Image conversion device and image decoding device

The image conversion and decoding devices optimize the relationship between slices and rectangular areas in VVC by adjusting loop filter application and using entry points, enhancing processing efficiency and flexibility in parallel video processing.

JP7755786B2Active Publication Date: 2025-10-17JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024176647
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-20
Filing Date
2024-10-08
Publication Date
2025-10-17
Estimated Expiration
2040-09-18

AI Technical Summary

Technical Problem

The processing efficiency of parallel processing tools in Versatile Video Coding (VVC) is inadequate, particularly in managing the relationship between slices and rectangular areas during image encoding and decoding.

Method used

An image conversion device and decoding device that divide images into blocks of predetermined size, adjusting the application of loop filters at slice and rectangular area boundaries to maintain or alter the relationship between slices and rectangular areas, and utilize entry points in NAL units to facilitate efficient parallel processing.

Benefits of technology

Enables high-efficiency parallel processing of video data by optimizing the relationship between slices and rectangular areas, allowing for flexible processing orders and reduced latency in image decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007755786000001
    Figure 0007755786000001
  • Figure 0007755786000002
    Figure 0007755786000002
  • Figure 0007755786000003
    Figure 0007755786000003
Patent Text Reader

Abstract

To provide an image decoding device that improves a processing efficiency of a parallel processing tool.SOLUTION: In an image decoding device, a conversion part that converts a coding picture containing a slice including an integer number of rectangular regions in each of which an image is divided into blocks of a predetermined size and one of more rows of blocks are collected, converts the coding picture so that a relation between the slice contained in the coding picture and the rectangular region is reconstructed if the coding picture is the same in a slice boundary and a rectangular region boundary in regard to an adaption of a loop filter, and converts the coding picture so that the relation between the slice contained in the coding picture and the rectangular region is not changed if the coding picture is not the same in the slice boundary and the rectangular region boundary in regard to the adaption of the loop filter.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image encoding and decoding technique that allows parallel processing. [Background technology]

[0002] HEVC defines tile and entropy synchronous coding as tools for parallel processing. Improvements to parallel processing tools are being considered for VVC (Versatile Video Coding), the next-generation video coding standard currently being formulated by the JVET (Joint Video Expert Team). [Prior art documents] [Non-patent literature]

[0003] Versatile Video Coding (Draft 6) Summary of the Invention

[0004] The technology in Non-Patent Document 1 is still in the process of being standardized, and there are issues with the processing efficiency of parallel processing tools.

[0005] In order to solve the above problem, an image conversion device of one embodiment of the present invention is an image conversion device that divides an image into blocks of a predetermined size and converts a coded picture including slices each including an integer number of rectangular areas each consisting of one or more rows of the blocks, and if whether or not a loop filter is applied to the coded picture is the same at slice boundaries and rectangular area boundaries, the image conversion device converts the coded picture so as to reconstruct the relationship between the slices and the rectangular areas included in the coded picture, and if whether or not a loop filter is applied to the coded picture is not the same at slice boundaries and rectangular area boundaries, the image conversion device converts the coded picture so as not to change the relationship between the slices and the rectangular areas included in the coded picture.

[0006] Another aspect of the present invention is an image decoding device that divides an image into blocks of a predetermined size and decodes a code string of slices, each slice including an integer number of rectangular areas, each of which is formed by grouping the blocks into one or more block rows, wherein the code string of the slices is included in an NAL unit, and entry points indicating start byte positions of the second and subsequent rectangular areas included in the slices are decoded from an NAL unit other than the NAL unit that includes the slice.

[0007] Yet another embodiment of the present invention is also an image decoding device that divides an image into blocks of a predetermined size and decodes a code sequence of slices, each slice including an integer number of rectangular areas, each of which is formed by grouping the blocks into one or more block rows, and when a flag indicating that the slice may include two or more rectangular areas indicates that the slice may include two or more rectangular areas, the device decodes, from a picture parameter set that sets picture parameters, a flag indicating whether a slice header includes entry points indicating start byte positions of the second and subsequent rectangular areas included in the slice.

[0008] According to the present invention, a parallel processing tool can be realized with high efficiency. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating a configuration of a first embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example in which a picture in a coded stream is divided into bricks. [Figure 3] A diagram explaining the syntax of a PPS regarding allocation information of CTUs to tiles and allocation information of CTU rows to bricks. [Figure 4] FIG. 10 is a diagram illustrating an example of a brick index. [Figure 5] FIG. 10 is a diagram illustrating a portion of the syntax of a PPS. [Figure 6] A diagram explaining the syntax of part of the slice header. [Figure 7] 10 is a flowchart illustrating the operation of a conversion unit. [Figure 8] A diagram showing the syntax of an SEI message. [Figure 9] A diagram showing the syntax of the SEI payload. [Figure 10] A diagram explaining the syntax of the entry point SEI message [Figure 11] FIG. 2 is a diagram illustrating an encoded stream output by a conversion unit. [Figure 12] FIG. 10 is a diagram illustrating a part of the syntax of a PPS according to a modification of the first embodiment. [Figure 13] FIG. 10 is a diagram illustrating a configuration of a second embodiment. [Figure 14] FIG. 2 is a diagram illustrating an example of a hardware configuration of a coding / decoding device according to a first embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] The embodiment of the present invention is based on the technology under consideration for VVC (Versatile Video Coding), a next-generation video coding standard currently being formulated by JVET (Joint Video Expert Team). VVC uses techniques such as block division, intra-prediction, inter-prediction, transform, quantization, entropy coding, and loop filtering, similar to conventional AVC and HEVC.

[0011] The following defines the techniques and technical terms used in the embodiments of the present invention. Hereinafter, unless otherwise specified, the terms "image" and "picture" are used interchangeably. In the embodiments of the present invention, encoding is performed in the image encoding unit, and decoding is performed in the image decoding unit.

[0012] <Coding tree units and coding blocks> A picture is divided equally into blocks of a predetermined size. These blocks of a predetermined size are defined as coding tree blocks (CTBs). Recursive division is possible within each coding tree block. The blocks to be coded and decoded after coding tree block division are defined as coding blocks (CUs). Intra prediction and inter prediction are coded (decoded) on a coding block basis. Furthermore, coding tree blocks and coding blocks are collectively defined as blocks. Appropriate block division enables efficient coding (decoding). The size of the coding tree block can be a fixed value agreed upon in advance between the encoder and decoder, or the size of the coding tree block determined by the encoder can be transmitted to the decoder.

[0013] In addition, in a picture, one luma coding tree block and two chroma coding tree blocks are defined as a coding tree unit (CTU) as a coding unit. There are CTUs that use one luma coding tree block as a unit and CTUs that use coding tree blocks for three color planes as a unit. The size of the CTU (CTB) is coded (decoded) using the SPS (Sequence Parameter Set).

[0014] <Tiles and Bricks> A tile is a rectangular area within a picture that consists of multiple CTUs. A brick is a rectangular area within a tile that consists of one or more CTU rows. A single CTU does not belong to multiple tiles; it belongs to only one tile. A single CTU does not belong to multiple bricks; it belongs to only one brick. A brick does not belong to multiple tiles; it belongs to only one tile. The allocation of CTUs to tiles and the allocation of CTU rows to bricks are set on a picture-by-picture basis. Information about the allocation of CTUs to tiles and the allocation of CTU rows to bricks is coded (decoded) using a Picture Parameter Set (PPS). As described above, a picture is divided into multiple tiles whose areas do not overlap. Furthermore, tiles are divided into multiple bricks whose areas do not overlap. Hereinafter, a rectangular area consisting of multiple CTUs in a picture will be referred to as a brick, but it is sufficient that this is a rectangular area consisting of one or more rows of CTUs in a tile, where a single CTU does not belong to multiple tiles but to only one tile, and a single CTU does not belong to multiple rectangular areas but to only one rectangular area.

[0015] With the parameter information of the SPS, which sets the parameters of the coded stream, and the PPS, which sets the parameters of the coded picture, and the slice header, each coded brick can be decoded independently. Therefore, with the SPS, PPS, and slice header, multiple coded bricks can be coded (decoded) simultaneously in parallel. A coded brick is defined as a code string of a brick generated by coding a brick. Hereinafter, a coded picture is defined as a code string of a picture generated by coding a picture, and a coded slice is defined as a code string of a slice generated by coding a slice.

[0016] Figure 3 is a diagram for explaining the syntax of PPS regarding the CTU allocation information to tiles and the CTU row allocation information to bricks. The brick_splitting_present_flag is a flag indicating whether to split a tile into two or more bricks. If the brick_splitting_present_flag is 1, one or more tiles are split into two or more bricks. If the brick_splitting_present_flag is 0, all tiles are not split into two or more bricks. If the brick_splitting_present_flag is 1, the splitting information to bricks for each tile is encoded (decoded).

[0017] When a tile is not split by bricks, one tile is treated as one brick. The tile index is an index assigned in raster scan order to the tiles in the picture. The brick index is an index assigned to the bricks in the picture. When the picture is split into multiple tiles, the brick index is an index assigned in raster scan order to each brick within the tile in tile index order for the tiles in the picture. The brick index is an integer value that is 0 or more and less than the number of bricks in the picture (NumBricksInPic). The raster scan order is the order of first scanning horizontally, then moving vertically and scanning horizontally again. Figure 4 is a diagram for explaining an example of the brick index. Figure 4(a) shows an example of 8 bricks in the picture, and Figure 4(b) shows an example of 5 bricks in the picture.

[0018] <CTU Processing Order> CTUs are processed in raster scan order within a brick, and bricks are processed in raster scan order within a tile.

[0019] <Slice and Slice Mode> A slice is an area that groups together an integer number of bricks. Slices have two modes for grouping blocks: rectangular slice mode and raster scan slice mode. Figure 5 is a diagram explaining part of the syntax of PPS. The rectangular slice mode and raster scan slice mode will be explained using Figure 5. If the flag (single_brick_per_slice_flag) indicating that a slice may contain two or more bricks indicates that the slice may contain two or more bricks (single_brick_per_slice_flag is 0), rect_slice_flag is coded (decoded). If a slice contains only one brick (single_brick_per_slice_flag is 1), rect_slice_flag is 0. If rect_slice_flag is 1, rectangular slice mode is used, and if rect_slice_flag is 0, raster scan slice mode is used.

[0020] <Rectangle slice mode> In rectangular slice mode, bricks within a rectangular area are grouped together as a slice. In rectangular slice mode, the order in which bricks are processed is not raster scan order, so the correspondence between slices and bricks cannot be uniquely determined. Therefore, the correspondence between slices and bricks can be obtained by encoding (decoding) information indicating the correspondence between slices and bricks using PPS.

[0021] The number of slices in a picture is obtained by subtracting 1 from the number of slices in the picture, and num_slices_in_pic_minus1 is coded (decoded) in the PPS to obtain the number of slices in the picture. The bottom_right_brick_idx_delta and brick_idx_delta_sign_flag, which indicate the bottom-right brick index included in each slice in the picture, are coded (decoded) in the PPS to derive the correspondence between slices and bricks in the picture.

[0022] As described above, in rectangular slice mode, by encoding (decoding) information indicating the correspondence between slices and bricks using PPS, it is possible to give flexibility to the order in which bricks are processed. For example, as shown in Figure 4(b), by assigning the large screen to the tile to which bricks 0 and 1 belong and the small screen to the other bricks, it is possible to prioritize the processing order of the large screen.

[0023] <Raster scan slice mode> In raster scan slice mode, bricks are sliced ​​in raster scan order. Therefore, in raster scan slice mode, the correspondence between the positions of bricks and slices does not need to be coded (decoded) using PPS. In raster scan slice mode, the information indicating the correspondence between slices and bricks is not coded (decoded) using PPS, allowing for flexibility in the correspondence between slices and bricks. For example, in Figure 4(b), the number of bricks included in a slice can be adjusted during encoding, such as whether the slice is composed of only brick 0 or two, brick 0 and brick 1. This makes it possible to adjust the code amount for a slice.

[0024] <Slice ID> The slice ID is an ID that uniquely identifies a slice. In raster scan slice mode, the slice ID of the ith slice[i] arranged in raster scan order is i, and the slice ID is a value greater than or equal to 0 and less than or equal to num_slices_in_pic_minus1. In rectangular slice mode, the slice ID can also be explicitly coded (decoded) as an arbitrary value within the PPS. In rectangular slice mode, if the slice ID is not explicitly coded (decoded) within the PPS, the slice ID of the ith slice[i] arranged in raster scan order is i, and the slice ID is a value greater than or equal to 0 and less than or equal to num_slices_in_pic_minus1.

[0025] <Loop filter boundary control flag> Here, the boundary control flags of the loop filter will be explained using the syntax in Figure 5. loop_filter_across_bricks_enabled_flag is a flag that indicates whether or not to apply the loop filter to brick boundaries; if it is 1, the loop filter is applied, and if it is 0, the loop filter is not applied. loop_filter_across_slices_enabled_flag is a flag that indicates whether or not to apply the loop filter to slice boundaries; if it is 1, the loop filter is applied, and if it is 0, the loop filter is not applied. When a brick boundary and a slice boundary overlap, if both flags are 1, the loop filter is applied, and if either or both flags are 0, the loop filter is not applied.

[0026] <slice header> Next, the slice header will be described. Figure 6 is a diagram explaining part of the syntax of the slice header. If the rectangular slice mode is selected or if NumBricksInPic, which indicates the number of bricks in a picture, is greater than 1, the slice address (slice_address) is coded (decoded). If the raster scan slice mode is selected and single_brick_per_slice_flag is 0, num_bricks_in_slice_minus1, which indicates the number of slices included in the slice, is coded (decoded).

[0027] In raster scan slice mode, the slice address indicates the first brick ID included in the slice. Brick IDs are assigned to CTUs included in each brick in a picture. CTUs belonging to the same brick have the same brick ID. Brick IDs are integer values ​​greater than or equal to 0 and less than the number of bricks in the picture (NumBricksInPic). In raster scan slice mode, the position of bricks in a picture is defined by the PPS, so the start position of a slice in a picture can be determined most quickly by setting the slice address to the first brick ID included in the slice. The bit length of slice_address is Ceil(Log2(NumBricksInPic)), and slice_address is a value greater than or equal to 0 and less than or equal to (NumBricksInPic-1). Here, in raster scan slice mode, information indicating the correspondence between slices and bricks is not encoded (decoded) in the PPS, so the slice address is set to the first brick ID included in the slice. In raster scan slice mode, bricks are included in a slice consecutively in raster scan order, so if you know the first brick ID and the number of slices included in the slice, you can uniquely determine the relationship between slices and bricks. Note that Ceil(x) is a function whose value is the smallest integer greater than or equal to x.

[0028] In rectangular slice mode, information indicating the correspondence between slices and bricks is coded (decoded) in the PPS. Therefore, the brick ID can be derived from the slice ID, and the slice address indicates the slice ID. Here, the slice address is used as the slice ID in rectangular slice mode, but it can also be the first brick ID included in the slice, as in raster scan slice mode. In rectangular slice mode, using the first brick ID included in the slice as the slice address, as in raster scan slice mode, eliminates the need for extra circuits and modules, thereby reducing the memory size and circuit scale of the encoder and decoder. Furthermore, if the signalled_slice_id_flag in the PPS is 0, slice ID management is not required, further reducing the memory size of the encoder and decoder.

[0029] <entry point> Next, the start byte position of an encoding brick in slice data will be described using Figure 6. As shown in Figure 6, the start byte position of an encoding brick in slice data is encoded (decoded) in the slice header as an entry point (entry_point_offset_minus1[i]). Therefore, a decoder can obtain the start byte position of the encoding brick by decoding the slice header. Here, NumEntryPoints is the number of bricks in a slice minus 1. The start byte position of the first encoding brick in slice data is 0, which is obvious, so it is not encoded (decoded). Instead, the start byte positions of the second and subsequent encoding bricks in the slice data are encoded (decoded). entry_point_offset_present_flag is a flag that determines whether to encode (decode) an entry point in the slice header. If entry_point_offset_present_flag is 1, the entry point is encoded (decoded). If entry_point_offset_present_flag is 0, the entry point is not encoded (decoded). offset_len_minus1 is a value that indicates the bit length of entry_point_offset_minus1. As shown in Figure 5, entry_point_offset_present_flag is coded (decoded) in PPS when single_tile_in_pic_flag is 0 or entropy_coding_sync_enabled_flag is 1. Here, single_tile_in_pic_flag is a flag indicating whether the picture is one tile. If single_tile_in_pic_flag is 1, the picture is composed of one tile, and if single_tile_in_pic_flag is 0, the picture is composed of multiple tiles. When a picture is composed of one tile, the tile is not allowed to be composed of multiple bricks, and one tile is composed of one brick. entropy_coding_sync_enabled_flag is a flag indicating whether parallel processing of CTU columns within a brick is permitted.If entropy_coding_sync_enabled_flag is 1, parallel processing of CTU columns within a brick is allowed, and if entropy_coding_sync_enabled_flag is 0, parallel processing within a brick is not allowed.

[0030] <sei> Supplemental enhancement information (SEI) is information that is not required for the process of generating pixels of a decoded image, but is necessary for the decoder to operate. For example, SEI includes a Buffering Period SEI, which is information for accurately operating a hypothetical reference decoder (HRD), and a Picture Timing SEI, which determines the output timing of a decoded image. SEI is encoded (decoded) as an SEI message.

[0031] <NALユニット> An NAL unit consists of a header and a payload. The payload type is coded (decoded) in the header, and the payload contains a code string of the type indicated by the payload type. Payload types include SPS, PPS, coded slice, and SEI message. For example, if the payload type indicates that a slice is included, the slice code string is stored in the payload, and if the payload type indicates that an SEI message is included, the SEI message code string is stored in the payload.

[0032] <Encoded stream> In VVC and HEVC, SPS, PPS, coded slices, SEI messages, etc. are stored in NAL units to form a coded stream. A coded slice consists of a slice header and slice data. A coding tree block is coded into the slice data. A coded stream contains one or more coded pictures.

[0033] <hrd> The HRD is a hypothetical reference decoder used to verify that a coded stream complies with the coding standard. The HRD defines a CPB (Coded Picture Buffer) for decoding the coded stream and a DPB (Decoded Picture Buffer) for outputting the decoded pictures. The HRD verifies that the CPB and DPB do not break when the coded stream is decoded based on the input / output timing of the CPB and DPB coded in the coded stream. If the CPB and DPB do not break, the coded stream complies with the coding standard. Type 1 HRD verifies that the coded slices of the coded stream conform to the standard. Type 1 HRD does not inspect NALs containing SPS, PPS, or SEI messages. The input / output timing of the CPB and DPB is encoded (decoded) in the SEI message.

[0034] (First embodiment) <Configuration of the first embodiment> A first embodiment of the present invention will be described. First, the configuration of this embodiment will be described. FIG. 1 is a diagram illustrating the configuration of the first embodiment. This embodiment is composed of a data server 2000 and an image analysis device 3000. The data server 2000 is composed of a storage unit 2001, a conversion unit 2002, and a transmission unit 2003. The data server 2000 has a function of converting encoded streams, and is therefore also called an image conversion device. The image analysis device 3000 is composed of a reception unit 3001, an image decoding unit 3002, and an image analysis unit 3003.

[0035] An encoded stream is stored in the storage unit 2001. The image decoding unit 3002 can decode the encoded stream stored in the storage unit 2001. The image decoding unit 3002 can process a maximum of eight bricks in parallel (hereinafter referred to as the maximum number of parallel brick processes).

[0036] <Operation of the First Embodiment> Next, the operation of this embodiment will be described. The operation of each component will be described with reference to Fig. 1. First, the operation of the data server 2000 and the image analysis device 3000 will be described.

[0037] Based on a request from the image analyzing device 3000, the data server 2000 reads out the coded stream and characteristic parameters of the image decoding unit 3002 from the storage unit 2001 and inputs them to the conversion unit 2002. The characteristic parameters are information indicating whether the image decoding unit 3002 supports the entry point SEI, and the characteristic parameters are assumed to be stored in advance in the storage unit 2001. The conversion unit 2002 converts the input coded stream and inputs the converted coded stream to the transmission unit 2003. The detailed operation of the conversion unit 2002 will be described later. The transmission unit 2003 transmits the input coded stream to the image analyzing device 3000.

[0038] The receiving unit 3001 inputs the coded stream input from the transmitting unit 2003 to the image decoding unit 3002. The image decoding unit 3002 decodes the input coded stream in parallel using the maximum number of bricks for parallel processing, outputs a decoded image, and inputs it to the image analysis unit 3003. Here, even if a frame rate for outputting the decoded image is set in the input coded stream, the image decoding unit 3002 outputs the decoded image regardless of the frame rate. The image analysis unit 3003 analyzes the input decoded image and outputs the analysis result. For example, the image analysis unit 3003 measures the image quality and recognizes faces, people, etc., of the input decoded image.

[0039] As described above, when analyzing a decoded image using the image analysis unit 3003, the analysis results output from the image analysis unit 3003 can be obtained in a short time by outputting the decoded image from the image decoding unit 3002 without being bound by the frame rate of the encoded stream.

[0040] <Encoded stream> The coded stream stored in the storage unit 2001 will now be described. Figure 2 is a diagram illustrating an example in which a picture in a coded stream is divided into bricks. For simplicity's sake, the picture is divided into tiles, but the tiles are not divided into bricks. In other words, one tile constitutes one brick. Figure 2(a) shows an example in which the screen is divided horizontally into eight bricks (tiles) B0 (T0) to B7 (T7) for each CTU row. Figure 2(b) shows an example in which the screen is divided vertically into eight bricks (tiles) B0 (T0) to B7 (T7) for each CTU column. Brick IDs 0 to 7 are assigned to bricks B0 to B7, respectively. In this example, the number of bricks in the picture, NumBricksInPic, is 8.

[0041] Here, one picture is assumed to consist of one slice. Therefore, slice 0 contains eight bricks, B0 to B7. Therefore, single_brick_per_slice_flag in the PPS is set to 0. The coded stream may be in raster scan slice mode or rectangular slice mode. However, entry_point_offset_present_flag, which is a flag indicating whether to code an entry point in the slice header, is set to 0.

[0042] Here, we will explain the effect of setting entry_point_offset_present_flag to 0. When entry_point_offset_present_flag is set to 1, it is necessary to encode the entry point obtained as a result of encoding the brick in the slice header. Even if the encoded stream is generated by encoding the bricks in a slice in parallel, the encoded stream cannot be transmitted until the number of bytes indicating the brick size is determined. Therefore, the encoded stream must be transmitted in slice units, and low latency can only be achieved in slice units. On the other hand, when entry_point_offset_present_flag is set to 0, it is not necessary to encode the brick entry point in the slice header, and the encoded stream can be transmitted in brick units, enabling low latency in brick units.

[0043] <Detailed operation of the conversion unit> Next, a detailed description will be given of the operation of the conversion unit 2002. Fig. 7 is a flowchart illustrating the operation of the conversion unit, which is executed for all coded pictures included in the coded stream.

[0044] It is checked whether the loop filter is applied on slice boundaries and brick boundaries (S100). Specifically, it is checked whether the values ​​of loop_filter_across_bricks_enabled_flag and loop_filter_across_slices_enabled_flag are the same. If the loop filter is applied on slice boundaries and brick boundaries and the value of loop_filter_across_bricks_enabled_flag are the same (YES in S100), it is checked whether the decoder is compatible with the entry point SEI (S101). If the loop filter is applied on slice boundaries and brick boundaries and the value of loop_filter_across_bricks_enabled_flag are the same (NO in S100), the slices are reconstructed (S104).

[0045] If the decoder is entry point SEI compatible (YES in S101), the entry point is added to the entry point SEI (S102). The detailed operation of adding the entry point to the entry point SEI will be described later. If the decoder is not entry point SEI compatible (NO in S101), the entry point is added to the slice header (S103).

[0046] Here, if the application of a loop filter is not the same between slice boundaries and brick boundaries, changing the bricks included in each slice will change the decoded image, requiring re-encoding. Therefore, if the application of a loop filter is not the same between slice boundaries and brick boundaries, the coded stream is converted so as not to change the configuration of bricks included in each slice.

[0047] On the other hand, if the application of a loop filter is the same for slice boundaries and brick boundaries, the decoded image does not change even if the bricks included in each slice are changed, and re-encoding is therefore not necessary. Therefore, if the application of a loop filter is the same for slice boundaries and brick boundaries, the coded stream is converted to change the configuration of bricks and slices.

[0048] Here, a step of checking whether the values ​​of loop_filter_across_bricks_enabled_flag and loop_filter_across_slices_enabled_flag are the same (S100) is added, taking into consideration decoders suitable for parallel processing of slices. When targeting general decoders, it is better to start with the process of checking whether the decoder is compatible with the entry point SEI (S101) without performing the process of checking whether the values ​​of loop_filter_across_bricks_enabled_flag and loop_filter_across_slices_enabled_flag are the same (S100).

[0049] <Adding entry points to slice headers> Here, the operation of adding entry points to slice headers will be described. To add entry points to slice headers, information on the entry points of bricks B1 to B7 is coded (decoded) into the coded stream according to the syntax of Fig. 6.

[0050] <Slice reconstruction> Next, the detailed operation of reconstructing a slice will be described. In the coded stream stored in the storage unit 2001, one slice contains eight bricks. This is reconstructed so that one slice contains one brick. Note that single_brick_per_slice_flag in the PPS is set to 1.

[0051] <Adding an entry point to the entry point SEI> Next, a detailed operation of adding an entry point to the entry point SEI will be described. To add an entry point to the entry point SEI, information on the entry points of bricks B1 to B7 is encoded (decoded) according to the syntax of Figures 8, 9, and 10.

[0052] Figure 8 shows the syntax of the SEI message. The payload type (payloadType) is derived from payload_type_byte. If the payload type is 128, the entry point SEI message is encoded (decoded) into the SEI payload (sei_payload). The entry point SEI message is the same as the entry point SEI.

[0053] Figure 9 shows the syntax of the SEI payload. In addition to the entry point SEI message, SEI messages such as the Buffering Period SEI message, Picture Timing, and Picture Timing SEI message are encoded (decoded) in the SEI payload. There are two NAL unit types that store SEI messages: PREFIX_SEI_NUT, which is encoded (decoded) before the coding slice, and SUFFIX_SEI_NUT, which is encoded (decoded) after the coding slice.

[0054] Figure 10 is a diagram explaining the syntax of the entry point SEI message. The syntax of the entry point SEI message will be explained using Figure 10. NumSlicesInPic is the number of slices in a picture, which can be obtained by PPS in rectangular slice mode, or by counting the number of slices in raster scan slice mode. offset_len_minus1[s] for the Sth slice and NumEntryPoints number of entry_point_offset_minus1 are coded (decoded). offset_len_minus1 and entry_point_offset_minus1 are the same as described above.

[0055] <Encoded stream after conversion> Fig. 11 is a diagram illustrating the coded stream output by the conversion unit 2002. Fig. 11(i) shows an example of an input coded stream. The stream is coded in the following order: SPS, PPS, coded slices constituting coded picture 0, coded slices 0, ... constituting coded picture 1, and coded slices constituting coded picture N.

[0056] Figure 11(a) shows an example of a coded stream in which an entry point has been added to a slice header. The following describes Figure 11(a). The NAL units of the SPS, PPS, and coded slice 0 form a coded stream. Coded slice 0 is composed of a slice header (Modified S-Header in Figure 11(a)) with an entry point added and slice data including eight coded bricks, from coded brick 0 (Brick 0 in Figure 11(a)) to coded brick 7 (Brick 7 in Figure 11(a)). The slice header with the entry point added is different from the slice header of the input coded stream. Note that the slice header with the entry point added is required to decode coded bricks 0 to 7.

[0057] Figure 11(b) shows an example of a coded stream in which slices have been reconstructed. The following explains Figure 11(b). The NAL units of the SPS, PPS, and coded slices 0 to 7 form a coded stream. Each coded slice i (i = 0, 1, ... 7) is composed of slice data including slice header i (New S-Header i in Figure 11(b)) and coded brick i (Brick i in Figure 11(b)). Here, slice header i is different from the slice header of the input coded stream. Note that slice header i is required to decode coded brick i.

[0058] In this way, parallelization in units of coded slices can be easily achieved by converting one coded slice into a coded picture including one coded brick and setting single_brick_per_slice_flag in the PPS to 1. In this case, an entry point is not required, and it is sufficient to be able to access in units of NAL units.

[0059] Here, the coded picture is converted so that one coding slice includes one brick, but this is not limiting and the bricks included in the slice may be rearranged according to the maximum number of brick parallel processes of the image decoding unit 3002. The coded picture may also be converted so that one coding slice includes multiple bricks. For example, if the maximum number of brick parallel processes of the image decoding unit 3002 is four, the coded picture is converted so that one coding slice includes two bricks.

[0060] Figure 11(c) shows an example of a coded stream in which an entry point is added to the entry point SEI. Figure 11(c) will be explained below. The NAL units of the SPS, PPS, coded slice 0, and SEI make up the coded stream. Coded slice 0 is composed of a slice header (S-Header in Figure 11(c)) and slice data including eight coded bricks, from coded brick 0 (Brick 0 in Figure 11(c)) to coded brick 7 (Brick 7 in Figure 11(c)). Here, the slice header is the same as the slice header of the input coded stream. Note that the slice header is required to decode coded bricks 0 to 7.

[0061] The size of the encoded stream after conversion when the entry point is added to the slice header is minimized compared to other methods. However, the slice header needs to be re-encoded with bit accuracy, which changes the size of the encoded slice.

[0062] The coded stream reconstructed from the coded slices consists of eight coded slices, so the size of the coded stream after conversion is larger than with other methods, but the structure of the coded stream is simpler because an entry point is no longer required. Therefore, although parallelization in coding brick units is not supported, it can be decoded in parallel by decoders that support parallelization in coding slice units.

[0063] For a coded stream in which an entry point has been added to the entry point SEI, the input coded stream remains unchanged; the entry point SEI simply needs to be inserted after the last coded slice in the coded picture, making it easy to generate a converted coded stream. In this way, by encoding (decoding) the entry point SEI after the coded slice, the entry point in the coded slice can be notified to the decoder while allowing flexibility in the relationship between slices and bricks in raster scan slice mode. Furthermore, because the coded stream is not modified, the coding stream conformance of Type 1 of the HRD is not affected. Therefore, there is no need to convert the coded stream while verifying HRD compliance, making it possible to reliably and easily convert the coded stream.

[0064] <Modification> A modification of the first embodiment will now be described. This modification differs from the present embodiment in the syntax of the PPS. Fig. 12 is a diagram explaining the syntax of part of the PPS of the modification of the first embodiment. The conditions for encoding (decoding) entry_point_offset_present_flag differ from those of Fig. 5 of the first embodiment. As shown in Fig. 12, entry_point_offset_present_flag is encoded (decoded) in the PPS when the flag indicating that a slice may contain two or more bricks indicates that a slice may contain two or more bricks (single_brick_per_slice_flag is 0) or when entropy_coding_sync_enabled_flag is 1.

[0065] As described above, by encoding (decoding) entry_point_offset_present_flag when single_brick_per_slice_flag is 0, even if the picture is composed of multiple tiles (single_tile_in_pic_flag is 0), if the slice contains only one brick, there is no need to encode entry_point_offset_present_flag. This eliminates the situation where entry_point_offset_present_flag is 1 when the slice contains only one brick, thereby reducing the amount of coding. Furthermore, by implicitly setting entry_point_offset_present_flag to 0 when entry_point_offset_present_flag does not exist, it is possible to define that no entry point exists when the slice contains only one brick. Furthermore, if entry_point_offset_present_flag is 0 in the slice header, there is no need to check the number of bricks included in the slice, thereby reducing the amount of processing.

[0066] (Second embodiment) <Configuration of the second embodiment> A second embodiment of the present invention will now be described. Fig. 13 is a diagram illustrating the configuration of the second embodiment. The second embodiment of the present invention differs from the first embodiment in that the data server 2000 is an image coding device 1000. Only the differences from the first embodiment will be described below.

[0067] The image decoding unit 3002 can decode the coded stream coded by the image coding unit 1002 .

[0068] Based on a request from the image analysis device 3000, the image encoding device 1000 reads image data and characteristic parameters of the image decoding unit 3002 from the storage unit 1001 and inputs them to the image encoding unit 1002. The image encoding unit 1002 encodes the input image data and inputs the encoded encoded stream to the transmission unit 1003.

[0069] The image encoding unit 1002 differs from the conversion unit 2002 in that it encodes the coded stream it outputs from image data. The coded stream output by the image encoding unit 1002 has the same structure as the coded stream output by the conversion unit 2002.

[0070] The following describes the operation of the image encoding unit 1002. The conversion unit 2002 obtains loop_filter_across_bricks_enabled_flag and loop_filter_across_slices_enabled_flag by decoding from the encoded stream, but the image encoding unit 1002 determines loop_filter_across_bricks_enabled_flag and loop_filter_across_slices_enabled_flag.

[0071] If whether to apply a loop filter is not the same on slice boundaries and brick boundaries, the image decoding unit 3002 checks whether it is an entry point SEI-compatible decoder. If whether to apply a loop filter is the same on slice boundaries and brick boundaries, the coded picture is coded so that one slice includes one brick.

[0072] If the image decoding unit 3002 is an entry point SEI-compatible decoder, it encodes the coded picture by using the entry point as the entry point SEI. If the image decoding unit 3002 is not an entry point SEI-compatible decoder, it encodes the coded picture by encoding the entry point into a slice header.

[0073] As described above, the image coding device 1000 can output a coded stream that can be decoded by the image decoding unit 3002 based on a request from the image analysis device 3000. Furthermore, the effects of the coded stream structure in this embodiment are the same as those in the first embodiment.

[0074] In all of the above-described embodiments, the coded stream output by the image coding device has a specific data format so that it can be decoded according to the coding method used in the embodiment. The coded stream may be provided by being recorded on a computer-readable recording medium such as an HDD, SSD, flash memory, or optical disk, or may be provided from a server via a wired or wireless network. Therefore, an image decoding device corresponding to this image coding device can decode a coded stream in this specific data format regardless of the providing means.

[0075] When a wired or wireless network is used to exchange coded streams between an image coding device and an image decoding device, the coded streams may be converted into a data format suitable for the transmission mode of the communication channel before being transmitted. In this case, a transmitting device is provided that converts the coded stream output by the image coding device into coded data in a data format suitable for the transmission mode of the communication channel and transmits the coded data to the network, and a receiving device is provided that receives the coded data from the network, restores the coded stream, and supplies it to the image decoding device. The transmitting device includes a memory that buffers the coded stream output by the image coding device, a packet processing unit that packetizes the coded stream, and a transmitting unit that transmits the packetized coded data via the network. The receiving device includes a receiving unit that receives the packetized coded data via the network, a memory that buffers the received coded data, and a packet processing unit that packetizes the coded data to generate a coded stream and provides it to the image decoding device.

[0076] When a wired or wireless network is used to exchange coded streams between an image coding device and an image decoding device, in addition to a transmitting device and a receiving device, a relay device may be provided that receives coded data transmitted by the transmitting device and supplies the coded data to the receiving device. The relay device includes a receiving unit that receives packetized coded data transmitted by the transmitting device, a memory that buffers the received coded data, and a transmitting unit that transmits the packetized coded data to the network. The relay device may further include a receiving packet processing unit that processes the packetized coded data into packets to generate a coded stream, a recording medium that stores the coded stream, and a transmitting packet processing unit that packetizes the coded stream.

[0077] Furthermore, a display unit for displaying images decoded by the image decoding device may be added to the configuration to form a display device.

[0078] Furthermore, an imaging unit may be added to the configuration, and the captured image may be input to the image encoding device, thereby forming an imaging device.

[0079] 14 shows an example of the hardware configuration of a coding / decoding device according to the present application. The coding / decoding device includes the configurations of the image coding device and image decoding device according to the embodiments of the present invention. The coding / decoding device 9000 includes a CPU 9001, a codec IC 9002, an I / O interface 9003, a memory 9004, an optical disk drive 9005, a network interface 9006, and a video interface 9009, and each unit is connected via a bus 9010.

[0080] The image encoding unit 9007 and the image decoding unit 9008 are typically implemented as a codec IC 9002. The image encoding process of the image encoding device according to the embodiment of the present invention is performed by the image encoding unit 9007, and the image decoding process of the image decoding device according to the embodiment of the present invention is performed by the image encoding unit 9007. The I / O interface 9003 is realized by, for example, a USB interface, and is connected to an external keyboard 9104, mouse 9105, etc. The CPU 9001 controls the encoding / decoding device 9000 so as to perform an operation desired by the user, based on user operations input via the I / O interface 9003. User operations via the keyboard 9104, mouse 9105, etc. include selecting whether to execute encoding or decoding functions, setting encoding quality, input / output destinations for the encoded stream, and input / output destinations for images.

[0081] When a user desires to play back images recorded on the disc recording medium 9100, the optical disc drive 9005 reads out an encoded bitstream from the inserted disc recording medium 9100 and sends the read encoded stream to an image decoding unit 9008 of the codec IC 9002 via a bus 9010. The image decoding unit 9008 performs image decoding processing on the input encoded bitstream in the image decoding device according to an embodiment of the present invention and sends the decoded image to an external monitor 9103 via a video interface 9009. The encoding / decoding device 9000 also has a network interface 9006 and is connectable to an external distribution server 9106 or a mobile terminal 9107 via a network 9101. When a user desires to play back images recorded on the distribution server 9106 or the mobile terminal 9107 instead of the images recorded on the disc recording medium 9100, the network interface 9006 acquires the encoded stream from the network 9101 instead of reading out the encoded bitstream from the input disc recording medium 9100. Furthermore, when the user wishes to play back the image recorded in memory 9004, image decoding processing is performed on the encoded stream recorded in memory 9004 in the image decoding device according to the embodiment of the present invention.

[0082] When the user desires to encode an image captured by an external camera 9102 and record the encoded stream in memory 9004, the video interface 9009 inputs the image from the camera 9102 and sends it to the image encoding unit 9007 of the codec IC 9002 via the bus 9010. The image encoding unit 9007 performs image encoding processing in the image encoding device according to an embodiment of the present invention on the image input via the video interface 9009 to create an encoded bitstream. The encoded bitstream is then sent to the memory 9004 via the bus 9010. When the user desires to record the encoded stream on a disc recording medium 9100 instead of the memory 9004, the optical disc drive 9005 writes the encoded stream to the inserted disc recording medium 9100.

[0083] It is also possible to realize a hardware configuration that has an image encoding device but not an image decoding device, or a hardware configuration that has an image decoding device but not an image encoding device. Such hardware configurations can be realized, for example, by replacing the codec IC 9002 with an image encoding unit 9007 or an image decoding unit 9008, respectively.

[0084] The above encoding and decoding processes may be realized not only as a transmission, storage, and receiving device using hardware, but also as firmware stored in a ROM (read-only memory) or flash memory, or as software for a computer, etc. The firmware program or software program may be provided by recording it on a computer-readable recording medium, or may be provided from a server via a wired or wireless network, or may be provided as data broadcasting on terrestrial or satellite digital broadcasting.

[0085] The present invention has been described above based on the embodiments. The embodiments are merely examples, and it will be understood by those skilled in the art that various modifications are possible in the combination of the components and treatment processes, and that such modifications are also within the scope of the present invention. [Industrial Applicability]

[0086] The present invention can be used in image encoding and decoding techniques. [Explanation of symbols]

[0087] 1000 image encoding device, 1001 storage unit, 1002 image encoding unit, 1003 transmission unit, 2000 image conversion device, 2001 storage unit, 2002 conversion unit, 2003 transmission unit, 3000 image analysis device, 3001 reception unit, 3002 image decoding unit, 3003 image analysis unit< / hrd> < / sei>

Claims

[Claim 1] An image decoding device that divides an image into blocks of a predetermined size and decodes a code sequence of a slice containing an integer number of rectangular areas each consisting of one or more block rows of the blocks, characterized in that when a flag indicating that the slice may contain two or more of the rectangular areas indicates that the slice may contain two or more of the rectangular areas, the image decoding device includes an entry point in the slice header indicating the starting byte position of the second or subsequent rectangular areas included in the slice, and when a flag allowing parallel processing of block rows within rectangular areas included in the slice indicates that parallel processing is allowed, the image decoding device includes an entry point in the slice header indicating the starting byte position of the second or subsequent rectangular areas included in the slice.

Citation Information

Patent Citations

  • Video coding with network abstraction layer units that include multiple encoded picture partitions

    US20130114735A1