Method and apparatus for transmitting slice segmentation information in image and video coding
By introducing new control syntax into the video bitstream, more flexible image segmentation and loop filtering across slice boundaries are achieved, solving the problem of inefficient image segmentation and filtering processing in the prior art, and improving the efficiency and image quality of video encoding and decoding.
Patent Information
- Application Number
- CN202510434244.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-01
- Filing Date
- 2021-04-01
- Publication Date
- 2025-08-12
AI Technical Summary
The existing video codec standards have problems such as inefficiency and insufficient flexibility in image segmentation and filtering processing, especially in the application of loop filtering in multi-slicing mode.
By introducing new control syntax into the video bitstream, image segmentation information is determined to allow for more flexible slicing and tile segmentation methods, and loop filtering across slice boundaries is enabled in multi-slicing situations.
It improves the efficiency and flexibility of video encoding and decoding, especially in multi-slicing mode, loop filtering can be applied more effectively, improving the quality of image reconstruction.
Smart Images

Figure CN120475150A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 003,362, filed April 1, 2020, which is hereby incorporated by reference in its entirety. Technical Field
[0003] The present invention generally relates to picture partitioning in video coding systems, and more particularly to signaling partitioning information associated with partitioning a picture into slices and tiles. Background Art
[0004] High-efficiency video coding (HEVC) is the latest international video codec standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) (Rec. ITU-T H.265 ISO / IEC 23008-2 version 3: High Efficiency Video Coding, April 2015). Figure 1 A block diagram of the HEVC codec system is provided. The input video signal is predicted using inter / intra prediction (110) by deriving a prediction signal (136) from an already coded image region. The prediction residual signal (116) is processed by a linear transform (transform, T) 118. The transform coefficients are quantized (quantized, Q) 120 and entropy encoded in the bitstream by an entropy encoder (122) along with other auxiliary information. After inverse transform (IT) 126 of the inverse quantization (IQ) 124 transform coefficients, a reconstructed signal (in Figure 1 The reconstructed signal is further processed by loop filtering (e.g., deblocking filter (DF) 130 and non-deblocking filter (NDF) 131) to remove coding artifacts. The decoded images are stored in a reference picture buffer (134) for use in predicting future images in the input video signal.
[0005] In HEVC, a coded image is partitioned into non-overlapping square block regions represented by associated coding tree units (CTUs). A coded image can be represented by a set of slices, each slice comprising an integer number of CTUs. Individual CTUs in a slice are processed in raster scanning order. Bi-predictive (B) slices can be decoded using intra prediction or inter prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Predictive (P) slices can be decoded using intra prediction or inter prediction, which uses up to one motion vector and reference indices to predict the sample values of each block. Intra (I) slices are decoded using only intra prediction.
[0006] A recursive quadtree (QT) structure can be used to partition a CTU into multiple non-overlapping coding units (CUs) to accommodate various local motion and texture characteristics. One or more prediction units (PUs) are designated for each CU. The prediction unit, together with the associated CU syntax, serves as the basic unit for sending prediction sub-information. The specified prediction process is used to predict the values of related pixel samples within the PU. The CU can be further partitioned using a residual quadtree (RQT) structure to represent the related prediction residual signal. The leaf nodes of the RQT correspond to transform units (TUs). A transform unit includes a transform block of luminance samples of size 8×8, 16×16, or 32×32, or a transform block of luminance samples of size 4×4, and two corresponding transform blocks of chrominance samples of an image in a 4:2:0 color format. An integer transform is applied to the transform block and the level values of the quantized coefficients and other auxiliary information are entropy encoded in the bitstream. Figure 2 An example of a block partition 210 (left) and its corresponding QT representation 220 (right) is shown. Solid lines represent CU boundaries and dashed lines represent TU boundaries.
[0007] The terms coding tree block (CTB), coding block (CB), prediction block (PB), and transform block (TB) are defined to specify a 2-D sample array of one color component associated with a CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. Similar relationships are valid for CUs, PUs, and TUs. Tree partitioning is generally applied to both luma and chroma, but some exceptions may apply when chroma reaches certain minimum sizes.
[0008] ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11's Joint Video Experts Team (JVET) are currently developing the next-generation video codec standard. The Versatile Video Coding (VVC) draft of JVET-Q2001 (B. Bross, J. Chen, S. Liu, "Versatile Video Coding (Draft 8)," Document of Joint Video Experts Team of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, JVET-Q2001, 17th Meeting: Brussels, BE, 7–17 January 2020) has adopted several promising new coding tools. Similar to HEVC, the VVC draft specified in JVET-Q2001 partitions the coded image into non-overlapping square block regions, represented by CTUs. Each CTU can be partitioned into one or more smaller CUs using a nested multi-type tree with quadtree splitting, binary splitting, and ternary splitting. The resulting CU partition can be square or rectangular.
[0009] In the VVC draft specified in JVET-Q2001, a tile is a sequence of CTUs covering a rectangular area of an image. The CTUs in a tile are scanned in raster scan order. An image is partitioned into one or more tile rows and one or more tile columns. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU columns in a tile of an image. Two slicing modes are supported, namely raster scan slicing mode and rectangular slicing mode, as indicated by the syntax element rect_slice_flag. In raster scan slicing mode, a slice includes a sequence of complete slices in a slice raster scan of an image. In rectangular slicing mode, a slice includes multiple complete tiles that together form a rectangular area of an image, or multiple consecutive complete columns of a tile that together form a rectangular area of an image. Within the rectangular area corresponding to the slice, the tiles in the rectangular slice are scanned in slice raster scan order. Figure 3 and Figure 4 Examples are provided for segmenting an image into tiles and slices in raster scan tiling mode and rectangular tiling mode, respectively. Figure 3 An example of segmenting an image with 18 x 12 luma CTUs into 12 tiles and 3 raster scan slices is shown. Each CTU is represented by a small rectangle surrounded by dashed lines, each tile is represented by a thin solid line, and each slice is represented by a gray area surrounded by thick lines. Figure 4 An example of partitioning an image with 18 x 12 luma CTUs into 24 tiles and 9 rectangular slices is shown. The syntax element slice_address specifies the raster scan tile index of the first tile in the slice. The value of the syntax element num_tiles_in_slice_minus1 plus 1 specifies the number of contiguous tiles in the current slice. Summary of the Invention
[0010] Disclosed are a method and apparatus for sending or parsing image segmentation information. According to the method, a current image is segmented into one or more slices and one or more tiles according to the image segmentation information. A control syntax is determined, and unless the image segmentation information indicates that a rectangular slice mode is selected and each sub-image is allowed to include more than one rectangular slice, and the current image includes only one rectangular slice in the current image, the control syntax is sent from the video bitstream on the encoder side or parsed from the video bitstream on the decoder side. If the image segmentation information indicates that multiple slices exist in the current image and the control syntax indicates that in-loop filtering is enabled, in-loop filtering is applied across slice boundaries. The syntax is signaled or parsed from the slice header level of the video bitstream corresponding to the target slice. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1An exemplary adaptive inter / intra video coding system is shown.
[0012] Figure 2 An example of block partitioning is shown, where the result of the block partitioning is shown on the left and the codec tree (also called partition tree structure) is shown on the right.
[0013] Figure 3 An example of segmenting an image with 18 x 12 luma CTUs into 12 tiles and 3 raster scan slices is shown.
[0014] Figure 4 An example of partitioning an image with 18 x 12 luma CTUs into 24 tiles and 9 rectangular slices is shown.
[0015] Figure 5 A flow chart of an exemplary video decoding system according to an embodiment of the present invention is shown, in which control syntax for controlling whether a loop filter is applied across slice boundaries is parsed if a set of conditions are met. DETAILED DESCRIPTION
[0016] The following description is of the best contemplated mode of implementing the present invention. The description is intended to illustrate the general principles of the present invention and should not be considered limiting. The scope of the present invention is best determined by reference to the appended claims.
[0017] In the VVC draft specified in JVET-Q2001, a tile is a sequence of CTUs covering a rectangular area of an image. The CTUs in a tile are scanned in raster scan order. An image is partitioned into one or more tile rows and one or more tile columns. A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU columns in a tile of an image. When a coded image is further partitioned into more than one slice or one tile using the syntax element no_pic_partition_flag equal to 0, the information used to derive the tile partitioning of the coded image is signaled by the syntax elements pps_log2_ctu_size_minus5, num_exp_tile_columns_minus1, num_exp_tile_rows_minus1 and tile_row_height_minus1 in the picture parameter set (PPS). When rectangular slice mode is used (i.e., rect_slice_flag = 1), the syntax element tile_idx_delta_present_flag may be signaled to specify whether the coded picture is partitioned into rectangular slice rows and rectangular slice columns in slice raster order. tile_idx_delta_present_flag equal to 0 indicates that the tile_idx_delta[i] syntax element is not in the PPS, and all pictures referencing the PPS are partitioned into rectangular slice rows and rectangular slice column rows in slice raster order. tile_idx_delta_present_flag equal to 1 indicates that the tile_idx_delta[i] syntax element (indicating the difference between the tile index of the current slice and the tile index of the next slice) may be present in the PPS, and all rectangular slices in the pictures referencing the PPS are specified in the order indicated by the value of tile_idx_delta[i], with i increasing in value. When the number of slices in the coded picture is equal to 1, the value of tile_idx_delta_present_flag is not signaled, and is inferred to be equal to 0.
[0018] According to one aspect of the present invention, when the number of tiles in a coded picture is equal to 1, there is only one tile index equal to 0 in the coded picture, and tile_idx_delta is not sent. In mode 1, the sending of tile_idx_delta_present_flag further depends on the number of tiles in the coded picture. In one embodiment, tile_idx_delta_present_flag is signaled only when the number of slices in the coded picture is greater than 1 and the number of tiles in the coded picture is greater than an integer threshold T, where T can be equal to 1, 2, or 3. When tile_idx_delta_present_flag is not sent, it is inferred to be equal to 0.
[0019] According to another aspect of the present invention, when the number of tile rows or tile columns in a coded image is equal to 1, the coded image is always partitioned into rectangular slice rows and rectangular slice columns (vertically or horizontally) in slice raster order.
[0020] In mode 2, sending tile_idx_delta_present_flag further depends on whether the number of tile rows or tile columns in the coded picture is greater than 1. In one embodiment, tile_idx_delta_present_flag is signaled only when the number of slices in the coded picture is greater than 1 and the number of tile rows or tile columns in the coded picture is greater than 1. When tile_idx_delta_present_flag is not sent, it is inferred to be equal to 0.
[0021] According to another aspect of the invention, a syntax control flag specifying whether loop filtering operations should be applied across slice boundaries is only relevant for coded pictures comprising more than one slice.In the proposed method, signaling of the control flag depends on whether the number of slices in the coded picture is greater than one.
[0022] In one example embodiment, according to the present invention, the video encoder specified in JVET-Q2001 is modified using Method 2 for sending the tile_idx_delta_present_flag. The modified syntax table for the PPS is provided in Table 1. In the proposed method, tile_idx_delta_present_flag is signaled only when the number of slices in the coded picture is greater than 1 and the number of tile rows or tile columns in the coded picture is greater than 1. When rectangular slicing mode is used and the number of coded slices is 1, the syntax element loop_filter_across_slices_enabled_flag, which specifies whether loop filtering operations can be performed across slice boundaries in pictures referencing the PPS, is not signaled.
[0023] Table 1 Syntax table modified based on JVET-Q2001 according to the proposed method
[0024]
[0025]
[0026]
[0027] Any of the methods proposed above can be implemented in an encoder and / or decoder. For example, any of the methods proposed above can be implemented in a high-level syntax encoding module or a high-level syntax decoding module of an encoder and / or decoder. Alternatively, any of the methods proposed above can be implemented as circuitry integrated into a high-level syntax encoding module of an encoder and / or a high-level syntax decoding module of a decoder. Any of the methods proposed above can also be implemented in an image encoder and / or decoder, where the resulting bitstream corresponds to a coded frame using only intra-image prediction.
[0028] Figure 5 A flow chart of an exemplary video decoding system according to an embodiment of the present invention is shown, in which control syntax for controlling whether a loop filter is applied across slice boundaries is parsed if a set of conditions are met. The steps shown in the flowchart can be implemented as program code that can be executed on one or more processors (e.g., one or more central processing units (CPUs)) on the encoder side and / or the decoder side. The steps shown in the flowchart can also be implemented based on hardware, such as one or more electronic devices or processors configured to perform the steps in the flowchart. According to the method, in step 510, a video bitstream including a current image is received. In step 520, image segmentation information is parsed from the video bitstream, wherein the current image is segmented into one or more slices and one or more tiles according to the image segmentation information. In step 530, a control syntax is determined. Unless the image segmentation information indicates that a rectangular slicing mode is selected and each sub-image is allowed to include more than one rectangular slice, and the current image includes only one rectangular slice in the current image, the control syntax is parsed from the video bitstream. In step 540, a reconstructed image is derived from the video bitstream, wherein the reconstructed image includes the image segmentation according to the image segmentation information. In step 550, if the image segmentation information indicates that multiple slices exist in the current image and the control syntax indicates that loop filtering is enabled, loop filtering is applied across slice boundaries.
[0029] It can be concluded accordingly that Figure 5 Flowchart of an exemplary video encoding system corresponding to the decoder in.
[0030] The flowchart shown is intended to illustrate an example of video encoding and decoding according to the present invention. Those skilled in the art may modify each step, rearrange the steps, split the steps, or combine the steps to practice the present invention without departing from the spirit of the present invention. In the present invention, specific syntax and semantics have been used to illustrate examples of implementations of the present invention. Those skilled in the art may implement the present invention by replacing the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0031] The above description is presented to enable those skilled in the art to implement the present invention provided in the context of a specific application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the specific embodiments shown and described, but rather to the widest range consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details are shown in order to provide a thorough understanding of the present invention. However, those skilled in the art will appreciate that the present invention can be implemented.
[0032] The embodiments of the present invention described above can be implemented in various hardware, software code, or a combination of both. For example, embodiments of the present invention can be integrated into one or more circuits of a video compression chip, or integrated into video compression software to perform the processing described herein. Embodiments of the present invention can also be implemented as program code to be executed on a digital signal processor (DSP) to perform the processing described herein. The present invention can also include many functions performed by computer processors, digital signal processors, microprocessors, or field programmable gate arrays (FPGAs). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines the specific methods embodied in the present invention. The software code or firmware code can be developed in different programming languages and in different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, software code styles and languages, as well as other means of configuring the code to perform the tasks according to the present invention, will not depart from the spirit and scope of the present invention.
[0033] The present invention may be embodied in other specific forms without departing from the spirit or essential characteristics of the present invention. The examples described are to be considered in all respects as illustrative only and not restrictive. The scope of the present invention is therefore indicated by the appended claims rather than the foregoing description. All changes that come within the meaning and range of equivalents of the claims are intended to be included within their scope.
Claims
1. A video sequence decoding method, the method comprising: receiving a video bitstream including a current image; Parsing image segmentation information from the video bitstream, wherein the current image is segmented into one or more slices and one or more tiles according to the image segmentation information; When at least one of the following conditions is met, a control syntax indicating whether to enable cross-slice boundary loop filtering is sent: rectangular slice mode is not used, each sub-image includes only one rectangular slice, and the current image includes more than one rectangular slice; deriving a reconstructed image from the video bitstream, wherein the reconstructed image includes an image segmentation according to the image segmentation information; and When the control syntax indicates that cross-slice boundary loop filtering is enabled, applying the loop filtering across the slice boundary is performed.
2. The video sequence decoding method according to claim 1, wherein: The control syntax is at the picture parameter set level of the video bitstream.
3. The video sequence decoding method according to claim 1, wherein: If the image segmentation information indicates that the rectangular slice mode is selected, it is inferred that the control syntax has a value indicating no loop filtering across slice boundaries, allowing each sub-image to include more than one rectangular slice, and the current image to include only one rectangular slice.
4. An apparatus for decoding a video sequence, the apparatus comprising one or more electronic circuits or processors configured to: receiving a video bitstream including a current image; Parsing image segmentation information from the video bitstream, wherein: dividing the current image into one or more slices and one or more tiles according to the image segmentation information; When at least one of the following conditions is met, a control syntax indicating whether to enable cross-slice boundary loop filtering is sent: rectangular slice mode is not used, each sub-image includes only one rectangular slice, and the current image includes more than one rectangular slice; deriving a reconstructed image from the video bitstream, wherein the reconstructed image includes image segmentation according to the image segmentation information; as well as When the control syntax indicates that cross-slice boundary loop filtering is enabled, applying the loop filtering across the slice boundary is performed.
5. A video sequence encoding method, the method comprising: receiving input data corresponding to the current image; Segmenting the current image into one or more slices and / or one or more tiles; Sending image segmentation information in the video bitstream; When at least one of the following conditions is met, a control syntax indicating whether to enable cross-slice boundary loop filtering is sent: rectangular slice mode is not used, each sub-image includes only one rectangular slice, and the current image includes more than one rectangular slice; as well as When the control syntax indicates that cross-slice boundary loop filtering is enabled, applying the loop filtering across the slice boundary is performed.
6. The video sequence encoding method according to claim 5, wherein: The control syntax is at the picture parameter set level of the video bitstream.
7. The video sequence encoding method according to claim 5, wherein: If the image segmentation information indicates that the rectangular slice mode is selected, it is inferred that the control syntax has a value indicating no loop filtering across slice boundaries, allowing each sub-image to include more than one rectangular slice, and the current image to include only one rectangular slice.
8. An apparatus for encoding a video sequence, the apparatus comprising one or more electronic circuits or processors configured to: receiving input data corresponding to the current image; Segmenting the current image into one or more slices and / or one or more tiles; Sending image segmentation information in the video bitstream; When at least one of the following conditions is met, a control syntax indicating whether to enable cross-slice boundary loop filtering is sent: rectangular slice mode is not used, each sub-image includes only one rectangular slice, and the current image includes more than one rectangular slice; as well as When the control syntax indicates that cross-slice boundary loop filtering is enabled, applying the loop filtering across the slice boundary is performed.