Video data processing method and apparatus for restricting block size in video coding

By introducing QTBT structure and RDO optimization in video encoding, limiting the size of the transform unit and segmenting the encoding tree unit and encoding unit, the high cost and low flexibility problems caused by large-size encoding units in HEVC are solved, and efficient video encoding is achieved.

CN114786009BActive Publication Date: 2025-07-08HFI INNOVATION INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210472310.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-01-12
Filing Date
2017-03-10
Publication Date
2025-07-08
Estimated Expiration
2037-03-10

AI Technical Summary

Technical Problem

When existing video encoding standards such as HEVC process high-resolution video data, the size of the encoding unit and the transformation unit is too large, resulting in high cost of realizing the transformation logic, and the existing block segmentation method is not flexible enough, which affects the encoding efficiency.

Method used

The quad-tree-binary tree (QTBT) structure is used to combine rate and distortion optimization (RDO) to segment the encoding tree unit and the encoding unit, limit the transform unit size, and divide it into smaller blocks through quad-tree or binary tree segmentation, reducing unnecessary segmentation flag transmission.

Benefits of technology

It improves the efficiency and flexibility of video encoding, reduces encoding complexity, and maintains encoding performance, and is suitable for effective encoding of high-resolution video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114786009B_ABST
    Figure CN114786009B_ABST
Patent Text Reader

Abstract

The present invention provides a video data processing method. The method includes: receiving input data associated with a current image, determining the size of a current coding tree unit or a current coding unit (CU) in the current coding tree unit, and if the size, width or height of the current coding tree unit or the current coding unit is greater than a threshold, the encoder or decoder divides the current coding tree unit or the current coding unit into multiple blocks until each block is not greater than the threshold. The current coding tree unit or the current coding unit is processed for prediction or compensation, and transformation or inverse transformation. The current coding tree unit is processed according to the current coding tree unit level syntax sent in the video bitstream. The encoder or decoder encodes or decodes the current coding tree unit. The threshold corresponds to the maximum supported transform unit size of the encoder or decoder.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This invention claims the priority of U.S. Provisional Patent Application No. 62 / 309,001, filed on March 16, 2016, with the title "Methods for pattern-based MV derivation for Video Coding", claims the priority of U.S. Provisional Patent Application No. 62 / 408,724, filed on October 15, 2016, with the title "Methods of coding unit coding", and claims the priority of U.S. Provisional Patent Application No. 62 / 445,284, filed on January 12, 2017, with the title "Methods of coding unit coding", the contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] This invention relates to block-based video data processing in video coding. Specifically, this invention relates to techniques for video coding and video decoding with restricted block sizes. Background Art

[0004] Video images encoded in the High Efficiency Video Coding (HEVC) standard can be divided into one or more slices, and each slice is divided into multiple coding tree units (CTUs). In the main profile of HEVC, the minimum and maximum sizes of the coding tree units are specified by syntax elements in the sequence parameter set (SPS). Each coding tree unit includes a luminance coding tree block (CTB), a corresponding chrominance coding tree block, and syntax elements. The size of the luminance coding tree block is selected from 16x16, 32x32, or 64x64, where the larger size can generally provide a better compression ratio to simplify or smooth texture regions. The coding tree units within a slice are processed in raster scan order.

[0005] HEVC supports dividing a coding tree unit into multiple coding units (CUs) to adapt to various local features using quadtree partitioning. For a 2Nx2N coding tree unit, it can be a single coding unit or can be divided into four smaller blocks of the same size (i.e., NxN). The quadtree partitioning process recursively divides each coding tree unit into smaller blocks until reaching the leaf nodes of the quadtree coding tree. The leaf nodes of the quadtree coding tree are called coding units in HEVC. The decision to encode an image region using inter (temporal) prediction or intra (spatial) prediction is made at the coding unit level. Since the minimum coding unit size can be 8x8, the minimum interval size (granularity) for switching between different basic prediction types is 8×8.

[0006] According to one of the prediction unit (PU) partitioning types as Figure 1 shown, one, two, or four prediction units are assigned to each coding unit. Among them, the prediction unit is the basic representative block for sharing prediction information. Figure 1 Eight different prediction unit partitioning types supported in HEVC are shown, including symmetric and asymmetric partitioning types. Inside each prediction unit, the same prediction process is applied, and relevant prediction information is sent to the decoder based on the prediction unit. After obtaining the prediction residual by applying the prediction process on the prediction unit, the coding unit is divided into transform units (TUs) according to another quadtree structure, which is similar to the quadtree coding tree used to obtain coding units from the largest coding unit. The transform unit is the basic representative block for the residual or transform coefficients to which the transform and quantization are applied. The transformed and quantized residual signal of the transform unit is encoded and transmitted to the decoder after being transformed and quantized on the basis of the transform unit.

[0007] Similar to the definition of the coding tree block, the coding block (CB), prediction block (PB), and transform block (TB) are defined as the sampling arrays specifying the luminance component or chrominance component related to the coding unit, prediction unit, and transform unit, respectively. The quadtree partitioning process is usually applied to both the luminance and chrominance components simultaneously, although there are exceptions when reaching certain minimum sizes of the chrominance component. Summary of the Invention

[0008] The present invention discloses a method and apparatus for processing video data in an encoding system and a decoding system. An embodiment of the encoding system receives input data associated with a current image of video data and determines a size of a current coding tree unit for transmitting one or more current coding tree unit level grammars. In some embodiments, the current coding tree unit includes one or more current coding units. A size, width, or height of the current coding tree unit or the current coding unit is checked using a threshold, and if the size, width, or height of the current coding tree unit or the current coding unit is greater than the threshold, the encoding system forces the current coding tree unit or the current coding unit to be divided into a plurality of blocks until a size, width, or height of each block in the current coding tree unit or the current coding unit is not greater than the threshold. The encoding system processes the current coding tree unit or the current coding unit through prediction and transformation. The encoding system also determines a coding tree unit level grammar of the current coding tree unit and processes the current coding tree unit according to the coding tree unit level grammar. The current coding tree unit is encoded to form a video bitstream, and the coding tree unit level grammar is transmitted in the video bitstream.

[0009] An embodiment of the video decoding system receives a video bitstream associated with a current coding tree unit in a current image and parses one or more coding tree unit level grammars and residuals of the current coding tree unit from the video bitstream. The current coding tree unit may include one or more current coding units. A size of the current coding tree unit or the current coding unit is determined and compared with a threshold, and if the size, width, or height of the current coding tree unit or the current coding unit is greater than the threshold, the decoding system infers that the current coding tree unit or the current coding unit is divided into a plurality of blocks until a size, width, or height of each block in the current coding tree unit or the current coding unit is not greater than the threshold. The current coding tree unit or the current coding unit is processed through prediction and inverse transformation, and the current coding tree unit is decoded according to one or more coding tree unit level grammars parsed from the video bitstream.

[0010] Embodiments for determining a threshold for further splitting of a current coding tree unit or a current coding unit correspond to the maximum supported transform unit size, and an example of the maximum supported transform unit size is 128×128. The maximum supported transform unit size may be sent at the sequence level, picture level, or slice level in a video bitstream to notify the decoder side. In some embodiments, if the size, width, or height of the current coding tree unit or the current coding unit is greater than the threshold, quadtree splitting or binary tree splitting is used to split the current coding tree unit or the current coding unit. In one embodiment, when splitting the current coding tree unit or the current coding unit using binary tree splitting, if the width or size of the current coding tree unit or the current coding unit is greater than the threshold, the current coding tree unit or the current coding unit is split by vertical splitting; if the height or size of the current coding tree unit or the current coding unit is greater than the threshold, the current coding tree unit or the current coding unit is split by horizontal splitting. In one embodiment, a split flag is sent in the video bitstream to indicate the splitting of the current coding tree unit or the current coding unit.

[0011] Some embodiments of the coding tree unit level syntax are various syntaxes for loop filtering processes such as sample adaptive offset (SAO), adaptive loop filtering (ALF), and deblocking filtering, and if the corresponding loop filtering process is applied to the current coding tree unit, the coding system may send one or more coding tree unit level syntaxes of the current coding tree unit. If the corresponding sample adaptive offset, adaptive loop filtering, and / or deblocking filter process is applied to the current coding tree unit, the decoding system may parse one or more coding tree unit level syntaxes of the current coding tree unit.

[0012] In one embodiment, the current picture is divided into coding tree unit groups, and the coding tree unit level syntax of the current coding tree unit is shared with other coding tree units in the same coding tree unit group. Examples of coding tree unit groups include MxN coding tree units, where M and N are positive integers, and the size of the coding tree unit group is a predefined value or is sent at the sequence level, picture level, or slice level in the video bitstream. The coding tree unit level syntax shared by the coding tree unit group may be one or a combination of the syntaxes of sample adaptive offset, adaptive loop filtering, and deblocking filtering.

[0013] In another embodiment, according to the position of the current coding tree unit, the width and height of the current coding tree unit or the current coding unit, and the width and height of the current picture, quadtree splitting, vertical binary tree splitting, or horizontal binary tree splitting is used to split the current coding tree unit or the current coding unit at the picture boundary.

[0014] In some embodiments of a video encoding or decoding system, if the size, width, or height is greater than a threshold, the current coding tree unit or the current coding unit is implicitly divided into multiple blocks, and then the video encoding or decoding system determines whether to further divide the current coding tree unit or the current coding unit according to the division decision. The division decision can be determined based on rate and distortion optimization (RDO) on the encoder side, and the video encoding system sends an information signal associated with the division decision in the video bitstream to indicate the division of the current coding tree unit or the current coding unit. The video decoding system decodes the division decision and divides the current coding tree unit or the current coding unit according to the decoded division decision. The blocks generated by the division decision are processed through prediction and transformation.

[0015] In some other embodiments, if the size, width, or height is not greater than the threshold, the current coding tree unit or the current coding unit is processed as a single block, and prediction processing and transformation processing are performed on each block in the current coding tree unit or the current coding unit without further division. In another embodiment, the video encoding system determines the size of the current coding unit according to the division decision, and if the size, width, or height is greater than the threshold, the current coding unit is divided into multiple blocks. If the current coding unit is divided into multiple blocks, prediction processing is performed on the current coding unit, and transformation processing is performed on each block in the current coding unit.

[0016] One aspect of the present disclosure further provides an apparatus for a video encoding system and an apparatus for a video decoding system to process video data with a restricted block size. The apparatus compares the size of the current coding tree unit or the current coding unit with the threshold, and if the size, width, or height of the current coding tree unit or the current coding unit is greater than the threshold, divides the current coding tree unit or the current coding unit.

[0017] One aspect of the present disclosure also provides a non-transitory computer-readable medium that stores program instructions for causing a processing circuit of an apparatus to execute a video encoding method or a video decoding method. The video encoding or decoding method includes checking whether the current coding tree unit or the current coding unit has a size, width, or height greater than the threshold, and dividing the current coding tree unit or the current coding unit according to the generated blocks with a size not greater than the threshold. For example, coding tree unit level syntax such as sample adaptive offset (SAO), adaptive loop filtering (ALF), and deblocking filtering is sent for the current coding tree unit or a coding tree unit group including multiple coding tree units. After reading the following detailed description of the embodiments of the present invention, other aspects and features of the present invention will be clear to those skilled in the art. Description of the Drawings

[0018] Various embodiments of the present disclosure presented as examples will be described in detail with reference to the following diagrams, where like diagrammatic references denote like components, and where:

[0019] Figure 1 Eight prediction unit (PU) partition types supported in High Efficiency Video Coding (HEVC) are marked out.

[0020] Figure 2A An exemplary block partition according to a Quad - Tree - Binary - Tree (QTBT) structure is shown.

[0021] Figure 2B Is shown corresponding to Figure 2A The coding tree structure of the block partition.

[0022] Figure 3 A flowchart of an exemplary video coding system according to an embodiment of the present invention is shown.

[0023] Figure 4 A flowchart of an exemplary video encoding system according to an embodiment of the present invention is shown.

[0024] Figure 5 A flowchart of an exemplary video decoding system according to an embodiment of the present invention is shown.

[0025] Figure 6 A block partitioning method for blocks at an image boundary according to an embodiment of the present invention is shown.

[0026] Figure 7 A schematic system block diagram for a video encoding system according to an embodiment of the present invention is shown.

[0027] Figure 8 A schematic system block diagram for a video decoding system according to an embodiment of the present invention is shown. Detailed Description of the Invention

[0028] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, can be arranged and designed in a wide variety of different configurations. Accordingly, the following more detailed description of the embodiments of the system and method of the present invention as shown in the figures is not intended to limit the scope of the present invention as claimed, but is merely representative of selected embodiments of the present invention.

[0029] References throughout this specification to "an embodiment", "some embodiments", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrases "in one embodiment" or "in some embodiments" throughout this specification are not necessarily all referring to the same embodiment, and these embodiments can be implemented individually or in combination with one or more other embodiments.

[0030] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the present invention can be implemented without one or more of the specific details, or with other methods, components, etc. In other cases, well-known structures or operations are not shown or described in detail to avoid obscuring the present invention.

[0031] Different from the quadtree structure supported by conventional video coding standards such as HEVC, a binary tree block partitioning structure can be used to partition a video image or a slice of video data. Depending on the partitioning type, such as symmetric horizontal partitioning and symmetric vertical partitioning, a block can be recursively partitioned into smaller blocks. For the block under consideration, a flag is used to indicate whether the block is partitioned into two smaller blocks. If the flag indicates partitioning, another syntax element is emitted to indicate which partitioning type is used, such as vertical partitioning or horizontal partitioning. The binary tree partitioning process can be repeated until the size (width or height) of the partitioned block reaches the minimum allowable block size (width or height).

[0032] The binary tree block partitioning structure is more flexible than the quadtree structure because more partitioning shapes are allowed. However, the coding complexity will also increase with the selection of the optimal partitioning shape. To balance the complexity and coding efficiency, the binary tree structure is combined with the quadtree structure, called the Quad-Tree-Binary-Tree (hereinafter simply referred to as QTBT) structure. Compared with the quadtree structure in HEVC, the QTBT structure shows satisfactory coding performance.

[0033] An exemplary QTBT structure is shown in Figure 2A . Wherein the larger block is first partitioned by the quadtree structure and then by the binary tree structure. Figure 2A Examples of block partitioning according to the QTBT structure are marked, Figure 2B and a tree diagram of the QTBT structure corresponding to the block partitioning shown in Figure 2A is shown. Figure 2A and 2B The solid lines in represent quadtree partitioning. Figure 2A and 2BThe dotted lines therein represent the binary tree partitioning. In each partitioning (i.e., non-leaf) node of the binary tree structure, a flag indicates which partitioning type (horizontal or vertical) is used, where 0 represents horizontal partitioning and 1 represents vertical partitioning. The QTBT structure partitioning Figure 2A The larger blocks therein are multiple smaller blocks, and these smaller blocks are further processed by prediction and transform coding. These smaller blocks are not further divided to form prediction units and transform units of different sizes. In other words, each block generated by the QTBT partitioning structure is the basic unit for prediction and transform processing. For example, Figure 2A the large block therein is a CTU of size 128x128, the minimum allowable size of a quadtree leaf node is 16x16, the maximum allowable size of a binary tree root node is 64x64, the minimum allowable width or height of a binary tree leaf node is 4, and the minimum allowable binary tree depth is 4. In this example, the CTU is first partitioned by the quadtree structure, and the leaf quadtree blocks can have sizes ranging from 16×16 to 128×128. If the leaf quadtree block is 128×128, it cannot be further partitioned by the binary tree structure because its size exceeds the maximum allowable size of a binary tree root node, which is 64x64. The leaf quadtree block is used as the root binary tree block with a binary tree depth equal to 0. When the binary tree depth reaches 4, non-partitioning is implicit; when the width of a binary tree node is equal to 4, non-vertical partitioning is implicit; and when the height of a binary tree node is equal to 4, non-horizontal partitioning is implicit. This QTBT structure can be applied to the luminance and chrominance components of I slices (i.e., intra-coded slices) separately, and to the luminance and chrominance components of P slices and B slices simultaneously. When encoding blocks in an I slice, the QTBT structure block partitioning of the luminance component and the two chrominance components may all be different.

[0034] To effectively encode and decode high-resolution video data, such as using 8K×4K video with 8000 pixels × 4000 pixels to represent each image, the next-generation video coding allows for larger basic representative blocks for video coding processing, such as prediction, transformation, and loop filters. HEVC recursively decomposes each coding tree unit into coding units, and prediction units and transformation units are respectively segmented from the corresponding coding units according to the prediction unit segmentation type and the quadtree segmentation structure. To simplify the block segmentation method of HEVC, each leaf node of the segmentation structure such as the quadtree structure or the QTBT segmentation structure is set as the basic representative block for prediction and transformation processing, so no further segmentation is required. In this case, the coding unit is equal to the prediction unit and also equal to the transformation unit. In some embodiments, the coding tree unit is equal to the coding unit, the prediction unit, and the transformation unit. Increasing the size of the coding tree unit and the coding unit means that the sizes supported by prediction and transformation also need to increase. For example, when the size of the coding tree unit is 256×256, the prediction unit size and the transformation unit size are also 256×256. However, as the transformation unit size increases linearly, the number of silicon gates of the transformation logic grows exponentially. The number of silicon gates of the transformation logic with a transformation unit size equal to 256x256 is too large and too costly to implement. The following embodiments provide some solutions to the problems caused by allowing large block sizes for video coding. Some embodiments of the present invention are applied to process video data with a unified block segmentation for prediction and transformation processing and video data with a restricted block size for transformation and inverse transformation. Some other embodiments are applied to process video data with a separate block segmentation for prediction and transformation processing but with a simplified block segmentation.

[0035] The first embodiment is an encoder that only limits the adjustment of the encoder to generate transform units with supported sizes for use in the transform logic hardware implementation. The encoder of the first embodiment determines the size of each coding unit (CU) or prediction unit (PU) for prediction processing and transform processing. Examples of coding unit or prediction unit size determination include performing rate-distortion optimization (RDO) using a splitting method such as quadtree splitting, binary tree splitting, or QTBT splitting to select the optimal size. According to an embodiment of the present invention, a splitting decision is determined to further split the current coding tree unit or current coding unit based on the rate and distortion optimization results, and information associated with the splitting decision in the video bitstream is sent to indicate the splitting of the current coding tree unit or current coding unit. The encoder then checks with a threshold whether the size MxN, width M, or height N of the current coding unit or prediction unit is greater than the threshold. For example, the threshold for the size is 128×128, and the threshold for M or N is 128. The threshold can be predefined or user-defined, and the threshold can be determined based on the maximum supported transform unit size. An example of the encoder sends the maximum supported transform unit size at the sequence-level, picture-level, or slice-level of the video bitstream, such as the maximum supported transform unit size sent in the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), or slice header. If the size, width M, or height N is greater than the threshold, the encoder forces the current coding unit or prediction unit to be divided into multiple transform units according to quadtree splitting or binary tree splitting. When the size, width M, or height N is greater than the threshold, the encoder can skip sending the transform unit splitting flag signal in the video bitstream to the decoder because the decoder also forces the splitting of the current coding unit or prediction unit. In another example, the encoder can send the transform unit splitting flag in the video bitstream, but the transform unit splitting flag is restricted to 1, indicating that the splitting is performed on the encoder side. If the size of the coding unit or prediction unit is greater than the threshold, the size of the transform unit is different from the size of the coding unit or prediction unit in this embodiment; if the size of the coding unit or prediction unit is not greater than the threshold, the size of the transform unit is the same as the size of the coding unit or prediction unit. In one example, if the size, M, or N of the current coding unit or prediction unit is greater than the threshold and quadtree splitting is used, the current coding unit or prediction unit is divided into four transform units, and the encoder may or may not send the transform unit splitting flag.In another example, if the width M is greater than a threshold and binary tree splitting is used, the current coding unit or prediction unit is split into two transform units by vertical splitting, and the encoder may or may not send a transform unit split flag indicating vertical splitting or the encoder may send a vertical transform unit split flag. If the height N is greater than a threshold and binary tree splitting is used, the current coding unit or prediction unit is divided into two transform units by horizontal splitting, and the encoder may or may not send a transform unit split flag indicating the use of horizontal splitting, or the encoder may send a horizontal transform unit split flag. Compared with the HEVC standard with two separate block splitting structures for coding units and transform units, the first embodiment of the present invention only sends one block splitting structure to the coding unit and the transform unit.

[0036] The second embodiment is a canonical solution for a decoder to decode video data with a larger block size. The decoder according to the second embodiment decodes the video bitstream and determines the block splitting of the current image based on information such as the splitting flags sent in the video bitstream. According to one embodiment of the present invention, each block in the current coding tree unit or the current coding unit is further split based on the decoded splitting flags sent in the video bitstream. When the decoder is processing the current coding unit or prediction unit, it checks the size MxN of the current coding unit or prediction unit. If the size is greater than a threshold, or if the width M or the height N is greater than a threshold, it is inferred that the transform unit is to be split without involving the splitting flag. Examples of the threshold are a size of 128x128, or a width or height of 128. The threshold can be determined by parsing relevant information in the video bitstream, or the threshold can be predefined or user-defined. In one example, if quadtree splitting is applied for splitting, if the size of the current coding unit or prediction unit, the width M or the height N, is greater than the threshold, since the transform unit split flag is presumed to be 1, the current coding unit or prediction unit is split into four transform units. In another example, if binary tree splitting is applied for splitting, if the width M or the height N is greater than the threshold, since the vertical or horizontal transform unit split flag is inferred to be 1, the current coding unit or prediction unit is divided into two transform units.

[0037] In the first and second embodiments, if the size, width, or height of the current coding unit or prediction unit is greater than a threshold, the encoder or decoder only forces the splitting of the transform unit rather than the current coding unit or prediction unit. In other words, if the size, width, or height of the current coding unit or prediction unit is not greater than the threshold, the size of the current coding unit or prediction unit is the same as the size of the current transform unit corresponding to the current coding unit or prediction unit. If the size, width, or height of the current coding unit or prediction unit is greater than the threshold, the size of the current coding unit or prediction unit is different from the size of the current transform unit corresponding to the current coding unit or prediction unit.

[0038] Figure 3 A flowchart of an exemplary video coding and decoding system incorporating the first or second embodiment of the present invention is shown. In step S31, the video coding system receives input data associated with a current block in a current image, and the current block can be a current coding unit or a current prediction unit. Step S32 determines the size of the current block, and step S33 checks whether the size, width, or height of the current block is greater than a threshold. An example of the threshold corresponds to the maximum supported transform unit size. If the size, width, or height of the current block is greater than the threshold, in step S34 the video coding system splits the current block into a plurality of blocks having sizes not greater than the threshold, width, or height. The blocks split from the current block are processed by transform or inverse transform in step S35, and the current block is processed by prediction. If the size, width, or height of the current block is not greater than the threshold, step S36 processes the current block by prediction, and transform or inverse transform. In step S37, the video coding system encodes or decodes the current block.

[0039] The third embodiment not only forces the encoder to split a larger coding unit or prediction unit into multiple transform units, but also forces the coding tree unit or coding unit to be split according to the comparison result. The encoder of the third embodiment first determines the size of each coding tree unit for transmitting the coding unit level syntax. One or more coding tree unit level syntaxes are determined for each coding tree unit and transmitted in the coding tree unit level of the video bitstream. Some examples of the coding tree unit level syntax include syntaxes specified for loop filters such as sample adaptive offset (SAO), adaptive loop filter (ALF), and deblocking filter. For example, the coding tree unit level syntax may include one or a combination of SAO type, SAO parameters, ALF parameters, and deblocking filter boundary strength. The video data in the coding tree unit shares the information in the coding tree unit level syntax. Then, the encoder checks the size MxN, width M, or height N of the current coding unit or the current coding tree unit in the current coding tree unit by comparing the size, width, or height with a threshold. If the size, width, or height is greater than the threshold, the video encoder splits the current coding tree unit or coding unit into smaller blocks until the size, width, or height of the smaller blocks is not greater than the threshold. The threshold can be determined by a predefined or user-defined maximum supported transform unit size. The threshold can be sent in the video bitstream to notify the decoder of forced splitting when the corresponding coding tree unit or coding unit is greater than the threshold. For example, the maximum supported transform unit size is sent in the sequence level, picture level, or slice level of the video bitstream. When the size, width, or height is greater than the threshold, an example of splitting the current coding tree unit or coding unit is to divide the current coding tree unit or coding unit into four blocks according to quadtree splitting. When the size or width is greater than the threshold, an example of splitting the current coding tree unit or coding unit is to split the current coding tree unit or coding unit into two blocks according to vertical splitting, or when the size or height is greater than the threshold, an example of splitting the current coding tree unit or coding unit is to split the current coding tree unit or coding unit into two blocks according to horizontal splitting. The video encoder determines whether to further split the current coding tree unit or coding unit or the smaller blocks by performing RDO to decide the best block size for prediction and transform processing. The video encoder performs prediction processing and transform processing on each block in the current coding tree unit or coding unit, and transmits the splitting decision in the video bitstream.

[0040] Figure 4A flowchart of an exemplary video coding system for processing video data with a restricted block size according to a third embodiment of the present invention is shown. In step S41, the video coding system receives input data associated with a current image, and in step S42, determines the size of a current coding tree unit for transmitting one or more coding tree unit level grammars. In step 43, the size, width, or height of a current coding unit or the current coding tree unit in the current coding tree unit is checked to determine whether it is greater than a threshold. If the check result in step S43 is affirmative, then in step S44, the current coding tree unit or the current coding unit is divided into blocks until the size, width, or height of each block is not greater than the threshold. If the check result in step S43 shows that the size, width, or height of the current coding tree unit or the current coding unit is not greater than the threshold, the video coding system proceeds to step S46. In step S46, the video coding system determines a splitting decision to further split the current coding tree unit or the current coding unit, and processes each block in the current coding tree unit or the coding unit through prediction and transformation. The splitting decision may be determined by the RDO result, which selects an optimal block splitting for prediction and transformation processing. In step S47, the video coding system determines the coding tree unit level grammar and applies it to the current coding tree unit, and in step S48, encodes the current coding tree unit to form a video bitstream and transmits the coding tree unit level grammar in the video bitstream.

[0041] When the size MxN, width M, or height N of the current coding tree unit or coding unit is greater than a threshold, the fourth embodiment forces the decoder to split the current coding tree unit or coding unit into smaller blocks. The decoder decodes information related to the split decision, such as a split flag from the video bitstream, and further splits the current coding tree unit or coding unit or the smaller blocks according to the split decision. Then each block in the current coding tree unit or coding unit is processed by prediction and inverse transformation in the decoder of this embodiment. The threshold can be predefined or user-defined, and in one embodiment, the threshold is determined by parsing relevant information in the video bitstream. The threshold corresponds to the maximum supported transform unit size, and some examples of the threshold are 128×128 for the transform unit size, 128 for the transform unit width or transform unit height, 64x64 for the transform unit size, and 64 for the transform unit width or transform unit height. When the decoder determines that the size MxN, width M, or height N of the current coding tree unit or coding unit is greater than the threshold, it infers that the current coding tree unit or coding unit is split without decoding the split flag. In other words, when the size, width, or height is greater than the threshold, when using quadtree splitting to split the current coding tree unit or coding unit into four smaller blocks, the quadtree split flag of the current coding tree unit or coding unit is inferred to be 1 (i.e., split). In another example, when the size or width is greater than the threshold, when using binary tree splitting to split the current coding tree unit or coding unit into two smaller blocks, the vertical split flag is inferred to be 1 (i.e., split). When the size or height is greater than the threshold, when using binary tree splitting to split the current coding tree unit or coding unit into two smaller blocks, the horizontal split flag is inferred to be 1 (i.e., split). Although the decoder performs decoding processes such as prediction and inverse transformation in each block of the current coding tree unit, the coding tree unit level syntax parsed from the video bitstream is used for all blocks in the current coding tree unit.

[0042] The third and fourth embodiments split the current coding tree unit at the root-level, or split the current coding unit, until the current coding tree unit or coding unit is not greater than the threshold. In these two embodiments, the basic representative blocks for applying prediction and transform processing are always the same. The encoder and decoder according to the third and fourth embodiments allow large coding tree units (e.g., 256×256) for sending the coding tree unit level syntax, while limiting the block size for transform or inverse transform processing to be less than or equal to the threshold (e.g., 128×128). In some embodiments, the split flag is not sent in the video bitstream to specify the split of coding tree units or coding units greater than the threshold, and the encoder and decoder infer the split of large coding tree units or coding units, thus saving the bits required for sending the split flag and achieving better coding efficiency.

[0043] Figure 5 A flowchart of an exemplary video decoding system including a video data processing method with a restricted block size according to a fourth embodiment of the present invention is shown. In step S51, the video decoding system receives a video bitstream including a current image, and in step S52, parses the coding tree unit level syntax and residuals of a current coding tree unit in the current image from the video bitstream. In step S53, the size of the current coding unit or the current coding tree unit in the current coding tree unit is determined, and in step S54, the size, width, or height of the current coding tree unit or the current coding unit is compared with a threshold. If the size, width, or height of the current coding tree unit or the current coding unit is greater than the threshold, then in step S55, the current coding tree unit or the current coding unit is divided into blocks until the size, width, or height of each block is not greater than the threshold. If the size, width, or height of the current coding tree unit or the current coding unit is not greater than the threshold, the video decoding system proceeds to step S57. In step S57, the video decoding system decodes information associated with the splitting decision from the video bitstream and further splits the current coding tree unit or the coding unit according to the splitting decision. Step S57 also includes processing each block of the current coding tree unit or the current coding unit through prediction and inverse transformation. In step S58, the video decoding system decodes the current coding tree unit according to the parsed coding tree unit level syntax.

[0044] The fifth embodiment shares one or more coding tree unit level grammars of the current coding tree unit with one or more other coding tree units. The coding tree unit level grammars associated with coding tools are sometimes very similar in multiple coding tree units. To further reduce the bits carried by the bitstream, the fifth embodiment introduces the concept of a coding tree unit group (CTU group). Coding tree units in the same coding tree unit group share some or all of the coding tree unit level grammars, which may include a combination of one or a set of grammars or parameters of SAO, ALF, and deblocking filtering. The size of the coding tree unit group can be a predefined value, or the size can be sent at the sequence level, picture level, or slice level of the video bitstream in, for example, the SPS, PPS, or slice file header. In one embodiment, the coding tree unit group sizes of any two of SAO, ALF, deblocking, or other loop filters can be the same or different. For example, the coding tree unit group size is represented by M×N coding tree units, where both M and N are positive integers and M + N>2. The pictures in the fifth embodiment are divided into coding tree unit groups, and the coding tree units in the same coding tree unit group share one or more coding tree unit level grammars. This idea of sharing coding tree unit level grammars can also be extended to share coding unit, prediction unit, or transform unit level grammars. For example, the coding unit level or transform unit level transform grammars in a coding tree unit are sent once, and these transform grammars are reused or shared by all coding units or transform units in the coding tree units belonging to the same coding tree unit group.

[0045] The sixth embodiment provides an alternative solution for a block splitting method in HEVC for splitting a current coding tree unit or a current coding unit at an image boundary. In HEVC, a coding unit at an image boundary is inferred to be split into four sub-coding units by a quadtree. The sixth embodiment of the present invention determines whether to apply quadtree splitting or binary tree splitting to the current coding tree unit or the current coding unit according to the boundary conditions of the current coding tree unit or the current coding unit. That is, according to the position of the current coding tree unit or the current coding unit, the width and height of the current coding tree unit or the current coding unit, and the width and height of the current image, the current coding tree unit or the current coding unit is split using quadtree splitting, vertical binary tree splitting, or horizontal binary tree splitting. Let (x0, y0) specify the upper-left sampled luminance position of the current coding tree unit or the current coding unit in the current image, cuWidth and cuHeight specify the width and height of the current coding tree unit or the current coding unit, and picWidth and picHeight specify the width and height of the current image. If the two values (x0 + cuWidth) and (y0 + cuHeight) are greater than picWidth and picHeight respectively, it is inferred that the current coding tree unit or the current coding unit is split by a quadtree. It is inferred that the current coding tree unit or the current coding unit is split into four sub-coding units of the same size cuWidth / 2 x cuHeight / 2 by a quadtree.

[0046] If only the value (x0 + cuWidth) is greater than picWidth, it is inferred that the current coding tree unit or the current coding unit is divided into two sub-coding units according to vertical binary tree splitting, and the sizes of the two sub-coding units are cuWidth / 2 x cuHeight. If only the value (y0 + cuHeight) is greater than picHeight, it is inferred that the current coding tree unit or the current coding unit is split into two sub-coding units according to horizontal binary tree splitting, and the sizes of the two sub-coding units are cuWidth x cuHeight / 2.

[0047] An encoder or a decoder according to the sixth embodiment splits a current coding tree unit or a current coding unit at an image boundary of a current image using one of quadtree splitting, vertical binary tree splitting, and horizontal binary tree splitting according to the position of the current block, the width and height of the current block, and the width and height of the current image. Figure 6Shows possible segmentation types of blocks at the image boundary according to the sixth embodiment. A video encoder or decoder receives input data associated with video image 60, and video image 60 is divided into a plurality of blocks including blocks 602, 604, and 606. Blocks 602, 604, and 606 may be coding tree units or coding units. The block 602 at the right boundary of video image 60 is forced to be segmented using vertical binary tree segmentation, and using horizontal binary tree segmentation, the block 604 at the bottom boundary of video image 60 is forced to be segmented. The block 606 at the bottom right boundary of video image 60 is segmented using quadtree segmentation.

[0048] Figure 7 Shows an exemplary system block diagram of a video encoder 700 implementing an embodiment of the present invention. Intra prediction 710 provides intra prediction based on the reconstructed video data of the current image, while motion prediction 712 performs motion estimation (ME) and motion compensation (MC) to provide a predictor based on video data from other images. The size, width, or height of the current coding tree unit or current coding unit in the current image is compared with a threshold, and if the size, width, or height is greater than the threshold, the current coding tree unit or current coding unit is divided into a plurality of blocks. In some embodiments, the current coding tree unit or current coding unit may be further divided according to the segmentation decision. Each block in the current coding tree unit or current coding unit can be predicted by intra prediction 710 or motion prediction 712. The block selected by motion prediction 712 is encoded in an inter prediction mode by inter prediction 7122 or in a merge mode by merge prediction 7124. Switch 714 selects one output from intra prediction 710 and motion prediction 712 and provides the selected predictor to adder 716 to form a prediction error, also known as a prediction residual.

[0049] The prediction residual of each block is further processed by a transform (T) 718, followed by quantization (Q) 720. Then, the transformed and quantized residual signal is encoded by entropy encoder 734 to form an encoded video bitstream. Then, the encoded video bitstream is wrapped with side information such as coding tree unit level syntax. Data associated with the side information is also provided to entropy encoder 734. The transformed and quantized residual signal of each block is processed by inverse quantization (IQ) 722 and inverse transform (IT) 724 to recover the prediction residual. As Figure 7As shown, the prediction residual is recovered by adding back the selected predictor at reconstruction (REC) 726 to produce the reconstructed video data. The reconstructed video data may be stored in a reference picture buffer (i.e., Ref. Pict. buffer) 732 and used to predict other pictures. Due to the encoding process, the reconstructed video data from reconstruction 726 may be subject to various impairments. Therefore, in-loop processing deblocking filter (DF) 728 and sample adaptive offset (SAO) 730 are applied to each coding tree unit of the reconstructed video data before being stored in the reference picture buffer 732 to further improve the picture quality. The syntax associated with the information for in-loop processing deblocking filter 728 and sample adaptive offset 730 is coding tree unit level syntax and is provided to entropy encoder 734 to be incorporated into the coded video bitstream.

[0050] For Figure 7 the corresponding video decoder 800 of video encoder 700 as Figure 8As shown. The encoded video bitstream is an input to video decoder 800 and is decoded by entropy decoder 810 to parse and recover the transformed and quantized residual signals, coding tree unit-level syntax such as DF and SAO information for each coding tree unit, and other system information. The decoding process of decoder 800 is similar to the reconstruction loop at encoder 700, except that decoder 800 only requires motion compensation prediction in motion prediction 814. Motion prediction 814 includes inter prediction 8142 and merge prediction 8144. Decoder 800 determines the size of the current coding tree unit or the size of the current coding unit in the current coding tree unit. If the size, width, or height of the current coding tree unit or the current coding unit is greater than a threshold, decoder 800 forces the current coding tree unit or the current coding unit to be divided into blocks until each block has a size, width, or height not greater than the threshold. According to some embodiments, the current coding tree unit or the current coding unit may be further divided according to information associated with the segmentation decision decoded from the video bitstream. Each block in the current coding tree unit or the current coding unit is decoded by intra prediction 812 or motion prediction 814. Blocks encoded in inter mode are decoded by inter prediction 8142, and blocks encoded in merge mode are decoded by merge prediction 8144. Switch 816 selects an intra predictor from intra prediction 812 or an inter predictor from inter prediction 814 according to the decoded mode information. The transformed and quantized residual signal associated with each block is recovered by inverse quantization (IQ) 820 and inverse transform (IT) 822. The recovered, transformed, and quantized residual signal is reconstructed by adding back the predictor in reconstruction 818 to produce the reconstructed video. The reconstructed video of each coding tree unit is further processed by deblocking filter 824 and sample adaptive offset 826 to produce the final decoded video according to the coding tree unit-level syntax. If the current decoded image is a reference image, the reconstructed video of the current decoded image is also stored in the reference image buffer for subsequent images in the decoding order.

[0051] Figure 7 and Figure 8Various components of the video encoder 700 and the video decoder 800 in [description] can be implemented by hardware components, one or more processors configured to execute program instructions stored in a memory, or a combination of hardware and processors. For example, the processor executes program instructions to control the reception of input data associated with the current image. The processor is equipped with a single or multiple processing cores. In some examples, the processor executes program instructions to perform functions in some components in the encoder 700 and the decoder 800, and the memory electrically coupled to the processor is used to store program instructions, information on the reconstructed images corresponding to the blocks, and / or intermediate data during the encoding or decoding process. In some embodiments, the memory includes a non-transitory computer-readable medium, such as semiconductor or solid-state memory, random access memory (RAM), read-only memory (ROM), hard disk, optical disc, or other suitable storage medium. The memory can also be a combination of two or more of the non-transitory computer-readable media listed above. As Figure 7 and 8 shown, the encoder 700 and the decoder 800 can be implemented in the same electronic device. Thus, if implemented in the same electronic device, various functional components of the encoder 700 and the decoder 800 can be shared or reused. For example, Figure 7 one or more of the reconstruction 726, transform 718, quantization 720, deblocking filter 728, sample adaptive offset 730, and reference image buffer 732 in [description] can also be used to implement respectively Figure 8 the functions of the reconstruction 818, transform 822, quantization 820, deblocking filter 824, sample adaptive offset 826, and reference image buffer 828 in [description].

[0052] Embodiments of the block segmentation method for a video coding system can be implemented in a circuit integrated into a video compression chip, or in program code integrated into video compression software to perform the above processing. For example, the decision of block partitioning can be implemented in program code executed on a computer processor, a digital signal processor (DSP), a microprocessor, or a field-programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the present invention by executing machine-readable software code or firmware code that defines the specific methods embodied in the present invention.

[0053] Without departing from the spirit or essential features of the present invention, the present invention can be implemented in other specific forms. The described examples are considered illustrative in all aspects and not restrictive. Thus, the scope of the present invention is indicated by the claims rather than the foregoing description. All changes within the meaning and scope equivalent to the claims will be included within its scope.

Claims

1. A block splitting method in an image or video coding system, comprising: Receiving input data associated with a current block in a current picture, where the current picture is divided into a plurality of non - overlapping blocks; Recursively splitting the current block into a plurality of leaf blocks according to the position of the current block in the current picture, including: Determining whether a first condition or a second condition is satisfied based on the position of the upper - left sampling of the current block, the widths of the current block and the current picture respectively, and the heights of the current block and the current picture respectively; When the position of the current block satisfies the first condition, recursively splitting the current block according to quadtree splitting; wherein the position of the current block satisfies the first condition only when the x - component of the upper - left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y - component of the upper - left sampling of the current block plus the height of the current block is greater than the height of the current picture; and When the position of the current block satisfies the second condition, recursively splitting the current block according to a selected splitting determined from the group consisting of vertical binary - tree splitting and horizontal binary - tree splitting; wherein if the x - component of the upper - left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y - component of the upper - left sampling of the current block plus the height of the current block is not greater than the height of the current picture, or the x - component of the upper - left sampling of the current block plus the width of the current block is not greater than the width of the current picture and the y - component of the upper - left sampling of the current block plus the height of the current block is greater than the height of the current picture, the position of the current block satisfies the second condition; and Encoding or decoding the current block by separately processing each leaf block in the current block for prediction and transform processing.

2. The block segmentation method in the image or video coding system according to claim 1, wherein When the x - component of the upper - left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y - component of the upper - left sampling of the current block plus the height of the current block is not greater than the height of the current picture, the selected splitting corresponds to vertical binary - tree splitting.

3. The block segmentation method in the image or video coding system according to claim 1, characterized in that, When the x - component of the upper - left sampling of the current block plus the width of the current block is not greater than the width of the current picture and the y - component of the upper - left sampling of the current block plus the height of the current block is greater than the height of the current picture, the selected splitting corresponds to horizontal binary - tree splitting.

4. The block segmentation method in the image or video coding system according to claim 1, wherein The first condition corresponds to one or a combination of the position of the current block in the current picture, the width of the current block, the height of the current block, the width of the current picture, and the height of the current picture.

5. The block segmentation method in the image or video coding system according to claim 1, characterized in that The second condition corresponds to one or a combination of the position of the current block in the current picture, the width of the current block, the height of the current block, the width of the current picture, and the height of the current picture.

6. The block segmentation method in the image or video coding system according to claim 1, wherein Each of the plurality of non - overlapping blocks corresponds to a coding tree unit, and each of the plurality of leaf blocks corresponds to a leaf coding unit.

7. A block splitting apparatus in an image or video coding system, the apparatus including one or more electronic circuits configured to: Receive input data associated with a current block in a current picture, where the current picture is divided into a plurality of non - overlapping blocks; Recursively divide the current block into multiple leaf blocks according to the position of the current block in the current picture, including: Determine whether the first condition or the second condition is satisfied based on the position of the upper-left sampling of the current block, the widths of the current block and the current picture respectively, and the heights of the current block and the current picture respectively; When the position of the current block satisfies the first condition, recursively divide the current block according to quadtree division; wherein the position of the current block satisfies the first condition only when the x component of the upper-left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y component of the upper-left sampling of the current block plus the height of the current block is greater than the height of the current picture; and When the position of the current block satisfies the second condition, recursively divide the current block according to the selected division determined from the group consisting of vertical binary tree division and horizontal binary tree division; wherein if the x component of the upper-left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y component of the upper-left sampling of the current block plus the height of the current block is not greater than the height of the current picture, or the x component of the upper-left sampling of the current block plus the width of the current block is not greater than the width of the current picture and the y component of the upper-left sampling of the current block plus the height of the current block is greater than the height of the current picture, the position of the current block satisfies the second condition; and Encode or decode the current block by processing each leaf block in the current block separately for prediction and transform processing.

8. The block segmentation device in the image or video coding system according to claim 7, wherein When the x component of the upper-left sampling of the current block plus the width of the current block is greater than the width of the current picture and the y component of the upper-left sampling of the current block plus the height of the current block is not greater than the height of the current picture, the selected division corresponds to vertical binary tree division.

9. The block segmentation device in the image or video coding system according to claim 7, characterized in that, When the x component of the upper-left sampling of the current block plus the width of the current block is not greater than the width of the current picture and the y component of the upper-left sampling of the current block plus the height of the current block is greater than the height of the current picture, the selected division corresponds to horizontal binary tree division.

10. The block segmentation device in the image or video coding system according to claim 7, characterized in that, The first condition corresponds to one or a combination of the position of the current block in the current picture, the width of the current block, the height of the current block, the width of the current picture, and the height of the current picture.

11. The block segmentation device in the image or video coding system according to claim 7, characterized in that, The second condition corresponds to one or a combination of the position of the current block in the current picture, the width of the current block, the height of the current block, the width of the current picture, and the height of the current picture.

12. The block segmentation device in the image or video coding system according to claim 7, characterized in that, Each of the multiple non-overlapping blocks corresponds to a coding tree unit, and each of the multiple leaf blocks corresponds to a leaf coding unit.

Citation Information

Patent Citations

  • Video encoding device, video decoding device, video encoding method, video decoding method, and program

    CN104904210A