Method and apparatus for video coding with intra sub-partition prediction and transform skipping

By applying the transform skip mode to the sub-blocks of intra-subpartition prediction in the video encoding and decoding method, the problem of inefficiency in the prior art is solved, and more efficient video encoding and decoding and better video quality are achieved.

CN120202671APending Publication Date: 2025-06-24HYUNDAI MOTOR CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380079650.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-22
Filing Date
2023-09-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art fails to effectively apply the transform skip mode to sub-blocks predicted according to intra-subpartitions, resulting in inefficient video encoding and decoding.

Method used

In the video encoding and decoding method, the current block is partitioned into a sub-block based on the size of the current block and the sub-block partition direction, and it is determined whether to apply a transform skip mode to the residual block of the sub-block, thereby generating a quantized residual block and determining the transform kernel implicitly or explicitly.

Benefits of technology

The video encoding and decoding efficiency is improved, the video quality is enhanced, and the problem of combining the transform skip mode and intra-subpartition prediction technology in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202671A_ABST
    Figure CN120202671A_ABST
Patent Text Reader

Abstract

The embodiment discloses a video coding and decoding method and device using intra sub-partition prediction and transform skipping. In this embodiment, an image decoding device divides a current block into sub-blocks, and then determines whether to use a transform skip on the sub-blocks of the current block. When transform skip is used, the image decoding device omits inverse transform for the sub-block. When the transform skip is not used, the image decoding device determines a transform kernel for the sub-block, and then inversely transforms the sub-block by using the determined transform kernel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video encoding and decoding method and apparatus using intra sub - partition prediction and transform skip. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit uncompressed video data.

[0004] Accordingly, an encoder is generally used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), and the Versatile Video Coding (VVC) has a coding and decoding efficiency improvement of about 30% or more compared to HEVC.

[0005] However, due to the gradual increase in image size, resolution, and frame rate, the amount of data to be encoded also increases. Accordingly, there is a need to provide a new compression technology with higher coding and decoding efficiency and improved image enhancement effects compared to existing compression technologies.

[0006] Intra prediction uses intra - picture pixel information to predict the pixel values of the current block to be encoded. Intra prediction can select one of multiple intra - prediction modes that is most suitable for the characteristics of the picture, and can use the selected mode to predict the current block. The encoder selects and uses one of the multiple intra - prediction modes to encode the current block. The encoder can then transmit information about the mode to the decoder.

[0007] The HEVC technology uses a total of 35 intra - prediction modes for intra - prediction, including 33 directional or angular modes with directionality and two non - directional or non - angular modes without directionality. However, as the spatial resolution of the image increases from 720×480 to 2048×1024 or 8192×4096, the size of the prediction block unit also increases, which requires adding more intra - prediction modes. As Figure 3a shown, the VVC technology uses 65 prediction modes with further sub - division for intra - prediction, which allows more types of prediction directions compared to the prior art.

[0008] Generally, an image to be encoded is partitioned into coding units (CUs) of various shapes and sizes, and thus encoded in units of CUs. At this time, the information specifying the partition is represented as a tree structure, which is sent to the decoder to indicate into which shapes and sizes of CUs the image is partitioned. On the other hand, when an image is partitioned into CUs and encoded in units of CUs, all pixels in a CU block can be intra-predicted according to a prediction mode. At this time, when the reference samples used for intra-prediction are far from the pixels in the CU block, the prediction efficiency decreases, and a large amount of energy may remain in the residual signal obtained through prediction. If the block is a horizontally (or vertically) elongated rectangular block or if the block size is large, the remaining energy in the residual signal may be more serious. Since some prediction directions will increase the distance between the pixels to be predicted and the reference samples, this will become serious. Sub-dividing the block into smaller CUs can be a solution to the problem. However, for each of the more sub-divided CU blocks, an increased overhead for transmitting intra-prediction modes is introduced.

[0009] On the other hand, there are techniques for solving the problem of increased overhead. In the prior art, a CU block is sub-divided into smaller blocks of equal size, which are called "sub-blocks", and intra-prediction is performed on each sub-block. By transmitting only one intra-prediction mode for the original CU block before sub-division and making the partitioned sub-blocks have a shared prediction mode, the prediction efficiency can be improved while reducing the overhead. This prior art is called the intra sub-partitions (ISP) technique.

[0010] By using transform techniques, the residual signal is transformed into a signal in the frequency domain. Depending on the transform, the energy in the block is concentrated in the low-frequency region, which can make it easier to encode the transformed residual signal. At this time, the encoder selects a transform technique more suitable for the residual signal, such as discrete cosine transform (DCT), discrete sine transform (DST), etc. The encoder uses the selected transform technique to transform the block to be encoded and transmits the information about the selected technique to the decoder.

[0011] According to the HEVC technology, by using the transformation of the Discrete Cosine Transform II (DCT-II) usually applied in the horizontal and vertical directions, the residual signal of the luminance channel is transformed into a frequency signal. For a block of size 4×4, quantization can be applied after applying the transformation of the Discrete Sine Transform VII (DST-VII) or after determining a transform skip mode in which no transformation is performed on the residual signal. However, with the development of image compression technology, various methods for generating predicted signals have been developed, thus providing various characteristics for the residual signals generated by applying these methods. In the recent VVC technology, new transformations such as DCT-VIII have been introduced to further diversify the transformations that can be applied to the residual signals. In addition, the transformation of DST-VII and the transform skip mode that are only applied to 4×4 blocks can be applied to blocks of other sizes.

[0012] However, the prior art does not consider applying the transform skip mode to sub-blocks according to the application of ISP technology. Therefore, in order to improve video coding and decoding efficiency and enhance video quality, a method for effectively associating the transform skip mode with sub-blocks predicted by ISP technology is needed. Summary of the Invention

[0013] Technical Problem

[0014] The present invention is dedicated to providing a video coding and decoding method and apparatus capable of applying a transform skip mode to sub-blocks predicted according to intra-frame sub-partition prediction (ISP prediction).

[0015] Technical Solution

[0016] At least one aspect of the present invention provides a method for a video decoding apparatus to reconstruct a current block. The method includes partitioning the current block into sub-blocks based on the size of the current block and the sub-block partitioning direction. The method further includes determining whether to use transform skip for the sub-blocks of the current block. The method further includes checking whether transform skip is used. The method further includes: when transform skip is not used, determining a transform kernel for the sub-blocks.

[0017] Another aspect of the present invention provides a method for a video encoding apparatus to encode a current block. The method includes partitioning the current block into sub-blocks based on the size of the current block and the sub-block partitioning direction. The method further includes generating a first quantized residual block by applying transform skip and quantization to the residual blocks of the sub-blocks. The method further includes implicitly deriving a transform kernel or explicitly determining a transform kernel. The method further includes generating a second quantized residual block by applying the transform kernel and quantization to the residual blocks of the sub-blocks.

[0018] Another aspect of the present invention provides a computer-readable recording medium that stores a bitstream generated by a video encoding method. The video encoding method includes partitioning a current block into sub-blocks based on the size of the current block and the sub-block partitioning direction. The video encoding method further includes generating a first quantized residual block by applying transform skip and quantization to the residual block of the sub-block. The video encoding method further includes implicitly deriving a transform kernel or explicitly determining a transform kernel. The video encoding method further includes generating a second quantized residual block by applying the transform kernel and quantization to the residual block of the sub-block.

[0019] Advantageous Effects

[0020] As described above, the present invention provides a video encoding / decoding method and apparatus that can apply a transform skip mode to sub-blocks predicted according to intra sub-partition prediction. Therefore, the video encoding / decoding method and apparatus improve video encoding / decoding efficiency and enhance video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a block diagram of a video encoding apparatus that can implement the technology of the present invention.

[0022] Figure 2 illustrates a method of partitioning a block using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.

[0023] Figure 3a and Figure 3b illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.

[0024] Figure 4 illustrates adjacent blocks of a current block.

[0025] Figure 5 is a block diagram of a video decoding apparatus that can implement the technology of the present invention.

[0026] Figure 6 is a schematic diagram showing rate-distortion estimation for determining prediction and transform.

[0027] Figure 7 is a diagram showing the proportion of coding units (CUs) that utilize new reference samples in coding units (CUs) that utilize intra sub-partition (ISP).

[0028] Figure 8 is a schematic diagram showing rate-distortion estimation for determining prediction and transform according to at least one embodiment of the present invention.

[0029] Figure 9It is a schematic diagram showing the use of transform skip and multiple transform selection (MTS) for corresponding ISP blocks according to at least one embodiment of the present invention.

[0030] Figure 10 It is a flowchart of a method for a video encoding device to encode a current block according to at least one embodiment of the present invention.

[0031] Figure 11 It is a flowchart of a method for a video decoding device to reconstruct a current block according to at least one embodiment of the present invention. Detailed Description

[0032] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. In addition, in the following description of some embodiments, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for clarity and conciseness.

[0033] Figure 1 It is a block diagram of a video encoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 1 the illustrated diagram, the video encoding device and the components of the device will be described.

[0034] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0035] Each component of the encoding device may be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0036] A video consists of one or more sequences including a plurality of images. Each image is segmented into a plurality of regions, and encoding is performed on each region. For example, an image is segmented into one or more tiles and / or slices. Here, one or more tiles may be defined as a tile group. Each tile and / or slice is segmented into one or more coding tree units (CTUs). In addition, each CTU is segmented into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information applied to the CUs included in one CTU is encoded as the syntax of the CTU. In addition, information applied to all blocks in a slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more images is encoded as the syntax of the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referred to by a plurality of images is encoded as the syntax of the Sequence Parameter Set (SPS). In addition, information commonly referred to by one or more SPSs is encoded as the syntax of the Video Parameter Set (VPS). In addition, information applied to one tile or tile group may also be encoded as the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header may be referred to as high-level syntax.

[0037] The image segmenter 110 determines the size of the coding tree unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and is transmitted to the video decoding device.

[0038] The image segmenter 110 segments each image constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively segments the CTUs by using a tree structure. A leaf node in the tree structure becomes a coding unit (CU), and the CU is a basic unit for encoding.

[0039] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree plus binarytree (QTBT) structure can be used, or a quadtree plus binarytreeternarytree (QTBTTT) structure can be used. Here, the binarytreeternarytree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).

[0040] Figure 2 is a schematic diagram for describing a method of dividing a block by using the QTBTTT structure.

[0041] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes and signals it to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or TT structure. For example, there can be two directions, namely, the direction of horizontally dividing the block of the corresponding node and the direction of vertically dividing the block of the corresponding node. As Figure 2 shown, when the MTT division starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is divided, and a flag indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) in the case where the node is divided, and signals it to the video decoding device.

[0042] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, the CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device starts encoding the first flag first according to the above scheme.

[0043] When QTBT is used as another example of the tree structure, there may be two types, that is, the type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal split) and the type that vertically splits the block of the corresponding node into two blocks of the same size (i.e., symmetric vertical split). The entropy encoder 155 encodes the split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and the split type information indicating the split type, and transmits them to the video decoding device. On the other hand, there may additionally be a type in which the block of the corresponding node is split into two asymmetric blocks. The asymmetric form may include the form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include the form in which the block of the corresponding node is split in the diagonal direction.

[0044] The CU may have various sizes according to the QTBT or QTBTTT split from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the "current block". When QTBTTT split is adopted, in addition to the square shape, the shape of the current block may also be a rectangular shape.

[0045] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0046] Generally, predictive encoding can be performed for each of the current blocks in the image. Generally, the prediction of the current block can be performed by using intra prediction technology (which uses data from the image including the current block) or inter prediction technology (which uses data from the image encoded before the image including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.

[0047] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) adjacent to the current block in the current image including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra prediction modes may include two non-directional modes including Planar mode and DC mode, and may include 65 directional modes. Adjacent pixels to be used and algorithm equations are defined differently according to each prediction mode.

[0048] For efficient directional prediction of a current block having a rectangular shape, the directional modes shown by the dashed arrows in Figure 3b (modes #67 to #80, intra prediction modes #-1 to #-14) may be additionally used. The directional modes may be referred to as "wide angle intra-prediction modes". In Figure 3b the arrows indicate the corresponding reference samples for prediction, rather than representing the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height smaller than the width, wide angle intra-prediction modes having an angle less than 45 degrees (intra prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width larger than the height, wide angle intra-prediction modes having an angle greater than -135 degrees are available.

[0049] The intra predictor 122 may determine the intra prediction to be used for encoding the current block. In some examples, the intra predictor 122 may encode the current block by using multiple intra prediction modes, and may also select an appropriate intra prediction mode to be used from test modes. For example, the intra predictor 122 may calculate rate-distortion values by using rate-distortion analysis of multiple tested intra prediction modes, and may also select an intra prediction mode having the best rate-distortion characteristics from the test modes.

[0050] The intra predictor 122 selects one intra prediction mode from multiple intra prediction modes, and predicts the current block by using adjacent pixels (reference pixels) and algorithm equations determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0051] The inter - frame predictor 124 generates a predicted block of the current block by using motion - compensation processing. The inter - frame predictor 124 searches for the block most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a predicted block of the current block by using the searched - for block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the predicted block in the reference image. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The entropy encoder 155 encodes the motion information including the information of the reference image and the information about the motion vector used for predicting the current block, and transmits it to the video decoding device.

[0052] The inter - frame predictor 124 can also perform interpolation of the reference image or reference blocks to increase the prediction accuracy. In other words, sub - samples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference image, the motion vector can represent fractional - unit precision rather than integer - sample - unit precision. For each target region to be encoded, such as units like slices, tiles, CTUs, CUs, etc., the precision or resolution of the motion vector can be set differently. When applying such adaptive motion vector resolution (AMVR), information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the precision of the motion - vector difference described below.

[0053] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference image to generate a first reference block and a second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Further, motion information including information about the two reference images used for predicting the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference picture list 0 may be composed of images in the pre-reconstructed images that are before the current image in the display order, and reference picture list 1 may be composed of images in the pre-reconstructed images that are after the current image in the display order. However, although not particularly limited thereto, pre-reconstructed images after the current image in the display order may be additionally included in reference picture list 0. Conversely, pre-reconstructed images before the current image may also be additionally included in reference picture list 1.

[0054] To minimize the amount of bits consumed for encoding motion information, various methods can be used.

[0055] For example, when the reference image and motion vector of the current block are the same as those of an adjacent block, information identifying the adjacent block is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.

[0056] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.

[0057] As the adjacent blocks for deriving the merge candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image can be used, as Figure 4 shown. In addition, in addition to the current image where the current block is located, blocks within the reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidates. For example, the co-located block of the current block within the reference image or a block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, zero vectors are added to the merge candidates.

[0058] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as motion information of a current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0059] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only the neighboring block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.

[0060] Thereafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.

[0061] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0062] In the AMVP mode, the inter-frame predictor 124 derives motion vector prediction candidates for the motion vector of the current block by using neighboring blocks of the current block. As neighboring blocks for deriving the motion vector prediction candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image shown in Figure 4 In addition, in addition to the current image where the current block is located, blocks in a reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as neighboring blocks for deriving the motion vector prediction candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0063] The inter-frame predictor 124 derives motion vector prediction candidates by using the motion vectors of neighboring blocks, and determines the motion vector prediction of the motion vector of the current block by using the motion vector prediction candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.

[0064] Motion vector prediction can be obtained by applying predefined functions (e.g., median and mean calculations, etc.) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined functions. In addition, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference image used to predict the current block are encoded.

[0065] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, the information for identifying the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference image used to predict the current block.

[0066] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0067] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as the transform unit, or the residual block can also be divided into multiple sub - blocks, and the transform can be performed by using the sub - blocks as the transform unit. Alternatively, the residual block is divided into two sub - blocks, i.e., a transform region and a non - transform region, to transform the residual signal by using only the transform region sub - block as the transform unit. Here, the transform region sub - block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating only the transform sub - block, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag), and signals it to the video decoding device. Additionally, the size of the transform region sub - block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding segmentation and signals it to the video decoding device.

[0068] On the other hand, the transformer 140 can perform the transformation of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transform set (MTS). The transformer 140 can select a pair of transformation functions in the MTS with the highest transformation efficiency, and can transform the residual block on each of the horizontal direction and the vertical direction. Information (mts_idx) about the pair of transformation functions in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.

[0069] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters, and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also quantize the relevant residual block immediately without transforming any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. A quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.

[0070] The rearrangement unit 150 can perform rearrangement of coefficient values on the quantized residual values.

[0071] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can use a zig-zag scan or a diagonal scan to scan the coefficients from the DC coefficient to the high-frequency region to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scan method to be used can be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.

[0072] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc., to generate a bitstream.

[0073] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, etc.) so that the video decoding device can partition blocks in the same way as the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (the merge index in the case of the merge mode, and information regarding the reference image index and the motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding the quantization parameter and information regarding the quantization matrix).

[0074] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0075] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.

[0076] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0077] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove blocking artifacts that occur due to block-based coding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The SAO filter 184 applies an offset in units of CTUs to enhance the subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the degree of the boundary and the amount of change of the corresponding block to compensate for distortion. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.

[0078] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter prediction of the blocks within the image to be encoded subsequently.

[0079] The video encoding device can store the bitstream of the encoded video data in a non-volatile storage medium or transmit the bitstream to the video decoding device through a communication network.

[0080] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 5 the video decoding device and the components of the device are described.

[0081] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transform unit 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0082] Similar to Figure 1 the video encoding device, each component of the video decoding device can be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component can be implemented as software, and the microprocessor can also be implemented to execute the functions of the software corresponding to each component.

[0083] The entropy decoder 510 extracts information related to block partitioning by decoding the bitstream generated by the video encoding device to determine the current block to be decoded, and extracts the prediction information and the information about the residual signal required to reconstruct the current block.

[0084] The entropy decoder 510 determines the size of a coding tree unit (CTU) by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and divides an image into CTUs having the determined size. In addition, a CTU is determined as the highest layer (i.e., the root node) of a tree structure, and split information of the CTU is extracted to divide the CTU by using the tree structure.

[0085] For example, when dividing a CTU by using a QTBTTT structure, first, a first flag (QT_split_flag) related to the split of a quantization tree (QT) is extracted to divide each node into four lower-layer nodes. In addition, a second flag (mtt_split_flag) related to the split of a multi-tree transform (MTT), a split direction (vertical / horizontal), and / or a split type (binary / trinary) are extracted with respect to a node corresponding to a leaf node of the QT to divide the corresponding leaf node into an MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a binary tree (BT) or a ternary tree (TT) structure.

[0086] As another example, when dividing a CTU by using a QTBTTT structure, a CU split flag (split_cu_flag) indicating whether to split a coding unit (CU) is extracted. When splitting the corresponding block, a first flag (QT_split_flag) may also be extracted. During the splitting process, for each node, zero or more recursive MTT splits may occur after zero or more recursive QT splits. For example, for a CTU, an MTT split may occur immediately, or conversely, only multiple QT splits may occur.

[0087] As another example, when dividing a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the split of a QT is extracted to divide each node into four lower-layer nodes. In addition, a split flag (split_flag) indicating whether to further split a node corresponding to a leaf node of the QT into a BT and split direction information are extracted.

[0088] On the other hand, when the entropy decoder 510 determines a current block to be decoded by using the split of a tree structure, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-frame predicted or inter-frame predicted. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts a syntax element for intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information about a syntax element representing inter-frame prediction information, i.e., a motion vector and a reference image to which the motion vector refers.

[0089] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on the quantized transform coefficients of the current block as information on the residual signal.

[0090] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video coding device.

[0091] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.

[0092] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.

[0093] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for inverse-transforming only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the non-inverse-transformed region with the value "0" as the residual signal to generate the final residual block of the current block.

[0094] In addition, when MTS is applied, the inverse transformer 530 determines the transform function or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.

[0095] The predictor 540 can include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.

[0096] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using the adjacent reference pixels of the current block according to the intra prediction mode.

[0097] The inter - frame predictor 544 determines the motion vector of the current block and the reference image for motion - vector reference by using the syntax element of the inter - frame prediction mode extracted from the entropy decoder 510.

[0098] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter - frame predictor 544 or the intra - frame predictor 542. When performing intra - frame prediction on a block to be decoded subsequently, the pixels within the reconstructed current block are used as reference pixels.

[0099] The loop filter unit 560 as an in - loop filter may include a de - blocking filter 562, a SAO filter 564, and an ALF 566. The de - blocking filter 562 performs de - blocking filtering on the boundaries between the reconstructed blocks to remove block artifacts that occur due to block - unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after de - blocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.

[0100] The reconstructed blocks filtered by the de - blocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter - frame prediction of the blocks within the image to be encoded subsequently.

[0101] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus that can apply a transform - skip mode to sub - blocks predicted according to intra sub - partitions (ISP) prediction techniques.

[0102] The following embodiments may be executed by the intra - frame predictor 122, the transformer 140, and the inverse transformer 165 in a video encoding apparatus. The following embodiments may also be executed by the inverse transformer 530 and the intra - frame predictor 542 in a video decoding apparatus.

[0103] When encoding the current block, the video encoding apparatus may generate signaling information associated with the present embodiment from the perspective of optimizing rate - distortion. The video encoding apparatus may encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding apparatus. The video decoding apparatus may decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.

[0104] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some regions of the coding unit.

[0105] In addition, a value of true for a flag indicates the case where the flag is set to 1. Further, a value of false for a flag indicates the case where the flag is set to 0.

[0106] I. Intra Prediction and Intra Sub-Partitioning (ISP)

[0107] In the VVC technology, the intra prediction mode of a luma block has sub-divided directional modes (i.e., modes - 14 to 80) in addition to non-directional modes (i.e., planar and DC modes), as Figure 3a and Figure 3b shown. Based on these prediction modes, there are several techniques for improving the coding and decoding efficiency of intra prediction. The ISP technique sub-partitions the current block into smaller blocks of the same size, and then allows sharing of the intra prediction mode across sub-blocks, but applies transforms to each sub-block. The sub-partitioning of the block can be performed in the horizontal or vertical direction.

[0108] In the following description, the larger block before sub-partitioning is referred to as the current block, and each of the smaller blocks generated by sub-partitioning is referred to as a sub-block.

[0109] The operation of the ISP technique is as follows.

[0110] The video coding device signals to the video decoding device the intra_subpartitions_mode_flag indicating whether to apply ISP and the intra_subpartitions_split_flag indicating the sub-partitioning method, as shown in Table 1.

[0111] [Table 1]

[0112]

[0113] Table 2 shows the IntraSubPartitionsSplitTyp of the sub-partitioning type according to the intra_subpartitions_mode_flag and the intra_subpartitions_split_flag.

[0114] [Table 2]

[0115] IntraSubPartitionsSplitType Name of IntraSubPartitionsSplitType 0 ISP_NO_SPLIT 1 ISP_HOR_SPLIT 2 ISP_VER_SPLIT

[0116] The ISP technique sets the partitioning type IntraSubPartitionsSplitType as follows.

[0117] If intra_subpartitions_mode_flag is 0, set IntraSubPartitionsSplitType to 0 and do not perform sub-block partitioning. That is, ISP is not applied.

[0118] If intra_subpartitions_mode_flag is not 0, apply ISP. In this case, set IntraSubPartitionsSplitType to the value of 1 + intra_subpartitions_split_flag and perform sub-block partitioning according to the partitioning type. If IntraSubPartitionsSplitType = 1, perform sub-block partitioning in the horizontal direction (ISP_HOR_SPLIT), and if IntraSubPartitionsSplitType = 2, perform sub-block partitioning in the vertical direction (ISP_VER_SPLIT). This means that intra_subpartitions_split_flag can indicate the sub-block partitioning direction.

[0119] For example, if the ISP mode of horizontal sub-partitioning is applied to the current block, IntraSubPartitionsSplitType is 1, intra_subpartitions_mode_flag is 1, and intra_subpartitions_split_flag is 0.

[0120] In the following description, intra_subpartitions_mode_flag is expressed as the sub-block partitioning application flag, intra_subpartitions_split_flag is expressed as the sub-block partitioning direction flag, and IntraSubPartitionsSplitType is expressed as the sub-block partitioning type.

[0121] In addition, ISP_HOR_SPLIT and horizontal partitioning can be used interchangeably, and ISP_VER_SPLIT and vertical partitioning can be used interchangeably.

[0122] The current block can be sub-partitioned in the horizontal or vertical direction. However, if the size of the current block is too small, the encoding / decoding efficiency of the sub-blocks generated by the sub-partitioning may be further reduced, or the sub-blocks may become smaller than the minimum transform unit, making it impossible to perform a transform on the sub-blocks. To prevent this from happening, the application of ISP can be restricted by referring to the size of the sub-blocks obtained from the partitioning. For example, the current block can be sub-partitioned to prevent the number of pixels in the sub-blocks of the partition from becoming less than 16. For example, if the size of the current block is 4×4, ISP is not applied. A block of size 4×8 or 8×4 can be partitioned into two sub-blocks of the same shape and size, which is a partition called Half_Split. Blocks of any other size can be partitioned into four sub-blocks of the same shape and size, which is a partition called Quarter_Split.

[0123] The video encoding device encodes the sub-blocks in sequential order from left to right or from top to bottom. Each sub-block shares the same intra prediction information. In the intra prediction for encoding each sub-block, the video encoding device can improve the compression efficiency by using the reconstructed pixels in the earlier encoded sub-blocks as the predicted pixel values for the subsequent sub-blocks.

[0124] II. Transform Skip Technique

[0125] When encoding / decoding an image, a transform is usually performed. However, in some cases, it may be beneficial not to perform the transform. Omitting the execution of the transform is called the transform skip technique. Additionally, the encoding information indicating whether to skip the transform for the current block is represented by the transform skip flag or transform_skip_flag. The transform_skip_flag is marked on the bitstream and sent from the video encoding device to the video decoding device. The video decoding device parses the transform_skip_flag from the bitstream. If transform_skip_flag = 0, the video decoding device performs the transform. If transform_skip_flag = 1, the video decoding device can omit the transform and perform the decoding process for the current block.

[0126] In the prior art, the video encoding device signals transform_skip_flag[x][y][compID] as shown in Table 3.

[0127] [Table 3]

[0128] [TU Level]

[0129]

[0130] Here, for each channel, x and y represent the coordinates of the upper left pixel of the residual block. compID indicates the channel. For example, if compID = 0, then compID indicates the luminance component, if compID = 1, then compID indicates the Cb component, and if compID = 2, then compID indicates the Cr component. In the prior art, when transform skip is not used (i.e., transform_skip_flag = 0), the video coding device passes mts_idx to the video decoding device, as shown in Table 4, to signal the transform kernel applied to the current block.

[0131] [Table 4]

[0132]

[0133] III. Deficiencies of the Prior Art and Embodiments of the Present Invention

[0134] The prior art cannot utilize both the ISP and transform skip techniques together. In the prior art, blocks using transform skip cannot benefit from the prediction according to the ISP technique, resulting in inefficiency.

[0135] Figure 6 is a schematic diagram showing rate-distortion estimation for determining prediction and transformation.

[0136] Figure 6 shows several cases where a video coding device according to the prior art performs rate-distortion estimation to determine how to predict and transform. In Figure 6 's example, transform skip is not tested for blocks performing ISP. In Figure 6 's example, rough mode decision (RMD), most probable mode (MPM), multiple reference line prediction (MRLP), matrix-weighted intra prediction (MIP), and transform skip mode (TSM) are shown. Pi (i = 1, 2, 3, 4, 5) represents the primary transform kernel, and Li (i = 1, 2) represents the low-frequency non-separable transform (LFNST) kernel for secondary transformation.

[0137] Hereinafter, from each angle of intra-frame sub-partitioning (ISP) and transform skip, the prior art can be analyzed as follows.

[0138] First to be described is the relationship between the performance of the ISP and the transformation. Generally, the performance improvement of the ISP can be attributed to the fact that when a block is partitioned into multiple sub-blocks, newly reconstructed reference samples located closer to each sub-block (i.e., the reconstructed samples from the previous sub-block) can be used for prediction. However, the above are just some factors contributing to the ISP performance improvement. Assuming the use of two additional bits, i.e., the use of a sub-block partitioning application flag and a sub-block partitioning direction flag, if the ISP performance improvement is mainly due to an increase in prediction accuracy with per-sub-block prediction, then most of the CUs utilizing the ISP need to use newly reconstructed reference samples. According to Figure 7 In the example of, depending on the video group, up to 41.4% of the CUs utilizing the ISP use reference samples reconstructed earlier than the new reference samples of the adjacent CUs for prediction, i.e., use the same reference samples as the CUs not utilizing the ISP. In Figure 7 In the example of, the categories represent the video groups used in the experiment. Figure 7 The example of shows that the per-sub-block transformation has a significant impact on the performance of the ISP.

[0139] Furthermore, based on the results of applying the transformation to the ISP-enabled CUs that use the ISP, one can expect that if the transformation is skipped for the sub-blocks, the prediction performance will be further improved. Table 5 shows the distribution of the coded block flag (CBF) in the CUs with the ISP divided by video group.

[0140] [Table 5]

[0141]

[0142] Here, "all 'CBF = 1'" indicates that all sub-blocks in the CU have CBF = 1. That is, all sub-blocks have non-zero transform coefficients. "The existence of 'CBF = 0'" indicates that at least one sub-block in the CU has CBF = 0. That is, at least one sub-block has zero transform coefficients. In this case, CBF = 0 indicates that the prediction is performed well and there are no non-zero transform coefficients in this block. For the CUs with CBF = 0, the encoding of the residual signal is not performed. Depending on the video group, for all sub-blocks within this CU, up to 38% of the CUs with Quarter_Split have CBF = 1. For all sub-blocks within this CU, up to 61% of the CUs with Half_Split have CBF = 1. These results show that the ISP is selected in those CUs even when the prediction is not performed well. Therefore, when it is difficult to achieve energy compression in sub-blocks with non-zero transform coefficients, it is expected that the coding performance can be improved by utilizing transformation skipping.

[0143] The benefits in terms of the application of transformation skipping are described below.

[0144] When effective energy compression cannot be performed after the transformation of a residual signal such as screen content, transform skip avoids performing the entire transformation process. This means that transform skip represents applying a transformation kernel in the form of an identity matrix to the residual signal. After performing transform skip, the residual signal remains unchanged. The percentage of CUs that utilize transform skip according to the video group can be expressed as shown in Table 6.

[0145] [Table 6]

[0146]

[0147]

[0148] According to Table 6, transform skip is used at a rate of 55% in slide editing of category F of screen content videos and at a similar large rate in other screen content videos. If blocks with transform skip can utilize ISP technology, it is expected that up to half of the blocks in the video group can improve prediction performance by applying ISP. However, the prior art cannot achieve the link between transform skip and ISP technology.

[0149] Figure 8 is a schematic diagram showing rate-distortion estimation for determining prediction and transformation according to at least one embodiment of the present invention.

[0150] The above problems of the prior art can be solved by allowing ISP to be used even in blocks using transform skip. Alternatively, the problems of the prior art can be solved by allowing transform skip to be used in blocks that are partitioned into sub-blocks and use ISP. For example, as Figure 8 shown, transform skip can be used to determine the transformation of blocks enabled with ISP, and the use or non-use of transform skip can be further tested during encoding. In this case, it can be equivalently determined whether to use transform skip for all sub-blocks within the current block, or it can be independently determined for each sub-block whether to use transform skip.

[0151] In addition, for blocks using ISP, in addition to transform skip, the use of explicit Multiple Transform Selection (MTS) can also be implemented. In this case, the same transformation kernel determined by explicit MTS can be used for all sub-blocks within the current block, or the transformation kernel can be determined by explicit MTS for each sub-block. For example, for blocks enabled with ISP, only transform skip can be used, transform skip or implicit MTS can be used, or transform skip or explicit MTS can be used. Each of the above cases can be common to all sub-blocks within the current block, or can be applied to each sub-block.

[0152] Here, the explicit MTS refers to a method of determining a transform kernel to be used by sending an MTS indicator signal from a video encoding device to a video decoding device. The implicit MTS refers to a method of deriving a transform kernel based on a prediction mode, the size of a transform unit, etc.

[0153] Figure 9 FIG. is a schematic diagram showing the use of transform skip and multiple transform selection (MTS) for blocks to which ISP is applied according to at least one embodiment of the present invention.

[0154] When there is a residual signal within the current block (such as within the block shown in Figure 9 ), if ISP is used for prediction, a transform can be performed on each sub-block. Therefore, by applying the solution of the present invention, a transform method for each sub-block can be determined. For example, the rightmost sub-block can use transform skip, and other sub-blocks can use MTS. Additionally, if transform skip is enabled at a high level, the present invention can implement a method that utilizes both ISP and transform skip.

[0155] Preferred embodiments for solving the above-mentioned defects of the prior art are presented below.

[0156] The following embodiments are described centering on a video encoding device, but can also be implemented in the same or similar manner by a video decoding device.

[0157] <Embodiment 1> Determine whether to use transform skip and a transform kernel for sub-blocks within a current block for which ISP is enabled

[0158] In this embodiment, a video encoding device can use a combined transform skip and ISP by determining whether to use transform skip and a transform kernel for each sub-block within a current block for which ISP is enabled. In this case, the sub-blocks within the current block can (1) use transform skip or implicit MTS, or (2) use transform skip or explicit MTS.

[0159] In this embodiment, it can be determined equivalently for all sub-blocks within the current block whether to use transform skip and a transform kernel (Embodiment 1-1), or it can be determined on a per-sub-block basis whether to use transform skip and a transform kernel (Embodiment 1-2).

[0160] <Embodiment 1-1> Determine equivalently for all sub-blocks whether to use transform skip and a transform kernel

[0161] In this embodiment, in order to use transform skip and ISP together, a video coding device can equally determine whether to use transform skip and a transform kernel for all sub-blocks within a current block enabled with ISP. That is, the type of transform to be shared or whether to use transform skip can be shared, which is similar to how all sub-blocks within a current block enabled with ISP share prediction modes. In this case, the sub-blocks within the current block can use transform skip or implicit MTS (Embodiment 1-1-1), or use transform skip or explicit MTS (Embodiment 1-1-2).

[0162] <Embodiment 1-1-1>When using transform skip or implicit MTS

[0163] In this embodiment, a video coding device equally determines whether to use transform skip for all sub-blocks within a current block enabled with ISP. That is, if transform skip is used in the current block enabled with ISP, then transform skip is applied to all sub-blocks within the current block. In the case where transform skip indicates that no transform is to be performed, this embodiment produces the same result as applying transform skip to the entire current block. If transform skip is not used in the current block enabled with ISP, then each sub-block within the current block performs a transform according to implicit MTS.

[0164] In this embodiment, a video coding device signals a flag indicating whether to use transform skip to signal the use or non-use of transform skip. An example syntax with an improvement over the prior art regarding the transmission of the flag is shown in Table 7.

[0165] [Table 7]

[0166]

[0167]

[0168] In Table 7, even when using ISP, that is, when IntraSubPartitionsSplitType is ISP_HOR_SPLIT or ISP_VER_SPLIT, transform_skip_flag can be signaled to indicate whether to use transform skip.

[0169] <Embodiment 1-1-2>When using transform skip or explicit MTS

[0170] In this embodiment, the video encoding device equally determines whether to use transform skip for all sub-blocks within the current ISP-enabled block. That is, if transform skip is used in the current ISP-enabled block, transform skip is applied to all sub-blocks within the current block. In the case where transform skip indicates that no transform is to be performed, this embodiment produces the same result as applying transform skip to the entire current block. If transform skip is not used in the current ISP-enabled block, each sub-block within the current block performs a transform according to explicit MTS.

[0171] In this embodiment, the video encoding device signals a flag indicating whether to use transform skip to signal the use or non-use of transform skip. Table 8 shows an example syntax that has an improvement over the prior art in signaling the flag.

[0172] [Table 8]

[0173]

[0174] In Table 8, even when using ISP, that is, when IntraSubPartitionsSplitType is ISP_HOR_SPLIT or ISP_VER_SPLIT, transform_skip_flag can be signaled to indicate whether to enable transform skip. Additionally, when signaling transform_skip_flag = 0, the video encoding device can signal mts_idx to signal the MTS transform kernel. Table 9 shows an example syntax that has an improvement over the prior art in signaling the MTS transform kernel.

[0175] [Table 9]

[0176]

[0177]

[0178] <Embodiment 1-2> Determine whether to use transform skip and transform kernel on a per-sub-block basis

[0179] In this embodiment, in order to use ISP together, the video encoding device can determine whether to use transform skip and transform kernel for the current ISP-enabled block on a per-sub-block basis. Each sub-block within the current block can use transform skip or implicit MTS (Embodiment 1-2-1), or use transform skip or explicit MTS (Embodiment 1-2-2).

[0180] <Embodiment 1-2-1> When using transform skip or implicit MTS

[0181] In this embodiment, the video coding device may determine whether to use transform skip on a per-subblock basis for the current ISP-enabled block. Transform skip may be applied to the subblocks determined to have transform skip enabled, and for the subblocks determined to have transform skip disabled, a transform may be applied according to implicit MTS.

[0182] In this embodiment, the video coding device signals a flag indicating whether to use transform skip to signal the use or non-use of transform skip on a per-subblock basis. An example syntax with an improvement over the prior art regarding the transmission of the flag is shown in Table 10.

[0183] [Table 10]

[0184]

[0185] The video coding device signals a transform_skip_flag indicating whether transform skip is enabled for the current ISP-enabled block, and signals a transform_skip_subpartitions indicating whether transform skip is enabled on a per-subblock basis for the current ISP-enabled block. transform_skip_subpartitions may be an array of 0s or 1s as many as the subblocks, thereby indicating whether transform skip is to be used for each subblock in the current ISP-enabled block. Alternatively, transform_skip_subpartitions may be an index indicating one of the sets that contain whether each subblock has transform skip enabled or disabled.

[0186] The example assumes that there are four subblocks in the current ISP-enabled block and transform_skip_subpartitions is an array. For example, if transform_skip_subpartitions is signaled as 0001, transform skip may be used for the rightmost subblock, as shown in the example in Figure 9 Each digit in the array can be determined by an arrangement between the video coding device and the video decoding device as to which subblock it represents as a flag. Another example assumes that there are four subblocks in the current ISP-enabled block and transform_skip_subpartitions is an index. For example, if the value represented by transform_skip_subpartitions is defined as shown in Table 11 and transform_skip_subpartitions is signaled as 0, transform skip may be used for the rightmost subblock, as shown in Figure 9As shown in the example. It can be determined whether each number in the set indicated by the index is a flag representing which sub-block through the arrangement between the video encoding device and the video decoding device.

[0187] [Table 11]

[0188] transform_skip_subpartitions Transform skip using each sub-block 0 (0,0,0,1) 1 (0,1,0,1) 2 (0,0,1,1) … …

[0189] <Embodiment 1-2-2>When using transform skip or explicit MTS

[0190] In this embodiment, the video encoding device can determine whether to use transform skip for the current ISP-enabled block on a per-sub-block basis. Transform skip can be applied to the sub-blocks determined to use transform skip, and for the sub-blocks determined not to use transform skip, a transform can be applied according to explicit MTS.

[0191] In this embodiment, the video encoding device signals a flag indicating whether to use transform skip to signal whether transform skip is enabled on a per-sub-block basis. An example syntax with improved transmission of the flag over the prior art is shown in Table 12.

[0192] [Table 12]

[0193]

[0194]

[0195] The video coding device signals the transform_skip_flag indicating whether to use transform skip for the currently enabled ISP block, and performs transform_skip_subpartitions to signal on a per-subblock basis whether to use transform skip for the currently enabled ISP block. In Table 12, the behavior of transform_skip_subpartitions is similar to a process. That is, by performing transform_skip_subpartitions, the video coding device can signal the transform_skip_subpartition_flag, which is a flag indicating whether to use transform skip on a per-subblock basis. If the transform_skip_subpartition_flag is 1, transform skip is applied to the relevant subblock. However, if the transform_skip_subpartition_flag is 0, the video coding device can further signal the mts_idx to convey the type of MTS transform kernel. The above process can be repeated for all subblocks within the currently enabled ISP block. Based on the above transform_skip_subpartitions, the pseudo-code of the signaling representing the syntax of each subblock on the video coding device side is shown in Table 13.

[0196] [Table 13]

[0197]

[0198] On the other hand, replacing the signaling steps with the parsing in Table 13 can generate the pseudo-code of the parsing representing the syntax on the video decoding device side.

[0199] <Embodiment 2> Determine at a higher level whether to apply both ISP and transform skip

[0200] In this embodiment, the video coding device can determine at a higher level whether to apply both ISP and transform skip. As determined at a higher level, it is determined whether transform skip is applied to the subblocks within the currently enabled ISP block. If transform skip is enabled at a higher level, transform skip is applied to the subblocks within the currently enabled ISP block (Embodiment 2-1), or if the use of both ISP and transform skip is enabled at a higher level, transform skip is applied to the subblocks within the currently enabled ISP block (Embodiment 2-2).

[0201] <Embodiment 2-1> If transform skip is enabled at a higher level, apply transform skip to the currently enabled ISP block

[0202] In this embodiment, in order to use both ISP and transform skip, if transform skip is enabled at a higher level, the video coding device may apply transform skip to the current block for which ISP is enabled. Since if transform skip is enabled at the SPS level, the sub-blocks within the current block for which ISP is enabled may each use transform skip, the method of Embodiment 1 may determine whether to use transform skip for each sub-block of the current block for which ISP is enabled. However, if transform skip is disabled at the SPS level, ISP is not used. An example syntax showing an improvement over the prior art for this embodiment is shown in Table 14.

[0203] [Table 14]

[0204]

[0205] <Embodiment 2-2>If the use of both ISP and transform skip is enabled at a higher level, apply transform skip to the current block for which ISP is enabled

[0206] In this embodiment, in order to use both ISP and transform skip, if the use of both ISP and transform skip is enabled at a higher level, the video coding device may apply transform skip to the current block for which ISP is enabled. In other words, the use of both ISP and transform is considered a technique. The video coding device may independently use high-level syntax to coordinate the use of the combined ISP and transform. In order to indicate whether ISP and transform skip are consistently enabled at the SPS level, the video coding device may signal sps_isp_transform_skip_enabled_flag. If both transform skip and ISP are enabled at the SPS level, i.e., sps_transform_skip_enabled_flag = 1 and sps_isp_enabled_flag = 1, the video coding device may signal sps_isp_transform_skip_enabled_flag as shown in Table 15.

[0207] [Table 15]

[0208]

[0209] Furthermore, if sps_isp_transform_skip_enabled_flag = 1, for each sub-block within the current block for which ISP is enabled, it may be determined whether transform skip and the transform kernel are enabled according to the method of Embodiment 1.

[0210] Now refer to Figure 10 and Figure 11, a method for applying transform skip is described.

[0211] Figure 10 is a flowchart of a method for a video encoding device to encode a current block according to at least one embodiment of the present invention.

[0212] The video encoding device partitions the current block into sub-blocks based on the size of the current block and the direction of sub-block partitioning (S1000). The video encoding device can determine the size of the current block and the direction of sub-block partitioning of the current block from the perspective of rate-distortion optimization.

[0213] The video encoding device can apply transform skip and quantization to the residual block of the sub-block to generate a first quantized residual block (S1002).

[0214] The video encoding device can generate the residual block of each sub-block by subtracting the predicted block of each sub-block generated according to intra prediction from the original sub-block.

[0215] The video encoding device determines a transform kernel (S1004).

[0216] The video encoding device can implicitly derive the transform kernel. Alternatively, the video encoding device can explicitly determine the transform kernel from the perspective of optimizing rate-distortion.

[0217] The video encoding device applies the transform kernel and quantization to the residual block of the sub-block to generate a second quantized residual block (S1006).

[0218] The video encoding device determines a transform skip flag based on the first quantized residual block and the second quantized residual block (S1008).

[0219] The video encoding device can equivalently determine the transform skip flag for all sub-blocks of the current block based on the first quantized residual block and the second quantized residual block. From the perspective of rate-distortion optimization, the video encoding device can determine the transform skip flag. For example, if the first quantized residual block is the best, the transform skip flag can be determined to be true. On the contrary, if the second quantized residual block is the best, the transform skip flag can be determined to be false.

[0220] As another example, the video encoding device can determine the transform skip flag for each sub-block of the current block based on the first quantized residual block and the second quantized residual block. For example, if the first quantized residual block is the best, the transform skip flag can be determined to be true for the relevant sub-block. On the other hand, if the second quantized residual block is the best, the transform skip flag can be determined to be false for the relevant sub-block.

[0221] The video encoding device encodes the transform skip flag (S1010).

[0222] When coding transform skip flags for respective sub - blocks, the video coding device may code an array generated by combining transform skip flags for respective sub - blocks. Alternatively, the video coding device may code an index indicating one of a set including transform skip flags for respective sub - blocks.

[0223] The video coding device checks the transform skip flag and the transform kernel (S1012).

[0224] If the transform skip flag is false and the transform kernel is explicitly determined (S1012 is yes), the video coding device codes an index indicating the transform kernel (S1014).

[0225] If the transform kernel is explicitly determined and equivalently applied to all sub - blocks of the current block, the video coding device may code an index indicating the transform kernel.

[0226] As another example, if the transform kernel for each sub - block is explicitly determined and applied to each sub - block, the video coding device may code an index indicating the transform kernel for each sub - block.

[0227] On the other hand, if the transform skip flag is true (S1012 is no), the transform is skipped, so the video coding device may skip the step of coding an index indicating the transform kernel.

[0228] Furthermore, if the transform skip flag is false and the transform kernel is implicitly derived (S1012 is no), the video coding device may skip the step of coding an index indicating the transform kernel.

[0229] Figure 11 is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention.

[0230] The video decoding device partitions the current block into sub - blocks based on the size of the current block and the direction of sub - block partitioning (S1100).

[0231] The video decoding device may decode the size of the current block and the sub - block partitioning direction of the current block from the bitstream.

[0232] The video decoding device determines whether to use transform skip for sub - blocks of the current block (S1102).

[0233] The video decoding device decodes the transform skip flag from the bitstream. Then, based on the decoded transform skip flag, the video decoding device may equivalently determine whether to perform transform skip for all sub - blocks of the current block.

[0234] As another example, a video decoding device decodes a transform skip flag for each sub-block from a bitstream. Then, the video decoding device may determine whether to use transform skip for each sub-block based on the decoded transform skip flag. To decode the transform skip flags of the respective sub-blocks, the video decoding device may decode an array generated by combining the transform skip flags of the respective sub-blocks. Alternatively, the video decoding device may decode an index indicating one of a set containing the transform skip flags of the respective sub-blocks.

[0235] The video decoding device determines whether to use transform skip (S1104).

[0236] If transform skip is used (Yes at S1104), the video decoding device may generate a residual block for each sub-block and skip the inverse transform for the sub-block. Then, the video decoding device may generate a reconstructed block for each sub-block by summing the prediction block of each sub-block and the residual block of the sub-block.

[0237] On the other hand, if transform skip is not used (No at S1104), the following steps are performed.

[0238] The video decoding device determines a transform kernel for the sub-block (S1106).

[0239] The video decoding device may implicitly derive the transform kernel and then determine the transform kernels of all sub-blocks as the derived transform kernels. Alternatively, the video decoding device may decode an index indicating the transform kernel from the bitstream and determine the transform kernels of all sub-blocks as the transform kernels indicated by the decoded index.

[0240] As another example, the video decoding device may implicitly derive the transform kernel for each sub-block and then determine the transform kernel of each sub-block as the derived transform kernel. Alternatively, the video decoding device may decode an index indicating the transform kernel of each sub-block from the bitstream and then determine the transform kernel of each sub-block as the transform kernel indicated by the decoded index.

[0241] The video decoding device performs an inverse transform on the sub-block using the determined transform kernel (S1108).

[0242] The video decoding device may perform an inverse transform on the sub-block to generate a residual block for each sub-block. Then, the video decoding device may generate a reconstructed block for each sub-block by summing the prediction block of each sub-block and the inversely transformed sub-block to generate a reconstructed block for each sub-block.

[0243] Although the steps in the respective flowcharts described are sequential, these steps merely illustrate the technical ideas of some embodiments of the present invention. Therefore, those of ordinary skill in the art to which the present invention pertains can perform the steps by changing the order described in the respective drawings or by performing two or more steps in parallel. Accordingly, the steps in the respective flowcharts are not limited to the order shown in the order of occurrence.

[0244] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present invention are labeled as "…… unit" to emphasize the possibility of their independent implementation.

[0245] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, which can be read and executed by one or more processors. The non-volatile recording medium can include various types of recording devices that store data in a form readable by a computer system. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.

[0246] Although the exemplary embodiments of the present invention have been described for illustrative purposes, those of ordinary skill in the art to which the present invention pertains should understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the present invention. Therefore, the embodiments of the present invention have been described for the sake of simplicity and clarity. The scope of the technical ideas of the embodiments of the present invention is not limited by the illustration. Accordingly, those of ordinary skill in the art to which the present invention pertains should understand that the scope of the present invention should not be limited by the embodiments clearly described above, but by the claims and their equivalents.

[0247] Reference Numerals

[0248] 122: Intra Predictor

[0249] 140: Transformer

[0250] 165: Inverse Transformer

[0251] 530: Inverse Transformer

[0252] 542: Intra Predictor.

[0253] Cross - Reference to Related Applications

[0254] This application claims the priority and benefit of Korean Patent Application No. 10-2022-0157432, filed on November 22, 2022, and Korean Patent Application No. 10-2023-0126757, filed on September 22, 2023, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for reconstructing a current block by a video decoding device, the method comprising: Partitioning the current block into sub - blocks based on the size of the current block and the sub - block partitioning direction; Determining whether to use transform skip for the sub - blocks of the current block; And Checking whether to use transform skip, wherein the method further comprises, when not using transform skip: Determining the transform kernel for the sub - blocks.

2. The method according to claim 1, further comprising, when not using transform skip: Performing inverse transform on the sub - blocks by utilizing the determined transform kernel, and Among them, The method further comprises, when using transform skip: Skipping the inverse transform of the sub - blocks.

3. The method according to claim 1, wherein Determining whether to use transform skip includes: Decoding the transform skip flag from the bitstream; and Based on the transform skip flag, equivalently determining whether to use transform skip for all sub - blocks of the current block.

4. The method according to claim 3, wherein, Determining the transform kernel includes: Implicitly deriving the transform kernel; and Determining the derived transform kernel as the transform kernel for all sub - blocks.

5. The method according to claim 3, wherein, Determining the transform kernel includes: Decoding the index indicating the transform kernel from the bitstream; and Determining the transform kernel indicated by the index as the transform kernel for all sub - blocks.

6. The method according to claim 1, wherein, Determining whether to use transform skip includes: Decoding the transform skip flag for each sub - block from the bitstream; and Based on the transform skip flag, determining whether to use transform skip for each sub - block.

7. The method according to claim 6, wherein Decoding the transform skip flag includes: Decoding the array of transform skip flags combined from the transform skip flags for each sub - block, or decoding the index indicating one of the sets containing the transform skip flags for each sub - block.

8. The method according to claim 6, wherein, Determining the transform kernel includes: Implicitly deriving the transform kernel for each sub - block; and Determining the derived transform kernel as the transform kernel for each sub - block.

9. The method according to claim 6, wherein Determining the transform kernel includes: Decoding the index indicating the transform kernel for each sub - block from the bitstream; and Determining the transform kernel indicated by the index as the transform kernel for each sub - block.

10. A method for encoding a current block by a video encoding device, the method comprising: Partitioning the current block into sub - blocks based on the size of the current block and the sub - block partitioning direction; Generating a first quantized residual block by applying transform skip and quantization to the residual block of the sub - blocks; Implicitly deriving the transform kernel or explicitly determining the transform kernel; And Generating a second quantized residual block by applying the transform kernel and quantization to the residual block of the sub - blocks.

11. The method according to claim 10, further comprising: Based on the first quantized residual block and the second quantized residual block, equivalently determining the transform skip flag for all sub - blocks of the current block; And Encoding the transform skip flag.

12. The method according to claim 11, further comprising: Checking the transform skip flag; And wherein the method further comprises, when the transform skip flag is false and the transform kernel is explicitly determined and equivalently applied to all sub - blocks of the current block: Encoding the index indicating the transform kernel.

13. The method according to claim 10, further comprising: Determining the transform skip flag for each sub - block of the current block based on the first quantized residual block and the second quantized residual block; And Encoding the transform skip flag for each sub - block.

14. The method according to claim 13, further comprising: Checking the transform skip flag; and wherein, the method further includes, when a transform skip flag is false and a transform kernel for each sub-block is explicitly determined and applied to each sub-block: encoding an index indicating the transform kernel for each sub-block.

15. The method according to claim 13, wherein, Encoding the transform skip flag includes: encoding an array of transform skip flags from a combination of transform skip flags for each sub-block, or encoding an index indicating one of a set containing the transform skip flags for each sub-block.

16. A computer-readable recording medium storing a bitstream generated by a video coding method, the video coding method including: partitioning a current block into sub-blocks based on the size of the current block and a sub-block partitioning direction; generating a first quantized residual block by applying transform skip and quantization to a residual block of the sub-blocks; implicitly deriving a transform kernel or explicitly determining a transform kernel; and generating a second quantized residual block by applying the transform kernel and quantization to the residual block of the sub-blocks.

Citation Information

Patent Citations

  • Operably connected data and power transmission equipment having medical guide wires and catheters with sensors

    KR1020220157432A

  • System for providing non face-to-face order service

    KR1020230126757A