Video coding and decoding method and apparatus
The SBT mode in video coding standards optimizes encoding and decoding by defining residual blocks for specific transform units within coding units, addressing inefficiencies in existing standards and enhancing compression efficiency.
Patent Information
- Application Number
- PCT/EP2025/077904
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-23
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-30
AI Technical Summary
Existing video coding standards like VVC and ECM have limitations in transform unit configurations for sub-block transforms (SBT), which result in inefficient encoding and decoding processes due to the need to encode transform units with and without residuals, leading to suboptimal compression efficiency.
The introduction of a sub-block transform (SBT) mode that defines a residual block for a first transform unit extending across a medial portion of a coding unit, while omitting residual blocks for other units, allowing efficient encoding and decoding by associating motion information and prediction blocks only with the residual unit.
Enhances encoding and decoding efficiency by reducing redundant data transmission and storage, improving compression performance in video coding standards like VVC and ECM.
Smart Images

Figure EP2025077904_30042026_PF_FP_ABST
Abstract
Description
VIDEO CODING AND DECODING METHOD AND APPARATUSBACKGROUND
[0001] In order to facilitate storage and / or transmission of video images in an efficient manner, frames of a video image may be encoded in order to compress the video information prior to storage and / or transmission. The encoded information may then be decoded and, as a result, decompressed by a decoder in order to reconstruct the video images for display or the like.
[0002] Video images may be encoded and decoded in accordance with various video coding standards. For example, versatile video coding (VVC) is an international video coding standard. An enhanced compression model (ECM) that is built on top of VVC is also being developed as a future video coding standard. Both VVC and ECM are block-based video coding standards in which an input picture is divided into coding tree units (CTUs). Each CTU may be further divided into coding units (CUs) or blocks. A CU may be encoded in accordance with either an inter-coding mode or an intra-coding mode. In inter-coding, a CU is encoded by reference to a corresponding block in another picture, while in intra-coding mode a CU is encoded by reference to another block within the same picture.
[0003] If a CU is encoded in an inter-coding mode, the encoder identifies a temporal prediction block from among one or more reference pictures and motion information is then provided to the decoder to enable the decoder to identify the same temporal prediction block in the one or more reference pictures. Thus, the CU may be encoded by reference to the temporal prediction block that may then be referenced by the decoder in order to reconstruct the original CU. In an intra-coding mode, a spatial prediction block from the same picture may be derived from spatially neighboring pixels with the spatial prediction block then used by the encoder to encode the CU and also by the decoder to reconstruct the original CU. By referencing a prediction block as opposed to transmitting the CU itself, a CU may be encoded in either an inter-coding mode or an intra-coding mode in a compressed manner in order to increase the efficiency with which the CU is stored and / or transmitted.
[0004] Various types of motion information may be signaled by an encoder operating in an inter-coding mode including motion vectors, reference pictures, reference lists or the like to permit the temporal prediction block to be identified. In addition to identifying the temporal prediction block with motion information, an encoder also identifies the difference between the prediction block and the current CU with the difference termed a residual. Theencoder may encode the residual in another block structure termed a transform unit (TU) in VVC and ECM. A coding unit generally has one or more transform units associated with the coding unit. Additionally, a transform unit may include a luma transform block (TB) and two corresponding chroma transform blocks. In an instance in which a coding unit is coded with a skip mode such that the coded block flag (CBF) is set to 0, the residual of the associated transform blocks is assumed to also be 0.
[0005] Sub-block transform (SBT) is an inter-coding mode in VVC and ECM. A CU that is encoded using SBT includes a plurality of transform units. However, only one of the plurality of transform units has a residual and, as a result, is encoded, while the other transform units of the CU have no residual in, as a result, are not encoded. VVC supports a number of SBT modes with each mode having a different size and / or position of the transform unit relative to the coding unit. However, each SBT mode only supports a transform unit of a single size. For example, for a vertically oriented transform unit, one SBT mode only supports the transform unit having a width that is one half the overall width of the CU and another SBT mode only supports the transform unit having a width that is one quarter the overall width of the CU. Similarly, for a horizontally oriented transform unit, one SBT mode only supports the transform unit having a height that is one half the overall height of the CU and another SBT mode only supports the transform unit having a height that is one quarter the overall height of the CU.
[0006] As shown in Figures 1 A and IB, the transform units may extend vertically. The transform unit designated A in Figures 1 A and IB has a residual associated therewith while the other transform units designated B in Figures 1 A and IB has no residual. Thus, the residual represented by transform unit A is encoded, while the other transform unit need not be encoded as there is no residual associated therewith. Transform unit A may encompass half of the width w of the coding unit such that wl = U w. Alternatively, transform unit A may encompass a quarter of the width w of the coding unit such that wl = % w. As shown in Figures 1C and ID, the transform units may alternatively extend horizontally. The transform unit designated A in Figures 1C and ID has a residual associated therewith while the other transform units designated B in Figures 1C and 1C has no residual. Thus, the residual represented by transform unit A is encoded, while the other transform unit need not be encoded as there is no residual associated therewith. Transform unit A may encompass half of the height h of the coding unit such that hl = U h.
[0007] An extension to SBT has been proposed by JVET-AI0282 that introduces two additional SBT configurations, namely, a corner configuration as shown in Figure 2A and a center configuration as shown in Figure 2B. In these configurations, the transform unit that is in the corner of the coding unit 20, such as the upper left corner of the coding unit of Figure 2A, has an associated residual and, as a result, is encoded, while the transform unit representative of the remainder of the coding unit does not have a residual associated therewith and, as a result, is not encoded. The dimensions of the transform unit in the comer of the coding unit may be one quarter of the width of the coding unit for transform unit 22 or one half of the width of the coding unit for transform unit 24 as shown in solid and dashed lines, respectively, in Figure 2A. Similarly, the transform unit 26 that is in the center of the coding unit of Figure 2B has an associated residual and, as a result, be encoded, while the transform unit representative of the remainder of the coding unit does not have a residual associated therewith, as a result, is not encoded.BRIEF SUMMARY
[0008] In one aspect of the present disclosure, an apparatus is provided that includes at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to encode a coding unit using a sub-block transform that includes a plurality of transform units. Encoding the coding unit includes defining a residual block for a first transform unit of the plurality of transform units. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit. The instructions, when executed by the at least one processor, cause the apparatus to cause at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
[0009] The instructions, when executed by the at least one processor, cause the apparatus of an example embodiment to define the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. In one embodiment, the instructions, when executed by the at least one processor, cause the apparatus to encode the coding unit without defining a residual block for the second and thirdtransform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The instructions, when executed by the at least one processor, cause the apparatus of an example embodiment to define the residual block for the first transform unit by defining comers of the residual block for the first transform unit. In one embodiment, the instructions, when executed by the at least one processor, further cause the apparatus to cause information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
[0010] In another aspect, a method is provided that includes encoding a coding unit using a sub-block transform that includes a plurality of transform units. Encoding the coding unit includes defining a residual block for a first transform unit of the plurality of transform units. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit. The method also includes causing at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
[0011] In one embodiment, defining the residual block for the first transform unit includes defining the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. Encoding the coding unit in accordance with an example embodiment includes encoding the coding unit without defining a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. Defining the residual block for the first transform unit in accordance with an example embodiment includes defining corners of the residual block for the first transform unit. The method of an example embodiment also includes causing information regarding asub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
[0012] In a further aspect of the present disclosure, a computer program product is provided that includes at least one non-transitory computer-readable storage medium having computer-executable program code portions stored therein with the computer-executable program code portions including program code instructions configured to encode a coding unit using a sub-block transform that includes a plurality of transform units. The program code instructions are configured to encode the coding unit by defining a residual block for a first transform unit of the plurality of transform units. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit. The program code instructions are also configured to cause at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
[0013] The program code instructions are configured in accordance with an example embodiment to define the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. In one embodiment, the program code instructions are configured to encode the coding unit without defining a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The program code instructions are configured in accordance an example embodiment to define the residual block for the first transform unit by defining corners of the residual block for the first transform unit. In one embodiment, the program code instructions are also configured to cause information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
[0014] In yet another aspect, an apparatus is provided that includes means for encoding a coding unit using a sub-block transform that includes a plurality of transform units. The means for encoding the coding unit includes means for defining a residual block for a first transform unit of the plurality of transform units. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit. The apparatus also includes means for causing at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
[0015] In one embodiment, the means for defining the residual block for the first transform unit includes means for defining the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. The means for encoding the coding unit in accordance with an example embodiment includes means for encoding the coding unit without defining a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The means for defining the residual block for the first transform unit in accordance with an example embodiment includes means for defining comers of the residual block for the first transform unit. The apparatus of an example embodiment also includes means for causing information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
[0016] In one aspect of the present disclosure, an apparatus is provided that includes at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to receive motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a sub-block transform of the coding unit. The instructions, when executed by the at least one processor, also cause the apparatus to decode the coding unit using the prediction block for the coding unit and the residual blockfor the first transform unit. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
[0017] In one embodiment, the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. The instructions, when executed by the at least one processor, cause the apparatus of an example embodiment to receive the motion information, the prediction block and the information regarding the residual block for the first transform unit without receiving a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The instructions, when executed by the at least one processor, cause the apparatus of an example embodiment to decode the coding unit using the residual block for the first transform unit as defined by corners of the residual block for the first transform unit. In one embodiment, the instructions, when executed by the at least one processor, further cause the apparatus to receive information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit. The instructions, when executed by the at least one processor, further cause the apparatus of this embodiment to decode the coding unit using the residual block for the first transform unit based, at least in part, upon the sub-block transform mode.
[0018] In another aspect, a method is provided that includes receiving motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a subblock transform of the coding unit. The method also includes decoding the coding unit using the prediction block for the coding unit and the residual block for the first transform unit. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
[0019] In one embodiment, the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. The motion information, the prediction block and the information regarding the residual block for the first transform unit are received in one embodiment without receiving a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. Decoding the coding unit using the residual block for the first transform unit may include, in accordance with an example embodiment, decoding the coding unit using the residual block as defined by corners of the residual block for the first transform unit. The method of an example embodiment also includes receiving information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit. In this embodiment, decoding the coding unit using the residual block for the first transform unit is based, at least in part, upon the sub-block transform mode.
[0020] In a further aspect of the present disclosure, a computer program product is provided that includes at least one non-transitory computer-readable storage medium having computer-executable program code portions stored therein with the computer-executable program code portions including program code instructions configured to receive motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a subblock transform of the coding unit. The program code instructions are also configured to decode the coding unit using the prediction block for the coding unit and the residual block for the first transform unit. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
[0021] In one embodiment, the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. The program code instructions of an example embodiment are configured to receive the motioninformation, the prediction block and the information regarding the residual block for the first transform unit without receiving a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically. Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The program code instructions of an example embodiment are configured to decode the coding unit using the residual block for the first transform unit as defined by corners of the residual block for the first transform unit. In one embodiment, the program code instructions are also configured to receive information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit. The program code instructions of this example embodiment are also configured to decode the coding unit using the residual block for the first transform unit based, at least in part, upon the sub-block transform mode.
[0022] In yet another aspect, an apparatus is provided that includes means for receiving motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a sub-block transform of the coding unit. The apparatus also includes means for decoding the coding unit using the prediction block for the coding unit and the residual block for the first transform unit. The first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
[0023] In one embodiment, the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit. The motion information, the prediction block and the information regarding the residual block for the first transform unit are received in one embodiment without receiving a residual block for the second and third transform units. The first pair of opposed edges may include an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.Alternatively, the first pair of opposed edges may include a left edge and a right edge of the coding unit such that the first transform unit extends horizontally. In one embodiment, a size of the first transform unit is dependent upon a size of the coding unit. The means fordecoding the coding unit using the residual block for the first transform unit may include, in accordance with an example embodiment, means for decoding the coding unit using the residual block as defined by corners of the residual block for the first transform unit. The apparatus of an example embodiment also includes means for receiving information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit. In this embodiment, the means for decoding the coding unit using the residual block for the first transform unit includes means for decoding the coding unit based, at least in part, upon the sub-block transform mode.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In the following, certain embodiments of the present disclosure will be described in greater detail with reference to the accompanying drawings, in which:
[0025] Figures 1 A-1D are representations of the transform unit configurations of several different sub-block transform (SBT) modes;
[0026] Figures 2A and 2B are representations of corner and center configurations, respectively, of two additional SBT modes;
[0027] Figure 3 is a block diagram illustrating an encoder and a decoder that may be configured in accordance with an example embodiment of the present disclosure;
[0028] Figure 4 illustrates an apparatus that may be embodied by an encoder or a decoder in order to encode and / or decode blocks, respectively, in accordance with an example embodiment of the present disclosure;
[0029] Figures 5A and 5B depict transform units having residual blocks that extend across a medial portion of a coding unit in accordance with an example embodiment to the present disclosure;
[0030] Figure 6 is a flow chart illustrating encoding operations performed in accordance with an example embodiment at the present disclosure; and
[0031] Figure 7 is a flow chart illustrating encoding operations performed in accordance with an example embodiment at the present disclosure.DETAILED DESCRIPTION
[0032] The following embodiments are exemplary. Although the specification may refer to “an”, “one”, or “some” embodiment s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment s), or that a particularfeature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments. Further, when a particular feature, structure, or characteristic is described in connection of an embodiment, it is within the knowledge of one skilled in the art to apply such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. It shall be understood that although the terms “first,” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0033] For the purposes of the present disclosure, the phrases “at least one of A or B”, “at least one of A and B”, and “A and / or B” means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0034] Video images are generally encoded and therefore compressed by an encode prior to being stored and / or transmitted, such as to a decoder that then decodes and decompresses the video images. In this regard, Figure 3 depicts an encoder 30 configured to encode frames of a video image so as to compress the frames prior to transmission to, for example, a decoder 32 and / or storage, such as in a memory device 34. The decoder of Figure 1 is configured to receive the encoded frames of the video image from the memory device, such as following storage of the encoded frame, or from the encoder, either directly or via one or more intermediary devices. The decoder is then configured to decode and, therefore, decompress the frames in order to reconstruct a representation of the video image.
[0035] The encoding and decoding of video images may be done in accordance with one or more video coding standards. For example, versatile video coding (VVC) is an international video coding standard and an enhanced compression model (ECM) that is built on top of VVC is being developed as a future video coding standard. Both VVC and ECM are block-based video coding standards in which an input picture is divided into coding tree units (CTUs) and each CTU may be further divided into coding units (CUs) or blocks. A CU may be coded in accordance with either an inter-coding mode or an intra-coding mode. For intercoding in VVC and ECM, a sub-block transform (SBT) is a coding mode in which a coding unit that is to be encoded includes multiple transform units (TUs). Pursuant to the SBT, only one of the TUs is associated with a residual block and, as a result, is encoded, while the other TUs of the coding unit have no residual and, as a result, need not be encoded.
[0036] In the inter-coding mode, an encoder identifies, for a coding unit, a temporal prediction block in a reference picture of the video image. In the SBT coding mode, the encoder determines the residual block for a first TU of the coding unit in order to define the difference between that portion of the original block that corresponds positionally to the first TU and the corresponding portion of the temporal prediction block. However, in the SBT coding mode, the other TUs of the coding unit do not have an associated residual and, as a result, are not considered to be different than the corresponding portions of the temporal prediction block associated with the coding unit. In addition to determining the temporal prediction block associated with the coding unit and the residual block for the first TU of the coding unit, the encoder also determines motion information, such as motion vectors, reference pictures, reference lists or the like, in order to identify the temporal prediction block relative to the coding unit of the current image being encoded.
[0037] As shown in Figure 3, the encoder 30 is then configured to store and / or transmit the temporal prediction block, if not already stored and / or transmitted, as well as the motion information and the residual block associated with the first TU of the coding unit. In one embodiment, the encoder may be configured to store the temporal prediction block, the motion information and the residual block in memory 34 for subsequent access and retrieval by a decoder 32. Alternatively, the encoder may be configured to transmit the temporal prediction block, the motion information and the residual block to the decoder, either directly or indirectly. In order to decode and therefore decompress a block from an image frame, the decoder identifies the temporal prediction block associated with the coding unit based upon the motion information, such as the motion vectors that point from the coding unit to the temporal prediction block, and then decodes the block utilizing the residual block of the first TU in order to define a difference between the portion of the reconstructed block that corresponds positionally to the first TU and the corresponding portion of the temporal prediction block associated with the coding unit. As described below, the other TUs of the coding unit do not have residual blocks associated therewith such that no difference is identified between the portions of the reconstructed block that correspond positionally to the other TUs of the coding unit and the corresponding portions of the temporal prediction block associated with the coding unit. As such, the encoder is configured to encode the coding unit using a sub-block transform by defining a residual block for the first transform unit of the plurality of transform units without similarly defining residual blocks for the other transformunits, e.g., second and third transform units. Thus, the efficiency with which the encoder is configured to encode the coding unit is enhanced.
[0038] Figure 4 shows, by way of example, a block diagram of an apparatus 40. The apparatus 40 comprises, for example, at least one processor 42 and at least one memory 44 storing instructions 46 that, when executed by the at least one processor, cause the apparatus 40 at least to perform the method or methods as disclosed herein, and any of the embodiments thereof. In an example, the at least one memory and the instructions (e.g. a computer program code, software), are configured, with the at least one processor, to cause the apparatus 40 to perform the method or methods as disclosed herein, and any of the embodiments thereof.
[0039] A processor 42 may comprise circuitry, or be constituted as circuitry or circuitries, the circuitry or circuitries being configured to perform phases of methods in accordance with example embodiments described herein. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations, such as implementations in only analog and / or digital circuitry, and (b) combinations of hardware circuits and software, such as, as applicable: (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a user equipment, to perform various functions) and (c) hardware circuit(s) and or processor(s), such as a microprocessor s) or a portion of a microprocessor s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0040] The memory 44 may be implemented using any suitable data storage technology. The memory may comprise a database for storing data. The memory 44 may be at least in part external to apparatus 40 but accessible to apparatus 40.
[0041] The instructions 46 may be comprised in a computer readable medium or a non-transitory computer readable medium. A term non-transitory, as used herein, is a limitation ofthe medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., random access memory, RAM, vs. read only memory, ROM).
[0042] For example, the apparatus 40 is embodied by an encoder 30 and / or a decoder 32 as shown in Figure 3. The apparatus 40 as embodied by an encoder 30 and / or a decoder 32, e.g., as a chipset configured to control the encoder and / or decoder, respectively, may be caused or configured to perform at least the methods of Figures 6 and 7, respectively.
[0043] The apparatus 40 comprises a radio interface 49. The radio interface 49 may provide the apparatus 40 with communication capabilities. The radio interface 49 may comprise a receiver configured to receive information in accordance with at least one cellular or non-cellular standard. The radio interface 49 may comprise a transmitter configured to transmit information in accordance with at least one cellular or non-cellular standard. The receiver may comprise more than one receiver. The transmitter may comprise more than one transmitter. The radio interface 49 may comprise a transceiver configured to receive and transmit information in accordance with at least one cellular or non-cellular standard. The transceiver may comprise more than one transceiver.
[0044] The apparatus 40 may comprise a user interface 48 comprising, for example, at least one of a keypad, a microphone, a touch display, a display, a speaker, etc. The user interface 48 may be used to control the apparatus by the user. The user interface 48 may be external to the apparatus 10. For example, the apparatus 40 may be connected to another device, such as a computer, either via wireless or wired connection, and the apparatus 40 is controlled by the user via the computer.
[0045] In accordance with an example embodiment, an SBT coding mode having different SBT configurations is defined. As shown in Figures 5A and 5B, an SBT coding mode defines a plurality of transform units including a first transform unit TUI that extends across a medial portion of the coding unit and second and third transform units TU2, TU3 being on opposed sides of the first transform unit TUI. The first transform unit TUI has a residual block associated therewith that defines a difference between the portion of the reconstructed block that corresponds positionally to the first transform unit TU 1 and the prediction block of a reference picture. However, the second and third transform units TU2, TU3 are not associated with corresponding residual blocks and, as a result, do not have any difference defined relative to the prediction block of the reference picture. The first transform unit TUI extends across the medial portion of the coding unit to at least one of the first pair of opposed edges of the coding unit. For example, the first transform unit TUI may extendacross the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of the opposed edges of the coding unit, as shown in Figures 5A and 5B.
[0046] The first transform unit may extend either vertically or horizontally. As shown in Figure 5 A, the first transform unit TUI extends vertically such that the first pair of opposed edges includes upper and lower edges of the coding unit. In this embodiment, the first transform unit TUI extends vertically across the medial portion of the coding unit to either the upper edge or the lower edge of the coding unit and, in the illustrated embodiment, to each of the upper edge and lower edge of the coding unit. In another embodiment depicted in Figure 5B, the first transform unit TUI extends horizontally such that the first pair of opposed edges includes left and right edges of the coding unit. In this embodiment, the first transform unit TUI extends horizontally across the medial portion of the coding unit to either the left edge or the right edge of the coding unit and, in the illustrated embodiment, to each of the left edge and right edge of the coding unit.
[0047] In the illustrated embodiment, the first transform unit TUI has a rectangular shape that is defined by four comers. In the embodiment of Figure 5 A in which the first transform unit TUI extends vertically, the four corners are defined as (Aw / C,0), ((Bw / C)-l,0),(Aw / C,h-1) and ((Bw / C)-l,h-l) for a coding unit having a width of w and a height of h. In the embodiment illustrated in Figure 5A, the corners of the coding unit are also defined as (0,0), (w-1,0), (0,h-l) and (w-l,h-l). The parameters of A, B and C define the position of the first transform unit TUI relative to the coding unit as well as the dimensions of the first transform unit TUI and, as a result, of the corresponding residual block. Although the parameters A, B and C may have any of a variety of different values, the A, B and C parameters of one embodiment are 1, 3 and 4, respectively. In another embodiment, the A, B and C parameters are 3, 5 and 8, respectively.
[0048] In the embodiment of Figure 5B in which the first transform unit TUI extends horizontally, the four corners are defined as (0,Ah / C), (0, (Bh / C)-1), (w-1, Ah / C) and (w-1, (Bh / C)-1) for a coding unit having a width of w and a height of h. In the embodiment illustrated in Figure 5B, the comers of the coding unit are again defined as (0,0), (w-1,0), (0,h-l) and (w-l,h-l). The parameters of A, B and C define the position of the first transform unit TUI relative to the coding unit as well as the dimensions of the first transform unit TUI and, as a result, of the corresponding residual block. Although the parameters A, B and C may have any of a variety of different values, the A, B and C parameters of one embodiment are 1,3 and 4, respectively. In another embodiment, the A, B and C parameters are 3, 5 and 8, respectively.
[0049] Regardless of the horizontal or vertical orientation of the first transform unit TUI, the size of the first transform unit TU 1 may be dependent upon the size of the coding unit and, in one embodiment, the size of the first transform unit TU 1 varies inversely to the size of the coding unit. In this embodiment, the first transform unit TUI is smaller in an instance in which the coding unit is larger and, conversely, the first transform unit TUI is larger in an instance so the coding unit is smaller. With respect to a vertically extending first transform unit TUI as shown in Figure 5 A, the width w of the first transform unit varies inversely to the size of the coding unit. Conversely, with respect to a horizontally extending first transform unit TUI as shown in Figure 5 A, the height of the first transform unit varies inversely to the size of the coding unit. By way of example for a coding unit having a width of w and a height of h, the size of a vertically extending first transform unit TUI of one embodiment is w / 2 x h (with the top left corner of the first transform unit being located at (w / 4, 0)) in an instance in which the width of the coding unit is less than 32 pixels, but the size of the vertically extending first transform unit may be w / 4 x h (with the top left corner of the first transform unit being located at (3w / 8, 0)) in an instance in which the width of the coding unit is at least 32 pixels.
[0050] Referring now to Figure 6, the operations performed by an apparatus 40, such as illustrated in Figure 4 as embodied by an encoder 30 are depicted. As shown in block 60, the apparatus embodied by the encoder includes means, such as at least one processor 42 or the like, for encoding a coding unit using a sub-block transform including a plurality of transform units. The apparatus embodied by the encoder includes means, such as the processing circuitry or the like, for encoding the coding unit by defining a residual block for a first transform unit TU 1 of the plurality of transform units, such as by defining the corners of the residual block for the first transform unit TU 1. Although a residual block for the first transform unit TUI is defined in relation to encoding the coding unit, the apparatus embodied by the encoder includes means, such as the processing circuitry or the like, for encoding the coding unit without defining a residual block in the second and third transform units TU2, TU3. As a result, no difference is defined between the portions of the reconstructed block that correspond positionally to the second and third transform units TU2, TU3 and the corresponding portions of the prediction block.
[0051] The first transform unit TUI extends across the medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by second and third transform units TU2, TU3 of the plurality of transform units from a second pair of opposed edges of the coding unit. In one embodiment, the apparatus 40 embodied by the encoder 30 includes means, such as the at least one processor 42 or the like, for defining the residual block for the first transform unit TUI so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit.
[0052] As shown in block 62 of Figure 6, the apparatus 40 embodied by the encoder 30 also includes means, such as the at least one processor 42, the radio interface 49 or the like, for causing at least one of storage or transmission of the motion information, the prediction block identified by the motion information to be associated with the coding unit and information regarding the residual block (such as the values of the pixels that comprise the residual block) for the first transform unit TUI. In an example embodiment, the apparatus embodied by the encoder includes means, such as the at least one processor, the radio interface or the like, for causing at least one of storage or transmission not only of the motion information, the prediction block and the information regarding the residual block for the first transform unit TUI but also information regarding the sub-block transform mode that specifies that the first transform unit TUI extends across medial portion of the coding unit and that further specifies whether the first transform unit TUI has a vertical or horizontal configuration. In some embodiments, the apparatus embodied by the encoder may also be configured to store and / or transmit information defining the size and / or location of the first transform unit TUI, at least relative to the coding unit. By causing at least one of storage or transmission, the apparatus embodied by the encoder may be configured to provide the motion information, the prediction block and the information regarding the residual block (as well as any additional information) to a storage device for storage and / or to a transmitter for transmission. In other embodiments, the apparatus embodied by the encoder may be configured to cause the motion information, the predication block and the information regarding the residual block (as well as any additional information) to be stored and / or transmitted by storing and / or transmitting the motion information, the predication block and the information regarding the residual block (as well as any additional information) itself.
[0053] With reference to Figure 7, the operations performed by an apparatus 40 embodied by a decoder 32 are depicted. As shown in block 70, the apparatus embodied by the decoderincludes means, such as the radio interface 49, at least one processor 42 or the like, for receiving motion information, a prediction block identified by the motion information to be associated with the coding unit and information regarding a residual block for a first transform unit TUI (such as the values of the pixels that comprise the residual block) of a sub-block transform of the coding unit. In this regard, the apparatus embodied by the decoder may include means, such as the radio interface, the at least one processor or the like, for receiving the motion information, the prediction block and the information regarding the residual block for the first transform unit TUI without receiving a residual block for the second or third transform units TU2, TU3. As a result, no difference is defined between the portions of the reconstructed block that correspond positionally to the second and third transform units TU2, TU3 and the corresponding portions of the prediction block.
[0054] In one example embodiment, the apparatus 40 embodied by the decoder 32 also includes means, such as the radio interface 49, one or more processors 42 or the like, for receiving information regarding a sub-block transform mode in which the first transform unit TUI extends across the medial portions of the coding unit. For example, the SBT mode may specify whether the first transform unit TUI extends horizontally or vertically. In some embodiments, the apparatus embodied by the decoder may also be configured to receive information defining the size and / or location of the first transform unit TUI, at least relative to the coding unit.
[0055] As shown in block 72, the apparatus 40 embodied by the decoder 32 also includes means, such as at least one processor 42 or the like, for decoding the coding unit using the prediction block for the coding unit and the residual block for the first transform unit TU 1. In an embodiment in which information regarding the sub-block transform mode is also received, the apparatus embodied by the decoder includes means, such as the at least one processor or the like, for decoding the coding unit using the residual block for the first transform unit TUI in a manner based, at least in part, upon the sub-block transform mode. In this regard, the sub-block transform mode identifies to the decoder the type and orientation of the first transform unit TUI, such as by extending either horizontally or vertically through a medial portion of the coding unit, to permit the first transform unit TUI to be properly identified. In addition, in an embodiment in which information defining the size and / or location of the first transform unit TUI is also received, the apparatus embodied by the decoder may include means, such as the at least one processor or the like, for decoding the coding unit using the residual block for the first transform unit TUI in a manner based, atleast in part, upon the information defining the size and / or location of the first transform unit TUI.
[0056] By transmitting the residual block of the first transform unit TUI without transmitting residual blocks for the second and third transform units TU2, TU3, the encoding and decoding process as well as the storage and transmission of the compressed representation of the frame of the video image may be performed efficiently. Additionally, by defining the first transform unit TUI to extend across a medial portion of the coding unit, the likelihood that the most salient portions of an image are included within the first transform unit TUI and, as a result, are encoded and decoded with greater fidelity as a result of having a residual block associated therewith are enhanced. In this regard, the most salient objects of an image may be more likely to be within the medial portion of an image as opposed to lateral portions of the image. Thus, by defining a residual block associated with a first transform unit TUI that extends through the medial portion of a coding unit without correspondingly defining residual blocks associated with second and third transform units TU2, TU3 that are positioned laterally relative to the first transform unit TUI, the portion of a frame corresponding to the medial portion of the coding unit may be reproduced with more fidelity than the lateral portions, thereby serving to increase the efficiency of the encoding and decoding process by not expending the processing and communication resources to define residual blocks for the second and third transform units TU2, TU 3 that correspond to lateral portions of the frame while still reproducing the portions of the frame that are likely most salient by defining a residual block associated of the first transform unit TUI that extends through the medial portion of a coding unit.
[0057] Figures 6 and 7 are flowcharts illustrating a method according to an example embodiment. It will be understood that each block or signal and combination of blocks and signals may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other communication devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by the memory 44 of an apparatus 40 employing an example embodiment and executed by at least one process or 42. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (for example, hardware) to produce a machine, such that the resulting computer or other programmable apparatusimplements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0058] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, can be implemented by special purpose hardwarebased computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0059] As described above, at least some of the processes described herein may be carried out by an apparatus comprising means for carrying out at least some of the described processes. Means for performing method steps as disclosed herein may include software and / or hardware components of the apparatus 40. For example, the at least one processor 42, the memory 44, and the instructions 46 form means for carrying out the method or methods as disclosed herein, and any of the embodiments thereof. As used herein the term “means” is to be construed in singular form, i.e. referring to a single element, or in plural form, i.e. referring to a combination of single elements. Therefore, terminology “means for [performing A, B, C]”, is to be interpreted to cover an apparatus in which there is only one means for performing A, B and C, or where there are separate means for performing A, B and C, or partially or fully overlapping means for performing A, B, C. Further, terminology “means for performing A, means for performing B, means for performing C” is to be interpreted to cover an apparatus in which there is only one means for performing A, B and C, or where there are separate means for performing A, B and C, or partially or fully overlapping means for performing A, B, C.
[0060] Even though the present disclosure has been described above with reference to an example according to the accompanying drawings, it is clear that the present disclosure is not restricted thereto but can be modified in several ways within the scope of the appendedclaims. Therefore, all words and expressions should be interpreted broadly and they are intended to illustrate, not to restrict, the embodiment. It will be obvious to a person skilled in the art that, as technology advances, the inventive concept can be implemented in various ways. Further, it is clear to a person skilled in the art that the described embodiments may, but are not required to, be combined with other embodiments in various ways.
Claims
THAT WHICH IS CLAIMED:
1. An apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least:encoding a coding unit using a sub-block transform comprising a plurality of transform units, wherein encoding the coding unit comprises defining a residual block for a first transform unit of the plurality of transform units, wherein the first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit; andcausing at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
2. An apparatus according to Claim 1, wherein the instructions, when executed by the at least one processor, cause the apparatus to define the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit.
3. An apparatus according to any one of Claims 1 or 2, wherein the instructions, when executed by the at least one processor, cause the apparatus to encode the coding unit without defining a residual block for the second and third transform units.
4. An apparatus according to any one of Claims 1 to 3, wherein the first pair of opposed edges comprise an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.
5. An apparatus according to any one of Claim 1 to 3, wherein the first pair of opposed edges comprise a left edge and a right edge of the coding unit such that the first transform unit extends horizontally.
6. An apparatus according to any one of Claims 1 to 5, wherein a size of the first transform unit is dependent upon a size of the coding unit.
7. An apparatus according to any one of Claims 1 to 6, wherein the instructions, when executed by the at least one processor, cause the apparatus to define the residual block for the first transform unit by defining corners of the residual block for the first transform unit.
8. An apparatus according to any one of Claims 1 to 7, wherein the instructions, when executed by the at least one processor, further cause the apparatus to cause information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
9. A method comprising:encoding a coding unit using a sub-block transform comprising a plurality of transform units, wherein encoding the coding unit comprises defining a residual block for a first transform unit of the plurality of transform units, wherein the first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the plurality of transform units from a second pair of opposed edges of the coding unit; andcausing at least one of storage or transmission of motion information, a prediction block identified by the motion information to be associated with the coding unit, and information regarding the residual block for the first transform unit.
10. A method according to Claim 9, wherein defining the residual block for the first transform unit comprises defining the residual block for the first transform unit so as to extend across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit.
11. A method according to any one of Claims 9 or 10, wherein encoding the coding unit comprises encoding the coding unit without defining a residual block for the second and third transform units.
12. A method according to any one of Claims 9 to 11, wherein the first pair of opposed edges comprise an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.
13. A method according to any one of Claims 9 to 11, wherein the first pair of opposed edges comprise a left edge and a right edge of the coding unit such that the first transform unit extends horizontally.
14. A method according to any one of Claims 9 to 13, wherein a size of the first transform unit is dependent upon a size of the coding unit.
15. A method according to any one of Claims 9 to 14, wherein defining the residual block for the first transform unit comprises defining corners of the residual block for the first transform unit.
16. A method according to any one of Claims 9 to 15, further comprising causing information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit to be stored or transmitted.
17. An apparatus comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to perform at least:receiving motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a sub-block transform of the coding unit; anddecoding the coding unit using the prediction block for the coding unit and the residual block for the first transform unit, wherein the first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
18. An apparatus according to Claim 17, wherein the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the firstpair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit.
19. An apparatus according to any one of Claims 17 or 18, wherein the instructions, when executed by the at least one processor, cause the apparatus to receive the motion information, the prediction block and the information regarding the residual block for the first transform unit without receiving a residual block for the second and third transform units.
20. An apparatus according to any one of Claims 17 to 19, wherein the first pair of opposed edges comprise an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.
21. An apparatus according to any one of Claims 17 to 19, wherein the first pair of opposed edges comprise a left edge and a right edge of the coding unit such that the first transform unit extends horizontally.
22. An apparatus according to any one of Claims 17 to 21, wherein a size of the first transform unit is dependent upon a size of the coding unit.
23. An apparatus according to any one of Claims 17 to 22, wherein the instructions, when executed by the at least one processor, cause the apparatus to decode the coding unit using the residual block for the first transform unit as defined by corners of the residual block for the first transform unit.
24. An apparatus according to any one of Claims 17 to 23, wherein the instructions, when executed by the at least one processor, further cause the apparatus to receive information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit, and wherein the instructions, when executed by the at least one processor, further cause the apparatus to decode the coding unit using the residual block for the first transform unit based, at least in part, upon the sub-block transform mode.
25. A method comprising:receiving motion information, a prediction block identified by the motion information to be associated with a coding unit, and information regarding a residual block for a first transform unit of a sub-block transform of the coding unit; anddecoding the coding unit using the prediction block for the coding unit and the residual block for the first transform unit, wherein the first transform unit extends across a medial portion of the coding unit to at least one of a first pair of opposed edges of the coding unit while being spaced by transform blocks of second and third transform units of the sub-block transform from a second pair of opposed edges of the coding unit.
26. A method according to Claim 25, wherein the residual block for the first transform unit extends across the medial portion of the coding unit from one edge of the first pair of opposed edges of the coding unit to another edge of the first pair of opposed edges of the coding unit.
27. A method according to any one of Claims 25 or 26, wherein the motion information, the prediction block and the information regarding the residual block for the first transform unit are received without receiving a residual block for the second and third transform units.
28. A method according to any one of Claims 25 to 27, wherein the first pair of opposed edges comprise an upper edge and a lower edge of the coding unit such that the first transform unit extends vertically.
29. A method according to any one of Claims 25 to 27, wherein the first pair of opposed edges comprise a left edge and a right edge of the coding unit such that the first transform unit extends horizontally.
30. A method according to any one of Claims 25 to 29, wherein a size of the first transform unit is dependent upon a size of the coding unit.
31. A method according to any one of Claims 25 to 30, wherein decoding the coding unit using the residual block for the first transform unit comprises decoding the coding unit using the residual block as defined by corners of the residual block for the first transform unit.
32. A method according to any one of Claims 25 to 31, further comprising receiving information regarding a sub-block transform mode in which the first transform unit extends across the medial portion of the coding unit, and wherein decoding the coding unit using the residual block for the first transform unit is based, at least in part, upon the sub-block transform mode.
Citation Information
Patent Citations
One-level transform split
US20210352327A1
Methods and apparatus for implicit sub-block transform coding
WO2023246901A1