Position-dependent spatial variation transform for video coding

The use of spatial-varying transforms in video coding optimizes encoding efficiency by adapting transform types and positions within residual blocks, addressing the challenge of high-resolution media delivery without delay.

JP2025102804AActive Publication Date: 2025-07-08HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025040052
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-02-23
Filing Date
2025-03-13
Publication Date
2025-07-08
Estimated Expiration
2039-02-04

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing video data, particularly in providing high-resolution media on demand without significant transmission delays, and there is a need for improved encoding methods to enhance coding efficiency.

Method used

The implementation of a spatial-varying transform (SVT) that adapts the type and position of transform blocks based on candidate positions within residual blocks, using different transforms for different block locations to optimize encoding efficiency.

Benefits of technology

This approach increases coding efficiency by effectively utilizing different transforms based on block positions, reducing the size of video files and improving transmission speed for high-resolution media delivery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025102804000001_ABST
    Figure 2025102804000001_ABST
Patent Text Reader

Abstract

To provide a mechanism for position-dependent spatial variation transform for video coding (spatial varying transform, SVT).SOLUTION: A prediction block and a corresponding transform residual block are received at a decoder. The type of spatial variation transform (SVT) used to generate the transform residual block is determined. The position of the SVT relative to the transform residual block is also determined. The inverse of the SVT is applied to the transformed residual block to reconstruct the reconstructed residual block. The reconstructed residual block is then combined with the prediction block to reconstruct the image block.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] [Cross - Reference to Related Applications] This application claims the priority of U.S. Provisional Patent Application No. 62 / 634,613, entitled "Position Dependent Spatial Varying Transform for Video Coding", filed on Feb. 23, 2018 by Yin Zhao et al., and incorporates by reference the entire teachings and disclosures thereof.

[0002] [Description of Research or Development Sponsored by the Federal Government] Not applicable.

[0003] [Reference to Microfiche Appendix] Not applicable.

[0004] [Background Art] Video coding is a process of compressing video images into a smaller format. Video coding enables the encoded video to occupy less space when stored on a medium. Further, video coding supports streaming media. Specifically, content providers desire to provide media to end - users at even higher resolutions. Further, content providers desire to provide media on demand without making the user wait for a long time for such media to be transmitted to end - user devices such as televisions, computers, tablets, phones, etc. Advances in video coding compression support the reduction of the size of video files and thus support both of the above - mentioned goals when applied with corresponding content delivery systems.

Summary of the Invention

[0005] The first aspect relates to a method implemented in a computing device. The method includes steps of analyzing a bitstream by a processor of the computing device to obtain a prediction block and a transform residual block corresponding to the prediction block, determining by the processor a type of spatial-varying transform (SVT) used to generate the transform residual block, determining by the processor a position of the SVT with respect to the transform residual block, determining by the processor an inverse of the SVT based on the position of the SVT, applying by the processor the inverse of the SVT to the transform residual block to generate a reconstructed residual block, and combining by the processor the reconstructed residual block with the prediction block to reconstruct an image block.

[0006] The method facilitates an increase in the coding efficiency of SVT. In this regard, the transform block is located at various candidate positions with respect to the corresponding residual block. Accordingly, the disclosed mechanism uses different transforms for the transform block based on the candidate positions.

[0007] In a first implementation manner of the method according to the first aspect itself, the type of SVT is an SVT vertical (SVT-V) type or an SVT horizontal (SVT-H) type.

[0008] In a second implementation manner of the method according to the first aspect itself or any of the above implementation manners of the first aspect, the SVT-V type includes a height equal to the height of the transform residual block and a width equal to half of the width of the transform residual block, and the SVT-H type includes a height equal to half of the height of the transform residual block and a width equal to the width of the transform residual block.

[0009] In a third implementation manner of the method according to the first aspect itself or any of the above implementation manners of the first aspect, the svt_type_flag is parsed from the bitstream to determine the type of SVT.

[0010] In a fourth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, when only one type of SVT is allowed for the residual block, the type of SVT is determined by speculation.

[0011] In a fifth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the position index is analyzed from the bitstream to determine the position of the SVT.

[0012] In a sixth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the position index includes a binary code indicating a position from a set of candidate positions determined according to a candidate position step size (CPSS).

[0013] In a seventh implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the most likely position of the SVT is assigned the fewest number of bits in the binary code indicating the position index.

[0014] In an eighth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, when a single candidate position is available for SVT conversion, the position of the SVT is speculated by the processor.

[0015] In a ninth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, when the residual block is generated by template matching in the inter-prediction mode, the position of the SVT is speculated by the processor.

[0016] In a tenth implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the inverse discrete sine transform (DST) is used for the SVT vertical (SVT-V) type conversion located at the left boundary of the residual block.

[0017] In the 11th implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the inverse DST is used for the SVT horizontal (SVT-H) type conversion located at the upper boundary of the residual block.

[0018] In the 12th implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the inverse discrete cosine transform (DCT) is used for the SVT-V type conversion located at the right boundary of the residual block.

[0019] In the 13th implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, the inverse DCT is used for the SVT-H type conversion located at the lower boundary of the residual block.

[0020] In the 14th implementation form of the method in the form of the first aspect itself or any of the above implementation forms of the first aspect, when the right adjacent of the coding unit related to the reconstructed residual block is reconstructed and the left adjacent of the coding unit is not reconstructed, the samples in the reconstructed residual block are horizontally flipped before combining the reconstructed residual block with the prediction block.

[0021] The second aspect relates to a method implemented in a computing device. The method includes receiving a video signal from a video capture device, the video signal including image blocks; generating a prediction block and a residual block by a processor of the computing device to represent the image blocks; selecting a conversion algorithm for the SVT based on the position of the spatial variance transform (SVT) for the residual block by the processor; converting the residual block into a transformed residual block using the selected SVT by the processor; encoding the type of the SVT into a bitstream by the processor; encoding the position of the SVT into the bitstream by the processor; and encoding the prediction block and the transformed residual block into the bitstream for transmission to a decoder by the processor.

[0022] The method facilitates an increase in the encoding efficiency of the SVT. In this regard, the transform block is located at various candidate positions with respect to the corresponding residual block. Accordingly, the disclosed mechanism uses different transforms for the transform block based on the candidate positions.

[0023] In the form of a first implementation manner of the method according to the first aspect itself, the type of SVT is an SVT vertical (SVT-V) type or an SVT horizontal (SVT-H) type.

[0024] In the form of a second implementation manner of the method according to the second aspect itself or in the form of any of the above implementation manners of the second aspect, the SVT-V type includes a height equal to the height of the residual block and a width equal to half of the width of the residual block.

[0025] In the form of a third implementation manner of the method according to the second aspect itself or in the form of any of the above implementation manners of the second aspect, the SVT-H type includes a height equal to half of the height of the residual block and a width equal to the width of the residual block.

[0026] In the form of a fourth implementation manner of the method according to the second aspect itself or in the form of any of the above implementation manners of the second aspect, the position of the SVT is encoded with a position index.

[0027] In the form of a fifth implementation manner of the method according to the second aspect itself or in the form of any of the above implementation manners of the second aspect, the position index includes a binary code indicating a position from a set of candidate positions determined according to a candidate position step size (CPSS).

[0028] In the form of a sixth implementation manner of the method according to the second aspect itself or in the form of any of the above implementation manners of the second aspect, the most likely position of the SVT is assigned as the least number of bits in the binary code indicating the position index.

[0029] In a seventh implementation form of the method in the form of the second aspect itself or any of the above implementation forms of the second aspect, the discrete sine transform (DST) algorithm is used by the processor for the SVT vertical (SVT-V) type transform located at the left boundary of the residual block.

[0030] In an eighth implementation form of the method in the form of the second aspect itself or any of the above implementation forms of the second aspect, the DST algorithm is selected by the processor for the SVT horizontal (SVT-H) type transform located at the upper boundary of the residual block.

[0031] In a ninth implementation form of the method in the form of the second aspect itself or any of the above implementation forms of the second aspect, the discrete cosine transform (DCT) algorithm is selected by the processor for the SVT-V type transform located at the right boundary of the residual block.

[0032] In a tenth implementation form of the method in the form of the second aspect itself or any of the above implementation forms of the second aspect, the DCT algorithm is selected by the processor for the SVT-H type transform located at the lower boundary of the residual block.

[0033] In an eleventh implementation form of the method in the form of the second aspect itself or any of the above implementation forms of the second aspect, when the right adjacent of the coding unit related to the residual block is coded and the left adjacent of the coding unit is not coded, the method further includes a step of horizontally inverting the samples in the residual block by the processor before the processor converts the residual block into a transformed residual block.

[0034] A third aspect relates to an encoding apparatus including a receiver configured to receive an encoded picture or to receive a bitstream to be decoded, a transmitter coupled to the receiver, the transmitter being configured to transmit the bitstream to a decoder or to transmit the decoded picture to a display, a memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions, and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory and to execute any of the methods of the above aspects or implementation manners.

[0035] The encoding apparatus facilitates an increase in the encoding efficiency of SVT. In this regard, the transform block is located at various candidate positions with respect to the corresponding residual block. Accordingly, the disclosed mechanism uses different transforms for the transform block based on the candidate positions.

[0036] In a first implementation manner of the apparatus according to the third aspect itself, the apparatus further includes a display configured to display an image.

[0037] A fourth aspect relates to a system including an encoder and a decoder communicating with the encoder. The encoder or the decoder includes any of the encoding apparatuses of the above aspects or implementation manners.

[0038] The system facilitates an increase in the encoding efficiency of SVT. In this regard, the transform block is located at various candidate positions with respect to the corresponding residual block. Accordingly, the disclosed mechanism uses different transforms for the transform block based on the candidate positions.

[0039] A fifth aspect relates to an encoding means including a receiving means configured to receive an encoded picture or to receive a bitstream to be decoded, a transmitting means coupled to the receiving means, the transmitting means being configured to transmit the bitstream to a decoder or to transmit the decoded picture to a display means, a storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions, and a processing means coupled to the storage means, the processing means being configured to execute the instructions stored in the storage means and to execute a method in any of the above aspects or implementation manners.

[0040] The encoding means facilitates an increase in the encoding efficiency of SVT. In this regard, the transform block is located at various candidate positions with respect to the corresponding residual block. Accordingly, the disclosed mechanism uses different transforms for the transform block based on the candidate positions.

[0041] For the purpose of clarity, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present disclosure.

[0042] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.

Brief Description of the Drawings

[0043] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0044] First, exemplary implementation manners of one or more embodiments are provided below. However, it should be understood that the disclosed system and / or method may be implemented using any number of techniques, whether currently known or existing. The disclosure should in no way be limited to the exemplary implementation manners, drawings, and techniques shown and described herein, but may be modified within the scope of the appended claims, together with the full scope of their equivalents.

[0045] The standard currently known as High Efficiency Video Coding (HEVC) is an advanced video coding system developed under the Joint Collaborative Team on Video Coding (JCT-VC) group of video coding experts from the International Telecommunication Union - Telecommunication Standardization Sector (ITU-T) study group. Details regarding the HEVC standard can be found in ITU-T Rec. H.265 and International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) 23008-2 (2013), High efficiency video coding, final draft approval Jan. 2013 (officially published by ITU-T in June 2013 and by ISO / IEC in November 2013), which is incorporated herein by reference. An overview of HEVC can be found in G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, 「Overview of the High Efficiency Video Coding (HEVC) Standard」, IEEE Trans. Circuits and Systems for Video Technology, Vol. 22, No. 12, pp. 1649-1668, December 2012, which is incorporated herein by reference.

[0046] FIG. 1 is a block diagram showing an exemplary encoding system 10 that can utilize video encoding techniques such as encoding using an SVT mechanism. As shown in FIG. 1, encoding system 10 includes a source device 12 that provides encoded video data to be decoded by a destination device 14 at a later time. In particular, source device 12 may provide the video data to destination device 14 via a computer-readable medium 16. Source device 12 and destination device 14 may include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephones such as so-called "smart" phones, so-called "smart" pads, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, source device 12 and destination device 14 may be equipped for wireless communication.

[0047] Destination device 14 may receive the encoded video data to be decoded via computer-readable medium 16. Computer-readable medium 16 may include any type of medium or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, computer-readable medium 16 may include a communication medium that enables source device 12 to transmit the encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be useful in facilitating communication from source device 12 to destination device 14.

[0048] In some examples, the encoded data may be output from the output interface 22 to a storage device. Similarly, the encoded data may be accessed from the storage device by an input interface. The storage device may include any of a variety of distributed or locally accessible data storage media such as a hard drive, Blu-ray disk, digital video disk (DVD), Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device may correspond to a file server or other intermediate storage device that can store the encoded video generated by the source device 12. The destination device 14 may access the video data stored in the storage device via streaming or downloading. The file server may be any type of server that can store the encoded video data and transmit the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 may access the encoded video data through any standard data connection including an Internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.

[0049] The technology disclosed herein is not necessarily limited to wireless applications or settings. The technology may be applied to video encoding that supports any of a variety of multimedia applications, such as wireless television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions such as dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the encoding system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0050] In the example of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to this disclosure, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 may be configured to apply techniques for video encoding. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 12 may receive video data from an external video source such as an external camera. Similarly, the destination device 14 may interface with an external display device rather than including an integrated display device.

[0051] The illustrated encoding system 10 of FIG. 1 is merely an example. The techniques for video encoding may be performed by any digital video encoding and / or decoding device. The techniques of this disclosure are generally performed by a video encoding device, but the techniques may also typically be performed by a video encoder / decoder, typically referred to as a “codec”. Further, the techniques of this disclosure may also be performed by a video processor. The video encoder and / or decoder may be a graphics processing unit (GPU) or similar device.

[0052] The source device 12 and the destination device 14 are merely examples of an encoding device such that the source device 12 generates encoded video data for transmission to the destination device 14. In some examples, the source device 12 and the destination device 14 may operate in a substantially symmetric manner such that each of the source and destination devices 12, 14 includes video encoding and decoding components. Accordingly, the encoding system 10 may support one-way or two-way video transmission between the video devices 12, 14, for example, for video streaming, video playback, video broadcast, or video telephony.

[0053] The video source 18 of the source device 12 may include a video capture device such as a video camera, a video archive including previously captured video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 18 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video.

[0054] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 may form a so-called camera phone or video phone. However, as described above, the techniques described in this disclosure may generally be applicable to video encoding and may also be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video information may then be output onto the computer-readable medium 16 by the output interface 22.

[0055] The computer-readable medium 16 may include a transient medium such as a wireless broadcast or wired network transmission, or a storage medium (i.e., a non-transient storage medium) such as a hard disk, flash drive, compact disk, digital video disk, Blu-ray disk, or other computer-readable medium. In some examples, a network server (not shown) may receive the encoded video data from the source device 12 and provide the encoded video data to the destination device 14, for example, via a network transmission. Similarly, a computing device of a media manufacturing facility such as a disk stamping facility may receive the encoded video data from the source device 12 and manufacture a disk containing the encoded video data. Thus, the computer-readable medium 16 may be understood to include one or more computer-readable media in various forms in various examples.

[0056] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information on the computer-readable medium 16 may include syntax information defined by the video encoder 20, which is also used by the video decoder 30 and includes syntax elements that describe the characteristics and / or processing of blocks and other encoded units, such as groups of pictures (GOPs). The display device 32 displays the decoded video data to the user and may include any of various display devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0057] Video encoder 20 and video decoder 30 may operate according to a video coding standard such as the currently under - development High Efficiency Video Coding (HEVC) standard, and may conform to the HEVC Test Model (HM). Alternatively, video encoder 20 and video decoder 30 may operate according to other proprietary or industry standards such as, alternatively, the International Telecommunication Union - Telecommunication Standardization Sector (ITU - T) H.264 standard, also known as Moving Picture Expert Group (MPEG) - 4, Part 10, Advanced Video Coding (AVC), H.265 / HEVC, or extensions of such standards. However, the technology of this disclosure is not limited to any particular coding standard. Other examples of video coding standards include MPEG - 2 and ITU - T H.263. Although not illustrated in FIG. 1, in some aspects, video encoder 20 and video decoder 30 may be integrated with an audio encoder and decoder, respectively, and may include an appropriate multiplexer - demultiplexer (MUX - DEMUX) unit or other hardware and software to process the encoding of both audio and video in a common data stream or separate data streams. When applicable, the MUX - DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0058] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) within the respective device. Devices including the video encoder 20 and / or the video decoder 30 may include integrated circuits, microprocessors, and / or wireless communication devices such as cellular telephones.

[0059] FIG. 2 is a block diagram illustrating an example of a video encoder 20 that may implement video encoding techniques. The video encoder 20 may perform intra and inter encoding of video blocks within a video slice. Intra encoding relies on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame or picture. Inter encoding relies on temporal prediction to reduce or remove temporal redundancy in the video within adjacent frames or pictures of a video sequence. The intra mode (I mode) may indicate any of several spatial-based encoding modes. Inter modes, such as uni-directional (also known as uni prediction) prediction (P mode) or bi-directional (also known as bi prediction) (B mode), may indicate any of several temporal-based encoding modes.

[0060] As shown in FIG. 2, video encoder 20 receives the current video block within the video frame to be encoded. In the example of FIG. 2, video encoder 20 includes a mode selection unit 40, a reference frame memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. Next, mode selection unit 40 includes a motion compensation unit 44, a motion estimation unit 42, an intra prediction unit (also known as intra prediction) unit 46, and a partitioning unit 48. For video block reconstruction, video encoder 20 also includes an inverse quantization unit 58, an inverse transform unit 60, and an adder 62. A deblocking filter (not shown in FIG. 2) may also be included to filter the block boundaries to remove blocky artifacts from the reconstructed video. Optionally, the deblocking filter typically filters the output of adder 62. Additional filters (in-loop or post-loop) may also be used in addition to the deblocking filter. Such filters are not shown for simplicity, but may optionally filter the output of adder 50 (as an in-loop filter).

[0061] During the encoding process, video encoder 20 receives the video frame or slice to be encoded. The frame or slice may be divided into a plurality of video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-prediction encoding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. Alternatively, intra prediction unit 46 may perform intra-prediction encoding of the received video block with respect to one or more adjacent blocks in the same frame or slice as the block to be encoded to provide spatial prediction. Video encoder 20 may execute a plurality of encoding paths, for example, to select an appropriate encoding mode for each block of video data.

[0062] Furthermore, the partitioning unit 48 may partition a block of video data into sub-blocks based on an evaluation of a previous partitioning method in a previous coding path. For example, the partitioning unit 48 may first partition a frame or slice into largest coding units (LCUs), and then partition each of the LCUs into sub-coding units (sub-CUs) based on rate-distortion analysis (e.g., rate-distortion optimization). The mode selection unit 40 may further generate a quadtree data structure indicating the partitioning of the LCUs into sub-CUs. A leaf node CU of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).

[0063] The present disclosure uses the term "block" to refer to any of a CU, a PU, or a TU in the context of HEVC, or a similar data structure in the context of other standards (e.g., macroblocks and their sub-blocks in H.264 / AVC). A CU includes a coding node, a PU, and a TU associated with the coding node. The size of a CU corresponds to the size of the coding node and is square in terms of shape. The size of a CU may range from 8×8 pixels to a size of a tree block having a maximum of 64×64 pixels or more. Each CU may include one or more PUs and one or more TUs. Syntax data associated with a CU may describe, for example, the partitioning of the CU into one or more PUs. The partitioning mode may vary depending on whether the CU is coded in skip or direct mode, intra prediction mode, or inter prediction (also known as inter prediction) mode. A PU may be partitioned non-squarely in terms of shape. Syntax data associated with a CU may also describe, for example, the partitioning of the CU into one or more TUs according to a quadtree. A TU may be square or non-square (e.g., rectangular) in terms of shape.

[0064] The mode selection unit 40 may select, for example, one of the encoding modes, i.e., intra or inter, based on the error result, and provide the resulting intra or inter encoded block to the adder 50 to generate residual block data, and provide the encoded block for use as a reference frame to the adder 62 to reconstruct it. The mode selection unit 40 also provides syntax elements such as motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy encoding unit 56.

[0065] The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion for a video block. The motion vector may indicate, for example, the displacement of the PU of the video block within the current video frame or picture with respect to the predicted block within the reference frame (or other encoded unit) with respect to the current block (or other encoded unit) encoded within the current frame. The predicted block is a block that is found to closely match the block to be encoded with respect to the pixel difference determined by the sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. In some examples, the video encoder 20 may calculate values for sub-integer pixel positions of the reference picture stored in the reference frame memory 64. For example, the video encoder 20 may interpolate values at 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 may perform a motion search for both full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0066] The motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a prediction block in a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 transmits the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0067] The motion compensation performed by the motion compensation unit 44 may include retrieving or generating a prediction block based on the motion vector determined by the motion estimation unit 42. Similarly, in some examples, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated. Upon receiving a motion vector for a PU of the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference picture lists. The adder 50 forms a residual video block and a pixel difference value by subtracting the pixel values of the prediction block from the pixel values of the encoded current video block, as described below. Generally, the motion estimation unit 42 performs motion estimation on the luma component, and the motion compensation unit 44 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The mode selection unit 40 may also generate syntax elements related to the video block and the video slice for use by the video decoder 30 when decoding the video block of the video slice.

[0068] As described above, the intra prediction unit 46 may perform intra prediction on the current block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44. In particular, the intra prediction unit 46 may determine an intra prediction mode for use in encoding the current block. In some examples, the intra prediction unit 46 may encode the current block using various intra prediction modes, for example, between separate encoding paths, and the intra prediction unit 46 (or in some examples, the mode selection unit 40) may select an appropriate intra prediction mode for use from among the tested modes.

[0069] For example, the intra prediction unit 46 may calculate rate-distortion values using rate-distortion analysis for various tested intra prediction modes and select an intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to generate the encoded block, and the bit rate (i.e., the number of bits) used to generate the encoded block. The intra prediction unit 46 may calculate a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0070] Furthermore, the intra prediction unit 46 may be configured to encode depth blocks of a depth map using a depth modeling mode (DMM). The mode selection unit 40 may determine, for example, using rate-distortion optimization (RDO), whether available DMM modes produce better encoding results than intra prediction modes and other DMM modes. The data of the texture image corresponding to the depth map may be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may also be configured to perform inter prediction on depth blocks of the depth map.

[0071] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode. The video encoder 20 may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of encoding contexts for various blocks, and indications of the most likely intra prediction mode, intra prediction mode index table, and modified intra prediction mode index table to be used for each context in the transmitted bitstream.

[0072] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the original encoded video block. The adder 50 represents a component or components that perform this subtraction operation.

[0073] The transform processing unit 52 applies a transform such as a discrete cosine transform (DCT) or a conceptually similar transform to the residual block to generate a video block including residual transform coefficient values. The transform processing unit 52 may perform other transforms conceptually similar to the DCT. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used.

[0074] The conversion processing unit 52 applies the conversion to the residual blocks and generates a block of residual conversion coefficients. The conversion may convert the residual information from the pixel value domain to a conversion domain such as the frequency domain. The conversion processing unit 52 may send the resulting conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0075] Following quantization, the entropy coding unit 56 entropy codes the quantized conversion coefficients. For example, the entropy coding unit 56 may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context may be based on neighboring blocks. Following the entropy coding by the entropy coding unit 56, the coded bit stream may be sent to other devices (e.g., the video decoder 30), or archived for later transmission or retrieval.

[0076] The inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct a residual block in the pixel domain for later use as a reference block, for example. The motion compensation unit 44 may calculate a reference block by adding the residual block to a predicted block of one of the frames in the reference frame memory 64. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-pixel values for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. The reconstructed video block may be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-coding blocks within subsequent video frames.

[0077] FIG. 3 is a block diagram showing an example of a video decoder 30 that can implement video encoding techniques. In the example of FIG. 3, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. In some examples, the video decoder 30 may execute a decoding path that generally opposes the encoding path described with respect to the video encoder 20 (FIG. 2). The motion compensation unit 72 may generate prediction data based on the motion vectors received from the entropy decoding unit 70, while the intra prediction unit 74 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 70.

[0078] During the decoding process, video decoder 30 receives from video encoder 20 an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. Entropy decoding unit 70 of video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 70 transfers the motion vectors and other syntax elements to motion compensation unit 72. Video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.

[0079] When a video slice is encoded as an intra-coded (I) slice, intra prediction unit 74 may generate prediction data for video blocks of the current video slice based on the signaled intra prediction mode and data from previously decoded blocks of the current frame or picture. When a video frame is encoded as an inter-coded (e.g., B, P, or GPB) slice, motion compensation unit 72 generates a prediction block for video blocks of the current video slice based on the motion vectors and other syntax elements received from entropy decoding unit 70. The prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, namely, list 0 and list 1, using default construction techniques based on the reference pictures stored in reference frame memory 82.

[0080] The motion compensation unit 72 analyzes the motion vector and other syntax elements to determine prediction information about the video blocks of the current video slice, and uses the prediction information to generate a prediction block for the decoded current video block. For example, the motion compensation unit 72 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to encode the video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more of the reference picture lists for the slice, the motion vector for each inter-encoded video block of the slice, the inter prediction state for each inter-encoded video block of the slice, and other information for decoding the video blocks within the current video slice.

[0081] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 may use the interpolation filter used by the video encoder 20 during the encoding of the video block to calculate the interpolation value for the sub-integer pixels of the reference block. In this case, the motion compensation unit 72 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the prediction block.

[0082] The data of the texture image corresponding to the depth map may be stored in the reference frame memory 82. The motion compensation unit 72 may also be configured to perform inter prediction on the depth blocks of the depth map.

[0083] Disclosed herein are various mechanisms for increasing the encoding efficiency of SVT. As described above, the transform block may be located at various candidate positions with respect to the corresponding residual block. The disclosed mechanisms use different transforms for the transform block based on the candidate positions. For example, the inverse discrete sine transform (DST) can be applied to a candidate position covering the lower right corner of the residual block. Also, the inverse DCT can be applied to a candidate position covering the upper left corner of the residual block. This mechanism can be advantageous because generally DST is more efficient than DCT for transforming residual blocks having more residual information distributed in the lower right corner, while generally DCT is more efficient than DST for transforming residual blocks having more residual information distributed in the upper left corner. Also, it should be noted that the lower right corner of the residual block statistically contains more residual information in most cases. The disclosed mechanisms also support horizontally reversing the residual samples of the transform block in some cases. For example, the residual samples may be horizontally reversed / inverted after the inverse transform block is applied. This may occur when the adjacent block on the right side of the current residual block has already been reconstructed and the adjacent block on the left side of the current residual block has not been reconstructed. This may also occur when the inverse DST is used as part of the corresponding transport block. This approach provides greater flexibility for encoding the information closest to the reconstructed block, which results in a reduction of the corresponding residual information. The disclosed mechanisms also support context encoding of the candidate position information for the transform block. The prediction information corresponding to the residual block may be used to encode the candidate position information. For example, in some cases, the residual block may correspond to a predicted block generated in the template matching mode. Further, the template used in the template matching mode may be selected based on the spatially adjacent reconstructed regions of the residual block. In such cases, the lower right portion of the residual block may contain more residual information than other portions of the residual block. Therefore, the candidate position covering the lower right portion of the residual block is most likely to be selected as the best position for transformation.Therefore, when the residual block is related to template matching-based inter prediction, only one candidate position may be made available for the residual block, and / or other context encoding techniques may be used for encoding the position information for the conversion.

[0084] FIG. 4 is a schematic diagram of an exemplary intra prediction mode 400 used for video encoding in the HEVC model. The video compression method utilizes data redundancy. For example, most images contain groups of pixels that have the same or similar color and / or light as adjacent pixels. As a specific example, an image of the night sky may include a large area of black pixels and a cluster of white pixels representing stars. Intra prediction mode 400 utilizes these spatial relationships. Specifically, a frame can be divided into a series of blocks containing samples. Instead of transmitting each block, the light / color of a block can be predicted based on its spatial relationship to reference samples within adjacent blocks. For example, the encoder may indicate that the current block contains the same data as the reference samples within a previously encoded block located at the upper left corner of the current block. The encoder may then encode the prediction mode instead of the values of the block. This significantly reduces the size of the encoding. As shown by intra prediction mode 400, the upper left corner corresponds to prediction mode 18 in HEVC. Thus, the encoder can simply store prediction mode 18 for the block instead of encoding the pixel information for the block. As shown, intra prediction mode 400 in HEVC includes 33 angular prediction modes from prediction mode 2 to prediction mode 34. Intra prediction mode 400 also includes intra - planar mode 0 for predicting flat regions and intra - direct current (DC) mode 1. Intra - planar mode 0 predicts the block as an amplitude plane having vertical and horizontal gradients derived from adjacent reference samples. Intra - DC mode 1 predicts the block as the average value of adjacent reference samples. Intra prediction mode 400 may be used to signal the luma (e.g., light) component of a block. Intra prediction can also be applied to chroma (e.g., color) values. In HEVC, chroma values are predicted by using the planar mode, angle 26 mode (e.g., vertical), angle 10 mode (e.g., horizontal), intra - DC, and derived mode, and the derived mode predicts the correlation between the chroma component and the luma component encoded by intra prediction mode 400.

[0085] The intra prediction mode 400 for an image block is stored by an encoder as prediction information. It should be noted that although the intra prediction mode 400 is used for prediction in a single frame, inter prediction may also be used. Inter prediction utilizes the temporal redundancy across multiple frames. As an example, a scene in a movie may include a relatively static background such as a stationary desk. Thus, the desk is shown as substantially the same set of pixels across multiple frames. Inter prediction uses a block in the first frame to predict a block in the second frame, which, in this example, avoids the need to encode the desk in each frame. Inter prediction uses a block matching algorithm to match and compare a block in the current frame with blocks in a reference frame. Then, the motion vector can be encoded to indicate the position of the best-matching block in the reference frame and the position of the same location in the current / target block. Thus, a series of image frames can be represented as a series of blocks, which can then be represented as prediction blocks including prediction modes and / or motion vectors.

[0086] FIG. 5 shows an example of intra prediction 500 in video encoding using an intra prediction mode such as intra prediction mode 400. As shown, the current block 501 can be predicted by samples in the adjacent block 510. The encoder may generally encode the image from top left to bottom right. However, in some cases, the encoder may encode from right to left as described below. It should be noted that when used here, right indicates the right side of the encoded image, left indicates the left side of the encoded image, up indicates the upper side of the encoded image, and down indicates the lower side of the encoded image.

[0087] It should be noted that the current block 501 does not necessarily exactly match the samples from the adjacent block 510. In such cases, the prediction mode is encoded from the adjacent block 510 that most closely matches. To enable the decoder to determine the appropriate value, the difference between the predicted value and the actual value is retained. This is called residual information. Residual information occurs in both intra prediction 500 and inter prediction.

[0088] FIG. 6 is a schematic diagram of an exemplary video encoding mechanism 600 based on intra prediction 500 and / or inter prediction. The image block 601 can be obtained by the encoder from one or more frames. For example, the image may be divided into a plurality of rectangular image regions. Each region of the image corresponds to a coding tree unit (CTU). The CTU is divided into a plurality of blocks such as coding units in HEVC. Then, the block partition information is encoded into the bitstream 611. Therefore, the image block 601 is a partitioned part of the image and includes pixels representing the luma component and / or the chroma component in the corresponding part of the image. During encoding, the image block 601 is encoded as a prediction block 603 that includes prediction information such as a prediction mode (e.g., an intra prediction mode) for intra prediction and / or a motion vector for inter prediction. Then, encoding the image block 601 as the prediction block 603 may leave a residual block 605 that includes residual information indicating the difference between the prediction block 603 and the image block 601.

[0089] It should be noted that the image block 601 may be divided as an encoding unit including one prediction block 603 and one residual block 605. The prediction block 603 may include all the prediction samples of the encoding unit, and the residual block 605 may include all the residual samples of the encoding unit. In such a case, the prediction block 603 is the same size as the residual block 605. In other examples, the image block 601 may be divided as an encoding unit including two prediction blocks 603 and one residual block 605. In such a case, each prediction block 603 includes a part of the prediction samples of the encoding unit, and the residual block 605 includes all of the residual samples of the encoding unit. In still other examples, the image block 601 is divided into an encoding unit including two prediction blocks 603 and four residual blocks 605. The division pattern of the residual blocks 605 within the encoding unit may be signaled in the bitstream 611. Such a position pattern may include a Residual Quad-Tree (RQT) in HEVC. Further, the image block 601 may include only the luma component (e.g., light) shown as the Y component of the image samples (or pixels). In other cases, the image block 601 may include the Y, U, and V components of the image samples, where U and V represent the chrominance components (e.g., color) in the blue luminance and red luminance (UV) color space.

[0090] SVT may be used to further compress the information. Specifically, SVT uses the transform block 607 to further compress the residual block 605. The transform block 607 includes transforms such as inverse DCT and / or inverse DST. The difference between the prediction block 603 and the image block 601 is adapted to the transform by using transform coefficients. By indicating the transform mode (e.g., inverse DCT and / or inverse DST) of the transform block 607 and the corresponding transform coefficients, the decoder can reconstruct the residual block 605. When exact reproduction is not required, the transform coefficients can be further compressed by rounding to specific values to create a better fit by the transform. This process is known as quantization and is performed according to quantization parameters that describe the acceptable quantization. Therefore, the transform mode, transform coefficients, and quantization parameters of the transform block 607 are stored as transform residual information in the transform residual block 609, which may also be simply called the residual block in some cases.

[0091] Next, the prediction information of the prediction block 603 and the transform residual information of the transform residual block 609 can be encoded into the bitstream 611. The bitstream 611 can be stored and / or transmitted to the decoder. Then, the decoder can reverse the process to recover the image block 601. Specifically, the decoder can use the transform residual information to determine the transform block 607. Then, the transform block 607 can be used with the transform residual block 609 to determine the residual block 605. Then, the residual block 605 and the prediction block 603 can be used to reconstruct the image block 601. Then, the image block 601 can be positioned relative to other decoded image blocks 601 to reconstruct the frame and position such a frame to recover the encoded video.

[0092] Regarding SVT, it will be described in more detail here. To perform SVT, the transform block 607 is selected to be smaller than the residual block 605. The transform block 607 is used to transform the corresponding part of the residual block 605 and leave the rest of the residual block 605 without further encoding / compression. This is because the residual information is generally not evenly distributed across the residual block 605. SVT uses smaller transform blocks 607 with adaptive positions to capture most of the residual information within the residual block 605 without requiring the entire residual block 605 to be transformed. This approach can achieve better encoding efficiency than transforming all the residual information within the residual block 605. Since the transform block 607 is smaller than the residual block 605, SVT uses a mechanism to signal the position of the transform with respect to the residual block 605. For example, when SVT is applied to a residual block 605 of size w×h (e.g., width×height), the size and position information of the transform block 607 may be encoded in the bitstream 611. This enables the decoder to reconstruct the transform block 607 and configure the transform block 607 at the correct position with respect to the transform residual block 609 for the reconstruction of the residual block 605.

[0093] It should be noted that some prediction blocks 603 can be encoded without generating a residual block 605. However, in such cases, the use of SVT does not occur, so it will not be described further. As described above, SVT may be used for inter prediction blocks or intra prediction blocks. Further, SVT may be used for the residual block 605 generated by a specified inter prediction mechanism (e.g., transform model-based motion compensation), but may not be used for the residual block 605 generated by other specified inter prediction mechanisms (e.g., affine model-based motion compensation).

[0094] FIG. 7 shows an example 700 of SVT including a conversion block 707 and a residual block 705. The conversion block 707 and the residual block 705 in FIG. 7 are similar to the conversion block 607 and the residual block 605 in FIG. 6, respectively. For ease of reference, the example 700 of SVT is referred to as SVT-I, SVT-II, and SVT-III.

[0095] SVT-I is described as w_t = w / 2, h_t = h / 2, where w_t and h_t represent the width and height of the conversion block 707, respectively, and w and h represent the width and height of the residual block 705, respectively. For example, both the width and height of the conversion block 707 are half of the width and height of the residual block 705. SVT-II is described as w_t = w / 4, h_t = h, and the variables are as described above. For example, the width of the conversion block 707 is one-fourth of the width of the residual block 705, and the height of the conversion block 707 is equal to the height of the residual block 705. SVT-III is described as w_t = w, h_t = h / 4, and the variables are as described above. For example, the width of the conversion block 707 is equal to the width of the residual block 705, and the height of the conversion block 707 is one-fourth of the height of the residual block 705. The type information indicating the type of SVT (e.g., SVT-I, SVT-II, or SVT-III) is encoded in the bitstream to support reconstruction by the decoder.

[0096] As can be seen from FIG. 7, each conversion block 707 can be located at various positions with respect to the residual block 705. The position of the conversion block 707 is represented by a position offset (x, y) with respect to the upper left corner of the residual block 705. x indicates the horizontal distance between the upper left corner of the conversion block 707 and the upper left corner of the residual block 705 in units of pixels, and y indicates the vertical distance between the upper left corner of the conversion block 707 and the upper left corner of the residual block 705 in units of pixels. Each potential position of the conversion block 707 within the residual block 705 is called a candidate position. For the residual block 705, the number of candidate positions is (w - w_t + 1) × (h - h_t + 1) for the type of SVT. More specifically, for a 16×16 residual block 705, when SVT-I is used, there are 81 candidate positions. When SVT-II or SVT-III is used, there are 13 candidate positions. Once determined, the x and y values of the position offset are encoded into the bitstream along with the type of SVT block used. To reduce the complexity for SVT-I, a subset of 32 positions can be selected from the 81 possible candidate positions. This subset then acts as the allowed candidate positions for SVT-I.

[0097] One drawback of one SVT method using one of the examples 700 of SVT is that encoding the SVT position information as residual information results in a significant signaling overhead. Further, the complexity of the encoder can increase significantly as the number of positions to be tested by a compression quality process such as Rate-Distortion Optimization (RDO) increases. Since the number of candidate positions increases with the size of the residual block 705, for larger residual blocks 705 such as 32×32 or 64×128, the signaling overhead can become even larger. Another drawback of using one of the examples 700 of SVT is that the size of the conversion block 707 is one-fourth the size of the residual block 705. Such a conversion block 707 of this size may often not be large enough to cover the main residual information within the residual block 705.

[0098] FIG. 8 shows an example 800 of a further SVT including a conversion block 807 and a residual block 805. The conversion block 807 and the residual block 805 in FIG. 8 are similar to the conversion blocks 607, 707 and the residual blocks 605, 705 in FIGS. 6-7, respectively. For ease of reference, the example 800 of the SVT is referred to as SVT vertical (SVT-V) and SVT horizontal (SVT-H). The example 800 of the SVT is similar to the example 700 of the SVT, but is designed to support reduced signaling overhead and less complex processing requirements for the encoder.

[0099] SVT-V is described as w_t = w / 2 and h_t = h, where the variables are as described above. The width of the conversion block 807 is half the width of the residual block 805, and the height of the conversion block 807 is equal to the height of the residual block 805. SVT-H is described as w_t = w and h_t = h / 2, where the variables are as described above. For example, the width of the conversion block 807 is equal to the width of the residual block 805, and the height of the conversion block 807 is half the height of the residual block 805. SVT-V is similar to SVT-II, and SVT-H is similar to SVT-III. Compared with SVT-II and SVT-III, the conversion block 807 in SVT-V and SVT-H is enlarged to half of the residual block 805 so that the conversion block 807 covers more residual information within the residual block 805.

[0100] Similar to Example 700 of SVT, Example 800 of SVT can include several candidate positions, which are possible acceptable positions of the conversion block (e.g., conversion block 807) with respect to the residual block (e.g., residual block 805). The candidate positions are determined according to the candidate position step size (CPSS). The candidate positions may be separated at equal intervals specified by the CPSS. In such a case, the number of candidate positions is reduced to 5 or less. The reduced number of candidate positions reduces the signaling overhead related to the position information because the selected positions for conversion can be signaled with fewer bits. Furthermore, reducing the number of candidate positions makes the selection of the conversion position algorithmically easier, which allows for a reduction in the complexity of the encoder (e.g., results in fewer computational resources used for encoding).

[0101] FIG. 9 shows an example 900 of SVT including a conversion block 907 and a residual block 905. The conversion block 907 and the residual block 905 in FIG. 9 are similar to the conversion blocks 607, 707, 807 and the residual blocks 605, 705, 805 in FIGS. 6-8, respectively. FIG. 9 shows various candidate positions, which are possible acceptable positions of the conversion block (e.g., conversion block 907) with respect to the residual block (e.g., residual block 905). Specifically, the examples of SVT in FIGS. 9A-9E use SVT-V, and the examples of SVT in FIGS. 9F-9J use SVT-H. The acceptable candidate positions for conversion depend on the CPSS, which further depends on the portion of the residual block 905 that the conversion block 907 should cover and / or the step size between the candidate positions. For example, the CPSS may be calculated as s = w / M1 for SVT-V or s = h / M2 for SVT-H, where w and h are the width and height of the residual block, respectively, and M1 and M2 are predetermined integers in the range of 2 to 8. More candidate positions are allowed by larger values of M1 or M2. For example, both M1 and M2 may be set to 8. In this case, the value of the position index (P) that describes the position of the conversion block 907 with respect to the residual block 905 is between 0 and 4.

[0102] In another example, the CPSS is calculated as s = max(w / M1, Th1) for SVT-V, or s = max(h / M2, Th2) for SVT-H, where Th1 and Th2 are predefined integers that specify the minimum step size. Th1 and Th2 may be integers greater than or equal to 2. In this example, Th1 and Th2 are set to 4, M1 and M2 are set to 8, and different block sizes may have different numbers of candidate positions. For example, when the width of the residual block 905 is 8, two candidate positions, particularly the candidate positions in FIGS. 9A and 9E, are available for SVT-V. For example, when the step size indicated by Th1 is large and the portion of the residual block 905 covered by the conversion block 907 indicated by w / M1 is also large, only two candidate positions satisfy the CPSS. However, when w is set to 16, the portion of the residual block 905 covered by the conversion block 907 decreases due to the change in w / M1. This results in more candidate positions, in this case, three candidate positions as shown in FIGS. 9A, 9C, and 9E. All five candidate positions shown in FIGS. 9A-9E are available when the width of the residual block 905 is greater than 16 and the values of Th1 and M1 are as described above.

[0103] Other examples may also be found when the CPSS is calculated according to other mechanisms. Specifically, the CPSS may be calculated as s = w / M1 for SVT-V, or s = h / M2 for SVT-H. In this case, when M1 and M2 are set to 4, three candidate positions are allowed for SVT-V (e.g., the candidate positions in FIGS. 9A, 9C, and 9E), and three candidate positions are allowed for SVT-H (e.g., the candidate positions in FIGS. 9F, 9H, and 9J). Further, when M1 and M2 are set to 4, the portion of the residual block 905 covered by the conversion block 907 increases, resulting in two acceptable candidate positions for SVT-V (e.g., the candidate positions in FIGS. 9A and 9E) and two acceptable candidate positions for SVT-H (e.g., the candidate positions in FIGS. 9F and 9J).

[0104] In another example, the CPSS is calculated as s = max(w / M1, Th1) for SVT-V or s = max(h / M2, Th2) for SVT-H as described above. In this case, T1 and T2 are set as predefined integers, for example, 2, M1 is set as 8 when w ≥ h, or set as 4 when w < h, M2 is set as 8 when h ≥ w, or set as 4 when h < w. For example, the portion of the residual block 905 covered by the conversion block 907 depends on whether the height of the residual block 905 is greater than the width of the residual block 905 or vice versa. Therefore, the number of candidate positions for SVT-H or SVT-V further depends on the aspect ratio of the residual block 905.

[0105] In another example, the CPSS is calculated as s = max(w / M1, Th1) for SVT-V or s = max(h / M2, Th2) for SVT-H as described above. In this case, the values of M1, M2, Th1, and Th2 are derived from the high-level syntax structure of the bitstream (e.g., the sequence parameter set). For example, the values used to derive the CPSS can be signaled within the bitstream. M1 and M2 may share the same value parsed from the syntax element, and Th1 and Th2 may share the same value parsed from other syntax elements.

[0106] FIG. 10 shows an exemplary SVT position 1000 indicating the position of the transform block 1007 relative to the residual block 1005. Although six different positions (e.g., three vertical positions and three horizontal positions) are shown, it should be recognized that a different number of positions may be used in actual applications. The SVT transform position 1000 is selected from the candidate positions in the example 900 of the SVT in FIG. 9. Specifically, the selected SVT transform position 1000 may be encoded with a position index (P). The position index P can be used to determine the position offset (Z) of the upper left corner of the transform block relative to the upper left corner of the residual block. For example, this position correlation can be determined according to Z = s×P, where s is the CPSS for the transform block based on the SVT type and is calculated as described with respect to FIG. 9. When the transform block is of the SVT-V type, the value of P may be

Number

Number

[0107] As will be described in more detail below, the encoder may encode the SVT transform type (e.g., SVT-H or SVT-T) and the residual block size within the bitstream by using flags. The decoder may then determine the SVT transform size based on the SVT transform size and the residual block size. Once the SVT transform size is determined, the decoder can determine an acceptable candidate position for the SVT transform, such as the candidate positions in example 900 of SVT in FIG. 9, according to the CPSS function. Since the decoder can determine the candidate position of the SVT transform, the encoder may not need to signal the coordinates of the position offset. Instead, a code can be used to indicate which candidate position is used for the corresponding transform. For example, the position index P may be binarized into one or more bins using a truncated unary code for further compression. As a specific example, when the P value is within the range of 0 to 4, the P values 0, 4, 2, 3, and 1 can be binarized as 0, 01, 001, 0001, and 0000, respectively. This binary code is more compressed than representing the decimal value of the position index. As another example, when the P value is within the range of 0 to 1, the P values 0 and 1 can be binarized as 0 and 1, respectively. Therefore, the position index can be increased or decreased to a desired size to signal a specific transform block position considering the possible candidate positions of the transform block.

[0108] The position index P may be binarized into one or more bins by using the most likely position and the remaining less likely positions. For example, when the left and upper adjacent blocks have already been decoded in the decoder and are thus available for prediction, the most likely position may be set as the position covering the lower right corner of the residual block. In one example, when the P value is in the range from 0 to 4 and position 4 is set as the most likely position, the P values 4, 0, 1, 2, and 3 are binarized as 1, 000, 001, 010, and 011, respectively. Further, when the P value is in the range from 0 to 2 and position 2 is set as the most likely position, the P values 2, 0, and 1 are binarized as 1, 01, and 00, respectively. Thus, the most likely position index of the candidate positions is indicated by the minimum number of bits in order to reduce the most common signaling overhead. The probability can be determined based on the coding order of the adjacent reconstructed blocks. Thus, the decoder can infer the codeword method to be used for the corresponding block based on the decoding method used.

[0109] For example, in HEVC, the coding order of coding units is generally top - to - bottom and left - to - right. In such a case, the right side of the current coding / decoding coding unit is not available, and the upper right corner is set as the more likely transformation position. However, the motion vector predictor is derived from the left and upper spatial adjacencies. In such a case, the residual information becomes statistically stronger towards the lower right corner. In this case, the candidate position covering the lower right portion is the most likely position. Further, when an adaptive coding order of coding units is used, one node may be vertically split into two child nodes, and the right child node may be coded before the left child node. In this case, the right adjacent of the left child node is reconstructed before the decoding / coding of the left child node. Further, in such a case, the left adjacent pixels cannot be used. When the right adjacent is available and the left adjacent is not available, the lower left portion of the residual block is likely to contain a large amount of residual information, and thus the candidate position covering the lower left portion of the residual block becomes the most likely position.

[0110] Therefore, the position index P may be binarized into one or more bins according to whether the right side adjacent to the residual block is reconstructed. In one example, the P value is in the range of 0 to 2, as indicated by the SVT transform position 1000. When the right side adjacent to the residual block is reconstructed, the P values 0, 2, and 1 are binarized as 0, 01, and 00. Otherwise, the P values 2, 0, and 1 are binarized as 0, 01, and 00. In other examples, when the right side adjacent to the residual block is reconstructed but the left side adjacent to the residual block is not reconstructed, the P values 0, 2, and 1 are binarized as 0, 00, and 01. Otherwise, the P values 2, 0, and 1 are binarized as 0, 00, and 01. In these examples, the position corresponding to a single bin is the most likely position, and the other two positions are the remaining positions. For example, the most likely position depends on the availability of the right adjacent.

[0111] The probability distribution of the best position in terms of rate distortion performance may be quite different during the inter prediction mode. For example, when the residual block corresponds to a predicted block generated by template matching with spatially adjacent reconstructed pixels as a template, the best position is the most probable position 2. In other inter prediction modes, the probability that position 2 (or position 0 when the right adjacent is available and the left adjacent is not available) is the best position is lower than the probability of the template matching mode. In view of this, the context model for the first bin of the position index P may be determined according to the inter prediction mode related to the residual block. More specifically, when the residual block is related to the template matching based inter prediction, the first bin of the position index P uses the first context model. Otherwise, the second context model is used to encode / decode this bin.

[0112] In other examples, when the residual block is related to template matching-based inter prediction, the most likely position (e.g., position 2, or position 0 when the right neighbor is available but the left neighbor is not) is directly set as the transform block position, and the position information is not signaled within the bitstream. Otherwise, the position index is explicitly signaled within the bitstream.

[0113] It should also be noted that different transforms can be used depending on the position of the transform block relative to the residual block. For example, the left side of the residual block is reconstructed and the right side of the residual block is not reconstructed, which occurs for video encoding having a fixed encoding order of encoding units from left to right and top to bottom (e.g., the encoding order in HEVC). In this case, for the candidate position covering the bottom right corner of the residual block, DST (e.g., DST version 7 (DST-7) or DST version 1 (DST-1)) may be used for the transform in the transform block during encoding. Thus, the inverse DST transform is used at the decoder for the corresponding candidate position. Further, for the candidate position covering the top left corner of the residual block, DCT (e.g., DCT version 8 (DCT-8) or DCT version 2 (DCT-2)) may be used for the transform in the transform block during encoding. Thus, the inverse DCT transform is used at the decoder for the corresponding candidate position. This is because in this case, the bottom right corner is the farthest from the spatially reconstructed region among the four corners. Further, when the transform block covers the bottom right corner of the residual block, DST is more effective than DCT for transforming the residual information distribution. However, when the transform block covers the top left corner of the residual block, DCT is more effective than DST for transforming the residual information distribution. For the rest of the candidate positions, the transform type can be either inverse DST or DCT. For example, when the candidate position is closer to the bottom right corner than the top left corner, inverse DST is used as the transform type. Otherwise, inverse DCT is used as the transform type.

[0114] As a specific example, as shown in FIG. 10, three candidate positions for the conversion block 1007 may be allowed. In this case, position 0 covers the upper left corner, and position 2 covers the lower right corner. Position 1 is at the center of the residual block 1005 and is equidistant from both the left and right corners. The conversion types can be selected as DCT-8, DST-7, and DST-7 for positions 0, 1, and 2, respectively, in the encoder. Then, the inverse conversions DCT-8, DST-7, and DST-7 can be used for positions 0, 1, and 2, respectively, in the decoder. In other examples, the conversion types for positions 0, 1, and 2 are DCT-2, DCT-2, and DST-7, respectively, in the encoder. Then, the inverse conversions DCT-2, DCT-2, and DST-7 can be used for positions 0, 1, and 2, respectively, in the decoder. Therefore, the conversion types for the corresponding candidate positions can be determined in advance.

[0115] In some cases, the above-mentioned multiple position-dependent conversions may be applied only to the luma conversion block. The corresponding chroma conversion block may always use inverse DCT-2 in the conversion / inverse conversion process.

[0116] FIG. 11 shows an example 1100 of horizontal flipping of residual samples. In some cases, advantageous residual compression can be achieved by horizontally flipping the residual information in the residual block (e.g., residual block 605) before applying the conversion block (e.g., conversion block 607) in the encoder. Example 1100 shows such a horizontal flip. In this regard, horizontal flipping indicates rotating the residual samples in the residual block around an axis near the middle between the left side and the right side of the residual block. Such horizontal flipping occurs before applying the conversion (e.g., conversion block) in the encoder and after applying the inverse conversion (e.g., conversion block) in the decoder. Such flipping may be used when specified predefined conditions occur.

[0117] In one example, horizontal flipping occurs when the transform block uses DST / inverse DST in the transform process. In this case, the right neighbor of the residual block is encoded / reconstructed before the current block, and the left neighbor is not encoded / reconstructed before the current block. The horizontal flipping process exchanges the residual samples in column i of the residual block with the residual samples in column w-1-i of the residual block. In this context, w is the width of the transform block, and i = 0, 1, ..., (w / 2)-1. Horizontal flipping of the residual samples can increase the encoding efficiency by better adapting the residual distribution to the DST transform.

[0118] FIG. 12 is a flowchart of an exemplary method 1200 for video decoding by position-dependent SVT using the above mechanism. Method 1200 may be initiated in a decoder when receiving a bitstream such as bitstream 611. Method 1200 uses the bitstream to determine prediction blocks and transform residual blocks such as prediction block 603 and transform residual block 609. Method 1200 also determines a transform block such as transform block 607, which is used to determine a residual block such as residual block 605. Next, residual block 605 and prediction block 603 are used to reconstruct an image block such as image block 601. It should be noted that method 1200 is described from the perspective of the decoder, but a similar method may be used (e.g., conversely) to encode video by using SVT.

[0119] In block 1201, a bitstream is obtained in the decoder. The bitstream may be received from a memory or a streaming source. The bitstream includes data that can be decoded into at least one image corresponding to the video data from the encoder. Specifically, the bitstream includes block partitioning information that can be used to determine an encoded unit including a prediction block and a residual block from the bitstream, as described in mechanism 600. Therefore, the encoding information related to the encoded unit can be analyzed from the bitstream, and the pixels of the encoded unit can be reconstructed based on the encoding information as described below.

[0120] In block 1203, the prediction block and the corresponding transform residual block are obtained from the bitstream based on the block partitioning information. In this example, the transform residual block is encoded according to SVT as described with respect to mechanism 600 above. Then, method 1200 reconstructs a residual block of size w×h from the transform residual block, as described below.

[0121] In block 1205, the use of SVT, the type of SVT, and the transform block size are determined. For example, the decoder first determines whether SVT is used in encoding. This is because some encodings use a transform whose size is that of the residual block. The use of SVT can be signaled by a syntax element in the bitstream. Specifically, when a residual block is allowed to use SVT, a flag such as svt_flag is parsed from the bitstream. When the transform residual block has non-zero transform coefficients (e.g., corresponding to any luma or chroma component), the residual block is allowed to use SVT. For example, when the residual block contains any residual data, the residual block may use SVT. The SVT flag indicates whether the residual block is encoded using a transform block of the same size as the residual block (e.g., the svt_flag is set to 0), or whether the residual block is encoded using a transform block of a size smaller than the residual block (e.g., the svt_flag is set to 1). The coded block flag (cbf) can be used to indicate whether the residual block contains non-zero transform coefficients for the color components, as used in HEVC. Also, the root coded block (root cbf) flag can indicate whether the residual block contains non-zero transform coefficients for any of the color components, as used in HEVC. As a specific example, when an image block is predicted using inter prediction and either the residual block width or the residual block height falls within a predetermined range of [a1, a2], the residual block is allowed to use SVT, where a1 = 16 and a2 = 64, a1 = 8 and a2 = 64, or a1 = 16 and a2 = 128. The values of a1 and a2 can be predetermined fixed values. The values may also be derived from the sequence parameter set (SPS) or the slice header in the bitstream. When the residual block does not use SVT, the transform block size is set as the width and height of the residual block size.Otherwise, the conversion size is determined based on the SVT conversion type.

[0122] Once the decoder determines that SVT is being used for the residual block, the decoder determines the type of the SVT conversion block used and derives the conversion block size according to the SVT type. The allowable SVT types for the residual block are determined based on the width and height of the residual block. The SVT-V conversion as shown in FIG. 8 is allowed when the width of the residual block is within the range [a1, a2], such values being defined above. The SVT-H conversion as shown in FIG. 8 is allowed when the height of the residual block is within the range [a1, a2], such values being defined above. SVT may be used only for the luma component within the residual block, or SVT may be used for both the luma and chroma components within the residual block. When SVT is used only for the luma component, the residual information of the luma component is converted by SVT, and the chroma component is converted by converting the size of the residual block. When both SVT-V and SVT-H are allowed, a flag such as svt_type_flag may be encoded in the bitstream. The svt_type_flag indicates whether SVT-V is being used for the residual block (e.g., svt_type_flag is set to 0) or whether SVT-H is being used for the residual block (e.g., svt_type_flag is set to 1). Once the type of SVT conversion is determined, the conversion block size is set according to the signaled SVT type (e.g., for SVT-V, w_t = w / 2 and h_t = h; for SVT-H, w_t = w and h_t = h / 2). When only SVT-V is allowed or only SVT-H is allowed, the svt_type_flag may not be encoded in the bitstream. In such a case, the decoder can infer the conversion block size based on the allowed SVT type.

[0123] Once the SVT type and size are determined, the decoder proceeds to block 1207. In block 1207, the decoder determines the position of the transform for the residual block and the type of transform (e.g., either DST or DCT). The position of the transform block can be determined according to the syntax elements in the bitstream. For example, the position index is, in some examples, directly signaled and thus can be parsed from the bitstream. In other examples, the position can be inferred as described with respect to FIGS. 8 - 10. Specifically, the candidate positions for the transform can be determined according to the CPSS function. The CPSS function can determine the candidate positions by considering the width of the residual block, the height of the residual block, the SVT type determined by block 1205, the step size of the transform, and / or the portion of the residual block covered by the transform. Then, the decoder can determine the transform block position from the candidate positions by obtaining a p - index that includes code signaling the correct candidate position according to the candidate position selection probability as described with respect to FIG. 10 above. Once the transform block position is known, the decoder can infer the type of transform used by the transform block as described with respect to FIG. 10 above. Thus, the encoder can select the corresponding inverse transform.

[0124] In block 1209, the decoder analyzes the transform coefficients of the transform block based on the transform block size determined in block 1205. This process may be accomplished according to the transform coefficient analysis mechanism used in HEVC, H.264, and / or AVC. The transform coefficients may be encoded using run - length encoding and / or as a set of coefficient groups (CG). It should be noted that in some examples, block 1209 may be executed before block 1207.

[0125] In block 1211, the residual block is reconstructed based on the conversion position, conversion coefficient, and conversion type as determined above. Specifically, inverse quantization and inverse transformation of size w_t×h_t are applied to the conversion coefficients to recover the residual samples of the residual block. The size of the residual block with the residual samples is w_t×h_t. The inverse transformation may be an inverse DCT or an inverse DST according to the position-dependent conversion type determined in block 1207. The residual samples are assigned to the corresponding regions within the residual block according to the conversion block position. The residual samples either inside the residual block or outside the conversion block may be set to zero. For example, when SVT-V is used, the number of candidate positions is 5, and the position index indicates the fifth conversion block position. Therefore, the reconstructed residual samples are assigned to the region within the conversion candidate positions of example 900 of SVT in FIG. 9 (e.g., the shaded region in FIG. 9) and the region of size (w / 2)×h up to the region with zero residual samples (e.g., the unshaded region in FIG. 9).

[0126] In optional block 1213, the residual block information of the reconstructed block may be horizontally flipped as described with respect to FIG. 11. As described above, this may occur when the conversion block in the decoder uses an inverse DST, the right adjacent block has already been reconstructed, and the left adjacent block has not yet been reconstructed. Specifically, the encoder may horizontally flip the residual block before applying the DST conversion in the above case to increase the coding efficiency. Therefore, optional block 1213 may be used to correct such a horizontal flip in the encoder to generate an accurately reconstructed block.

[0127] In block 1215, the reconstructed residual block may be configured with a prediction block to generate a reconstructed image block that includes samples as part of an encoding unit. The filtering process may also be applied to the reconstructed samples, such as the deblocking filter and sample adaptive offset (SAO) processing in HEVC. The reconstructed image block may then be combined with other decoded image blocks in a similar manner to generate a frame of a media / video file. The reconstructed media file may then be displayed to a user on a monitor or other display device.

[0128] It should be noted that an equivalent implementation of method 1200 can be used to generate the samples reconstructed within the residual block. Specifically, the residual samples of the transform block can be directly configured with the prediction block at the position indicated by the transform block position information without first recovering the residual block.

[0129] FIG. 13 shows a method 1300 of video encoding. Method 1300 may be implemented in a decoder (e.g., video decoder 30). In particular, method 1300 may be implemented by a processor of the decoder. Method 1300 may be implemented when a bitstream is received directly or indirectly from an encoder (e.g., video encoder 20), or when obtained from memory. In block 1301, the bitstream is parsed to obtain a prediction block (e.g., prediction block 603) and a transform residual block corresponding to the prediction block (e.g., transform residual block 609). In block 1303, the type of SVT used to generate the transform residual block is determined. As described above, the type of SVT may be SVT-V or SVT-H. In an embodiment, the SVT-V type includes a height equal to the height of the transform residual block and a width equal to half the width of the transform residual block.

[0130] In an embodiment, the SVT-H type includes a height that is half the height of the transform residual block and a width equal to the width of the transform residual block. In an embodiment, the svt_type_flag is parsed from the bitstream to determine the type of SVT. In an embodiment, when only one type of SVT is allowed for the residual block, the type of SVT is determined by inference.

[0131] In block 1305, the position of the SVT for the transform residual block is determined. In an embodiment, the position index is parsed from the bitstream to determine the position of the SVT. In an embodiment, the position index includes a binary code indicating a position from a set of candidate positions determined according to CPSS. In an embodiment, the most likely position of the SVT is assigned the fewest number of bits in the binary code indicating the position index. In an embodiment, when a single candidate position is available for SVT transform, the position of the SVT is inferred by the processor. In an embodiment, when the residual block is generated by template matching in the inter prediction mode, the position of the SVT is inferred by the processor.

[0132] In block 1307, the inverse of the SVT is determined based on the position of the SVT. In block 1309, the inverse of the SVT is applied to the transform residual block to generate a reconstructed residual block (e.g., residual block 605). In an embodiment, the inverse DST is used for SVT-V type transform located at the left boundary of the residual block. In an embodiment, the inverse DST is used for SVT-H type transform located at the upper boundary of the residual block. In an embodiment, the inverse DCT is used for SVT-V type transform located at the right boundary of the residual block. In an embodiment, the inverse DCT is used for SVT-H type transform located at the lower boundary of the residual block.

[0133] In block 1311, a reconstructed residual block is combined with a prediction block to reconstruct an image block. In an embodiment, the image block is displayed on a display or monitor of an electronic device (e.g., a smartphone, a tablet, a laptop computer, a personal computer, etc.).

[0134] Optionally, method 1300 may also include horizontally flipping samples in the reconstructed residual block before combining the reconstructed residual block with the prediction block when the right adjacent of the encoding unit associated with the reconstructed residual block is reconstructed and the left adjacent of the encoding unit is not reconstructed.

[0135] FIG. 14 is a method 1400 for video encoding. Method 1400 may be implemented in an encoder (e.g., video encoder 20). In particular, method 1400 may be implemented by a processor of the encoder. Method 1400 may be implemented to encode a video signal. In block 1401, a video signal is received from a video capture device (e.g., a camera). In an embodiment, the video signal includes an image block (e.g., image block 601).

[0136] In block 1403, a prediction block (e.g., prediction block 603) and a residual block (e.g., residual block 605) are generated to represent the image block. In block 1405, a conversion algorithm is selected for the SVT based on the position of the SVT for the residual block. In block 1407, the residual block is converted to a transformed residual block using the selected SVT.

[0137] In block 1409, the type of the SVT is encoded in the bitstream. In an embodiment, the type of the SVT is an SVT-V type or an SVT-H type. In an embodiment, the SVT-V type includes a height equal to the height of the residual block and a width equal to half the width of the residual block. In an embodiment, the SVT-H type includes a height equal to half the height of the residual block and a width equal to the width of the residual block.

[0138] In block 1411, the position of the SVT is encoded in the bitstream. In an embodiment, the position of the SVT is encoded in a position index. In an embodiment, the position index includes a binary code indicating a position from a set of candidate positions determined according to CPSS. In an embodiment, the most likely position of the SVT is assigned as the least number of bits in the binary code indicating the position index.

[0139] In an embodiment, the DST algorithm is used by the processor for the SVT-V type conversion located at the left boundary of the residual block. In an embodiment, the DST algorithm is selected by the processor for the SVT-H type conversion located at the upper boundary of the residual block. In an embodiment, the DCT algorithm is selected by the processor for the SVT-V type conversion located at the right boundary of the residual block. In an embodiment, the DCT algorithm is selected by the processor for the SVT-H type conversion located at the lower boundary of the residual block.

[0140] Optionally, when the right adjacent of the coding unit related to the residual block is encoded and the left adjacent of the coding unit is not encoded, the processor may horizontally invert the samples in the residual block before converting the residual block into a transformed residual block.

[0141] In block 1413, the prediction block and the transformed residual block are encoded in the bitstream. In an embodiment, the bitstream is configured to be transmitted to the decoder and / or is transmitted to the decoder.

[0142] FIG. 15 is a schematic diagram of an exemplary computing device 1500 for video encoding according to an embodiment of the disclosure. The computing device 1500 is suitable for implementing the embodiments of the disclosure described herein. The computing device 1500 includes an inlet port 1520 and a receiver unit (Rx) 1510 for receiving data, a processor, logic unit or central processing unit (CPU) 1530 for processing data, a transmitter unit (Tx) 1540 and an outlet port 1550 for transmitting data, and a memory 1560 for storing data. The computing device 1500 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the inlet port 1520, the receiver unit 1510, the transmitter unit 1540, and the outlet port 1550 for the exit or entry of optical or electrical signals. The computing device 1500 may also include, in some examples, a wireless transmitter and / or receiver.

[0143] Processor 1530 is implemented by hardware and software. Processor 1530 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), a field-programmable gate array (FPGA), an application specific integrated circuit (ASIC), and a digital signal processor (DSP). Processor 1530 communicates with an ingress port 1520, a receiver unit 1510, a transmitter unit 1540, an egress port 1550, and a memory 1560. Processor 1530 includes an encoding / decoding module 1514. Encoding / decoding module 1514 implements the embodiments of the above disclosure such as method 1300 and method 1400, other mechanisms for encoding / reconstructing residual blocks based on the transform block positions when using SVT, and any of the other mechanisms above. For example, encoding / decoding module 1514 realizes, processes, prepares, or provides various encoding operations such as encoding of video data and / or decoding of video data as described above. Thus, the inclusion of encoding / decoding module 1514 provides a substantial improvement to the functionality of computing device 1500 and results in a transformation of computing device 1500 to different states. Alternatively, encoding / decoding module 1514 is realized as instructions stored in memory 1560 and executed by processor 1530 (e.g., as a computer program product stored on a non-transitory medium).

[0144] Memory 1560 includes one or more disks, tape drives, and solid state drives, and stores such programs when selected for execution, and may be used as an overflow data storage device to store instructions and data read during program execution. Memory 1560 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM). Computing device 1500 may also include input / output (IO) devices for interacting with an end user. For example, computing device 1500 may include a display such as a monitor for visual output, a speaker for audio output, and a keyboard / mouse / trackball, etc. for user input.

[0145] In summary, the above disclosure includes a mechanism for adaptively using multiple conversion types for conversion blocks at different positions. Further, the disclosure enables inverting residual samples within a residual block horizontally to support coding efficiency. This occurs when the conversion blocks use DST and inverse DST in the encoder and decoder, respectively, and when the right adjacent block is available and the left adjacent is not. Further, the disclosure includes a mechanism for supporting coded position information within a bitstream based on an inter prediction mode associated with a residual block.

[0146] FIG. 16 is a schematic diagram of an embodiment of an encoding means 1600. In the embodiment, the encoding means 1600 is implemented in a video encoding device 1602 (e.g., video encoder 20 or video decoder 30). The video encoding device 1602 includes a receiving means 1601. The receiving means 1601 is configured to receive an image to be encoded or to receive a bitstream to be decoded. The video encoding device 1602 includes a transmitting means 1607 coupled to the receiving means 1601. The transmitting means 1607 is configured to transmit the bitstream to a decoder or to transmit the decoded image to a display means (e.g., one of the I / O devices within the computing device 1500).

[0147] The video encoding device 1602 includes a storage means 1603. The storage means 1603 is coupled to at least one of the receiving means 1601 or the transmitting means 1607. The storage means 1603 is configured to store instructions. The video encoding device 1602 also includes a processing means 1605. The processing means 1605 is coupled to the storage means 1603. The processing means 1605 is configured to execute the instructions stored in the storage means 1603 and to execute the method disclosed herein.

[0148] A first component is directly coupled to a second component when there is no component intervening between the first component and the second component other than a line, trace, or other medium. A first component is indirectly coupled to a second component when there is a component intervening between the first component and the second component other than a line, trace, or other medium. The term "coupled" and its variations include both being directly coupled and being indirectly coupled. The use of the term "about" means a range including ±10% of the subsequent number unless otherwise specified.

[0149] Although several embodiments are provided in the present disclosure, it can be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. This example is to be considered illustrative and not restrictive, and its intention is not limited to the details given herein. For example, various elements or components may be combined with or integrated into other systems, or certain features may be omitted or not implemented.

[0150] Furthermore, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined with or integrated into other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other items shown or described as combined may be directly connected, or may be indirectly coupled or communicate through some interfaces, devices, or intermediate components, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and alternatives are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method implemented in a decoding device, comprising: analyzing a bitstream to obtain a prediction block and a transform residue block corresponding to the prediction block, wherein the transform residue block contains transform residue information; determining a type of spatial varying transform (SVT) used to generate the transform residue block, wherein the type of the SVT is an SVT vertical (SVT-V) type or an SVT horizontal (SVT-H) type, and when the type of the SVT is the SVT-V type, the SVT-V type includes a height equal to the height of the transform residue block and a width equal to half of the width of the transform residue block, or when the type of the SVT is the SVT-H type, the SVT-H type includes a height equal to half of the height of the transform residue block and a width equal to the width of the transform residue block; determining a position of the SVT with respect to the transform residue block; determining an inverse of the SVT based on the position of the SVT; applying the inverse of the SVT to the transform residue block to generate a reconstructed residue block based on the transform residue information; combining the reconstructed residue block with the prediction block to reconstruct an image block. A method comprising the above steps.

2. The method according to claim 1, further comprising analyzing an svt_type_flag from the bitstream to determine the type of the SVT.

3. The method according to claim 1 or 2, further comprising determining the type of the SVT by inference when only one type of SVT is allowed for the residue block.

4. The method according to claim 1 or 2, further comprising analyzing a position index from the bitstream to determine the position of the SVT.

5. A method implemented in an encoding device, comprising: generating a prediction block and a residue block to represent an image block; selecting a transform algorithm for the SVT based on a position of a spatial varying transform (SVT) for the residual block; transforming the residual block into a transform residual block using the selected SVT to obtain transform residual information; encoding a type of the SVT into a bitstream, the type of the SVT being an SVT vertical (SVT-V) type or an SVT horizontal (SVT-H) type, wherein when the type of the SVT is the SVT-V type, the SVT-V type includes a height equal to a height of the residual block and a width equal to half of a width of the residual block, or when the type of the SVT is the SVT-H type, the SVT-H type includes a height equal to half of the height of the residual block and a width equal to the width of the residual block; encoding a position of the SVT into the bitstream; encoding the prediction block and the transform residual block into the bitstream for transmission to a decoder, the transform residual block including the transform residual information; A method comprising.

6. The method according to claim 5, wherein the position of the SVT is encoded into a position index.

7. A decoding apparatus, comprising: a receiver configured to receive a bitstream to be decoded; a transmitter coupled to the receiver and configured to transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter and configured to store instructions; a processor coupled to the memory and configured to execute the instructions stored in the memory and execute the method according to claim 1; A decoding apparatus comprising.

8. The decoding apparatus according to claim 7, further comprising a display configured to display an image.

9. An encoding apparatus, comprising: a receiver configured to receive a picture to be encoded; a transmitter coupled to the receiver and configured to transmit the bitstream to a decoder; A memory coupled to at least one of the receiver or the transmitter, the memory being configured to store instructions; A processor coupled to the memory, the processor being configured to execute the instructions stored in the memory and to execute the method according to claim 5; An encoding device comprising the same. **Claim 10** An encoder; A decoder communicating with the encoder; A system comprising the same, wherein the encoder comprises the encoding device according to claim 9; and the decoder comprises the decoding device according to claim 7. **Claim 11** Means for decoding, comprising: Receiving means configured to receive a bitstream to be decoded; Transmitting means coupled to the receiving means, the transmitting means being configured to transmit the decoded image to a display means; Storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions; Processing means coupled to the storage means, the processing means being configured to execute the instructions stored in the storage means and to execute the method according to claim 1 or 2. Means for decoding comprising the same. **Claim 12** Means for encoding, comprising: Receiving means configured to receive a picture to be encoded; Transmitting means coupled to the receiving means, the transmitting means being configured to transmit the bitstream to a decoder; Storage means coupled to at least one of the receiving means or the transmitting means, the storage means being configured to store instructions; Processing means coupled to the storage means, the processing means being configured to execute the instructions stored in the storage means and to execute the method according to claim 5 or 6. Means for encoding comprising the same. **Claim 13** A method for storing a bitstream, comprising: Receiving the bitstream by at least one processor; Storing the bitstream in at least one memory. The bitstream is generated by the method according to claim 5 or 6. **Claim 14** A device for transmitting a bitstream, comprising: At least one processor configured to obtain the bitstream; At least one transmitter configured to transmit the bitstream. ** ​ The bitstream is a device generated by the method according to claim 5 or 6.

Citation Information

Patent Citations

  • Position-dependent space-varying transformations for video coding.

    JP7651669B2

  • Methods and apparatus for spatially varying residue coding

    WO2011005303A1

  • Spatial varying transforms for video coding

    WO2019076290A1