Position-related Spatial Variation Transform for Video Coding and Decoding
By analyzing the SVT type and position information in the bitstream and applying corresponding inverse transformation, the problem of low SVT encoding efficiency in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202410232196.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-23
- Filing Date
- 2019-02-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-02-04
AI Technical Summary
When existing video encoding and decoding technologies deal with spatial change transformation (SVT), it is difficult to effectively locate transformation blocks, resulting in low encoding efficiency.
By analyzing the bitstream, the type and position of the SVT are determined, and the inverse transformation of the SVT is determined based on this information, applied to transform the residual block to generate the reconstruction residual block, and finally combined with the prediction block to reconstruct the image block.
The encoding efficiency of SVT is improved, and the positioning and transformation process of the transformation block is optimized by adopting different transformations at each candidate position.
Smart Images

Figure CN118101966B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 201980015113.X, the original application date is February 4, 2019, and the entire content of the original application is incorporated herein by reference.
[0002] Cross - reference to related applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 634,613, filed on February 23, 2018, entitled "Position Dependent Spatial Varying Transform for Video Coding" by Yin Zhao et al., the teachings and disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0004] This application relates to the field of video coding and decoding, and in particular, to a method and related apparatus for position - dependent spatial - varying transform for video coding and decoding. Background Art
[0005] Video coding and decoding is the process of compressing video images into a smaller format. Video coding and decoding enables the encoded video to occupy less space when stored on a medium. In addition, video coding and decoding supports streaming media. Specifically, content providers desire to provide media to end - users in increasingly high definition. In addition, content providers desire to provide media on - demand without forcing users to wait for long periods of time for such media to be transmitted to end - user devices such as televisions, computers, tablets, phones, etc. Advancements in video coding and decoding compression can reduce the size of video files. Thus, when combined with corresponding content distribution systems, the above two goals are supported. Summary of the Invention
[0006] A first aspect relates to a method implemented in a computing device. The method includes: a processor of the computing device parsing a bitstream to obtain a prediction block and a transform residual block corresponding to the prediction block; the processor determining a type of spatial - varying transform (SVT) for generating the transform residual block; the processor determining a position of the SVT relative to the transform residual block; the processor determining an inverse of the SVT based on the position of the SVT; the processor applying the inverse of the SVT to the transform residual block to generate a reconstructed residual block; and the processor combining the reconstructed residual block with the prediction block to reconstruct an image block.
[0007] The method is beneficial to improving the coding efficiency of the SVT. In this regard, the transform block is located at various candidate positions relative to the corresponding residual block. Thus, the disclosed mechanism employs different transforms for the transform block based on the candidate positions.
[0008] In a first implementation of the method according to the first aspect, the type of the SVT is the SVT vertical (SVT-V) type or the SVT horizontal (SVT-H) type.
[0009] In a second implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, the height included in the SVT-V type is equal to the height of the transform residual block, and the width is half of the width of the transform residual block; the height included in the SVT-H type is half of the height of the transform residual block, and the width is equal to the width of the transform residual block.
[0010] In a third implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, parse the svt_type_flag in the bitstream to determine the type of the SVT.
[0011] In a fourth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, when only one type of SVT is allowed for the residual block, determine the type of the SVT by inference.
[0012] In a fifth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, parse the position index in the bitstream to determine the position of the SVT.
[0013] In a sixth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, the position index includes a binary code, where the binary code indicates a position in a set of candidate positions determined according to a candidate position step size (CPSS).
[0014] In a seventh implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, allocate the least number of bits in the binary code indicating the position index to the most likely position of the SVT.
[0015] In an eighth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, when a single candidate position is available for the SVT transform, the processor infers the position of the SVT.
[0016] In a ninth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, when the residual block is generated by template matching in an inter prediction mode, the processor infers the position of the SVT.
[0017] In a tenth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, an inverse discrete sine transform (DST) is employed for the SVT vertical (SVT-V) type transform located at the left boundary of the residual block.
[0018] In an eleventh implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, an inverse DST is employed for the SVT horizontal (SVT-H) type transform located at the top boundary of the residual block.
[0019] In a twelfth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, an inverse discrete cosine transform (DCT) is employed for the SVT-V type transform located at the right boundary of the residual block.
[0020] In a thirteenth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, an inverse DCT is employed for the SVT-H type transform located at the bottom boundary of the residual block.
[0021] In a fourteenth implementation of the method according to the first aspect or any of the foregoing implementations of the first aspect, when the right adjacent coding unit of the coding unit associated with the reconstructed residual block has been reconstructed and the left adjacent coding unit of the coding unit has not been reconstructed, samples in the reconstructed residual block are horizontally flipped before combining the reconstructed residual block with the prediction block.
[0022] A second aspect relates to a method implemented in a computing device. The method includes: receiving a video signal from a video capture device, where the video signal includes image blocks; a processor of the computing device generating a prediction block and a residual block to represent the image blocks; the processor selecting a transform algorithm for a spatially-variant transform (SVT) based on the position of the SVT relative to the residual block; the processor using the selected SVT to transform the residual block into a transformed residual block; the processor encoding the type of the SVT into a bitstream; the processor encoding the position of the SVT into the bitstream; and the processor encoding the prediction block and the transformed residual block into the bitstream for transmission to a decoder.
[0023] The method is conducive to improving the coding efficiency of the SVT. In this regard, the transform block is positioned at various candidate positions relative to the corresponding residual block. Accordingly, the disclosed mechanism employs different transforms for the transform block based on the candidate positions.
[0024] In a first implementation of the method according to the second aspect, the type of the SVT is the SVT vertical (SVT-V) type or the SVT horizontal (SVT-H) type.
[0025] In a second implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the height included in the SVT-V type is equal to the height of the transform residual block, and the width is half of the width of the transform residual block.
[0026] In a third implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the height included in the SVT-H type is half of the height of the transform residual block, and the width is equal to the width of the transform residual block.
[0027] In a fourth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the position of the SVT is encoded in a position index.
[0028] In a fifth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the position index includes a binary code, where the binary code indicates a position in a set of candidate positions determined according to a candidate position step size (CPSS).
[0029] In a sixth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the binary code indicating the least number of bits of the position index is assigned to the most likely position of the SVT.
[0030] In a seventh implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the processor employs a discrete sine transform (DST) algorithm for the SVT vertical (SVT-V) type transform located at the left boundary of the residual block.
[0031] In an eighth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the processor selects the DST algorithm for the SVT horizontal (SVT-H) type transform located at the top boundary of the residual block.
[0032] In a ninth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the processor selects a discrete cosine transform (DCT) algorithm for the SVT-V type transform located at the right boundary of the residual block.
[0033] In a tenth implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the processor selects the DCT algorithm for the SVT-H type transform located at the bottom boundary of the residual block.
[0034] In an eleventh implementation of the method according to the second aspect or any of the foregoing implementations of the second aspect, the method further includes: when the right adjacent coding unit of the coding unit associated with the residual block has been coded and the left adjacent coding unit of the coding unit has not been coded, before the processor converts the residual block into the transformed residual block, the processor horizontally flips the samples in the residual block.
[0035] A third aspect relates to an encoding / decoding device, including: a receiver for receiving an image for encoding or receiving a bitstream for decoding; a transmitter coupled to the receiver, wherein the transmitter is configured to transmit the bitstream to a decoder or transmit a decoded image to a display; a memory coupled to at least one of the receiver or the transmitter, wherein the memory is configured to store instructions; a processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory to perform the method according to any one of the foregoing aspects or implementations.
[0036] The encoding / decoding device is conducive to improving the encoding efficiency of SVT. In this regard, the transform block is positioned at various candidate positions relative to the corresponding residual block. Therefore, the disclosed mechanism employs different transforms for the transform block based on the candidate positions.
[0037] In a first implementation of the device according to the third aspect, the device further includes a display for displaying an image.
[0038] A fourth aspect relates to a system, wherein the system includes an encoder and a decoder in communication with the encoder. The encoder or the decoder includes the encoding / decoding device according to any one of the foregoing aspects or implementations.
[0039] The system is conducive to improving the encoding efficiency of SVT. In this regard, the transform block is positioned at various candidate positions relative to the corresponding residual block. Therefore, the disclosed mechanism employs different transforms for the transform block based on the candidate positions.
[0040] A fifth aspect relates to a component for encoding / decoding, including: a receiving component for receiving an image for encoding or receiving a bitstream for decoding; a transmitting component coupled to the receiving component, wherein the transmitting component is configured to transmit the bitstream to a decoder or transmit a decoded image to a display component; a storage component coupled to at least one of the receiving component or the transmitting component, wherein the storage component is configured to store instructions; a processing component coupled to the storage component, wherein the processing component is configured to execute the instructions stored in the storage component to perform the method according to any one of the foregoing aspects or implementations.
[0041] The components for encoding and decoding are conducive to improving the encoding efficiency of SVT. In this regard, the transform block is positioned at various candidate positions relative to the corresponding residual block. Therefore, the disclosed mechanism employs different transforms for the transform block based on the candidate positions.
[0042] For clarity, any of the above embodiments can be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present invention.
[0043] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. Brief Description of the Drawings
[0044] To understand the present invention more thoroughly, reference is now made to the following brief description, which is described in conjunction with the drawings and specific embodiments, in which the same reference numerals represent the same parts.
[0045] Figure 1 is a block diagram showing an exemplary encoding system that can utilize spatial-varying transform (SVT).
[0046] Figure 2 is a block diagram showing an exemplary encoding system that can utilize spatial SVT.
[0047] Figure 3 is a block diagram showing an example of a video decoder that can utilize spatial SVT.
[0048] Figure 4 is a schematic diagram of an intra prediction mode employed in video encoding and decoding.
[0049] Figure 5 shows an example of intra prediction in video encoding and decoding.
[0050] Figure 6 is a schematic diagram of an exemplary video encoding mechanism.
[0051] Figure 7 shows an exemplary SVT transform.
[0052] Figure 8 shows an exemplary SVT transform.
[0053] Figure 9 shows exemplary SVT transform candidate positions relative to the residual block.
[0054] Figure 10 shows exemplary SVT transform positions relative to the residual block.
[0055] Figure 11 shows an example of horizontal flipping of residual samples.
[0056] Figure 12 It is a flowchart of an exemplary method for video decoding using position-related SVT.
[0057] Figure 13 It is a flowchart of an exemplary method for video encoding and decoding.
[0058] Figure 14 It is a flowchart of an exemplary method for video encoding and decoding.
[0059] Figure 15 It is a schematic diagram of an exemplary computing device for video encoding and decoding.
[0060] Figure 16 It is a schematic diagram of an embodiment of a component for encoding and decoding. Detailed implementation
[0061] First, it should be understood that although the following provides illustrative implementations of one or more embodiments, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.
[0062] The standard now known as High Efficiency Video Coding (HEVC) is an advanced video coding system developed by the Joint Collaborative Team on Video Coding (JCT-VC), a joint collaboration team of video coding experts from the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) Study Group. Details of the HEVC standard can be found in ITU-T Rec. H.265 and International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) 23008-2 (2013), High Efficiency Video Coding, Final Draft Approval Date: January 2013 (officially released by ITU-T in June 2013 and by ISO / IEC in November 2013), which is incorporated herein by reference. An overview of HEVC can be found in G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEEE Trans. Circuits and Systems for Video Technology, Vol. 22, No. 12, pp. 1649-1668, December 2012, which is incorporated herein by reference.
[0063] Figure 1 is a block diagram showing an exemplary coding system 10, where the exemplary coding system 10 can utilize video coding techniques, such as encoding using the SVT mechanism. As Figure 1As shown, the encoding system 10 includes a source device 12, where the source device 12 provides encoded video data that is subsequently decoded by a destination device 14. In particular, the source device 12 can provide the video data to the destination device 14 via a computer-readable medium 16. The source device 12 and the destination device 14 can include any one of a variety of devices, including desktop computers, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones and "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device 12 and the destination device 14 can be equipped for wireless communication.
[0064] The destination device 14 can receive the encoded video data to be decoded via the computer-readable medium 16. The computer-readable medium 16 can include any type of medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the computer-readable medium 16 can include a communication medium such that the source device 12 can directly and in real time transmit the encoded video data to the destination device 14. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet network (such as a local area network, a wide area network, or a global network such as the Internet). The communication medium can include routers, switches, base stations, or any other device that can be used to facilitate communication between the source device 12 and the destination device 14.
[0065] In some examples, the encoded data can be output from the output interface 22 to a storage device. Similarly, the encoded data can be accessed from the storage device via the input interface. The storage device can include any kind of distributed or locally accessible data storage medium, such as a hard disk drive, a Blu-ray disc, a digital video disk (DVD), a Compact Disc Read-Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video data generated by the source device 12. The destination device 14 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 can access the encoded video data through any standard data connection, including an Internet connection. The standard data connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data in the storage device can be streaming, download transmission, or a combination thereof.
[0066] The techniques of the present invention are not necessarily limited to wireless applications or settings. These techniques can be applied to video coding and decoding to support any kind of multimedia application, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as HTTP dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored in a data storage medium), or other applications. In some examples, the system 10 can be used to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0067] In Figure 1In the example, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. According to the present invention, the video encoder 20 of the source device 12 and / or the video decoder 30 of the destination device 14 can be used to apply video encoding and decoding techniques. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 12 can receive video data from an external video source (such as an external camera). Similarly, the destination device 14 can be connected to an external display device instead of including an integrated display device.
[0068] Figure 1 The encoding system 10 shown is only one example. The video encoding and decoding techniques can be performed by any digital video encoding and / or decoding device. Although the techniques of the present invention are generally performed by video encoding and decoding devices, these techniques can also be performed by a video encoder / decoder (commonly referred to as a "codec"). In addition, the techniques of the present invention can also be performed by a video pre-processor. The video encoder and / or decoder can be a graphics processing unit (GPU) or a similar device.
[0069] The source device 12 and the destination device 14 are only examples of such encoding and decoding devices, where the source device 12 generates encoded video data for transmission to the destination device 14. In some examples, the source device 12 and the destination device 14 can operate substantially symmetrically, such that both the source device 12 and the destination device 14 include video encoding and decoding components. Thus, the encoding system 10 can support one-way or two-way video transmission between the video devices 12, 14 for video streaming, video playback, video broadcasting, video telephony, etc.
[0070] The video source 18 of the source device 12 can include a video capture device (such as a video camera), a video archive including previously captured video, and / or a video feed interface for receiving video from a video content provider. Optionally, the video source 18 can generate computer graphics-based data as the source video, or as a combination of live video, archived video, and computer-generated video.
[0071] In some cases, when the video source 18 is a video camera, the source device 12 and the destination device 14 can form a so-called camera phone or video phone. However, as described above, the techniques described in the present invention are generally applicable to video encoding and decoding and can also be used in wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. Then, the encoded video information can be output to the computer-readable medium 16 through the output interface 22.
[0072] The computer-readable medium 16 can include transient media, such as wireless broadcasts or wired network transmissions, or can include storage media (i.e., non-transitory storage media), such as hard disks, flash drives, optical discs, digital video discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device 12 and provide the encoded video data to the destination device 14 via network transmissions and the like. Similarly, a computing device in a media production facility (such as an optical disc stamping facility) can receive the encoded video data from the source device 12 and produce an optical disc that includes the encoded video data. Thus, in various examples, the computer-readable medium 16 can be understood to include one or more computer-readable media in various forms.
[0073] The input interface 28 of the destination device 14 receives information from the computer-readable medium 16. The information of the computer-readable medium 16 can include syntax information defined by the video encoder 20. This syntax information is also used by the video decoder 30 and includes syntax elements that describe the characteristics and / or processing of blocks and other coding units (e.g., group of pictures (GOP)). The display device 32 displays the decoded video data to the user and can include any type of display device, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light emitting diode (OLED) display, or other type of display device.
[0074] The video encoder 20 and the video decoder 30 can operate according to a video coding standard (such as the High Efficiency Video Coding (HEVC) standard currently under development) and can conform to the HEVC Test Model (HM). Alternatively, the video encoder 20 and the video decoder 30 can operate according to other proprietary or industry standards, such as the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.264 standard, or Motion Picture Expert Group (MPEG)-4, Part 10, Advanced Video Coding (AVC), H.265 / HEVC, and extensions of such standards. However, the techniques of the present invention are not limited to any particular coding standard. Other examples of video coding standards include MPEG-2 and ITU-T H.263. Although in Figure 1Although not shown in the figure, in some aspects, the video encoder 20 and the video decoder 30 can each be integrated with an audio encoder and an audio decoder, and can include appropriate multiplexer - demultiplexer (MUX - DEMUX) units or other hardware and software to perform encoding processing on both audio and video in a common data stream or separate data streams. If applicable, the MUX - DEMUX unit can conform to the ITU H.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).
[0075] The video encoder 20 and the video decoder 30 can each be implemented as any suitable encoder circuit, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented partially in software form, the device can store the instructions of the software in a suitable non - transitory computer - readable medium and use one or more processors to execute the instructions in hardware form to perform the technology of the present invention. Both the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders and can be integrated as part of a combined encoder / decoder (codec) in the corresponding device. Devices including the video encoder 20 and / or the video decoder 30 can include integrated circuits, microprocessors, and / or wireless communication devices such as cellular phones.
[0076] Figure 2 is a block diagram showing an example of a video encoder 20 in which video codec technology can be implemented. The video encoder 20 can perform intra - and inter - frame encoding of video blocks within a video slice. Intra - frame encoding relies on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame or image. Inter - frame encoding relies on temporal prediction to reduce or remove temporal redundancy in the video within adjacent frames or images of a video sequence. The intra - frame mode (I - mode) can refer to any one of several spatial - based encoding modes. Inter - frame modes, such as unidirectional (also known as uni - prediction) prediction (P - mode) or bidirectional prediction (also known as bi - prediction) (B - mode), can refer to any one of several temporal - based encoding modes.
[0077] As Figure 2 shown, the video encoder 20 receives a current video block within the video frame to be encoded. In Figure 2In the example, video encoder 20 includes mode selection unit 40, reference frame memory 64, adder 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Mode selection unit 40 sequentially includes motion compensation unit 44, motion estimation unit 42, intra prediction (also referred to as, intra-frame prediction) unit 46, and segmentation unit 48. For video block reconstruction, video encoder 20 further includes inverse quantization unit 58, inverse transform unit 60, and adder 62. A deblocking filter ( Figure 2 not shown in ) may be included to filter block boundaries so as to remove blocking artifacts from the reconstructed video. If desired, the deblocking filter will typically filter the output of adder 62. In addition to the deblocking filter, additional filters (in-loop or post-loop) may be used. Such filters are not shown for simplicity, but if desired, can filter the output of adder 50 (as an in-loop filter).
[0078] During the encoding process, video encoder 20 receives a video frame or slice to be encoded. The frame or slice may be segmented into multiple video blocks. Motion estimation unit 42 and motion compensation unit 44 perform inter-frame predictive coding on the received video block relative to one or more blocks in one or more reference frames to provide temporal prediction. Intra prediction unit 46 may also perform intra-frame predictive coding on the received video block relative to one or more adjacent blocks located in the same frame or slice as the block to be encoded to provide spatial prediction. Video encoder 20 may execute multiple encoding channels, e.g., select an appropriate encoding mode for each video data block.
[0079] In addition, segmentation unit 48 may segment a block of video data into sub-blocks based on an evaluation of a previous segmentation scheme in a previous encoding channel. For example, segmentation unit 48 may initially segment a frame or slice into largest coding units (LCUs), and based on rate-distortion analysis (e.g., rate-distortion optimization), segment each LCU into sub-coding units (sub-CUs). Mode selection unit 40 may further generate a quadtree data structure indicating the segmentation of the LCU into sub-CUs. The leaf node CUs of the quadtree may include one or more prediction units (PUs) and one or more transform units (TUs).
[0080] The present invention uses the term "block" to refer to any one of a CU, a PU, or a TU in the HEVC context, or a similar data structure in other standard contexts (e.g., a macroblock and its sub-blocks in H.264 / AVC). A CU includes an encoding node, a PU and a TU associated with the encoding node. The size of the CU corresponds to the size of the encoding node and is square. The size range of the CU can be from 8×8 pixels to a tree block size with a maximum of 64×64 pixels or larger. Each CU can include one or more PUs and one or more TUs. For example, the syntax data associated with the CU can describe splitting the CU into one or more PUs. The splitting pattern may vary depending on whether the CU is skipped or encoded in direct mode, intra prediction mode, or inter prediction (also known as, inter-frame prediction) mode. The PU can be split into non-square shapes. For example, the syntax data associated with the CU can also describe splitting the CU into one or more TUs according to a quadtree. The TU can be square or non-square (e.g., rectangular).
[0081] The mode selection unit 40 can select one of the intra or inter encoding modes based on error results, etc., provide the resulting intra or inter encoded block to the adder 50 to generate residual block data, and provide it to the adder 62 to reconstruct the encoded block for use as a reference frame. The mode selection unit 40 also provides syntax elements, such as motion vectors, intra mode indicators, splitting information, and other such syntax information, to the entropy encoding unit 56.
[0082] The motion estimation unit 42 and the motion compensation unit 44 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, where the motion vector is used to estimate the motion of a video block. For example, the motion vector can indicate the displacement of a PU of a video block within the current video frame or image relative to a predicted block within a reference frame (or other encoding unit), where the reference frame (or other encoding unit) is related to the current block of the current frame (or other encoding unit). The predicted block is a block found to highly match the block to be encoded in terms of pixel differences, where the pixel differences can be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. In some examples, the video encoder 20 can calculate values at sub-integer pixel positions of a reference image stored in the reference frame memory 64. For example, the video encoder 20 can insert values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation unit 42 can perform a motion search with respect to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0083] The motion estimation unit 42 calculates the motion vector of the PU of the video block in the inter-coded slice by comparing the position of the PU with the position of the predicted block of the reference image. The reference image can be selected from the first reference image list (list 0) or the second reference image list (list 1), and each list identifies one or more reference pictures stored in the reference frame memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.
[0084] The motion compensation performed by the motion compensation unit 44 may involve obtaining or generating a predicted block based on the motion vector determined by the motion estimation unit 42. Additionally, in some examples, the motion estimation unit 42 and the motion compensation unit 44 may be functionally integrated. Once the motion vector of the PU of the current video block is received, the motion compensation unit 44 can locate the predicted block pointed to by the motion vector in one of the reference image lists. The adder 50 forms a residual video block by subtracting the pixel values of the predicted block from the pixel values of the encoded current video block, thereby forming a pixel difference, as discussed below. Generally, the motion estimation unit 42 performs motion estimation with respect to the luminance component, and the motion compensation unit 44 uses the motion vector calculated based on the luminance component for both the chrominance component and the luminance component. The mode selection unit 40 may also generate syntax elements related to the video block and the video slice for use by the video decoder 30 to decode the video block of the video slice.
[0085] The intra prediction unit 46 can perform intra prediction on the current block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, as described above. In particular, the intra prediction unit 46 can determine the intra prediction mode for encoding the current block. In some examples, the intra prediction unit 46 can encode the current block using various intra prediction modes in a separate coding channel, and the intra prediction unit 46 (or in some examples, the mode selection unit 40) can select an appropriate intra prediction mode from the test modes for use.
[0086] For example, the intra prediction unit 46 can calculate the rate-distortion value for various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. The intra prediction unit 46 can calculate the ratio based on the distortion and rate of various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0087] In addition, the intra prediction unit 46 can be used to encode depth blocks of a depth image using a depth modeling mode (DMM). The mode selection unit 40 can determine whether the available DMM mode produces better encoding results than the intra prediction mode and other DMM modes (e.g., using rate-distortion optimization (RDO)). Data of the texture image corresponding to the depth image can be stored in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 can also be used to inter-predict depth blocks of the depth image.
[0088] After selecting an intra prediction mode for a block (e.g., one of a conventional intra prediction mode or a DMM mode), the intra prediction unit 46 can provide information indicating the intra prediction mode selected for the block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicating the selected intra prediction mode. The video encoder 20 can include configuration data in the transmitted bitstream, where the configuration data includes a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables), definitions of the encoding context for various blocks, and indications of the most probable intra prediction mode, the intra prediction mode index table, and the modified intra prediction mode index table for each context.
[0089] The video encoder 20 forms a residual video block by subtracting the prediction data from the mode selection unit 40 from the encoded original video block. The adder 50 represents one or more components that perform this subtraction operation.
[0090] The transform processing unit 52 applies a transform to the residual block, such as a discrete cosine transform (DCT) or a conceptually similar transform, thereby producing a video block including residual transform coefficient values. The transform processing unit 52 can perform other transforms conceptually similar to the DCT. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used.
[0091] The transform processing unit 52 applies a transform to the residual block, thereby producing a block with residual transform coefficients. The transform can convert the residual information from the pixel value domain to the transform domain, such as the frequency domain. The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 can subsequently scan the matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 can perform the scan.
[0092] After quantization, the entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, the entropy coding unit 56 may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. In the case of context-based entropy coding, the context may be based on neighboring blocks. After the entropy coding unit 56 performs entropy coding, the coded bitstream may be transmitted to another device (e.g., the video decoder 30) or archived for later transmission or retrieval.
[0093] The inverse quantization unit 58 and the inverse transform unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain, e.g., for later use as a reference block. The motion compensation unit 44 may calculate a reference block by adding the residual block to a predicted block of a frame in the reference frame memory 64. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for motion estimation. The adder 62 adds the reconstructed residual block to the motion compensation predicted block generated by the motion compensation unit 44 to generate a reconstructed video block for storage in the reference frame memory 64. The motion estimation unit 42 and the motion compensation unit 44 may use the reconstructed video block as a reference block for intra-coding blocks in subsequent video frames.
[0094] Figure 3 is a block diagram showing an example of a video decoder 30 in which video codec techniques may be implemented. In Figure 3 the example, the video decoder 30 includes an entropy decoding unit 70, a motion compensation unit 72, an intra prediction unit 74, an inverse quantization unit 76, an inverse transform unit 78, a reference frame memory 82, and an adder 80. In some examples, the video decoder 30 may perform a decoding channel that is generally opposite to the encoding channel described with respect to the video encoder 20 ( Figure 2 ). The motion compensation unit 72 may generate prediction data based on motion vectors received from the entropy decoding unit 70, while the intra prediction unit 74 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 70.
[0095] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements from the video encoder 20. The entropy decoding unit 70 of the video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators, as well as other syntax elements. The entropy decoding unit 70 forwards the motion vectors and other syntax elements to the motion compensation unit 72. The video decoder 30 may receive syntax elements at the video slice level and / or at the video block level.
[0096] When a video slice is encoded as an intra-coded (I) slice, the intra prediction unit 74 may generate prediction data for video blocks of the current video slice based on the intra prediction mode and data signaled from previously decoded blocks of the current frame or picture. When a video frame is encoded as an inter-coded (e.g., B, P, or GPB) slice, the motion compensation unit 72 generates a prediction block for a video block of the current video slice based on the motion vectors and other syntax elements received from the entropy decoding unit 70. These prediction blocks may be generated from one of the reference images within one of the reference image lists. The video decoder 30 may use a default construction technique to construct the reference frame lists: list 0 and list 1, based on the reference images stored in the reference frame memory 82.
[0097] The motion compensation unit 72 determines prediction information for video blocks in the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to generate a prediction block for the decoded current video block. For example, the motion compensation unit 72 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used for encoding a video block of a video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more reference image lists for the slice, a motion vector for each inter-coded video block of the slice, an inter prediction state for each inter-coded video block of the slice, and other information, to decode a video block within the current video slice.
[0098] The motion compensation unit 72 may also perform interpolation based on an interpolation filter. The motion compensation unit 72 may use the interpolation filter used by the video encoder 20 during video block encoding to calculate the interpolation of sub-integer pixels of a reference block. In this case, the motion compensation unit 72 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate a prediction block.
[0099] Data of a texture image corresponding to a depth image may be stored in the reference frame memory 82. The motion compensation unit 72 may also be used for inter prediction of depth blocks of a depth image.
[0100] This disclosure presents various mechanisms for improving the encoding efficiency of SVT. As described above, the transform block can be positioned at various candidate positions relative to the corresponding residual block. The disclosed mechanisms employ different transforms for the transform block based on the candidate positions. For example, the inverse discrete sine transform (DST) can be applied at the candidate position covering the lower right corner of the residual block. Additionally, the inverse DCT can be applied at the candidate position covering the upper left corner of the residual block. This mechanism can be beneficial because DST is generally more effective than DCT for transforming residual blocks with more residual information distributed in the lower right corner, while DCT is generally more effective than DST for transforming residual blocks with more residual information distributed in the upper left corner. It should also be noted that in most cases, the lower right corner of the residual block statistically includes more residual information. In some cases, the disclosed mechanism also supports reversing the residual samples of the transform block horizontally. For example, after applying the inverse transform block, the residual samples can be horizontally reversed / flipped. This may occur when the adjacent block on the right side of the current residual block has been reconstructed and the adjacent block on the left side of the current residual block has not been reconstructed. This may also occur when the inverse DST is employed as part of the corresponding transmission block. This method provides greater flexibility in encoding the information closest to the reconstructed block, which reduces the corresponding residual information. The disclosed mechanism also supports context encoding of the candidate position information of the transform block. The prediction information corresponding to the residual block can be used to encode the candidate position information. For example, in some cases, the residual block can correspond to a predicted block generated by a template matching pattern. Additionally, the template used in the template matching pattern can be selected based on the spatially adjacent reconstructed regions of the residual block. In this case, the lower right part of the residual block can include more residual information than other parts of the residual block. Therefore, the candidate position covering the lower right part of the residual block is most likely to be selected as the best position for the transform. Thus, when the residual block is associated with inter-frame prediction based on template matching, only one candidate position can be provided for the residual block, and / or other context encoding methods can be used to encode the position information for the transform.
[0101] Figure 4It is a schematic diagram of an exemplary intra prediction mode 400 adopted in video encoding and decoding in the HEVC model. Video compression schemes utilize data redundancy. For example, most images include groups of pixels that include the same or similar colors and / or light as adjacent pixels. As a particular example, a night sky image may include large areas of black pixels depicting stars and clusters of white pixels. Intra prediction mode 400 utilizes these spatial relationships. Specifically, a frame can be decomposed into a series of blocks that include samples. Then, instead of transmitting each block, the light / color of the block can be predicted based on the spatial relationship with reference samples in adjacent blocks. For example, the encoder can indicate that the current block includes the same data as the reference samples in the previously encoded block located at the upper left corner of the current block. Then, the encoder can encode the prediction mode rather than the value of the block. This significantly reduces the size of the encoding. As shown in intra prediction mode 400, the upper left corner corresponds to prediction mode 18 in HEVC. Therefore, the encoder can store only the prediction mode 18 of the block instead of encoding the pixel information of the block. As shown, intra prediction mode 400 in HEVC includes thirty-three angular prediction modes from prediction mode 2 to prediction mode 34. Intra prediction mode 400 also includes intra planar mode 0 for predicting smooth regions and intra direct current (DC) mode 1. Intra planar mode 0 predicts the block as an amplitude surface with vertical and horizontal slopes derived from adjacent reference samples. Intra DC mode 1 predicts the block as the average of adjacent reference samples. Intra prediction mode 400 can be used to signal the luminance (e.g., light) component of the block. Intra prediction can also be applied to chrominance (e.g., color) values. In HEVC, chrominance values are predicted by adopting planar mode, angular mode twenty-six (e.g., vertical), angular mode ten (e.g., horizontal), intra DC, and derived mode, where the derived mode predicts the correlation between the chrominance component encoded by intra prediction mode 400 and the luminance component.
[0102] The intra prediction mode 400 of the image block is stored by the encoder as prediction information. It should be noted that although intra prediction mode 400 is adopted for prediction in a single frame, inter prediction can also be adopted. Inter prediction utilizes the temporal redundancy between multiple frames. For example, a scene in a movie may include a relatively stationary background, such as an immobile table. Therefore, the table is described as a set of substantially the same pixels in multiple frames. Inter prediction uses a block matching algorithm to match and compare a block in the current frame with a block in a reference frame. Then, the motion vector can be encoded to indicate the position of the best-matching block in the reference frame and the co-located position of the current / target block. Therefore, a series of image frames can be represented as a series of blocks, which can then be represented as prediction blocks including prediction modes and / or motion vectors.
[0103] Figure 5 An example of intra prediction 500 in video coding and decoding using an intra prediction mode such as intra prediction mode 400 is shown. As shown, the current block 501 can be predicted by samples in the neighboring blocks 510. The encoder can typically encode an image from top left to bottom right. However, in some cases, the encoder can encode from right to left, as discussed below. It should be noted that as used herein, the right side refers to the right side of the encoded image, the left side refers to the left side of the encoded image, the top refers to the top side of the encoded image, and the bottom refers to the bottom side of the encoded image.
[0104] It should be noted that the current block 501 may not always exactly match the samples from the neighboring blocks 510. In such cases, the prediction mode is encoded from the most matching neighboring block 510. To allow the decoder to determine the appropriate value, the difference between the predicted value and the actual value is retained. This is called residual information. Residual information exists both in intra prediction 500 and in inter prediction.
[0105] Figure 6 is a schematic diagram of an exemplary video coding and decoding mechanism 600 based on intra prediction 500 and / or inter prediction. The image block 601 can be obtained by the encoder from one or more frames. For example, an image can be segmented into multiple rectangular image regions. Each region of the image corresponds to a Coding Tree Unit (CTU). The CTU is segmented into multiple blocks, such as coding units in HEVC. Then, the block segmentation information is encoded in the bitstream 611. Thus, the image block 601 is a segmented part of the image, including pixels representing the luminance component and / or chrominance component at the corresponding part of the image. During encoding, the image block 601 is encoded as a prediction block 603, where the prediction block 603 includes prediction information, such as a prediction mode for intra prediction (e.g., intra prediction mode 400) and / or a motion vector for inter prediction. Then, encoding the image block 601 as the prediction block 603 can leave a residual block 605, where the residual block 605 includes residual information indicating the difference between the prediction block 603 and the image block 601.
[0106] Note that the image block 601 can be partitioned into coding units including a prediction block 603 and a residual block 605. The prediction block 603 can include all the prediction samples of the coding unit, and the residual block 605 can include all the residual samples of the coding unit. In this case, the prediction block 603 has the same size as the residual block 605. In another example, the image block 601 can be partitioned into coding units including two prediction blocks 603 and a residual block 605. In this case, each prediction block 603 includes a part of the prediction samples of the coding unit, and the residual block 605 includes all the residual samples of the coding unit. In yet another example, the image block 601 is partitioned into coding units including two prediction blocks 603 and four residual blocks 605. The partitioning pattern of the residual blocks 605 in the coding unit can be signaled in the bitstream 611. Such location patterns can include the Residual Quad-Tree (RQT) in HEVC. Additionally, the image block 601 can include only the luminance component (e.g., light) of the image samples (or pixels), denoted as the Y component. In other cases, the image block 601 can include the Y, U, and V components of the image samples, where U and V indicate the chrominance components (e.g., colors) in the blue luminance and red luminance color spaces (UV).
[0107] SVT can be used to further compress the information. Specifically, SVT employs a transform block 607 to further compress the residual block 605. The transform block 607 includes transforms such as inverse DCT and / or inverse DST. The difference between the prediction block 603 and the image block 601 is adapted to the transform by employing transform coefficients. By indicating the transform mode (e.g., inverse DCT and / or inverse DST) of the transform block 607 and the corresponding transform coefficients, the decoder can reconstruct the residual block 605. When exact reproduction is not required, the transform coefficients can be further compressed by rounding to specific values to create a better fit to the transform. This process is called quantization and is performed according to a quantization parameter that describes the allowed quantization. Thus, the transform mode, transform coefficients, and quantization parameter of the transform block 607 are stored as transformed residual information in a transformed residual block 609, which in some cases can also simply be referred to as the residual block.
[0108] Then, the prediction information of the prediction block 603 and the transformed residual information of the transformed residual block 609 can be encoded in the bitstream 611. The bitstream 611 can be stored and / or transmitted to the decoder. Then, the decoder can perform the reverse process to recover the image block 601. Specifically, the decoder can use the transformed residual information to determine the transform block 607. Then, the transform block 607 can be used in combination with the transformed residual block 609 to determine the residual block 605. Then, the residual block 605 and the prediction block 603 can be used to reconstruct the image block 601. Then, the image block 601 can be positioned relative to other decoded image blocks 601 to reconstruct the frame and position such frames to recover the encoded video.
[0109] SVT will now be described in further detail. For SVT, a transform block 607 smaller than the residual block 605 is selected. The transform block 607 is used to transform the corresponding part of the residual block 605 and leave the remaining part of the residual block 605 without additional encoding / compression. This is because the residual information is typically not evenly distributed in the residual block 605. SVT uses a smaller transform block 607 with an adaptive position to capture most of the residual information in the residual block 605 without the need to transform the entire residual block 605. This method can achieve a higher coding efficiency than transforming all the residual information in the transformed residual block 605. Since the transform block 607 is smaller than the residual block 605, SVT employs a mechanism for signaling the position of the transform relative to the residual block 605. For example, when applying SVT to a residual block 605 of size w×h (e.g., width times height), the size and position information of the transform block 607 can be encoded into the bitstream 611. This allows the decoder to reconstruct the transform block 607 and combine the transform block 607 into the correct position relative to the transformed residual block 609 to reconstruct the residual block 605.
[0110] It should be noted that some prediction blocks 603 can be encoded without producing a residual block 605. However, this case does not result in the use of SVT and is not discussed further. As described above, SVT can be used for inter-prediction blocks or intra-prediction blocks. In addition, SVT can be used for residual blocks 605 generated by a specified inter-prediction mechanism (e.g., motion compensation based on a translational model), but not for residual blocks 605 generated by other specified inter-prediction mechanisms (e.g., affine model motion compensation).
[0111] Figure 7 An example 700 of SVT including a transform block 707 and a residual block 705 is shown. Figure 7 The transform block 707 and the residual block 705 of Figure 6 are respectively similar to the transform block 607 and the residual block 605 of
[0112] SVT-I is described as w_t = w / 2, h_t = h / 2, where w_t and h_t represent the width and height of transform block 707 respectively, and where w and h represent the width and height of residual block 705 respectively. For example, both the width and height of transform block 707 are half of the width and height of residual block 705. SVT-II is described as w_t = w / 4, h_t = h, where the variables are as described above. For example, the width of transform block 707 is one quarter of the width of residual block 705, and the height of transform block 707 is equal to the height of residual block 705. SVT-III is described as w_t = w, h_t = h / 4, where the variables are as described above. For example, the width of transform block 707 is equal to the width of residual block 705, and the height of transform block 707 is equal to one quarter of the height of residual block 705. Type information indicating the SVT type (e.g., SVT-I, SVT-II, or SVT-III) is encoded into the bitstream to support reconstruction by the decoder.
[0113] As Figure 7 shown, each transform block 707 can be positioned at various locations relative to residual block 705. The position of transform block 707 is represented by a position offset (x, y) to the upper left corner of residual block 705, where x indicates the horizontal distance in pixels between the upper left corner of transform block 707 and the upper left corner of residual block 705, and y indicates the vertical distance in pixels between the upper left corner of transform block 707 and the upper left corner of residual block 705. Each potential position of transform block 707 within residual block 705 is called a candidate position. For residual block 705, the number of candidate positions for one type of SVT is (w – w_t + 1) × (h – h_t + 1). More specifically, for a 16×16 residual block 705, when using SVT-I, there are eighty-one candidate positions. When using SVT-II or SVT-III, there are thirteen candidate positions. Once determined, the x value and y value of the position offset, together with the adopted SVT block type, are encoded into the bitstream. To reduce the complexity of SVT-I, a subset of thirty-two positions can be selected from the eighty-one possible candidate positions. Then, this subset serves as the allowed candidate positions for SVT-I.
[0114] One drawback of the SVT scheme using one of the SVT examples 700 is that encoding the SVT position information as residual information results in significant signaling overhead. Additionally, the encoder complexity can increase significantly as the number of positions tested in a compression quality process such as Rate-Distortion Optimization (RDO) increases. Since the number of candidate positions increases with the size of the residual block 705, for larger residual blocks 705, such as 32×32 or 64×128, the signaling overhead can be even greater. Another drawback of using one of the SVT examples 700 is that the size of the transform block 707 is one-fourth the size of the residual block 705. In many cases, the size of the transform block 707 of this size may not be sufficient to cover the main residual information in the residual block 705.
[0115] Figure 8 An additional SVT example 800 including a transform block 807 and a residual block 805 is shown. Figure 8 The transform block 807 and the residual block 805 of which are respectively similar to Figures 6 - 7 the transform blocks 607, 707 and the residual blocks 605, 705. For ease of reference, the SVT example 800 is referred to as SVT Vertical (SVT-V) and SVT Horizontal (SVT-H). The SVT example 800 is similar to the SVT example 700 but is designed to support reduced signaling overhead and lower complex processing requirements for the encoder.
[0116] SVT-V is described as w_t = w / 2 and h_t = h, where the variables are as described above. The width of the transform block 807 is half the width of the residual block 805, and the height of the transform block 807 is equal to the height of the residual block 805. SVT-H is described as w_t = w and h_t = h / 2, where the variables are as described above. For example, the width of the transform block 807 is equal to the width of the residual block 805, and the height of the transform block 807 is half the height of the residual block 805. SVT-V is similar to SVT-II, and SVT-H is similar to SVT-III. Compared with SVT-II and SVT-III, the transform block 807 in SVT-V and SVT-H is enlarged to half of the residual block 805, so that the transform block 807 covers more residual information in the residual block 805.
[0117] Similar to SVT example 700, SVT example 800 may include a number of candidate positions, where a candidate position is a possible allowed position of a transform block (e.g., transform block 807) relative to a residual block (e.g., residual block 805). The candidate positions are determined according to the candidate position step size (CPSS). Each candidate position may be separated by an equal spacing specified by the CPSS. In this case, the number of candidate positions is reduced to no more than 5. Reducing the number of candidate positions alleviates the signaling overhead associated with the position information, since the selected position for the transform can be signaled with fewer bits. In addition, reducing the number of candidate positions makes the selection of the transform position algorithmically simpler, thus allowing the complexity of the encoder to be reduced (e.g., resulting in fewer computational resources for encoding).
[0118] Figure 9 SVT example 900 including a transform block 907 and a residual block 905 is shown. Figure 9 The transform block 907 and the residual block 905 of Figures 6 - 8 are respectively similar to the transform blocks 607, 707, 807 and the residual blocks 605, 705, 805 of Figure 9 Various candidate positions are shown, where a candidate position is a possible allowed position of a transform block (e.g., transform block 907) relative to a residual block (e.g., residual block 905). Specifically, Figure 9 in (A)- Figure 9 the SVT example in (E) employs SVT-V, Figure 9 in (F)- Figure 9 the SVT example in (J) employs SVT-H. The allowed candidate positions for the transform depend on the CPSS, which further depends on the portion of the residual block 905 that the transform block 907 should cover and / or the step size between candidate positions. For example, for SVT-V, the CPSS may be calculated as s = w / M1, or for SVT-H, the CPSS may be calculated as s = h / M2, where w and h are the width and height of the residual block respectively, and M1 and M2 are predetermined integers in the range of 2 to 8. The larger the value of M1 or M2, the more candidate positions. For example, both M1 and M2 may be set to 8. In this case, the value of the position index (P) describing the position of the transform block 907 relative to the residual block 905 is between 0 and 4.
[0119] In another example, for SVT-V, the CPSS is calculated as s = max(w / M1, Th1), and for SVT-H, the CPSS is calculated as s = max(h / M2, Th2), where Th1 and Th2 are predefined integers that specify the minimum step size. Th1 and Th2 can be integers not less than 2. In this example, Th1 and Th2 are set to 4, and M1 and M2 are set to 8. Different block sizes can have different numbers of candidate positions. For example, when the width of the residual block 905 is 8, two candidate positions are available for SVT-V, specifically Figure 9 in (A) and Figure 9 in (E). For example, when the step size indicated by Th1 is large and the portion of the residual block 905 covered by the transform block 907 as indicated by w / M1 is also large, only two candidate positions satisfy the CPSS. However, when w is set to 16, since w / M1 changes, the portion of the residual block 905 covered by the transform block 907 decreases. This results in more candidate positions. In this case, as Figure 9 in (A), Figure 9 in (C), and Figure 9 in (E) shown, three candidate positions. When the width of the residual block 905 is greater than 16 and the values of Th1 and M1 are as discussed above, Figure 9 all five candidate positions shown in (A)- Figure 9 in (E) are available.
[0120] When calculating the CPSS according to other mechanisms, other examples can also be seen. Specifically, for SVT-V, the CPSS can be calculated as s = w / M1; or for SVT-H, the CPSS can be calculated as s = h / M2. In this case, when M1 and M2 are set to 4, SVT-V allows 3 candidate positions (for example, Figure 9 in (A), Figure 9 in (C), and Figure 9 in (E) candidate positions); SVT-H allows 3 candidate positions (for example, Figure 9 in (F), Figure 9 in (H), and Figure 9 in (J) candidate positions). In addition, when M1 and M2 are set to 4, the portion of the residual block 905 covered by the transform block 907 increases, resulting in SVT-V having two allowed candidate positions (for example, Figure 9 in (A) and Figure 9 in (E) candidate positions) and SVT-H having two allowed candidate positions (for example, Figure 9 in (F) and Figure 9 in (J) candidate positions).
[0121] In another example, as discussed above, for SVT-V, CPSS is calculated as s = max(w / M1, Th1), or for SVT-H, CPSS is calculated as s = max(h / M2, Th2). In this case, T1 and T2 are set to predefined integers (e.g., 2), M1 is set to 8 when w ≥ h; M1 is set to 4 when w < h. M2 is set to 8 when h ≥ w; or M2 is set to 4 when h < w. For example, the portion of the residual block 905 covered by the transform block 907 depends on whether the height of the residual block 905 is greater than the width of the residual block 905, or whether the width of the residual block 905 is greater than the height of the residual block 905. Therefore, the number of candidate positions for SVT-H or SVT-V also depends on the aspect ratio of the residual block 905.
[0122] In another example, as discussed above, for SVT-V, CPSS is calculated as s = max(w / M1, Th1), or for SVT-H, CPSS is calculated as s = max(h / M2, Th2). In this case, the values of M1, M2, Th1, and Th2 are derived from a high-level syntax structure (e.g., the sequence parameter set) in the bitstream. For example, the values used to derive CPSS can be signaled in the bitstream. M1 and M2 may share the same value parsed from a syntax element, and Th1 and Th2 may share the same value parsed from another syntax element.
[0123] Figure 10 An exemplary SVT position 1000 is shown that describes the position of the transform block 1007 relative to the residual block 1005. Although six different positions are shown (e.g., three vertical positions and three horizontal positions), it should be understood that in practical applications, a different number of positions may be used. The SVT transform position 1000 is selected from the candidate positions in the SVT example 900 of Figure 9 . Specifically, the selected SVT transform position 1000 can be encoded into a position index (P). The position index P can be used to determine the position offset (Z) of the upper left corner of the transform block relative to the upper left corner of the residual block. For example, this position correlation can be determined according to Z = s × P, where s is the CPSS of the transform block based on the SVT type and is calculated as discussed with respect to Figure 9 . When the transform block is of the SVT-V type, the value of P can be encoded as When the transform block is of the SVT-H type, the value of P can be encoded as More specifically, (0, 0) can represent the coordinates of the upper left corner of the residual block. In this case, the coordinates of the upper left corner of the transform block are (Z, 0) for SVT-V and (0, Z) for SVT-H.
[0124] As discussed further below in detail, the encoder can encode by adopting a flag to indicate the SVT transform type (e.g., SVT-H or SVT-T) and the residual block size in the bitstream. Then, the decoder can determine the SVT transform size based on the SVT transform size and the residual block size. Once the SVT transform size is determined, the decoder can determine the allowed candidate positions of the SVT transform according to the CPSS function, such as Figure 9 the candidate positions in SVT example 900 of Figure 9 . Since the decoder is able to determine the candidate positions of the SVT transform, the encoder may not signal the coordinates of the position offset. Instead, codes can be adopted to indicate which candidate positions are used for the corresponding transform. For example, the position index P can be binary-coded into one or more binary bits using a truncated unary code to increase compression. As a particular example, when the P value is in the range of 0 to 4, the P values 0, 4, 2, 3, and 1 can be binary-coded into 0, 01, 001, 0001, and 0000 respectively. This binary code is more compressed than the base decimal value representing the position index. As another example, when the P value is in the range of 0 to 1, the P values 0 and 1 can be binary-coded into 0 and 1 respectively. Therefore, the size of the position index can be increased or decreased as needed to signal the specific transform block position according to the possible candidate positions of the transform block.
[0125] The position index P can be binary-coded into one or more binary bits by adopting the most likely position and the remaining less likely positions. For example, when the left adjacent block and the upper adjacent block have been decoded at the decoder and are thus available for prediction, the most likely position can be set to the position covering the lower right corner of the residual block. In one example, when the P value is in the range of 0 to 4 and position 4 is set as the most likely position, the P values 4, 0, 1, 2, and 3 are binary-coded into 1, 000, 001, 010, and 011 respectively. Additionally, when the P value is in the range of 0 to 2 and position 2 is set as the most likely position, the P values 2, 0, and 1 are binary-coded into 1, 01, and 00 respectively. Thus, the most likely position index of the candidate positions is represented with the fewest bits to reduce the signaling overhead in the most common cases. The probability can be determined based on the encoding order of the adjacent reconstructed blocks. Therefore, the decoder can infer the codeword scheme to be used for the corresponding block based on the adopted decoding scheme.
[0126] For example, in HEVC, the coding unit coding order is usually top - down and left - to - right. In this case, the right side of the current coding / decoding coding unit is not available, making the upper - right corner more likely to be the transform position. However, the motion vector prediction values are derived from the left and upper spatial neighbors. In this case, the residual information is statistically denser towards the lower - right corner. In this case, the candidate position covering the lower - right part is the most likely position. Additionally, when using an adaptive coding unit coding order, a node can be vertically split into two child nodes, and the right child node can be coded before coding the left child node. In this case, the right neighbor of the left child node has been reconstructed before decoding / encoding the left child node. Additionally, in this case, the left adjacent pixels are not available. When the right adjacent pixels are available and the left adjacent pixels are not available, the lower - left part of the residual block may include a large amount of residual information. Therefore, the candidate position covering the lower - left part of the residual block becomes the most likely position.
[0127] Therefore, the position index P can be binary - valued into one or more bits according to whether the right side adjacent to the residual block has been reconstructed. In one example, the P value ranges from 0 to 2, as shown in SVT transform position 1000. When the right side adjacent to the residual block has been reconstructed, the P values 0, 2, and 1 are binary - valued into 0, 01, and 00. Otherwise, the P values 2, 0, and 1 are binary - valued into 0, 01, and 00. In another example, when the right side adjacent to the residual block has been reconstructed but the left side adjacent to the residual block has not been reconstructed, the P values 0, 2, and 1 are binary - valued into 0, 00, and 01. Otherwise, the P values 2, 0, and 1 are binary - valued into 0, 00, and 01. In these examples, the position corresponding to a single bit is the most likely position, and the other two positions are the remaining positions. For example, the most likely position depends on the availability of the right - hand adjacent residual block.
[0128] In the sense of rate - distortion performance, the probability distribution of the best position may vary greatly among different inter - frame prediction modes. For example, when the residual block corresponds to a predicted block generated by template matching with spatially adjacent reconstructed pixels, the best position is most likely to be position 2. For other inter - frame prediction modes, the probability that position 2 (or position 0 when the right neighbor is available and the left neighbor is not available) is the best position is lower than that in the template - matching mode. In view of this, the context model of the first binary bit of the position index P can be determined according to the inter - frame prediction mode associated with the residual block. More specifically, when the residual block is associated with inter - frame prediction based on template matching, the first binary bit of the position index P uses the first context model. Otherwise, the second context model is used to encode / decode this binary bit.
[0129] In another example, when the residual block is associated with inter-frame prediction based on template matching, the most likely position (e.g., position 2, or position 0 when the right neighbor is available but the left neighbor is not) is directly set as the transform block position, and the position information is not signaled in the bitstream. Otherwise, the position index is explicitly signaled in the bitstream.
[0130] It should also be noted that different transforms can be employed according to the position of the transform block relative to the residual block. For example, the left side of the reconstructed residual block is reconstructed without reconstructing the right side of the residual block, which occurs in video coding and decoding with a fixed coding unit coding order from left to right and top to bottom (e.g., the coding order in HEVC). In this case, for candidate positions covering the lower right corner of the residual block, DST (e.g., DST version 7 (DST-7) or DST version 1 (DST-1)) can be employed for the transform in the transform block during encoding. Accordingly, the inverse DST transform is employed at the decoder for the corresponding candidate positions. In addition, for candidate positions covering the upper left corner of the residual block, DCT (e.g., DCT version 8 (DCT-8) or DCT version 2 (DCT-2)) can be employed for the transform in the transform block during encoding. Accordingly, the inverse DCT transform is employed at the decoder for the corresponding candidate positions. This is because, in this case, among the four corners, the lower right corner is the farthest from the spatial reconstruction region. In addition, when the transform block covers the lower right corner of the residual block, DST transforms the residual information distribution more effectively than DCT. In addition, when the transform block covers the upper left corner of the residual block, DCT transforms the residual information distribution more effectively than DST. For the remaining candidate positions, the transform type can be inverse DST or inverse DCT. For example, when the candidate position is closer to the lower right corner than the upper left corner, inverse DST is employed as the transform type. Otherwise, inverse DCT is employed as the transform type.
[0131] As a specific example, it can be allowed that the transform block 1007 has three candidate positions, as Figure 10 shown. In this case, position 0 covers the upper left corner, and position 2 covers the lower right corner. Position 1 is in the middle of the residual block 1005, equidistant from the left and right corners. At the encoder, the transform types for positions 0, 1, and 2 can be selected as DCT-8, DST-7, and DST-7, respectively. Then, the inverse transforms DCT-8, DST-7, and DST-7 can be employed at the decoder for positions 0, 1, and 2, respectively. In another example, at the encoder, the transform types for positions 0, 1, and 2 are DCT-2, DCT-2, and DST-7, respectively. Then, the inverse transforms DCT-2, DCT-2, and DST-7 can be employed at the decoder for positions 0, 1, and 2, respectively. Accordingly, the transform types for the corresponding candidate positions can be predetermined.
[0132] In some cases, the above-mentioned position-related multiple transforms may be applied only to the luminance transform blocks. The corresponding chrominance transform blocks may always use the inverse DCT-2 during the transform / inverse transform process.
[0133] Figure 11 Example 1100 shows a horizontal flip of the residual samples. In some cases, beneficial residual compression can be achieved by horizontally flipping the residual information in a residual block (e.g., residual block 605) before applying a transform block (e.g., transform block 607) at the encoder. Example 1100 shows such a horizontal flip. In this context, a horizontal flip means that the residual samples in the residual block are rotated by half around an axis between the left side and the right side of the residual block. This horizontal flip occurs before applying the transform (e.g., transform block) at the encoder and after applying the inverse transform (e.g., transform block) at the decoder. This flip can be employed when a specified predefined condition occurs.
[0134] In one example, the horizontal flip occurs when the transform block uses DST / inverse DST during the transform process. In this case, the right adjacent residual block of the current block is encoded / reconstructed before the current block, and the left adjacent residual block is not encoded / reconstructed before the current block. The horizontal flip process exchanges the residual sample at the i-th column of the residual block with the residual sample at the w–1–i-th column of the residual block. In this context, w is the width of the transform block, and i = 0, 1, ……, (w / 2)–1. The horizontal flip of the residual samples can improve the coding efficiency by making the residual distribution better fit the DST transform.
[0135] Figure 12 FIG. 1200 is a flowchart of an exemplary method for video decoding using position-related SVT that employs the mechanism discussed above. After receiving a bitstream (such as bitstream 611), method 1200 can be initiated at the decoder. Method 1200 uses the bitstream to determine a prediction block and a transform residual block, such as prediction block 603 and transform residual block 609. Method 1200 also determines a transform block (such as transform block 607), where transform block 607 is used to determine a residual block (such as residual block 605). Then, image block 601 is reconstructed using residual block 605 and prediction block 603. It should be noted that although method 1200 is described from the perspective of the decoder, a similar method (e.g., in reverse) can be employed to encode video by using SVT.
[0136] In block 1201, a bitstream is obtained at a decoder. The bitstream can be received from a memory or from a streaming media source. The bitstream includes data that can be decoded into at least one image corresponding to video data from an encoder. Specifically, the bitstream includes block partitioning information, where the block partitioning information can be used to determine coding units that include prediction blocks and residual blocks in the bitstream as described in mechanism 600. Thus, coding information related to the coding units can be parsed from the bitstream, and pixels of the coding units can be reconstructed based on the coding information discussed below.
[0137] In block 1203, a prediction block and a corresponding transformed residual block are obtained from the bitstream based on the block partitioning information. For this example, the transformed residual block has been encoded according to SVT as discussed above with respect to mechanism 600. Then, method 1200 reconstructs a residual block of size w×h from the transformed residual block as discussed below.
[0138] In block 1205, SVT usage, SVT type, and transform block size are determined. For example, the decoder first determines whether SVT has been used in the coding. This is because the transform employed by some codings is the size of the residual block. The usage of SVT can be signaled by a syntax element in the bitstream. Specifically, when a residual block is allowed to employ SVT, a flag, such as svt_flag, is parsed from the bitstream. A residual block is allowed to employ SVT when the transformed residual block has non-zero transform coefficients (e.g., corresponding to any luminance component or chrominance component). For example, a residual block can employ SVT when the residual block includes any residual data. The SVT flag indicates whether the residual block is encoded using a transform block of the same size as the residual block (e.g., svt_flag set to 0) or using a transform block smaller than the residual block (e.g., svt_flag set to 1). A coded block flag (cbf) can be employed to indicate whether the residual block includes non-zero transform coefficients of color components, as used in HEVC. Additionally, a root coded block (root cbf) flag can indicate whether the residual block includes non-zero transform coefficients of any color components, as used in HEVC. As a specific example, when an image block is predicted using inter prediction and the width or height of the residual block falls within a predetermined range of [a1, a2], the residual block is allowed to use SVT, where a1 = 16 and a2 = 64, or a1 = 8 and a2 = 64, or a1 = 16 and a2 = 128. The values of a1 and a2 can be predetermined fixed values. These values can also be derived from a sequence parameter set (SPS) or a slice header in the bitstream. When the residual block does not employ SVT, the transform block size is set to the width and height of the residual block size. Otherwise, the transform size is determined based on the SVT transform type.
[0139] Once the decoder determines that SVT has been used for a residual block, the decoder determines the type of the SVT transform block used and derives the transform block size based on the SVT type. The allowed SVT types for the residual block are determined based on the width and height of the residual block. If the width of the residual block is within the range of [a1, a2] as defined above for such values, the SVT-V transform as shown in Figure 8 is allowed. When the height of the residual block is within the range of [a1, a2] as defined above for such values, the SVT-H transform as shown in Figure 8 is allowed. SVT can be used only for the luminance component in the residual block, or SVT can be used for both the luminance component and the chrominance component in the residual block. When SVT is used only for the luminance component, the luminance component residual information is transformed by SVT, and the chrominance component is transformed by converting the size of the residual block. When both SVT-V and SVT-H are allowed, a flag (such as svt_type_flag) can be encoded into the bitstream. The svt_type_flag indicates whether SVT-V is used for the residual block (e.g., svt_type_flag is set to 0) or SVT-H is used for the residual block (e.g., svt_type_flag is set to 1). Once the type of the SVT transform is determined, the transform block size is set according to the signaled SVT type (e.g., for SVT-V, w_t = w / 2 and h_t = h; for SVT-H, w_t = w and h_t = h / 2). When only SVT-V or only SVT-H is allowed, the svt_type_flag may not be encoded into the bitstream. In this case, the decoder can infer the transform block size based on the allowed SVT type.
[0140] Once the SVT type and size are determined, the decoder proceeds to block 1207. In block 1207, the decoder determines the position of the transform relative to the residual block and the transform type (e.g., DST or DCT). The position of the transform block can be determined according to a syntax element in the bitstream. For example, in some examples, the position index can be directly signaled and parsed from the bitstream. In other examples, the position can be inferred as discussed with respect to Figures 8 - 10 . Specifically, the candidate positions of the transform can be determined according to the CPSS function. The CPSS function can determine the candidate positions by considering the width of the residual block, the height of the residual block, the SVT type determined in block 1205, the step size of the transform, and / or the portion of the residual block covered by the transform. Then, the decoder can determine the transform block position according to the candidate positions by obtaining a p index, where the p index includes the code that signals the correct candidate position according to the candidate position selection probability as discussed above with respect to Figure 10 . Once the transform block position is known, the decoder can infer the transform type adopted by the transform block, as discussed above with respect toFigure 10 As discussed. Accordingly, the encoder can select the corresponding inverse transform.
[0141] In block 1209, the decoder parses the transform coefficients of the transform block based on the transform block size determined in block 1205. This process can be completed according to the transform coefficient parsing mechanism adopted in HEVC, H.264, and / or AVC. The transform coefficients can be encoded using run-length encoding and / or as a set of coefficient groups (CGs). It should be noted that in some examples, block 1209 can be performed before block 1207.
[0142] In block 1211, the residual block is reconstructed based on the transform position, transform coefficients, and transform type determined as above. Specifically, inverse quantization and inverse transform of size w_t×h_t are applied to the transform coefficients to recover the residual samples of the residual block. The size of the residual block with residual samples is w_t×h_t. According to the position-dependent transform type determined in block 1207, the inverse transform can be an inverse DCT or an inverse DST. According to the transform block position, the residual samples are assigned to the corresponding regions within the residual block. Any residual samples inside the residual block and outside the transform block can be set to zero. For example, when using SVT-V, the number of candidate positions is 5, and the position index indicates the fifth transform block position. Therefore, the reconstructed residual samples are assigned to Figure 9 the region in the transform candidate positions of the SVT example 900 (e.g., Figure 9 the shaded region in Figure 9 ), and the region of size (w / 2)×h with zero residual samples to the left of this region (
[0143] In optional block 1213, the residual block information of the reconstructed block can be horizontally flipped, as discussed with respect to Figure 11 As described above, this may occur when the inverse DST is used for the transform block at the decoder, the right adjacent block has been reconstructed, and the left adjacent block has not been reconstructed. Specifically, in the above case, the encoder can horizontally flip the residual block before applying the DST transform to improve the coding efficiency. Therefore, optional block 1213 can be used to correct this horizontal flip at the encoder to create an accurate reconstructed block.
[0144] In block 1215, the reconstructed residual block can be combined with the prediction block to generate a reconstructed image block of the sample, which is part of the coding unit. The filtering process can also be applied to the reconstructed samples, such as deblocking filtering and sample adaptive offset (SAO) processing in HEVC. Then, the reconstructed image blocks can be combined with other image blocks decoded in a similar manner to generate a frame of the media / video file. Then, the reconstructed media file can be displayed to the user on a monitor or other display device.
[0145] It should be noted that an equivalent implementation of method 1200 can be adopted to generate reconstructed samples in the residual block. Specifically, the residual samples of the transform block can be directly combined with the prediction block at the position indicated by the transform block position information without first restoring the residual block.
[0146] Figure 13 This is video coding and decoding method 1300. Method 1300 can be implemented in a decoder (e.g., video decoder 30). In particular, method 1300 can be implemented by the processor of the decoder. When a bitstream has been received directly or indirectly from an encoder (e.g., video encoder 20) or retrieved from a memory, method 1300 can be implemented. In block 1301, the bitstream is parsed to obtain a prediction block (e.g., prediction block 603) and a transform residual block corresponding to the prediction block (e.g., transform residual block 609). In block 1303, the SVT type used to generate the transform residual block is determined. As described above, the SVT type can be SVT-V or SVT-H. In one embodiment, the SVT-V type includes a height equal to the height of the transform residual block and a width that is half of the width of the transform residual block.
[0147] In one embodiment, the SVT-H type includes a height that is half of the height of the transform residual block and a width equal to the width of the transform residual block. In one embodiment, the svt_type_flag is parsed from the bitstream to determine the SVT type. In one embodiment, when only one type of SVT is allowed for the residual block, the SVT type is determined by inference.
[0148] In block 1305, the position of the SVT relative to the transform residual block is determined. In one embodiment, a position index is parsed from the bitstream to determine the SVT type. In one embodiment, the position index includes a binary code, where the binary code indicates a position in a set of candidate positions determined according to CPSS. In one embodiment, the least number of bits indicating the position index in the binary code is assigned to the most likely position of the SVT. In one embodiment, when a single candidate position is available for the SVT transform, the processor infers the position of the SVT. In one embodiment, when the residual block is generated by template matching in an inter-frame prediction mode, the processor infers the position of the SVT.
[0149] In block 1307, the inverse of the SVT is determined based on the position of the SVT. In block 1309, the inverse of the SVT is applied to the transform residual block to generate a reconstructed residual block (e.g., residual block 605). In one embodiment, an inverse DST is used for an SVT-V type transform located at the left boundary of the residual block. In one embodiment, an inverse DST is used for an SVT-H type transform located at the top boundary of the residual block. In one embodiment, an inverse DCT is used for an SVT-V type transform located at the right boundary of the residual block. In one embodiment, an inverse DCT is used for an SVT-H type transform located at the bottom boundary of the residual block.
[0150] In block 1311, the reconstructed residual block is combined with the prediction block to reconstruct the image block. In one embodiment, the image block is displayed on a display or monitor of an electronic device (e.g., a smartphone, a tablet, a laptop computer, a personal computer, etc.).
[0151] Optionally, method 1300 may further include: when the right adjacent coding unit of the coding unit associated with the reconstructed residual block has been reconstructed and the left adjacent coding unit of the coding unit has not been reconstructed, horizontally flipping the samples in the reconstructed residual block before combining the reconstructed residual block with the prediction block.
[0152] Figure 14 There is a video encoding and decoding method 1400. Method 1400 may be implemented in an encoder (e.g., video encoder 20). In particular, method 1400 may be implemented by a processor of the encoder. Method 1400 may be implemented to encode a video signal. In block 1401, a video signal is received from a video capture device (e.g., a camera). In one embodiment, the video signal includes an image block (e.g., image block 601).
[0153] In block 1403, a prediction block (e.g., prediction block 603) and a residual block (e.g., residual block 605) are generated to represent the image block. In block 1405, a transform algorithm is selected for the SVT based on the position of the SVT relative to the residual block. In block 1407, the residual block is transformed into a transform residual block using the selected SVT.
[0154] In block 1409, the type of the SVT is encoded into the bitstream. In one embodiment, the type of the SVT is an SVT-V type or an SVT-H type. In one embodiment, the SVT-V type includes a height equal to the height of the transform residual block and a width that is half of the width of the transform residual block. In one embodiment, the SVT-H type includes a height that is half of the height of the transform residual block and a width equal to the width of the transform residual block.
[0155] In block 1411, the position of the SVT is encoded into the bitstream. In one embodiment, the position of the SVT is encoded in a position index. In one embodiment, the position index includes a binary code, where the binary code indicates a position in a set of candidate positions determined according to CPSS. In one embodiment, the binary code indicating the position index with the fewest bits is assigned to the most likely position of the SVT.
[0156] In one embodiment, the processor employs the DST algorithm for the SVT-V type transform located at the left boundary of the residual block. In one embodiment, the processor selects the DST algorithm for the SVT-H type transform located at the top boundary of the residual block. In one embodiment, the processor selects the DCT algorithm for the SVT-V type transform located at the right boundary of the residual block. In one embodiment, the processor selects the DCT algorithm for the SVT-H type transform located at the bottom boundary of the residual block.
[0157] Optionally, when the right adjacent coding unit of the coding unit associated with the residual block has been encoded and the left adjacent coding unit of the coding unit has not been encoded, the processor horizontally flips the samples in the residual block before converting the residual block into a transformed residual block.
[0158] In block 1413, the prediction block and the transformed residual block are encoded into the bitstream. In one embodiment, the bitstream is configured to be transmitted to a decoder and / or transmitted to a decoder.
[0159] Figure 15 FIG. 13 is a schematic diagram of an exemplary computing device 1500 for video encoding and decoding according to an embodiment of the present invention. The computing device 1500 is suitable for implementing the disclosed embodiments described herein. The computing device 1500 includes: an input port 1520 and a receiver unit (Rx) 1510 for receiving data; a processor, a logic unit, or a central processing unit (CPU) 1530 for processing data; a transmitter unit (Tx) 1540 and an output port 1550 for transmitting data; and a memory 1560 for storing data. The computing device 1500 may further include an optical-to-electrical (OE) component and an electrical-to-optical (EO) component coupled to the input port 1520, the receiver unit 1510, the transmitter unit 1540, and an output port 1550 for optical or electrical signals to go in or out. In some examples, the computing device 1500 may further include a wireless transmitter and / or receiver.
[0160] The processor 1530 is implemented by hardware and software. The processor 1530 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 1530 communicates with the input port 1520, the receiver unit 1510, the transmitter unit 1540, the output port 1550, and the memory 1560. The processor 1530 includes an encoding / decoding module 1514. The encoding / decoding module 1514 implements the above-disclosed embodiments, such as method 1300 and method 1400, and other mechanisms for encoding / reconstructing residual blocks based on the transform block position when using SVT and any other mechanisms described above. For example, the encoding / decoding module 1514 implements, processes, prepares, or provides various encoding operations, such as encoding video data and / or decoding video data as discussed above. Thus, including the encoding / decoding module 1514 provides a significant improvement to the functionality of the computing device 1500 and enables the transition of the computing device 1500 to different states. Alternatively, the encoding / decoding module 1514 is implemented as instructions stored in the memory 1560 and executed by the processor 1530 (e.g., a computer program product stored on a non-transitory medium).
[0161] The memory 1560 includes one or more disks, tape drives, and solid state drives and can be used as an overflow data storage device to store such programs when selected for execution and to store the instructions and data read during program execution. The memory 1560 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM). The computing device 1500 may also include input / output (I / O) devices for interacting with an end user. For example, the computing device 1500 may include a display (such as a monitor) for visual output, speakers for audio output, and a keyboard / mouse / trackball, etc. for user input.
[0162] In summary, the above disclosure includes a mechanism for adaptively employing multiple transform types for transform blocks at different positions. Additionally, the present invention allows for horizontal flipping of residual samples in a residual block to improve coding efficiency. This occurs when the transform block uses DST and inverse DST at the encoder and decoder, respectively, and when the right adjacent block is available while the left adjacent block is not. Furthermore, the present invention includes a mechanism for supporting coding position information in a bitstream based on an inter-frame prediction mode associated with a residual block.
[0163] Figure 16 FIG. is a schematic diagram of an embodiment of an encoding component 1600. In an embodiment, the encoding component 1600 is implemented in a video codec device 1602 (e.g., video encoder 20 or video decoder 30). The video codec device 1602 includes a receiving component 1601. The receiving component 1601 is configured to receive an image for encoding or receive a bitstream for decoding. The video codec device 1602 includes a transmitting component 1607 coupled to the receiving component 1601. The transmitting component 1607 is configured to transmit the bitstream to a decoder or transmit the decoded image to a display component (e.g., one of the I / O devices in the computing device 1500).
[0164] The video codec device 1602 includes a storage component 1603. The storage component 1603 is coupled to at least one of the receiving component 1601 or the transmitting component 1607. The storage component 1603 is configured to store instructions. The video codec device 1602 further includes a processing component 1605. The processing component 1605 is coupled to the storage component 1603. The processing component 1605 is configured to execute the instructions stored in the storage component 1603 to perform the methods disclosed herein.
[0165] When there is no intermediate component between a first component and a second component other than a wire, trace, or other medium, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component in addition to a wire, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include both direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range of ±10% of the subsequent number.
[0166] Although several embodiments have been provided in the present invention, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention should be considered illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.
[0167] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques, or methods. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component in an electrical, mechanical, or other manner. Other instances of variations, substitutions, and alterations may be determined by those skilled in the art and may be exemplified without departing from the spirit and scope disclosed herein.
Claims
1. A method implemented in a decoding device, characterized in that, the method comprises: parsing a bitstream to obtain a prediction block and a transform residual block corresponding to the prediction block; determining a type of spatial-varying transform (SVT) for generating the transform residual block, the type of the SVT being a vertical SVT (SVT-V) type or a horizontal SVT (SVT-H) type; the height corresponding to the SVT-V type is equal to the height of the transform residual block, and the width is half of the width of the transform residual block; the height corresponding to the SVT-H type is half of the height of the transform residual block, and the width is equal to the width of the transform residual block; determining the position of the SVT relative to the transform residual block; determining an inverse transform of the SVT based on the position of the SVT; wherein, when the position of the SVT covers the lower right corner of the transform residual block, the inverse transform of the SVT includes an inverse discrete sine transform (DST) version 7; when the position of the SVT covers the upper left corner of the transform residual block, the inverse transform of the SVT includes an inverse discrete cosine transform (DCT) version 8; applying the inverse transform of the SVT to the transform residual block to generate a reconstructed residual block; combining the reconstructed residual block with the prediction block to reconstruct an image block.
2. The method according to claim 1, characterized in that, it further comprises: parsing an svt_type_flag in the bitstream to determine the type of the SVT.
3. The method according to claim 1 or 2, characterized in that, it further comprises: when only one type of SVT is allowed for the residual block, determining the type of the SVT by inference.
4. The method according to claim 1 or 2, characterized in that, it further comprises: parsing a position index in the bitstream to determine the position of the SVT.
5. The method according to claim 1 or 2, characterized in that, the position index includes a binary code, wherein the binary code indicates a position in a set of candidate positions determined according to a candidate position step size (CPSS).
6. The method according to claim 5, characterized in that, allocating the least number of bits in the binary code indicating the position index to the most likely position of the SVT.
7. The method according to claim 1 or 2, characterized in that, when a single candidate position is available for the SVT transform, the processor infers the position of the SVT.
8. The method according to claim 7, characterized in that, when the residual block is generated by template matching in an inter prediction mode, the processor infers the position of the SVT.
9. The method according to claim 1 or 2, characterized in that, an inverse discrete sine transform (DST) is adopted for a vertical SVT (SVT-V) type transform located at the left boundary of the residual block.
10. The method according to claim 1 or 2, characterized in that, an inverse DST is adopted for a horizontal SVT (SVT-H) type transform located at the top boundary of the residual block.
11. The method according to claim 1 or 2, characterized in that, An inverse discrete cosine transform (DCT) is applied to an SVT-V type transform located at the right boundary of the residual block.
12. The method according to claim 1 or 2, wherein, an inverse DCT is applied to an SVT-H type transform located at the bottom boundary of the residual block.
13. A method implemented in an encoding device, wherein, the method includes: receiving a video signal, wherein the video signal includes image blocks; generating a prediction block and a residual block to represent the image block; selecting a transform algorithm for the spatial varying transform (SVT) based on the position of the SVT relative to the residual block; using the selected SVT to transform the residual block into a transformed residual block, wherein when the position of the SVT covers the lower right corner of the transformed residual block, the SVT includes a discrete sine transform (DST) version 7; when the position of the SVT covers the upper left corner of the transformed residual block, the SVT includes a discrete cosine transform (DCT) version 8; encoding the type of the SVT into a bitstream; the type of the SVT is an SVT vertical (SVT-V) type or an SVT horizontal (SVT-H) type; the SVT-V type includes a height equal to the height of the transformed residual block and a width that is half of the width of the transformed residual block; the SVT-H type includes a height that is half of the height of the transformed residual block and a width equal to the width of the transformed residual block; encoding the position of the SVT into the bitstream; encoding the transformed residual block into the bitstream for transmission to a decoder.
14. The method according to claim 13, wherein, the position of the SVT is encoded in a position index.
15. The method according to claim 13 or 14, wherein, the position index includes a binary code, wherein the binary code indicates a position in a set of candidate positions determined according to a candidate position step size (CPSS).
16. The method according to claim 15, wherein, the least number of bits indicating the position index in the binary code is assigned to the most likely position of the SVT.
17. The method according to claim 13 or 14, wherein, a processor applies a discrete sine transform (DST) algorithm to an SVT vertical (SVT-V) type transform located at the left boundary of the residual block.
18. The method according to claim 17, wherein, the processor selects a DST algorithm for an SVT horizontal (SVT-H) type transform located at the top boundary of the residual block.
19. The method according to claim 17, wherein, the processor selects a discrete cosine transform (DST) algorithm for an SVT-V type transform located at the right boundary of the residual block.
20. The method according to claim 17, wherein, the processor selects a DCT algorithm for an SVT-H type transform located at the bottom boundary of the residual block.
21. An encoding and decoding device, wherein, comprising: A receiver for receiving an image for encoding or receiving a bitstream for decoding; A transmitter coupled to the receiver, wherein the transmitter is configured to transmit the bitstream to a decoder or transmit a decoded image to a display; A memory coupled to at least one of the receiver or the transmitter, wherein the memory is configured to store instructions; A processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory to perform the method according to any one of claims 1 to 20.
22. The encoding / decoding device according to claim 21, wherein, it further includes a display for displaying an image.
23. A system, wherein, it includes: An encoder; A decoder in communication with the encoder, wherein the encoder and the decoder include the encoding / decoding device according to claim 21 or 22.
24. A video encoding / decoding device, wherein, the video encoding / decoding device includes a receiving component, including: The receiving component for receiving an image for encoding or receiving a bitstream for decoding; A transmission component coupled to the receiving component, wherein the transmission component is configured to transmit the bitstream to a decoder or transmit a decoded image to a display component; A storage component coupled to at least one of the receiving component or the transmission component, wherein the storage component is configured to store instructions; A processing component coupled to the storage component, wherein the processing component is configured to execute the instructions stored in the storage component to perform the method according to any one of claims 1 to 20.
25. A non-volatile computer-readable medium, wherein, computer-executable instructions are stored on the non-volatile computer-readable medium; when the computer-executable instructions run on a video decoding device, the video decoding device is caused to execute the method according to any one of claims 1 to 20.
26. A computer program product, wherein, when the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 - 20.
27. A method for storing an encoded video bitstream, wherein, it includes: Receiving the video bitstream, wherein the video bitstream includes a video bitstream generated by the method according to any one of claims 13 to 20; Storing the video bitstream in a storage medium.
28. A system for storing an encoded video bitstream, wherein, it includes: A receiving unit for receiving the video bitstream, wherein the video bitstream includes a video bitstream generated by the method according to any one of claims 13 to 20; A storage unit for storing the video bitstream.
29. A method for transmitting an encoded video bitstream, wherein, it includes: Obtaining the video bitstream from a storage medium, wherein the video bitstream includes a video bitstream generated by the method according to any one of claims 13 to 20; Sending the video bitstream to a decoding device.
30. A system for transmitting an encoded video bitstream, wherein, it includes: An acquisition unit, configured to acquire the video bitstream from a storage medium, where the video bitstream includes a video bitstream generated by the method according to any one of claims 13 to 20; A sending unit, configured to send the video bitstream to a decoding device.
Citation Information
Patent Citations
Motion compensation prediction method based on motion vector restraint and weighting motion vector
CN103561263A
Alternative transforms for data compression
WO2015172337A1