Code Prediction for Block-Based Video Coding

A template-based hypothesis generation scheme for transform coefficient code prediction addresses decoding challenges, enhancing efficiency and quality in video coding by accurately predicting transform coefficient signs.

JP7719949B2Active Publication Date: 2025-08-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024507965
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-17
Filing Date
2022-08-16
Publication Date
2025-08-06
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently predicting transform coefficient signs during the decoding process, leading to suboptimal compression and quality trade-offs.

Method used

Implement a template-based hypothesis generation scheme for transform coefficient code prediction at the video decoder side, selecting a set of candidates, determining predictive codes, and updating dequantized coefficients based on received signaling bits.

Benefits of technology

Enhances the accuracy of transform coefficient sign prediction, improving video decoding efficiency and maintaining video quality during compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719949000021
    Figure 0007719949000021
  • Figure 0007719949000022
    Figure 0007719949000022
  • Figure 0007719949000023
    Figure 0007719949000023
Patent Text Reader

Abstract

Implementations of the present disclosure provide a video decoding apparatus and method for transform coefficient code prediction at a video decoder side. The method may include selecting a set of transform coefficient candidates from dequantized transform coefficients. The method may include applying a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of transform coefficient candidates. The method may include determining a combination of code candidates associated with the selected hypothesis to be a set of predictive codes for the set of transform coefficient candidates. The method may include estimating original codes of the set of transform coefficient candidates based on the set of predictive codes and a sequence of code signaling bits received from the video encoder. The method may include updating the dequantized transform coefficients based on the estimated original codes of the set of transform coefficient candidates.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Application No. 63 / 233,940, filed August 17, 2021, the entire contents of which are incorporated herein by reference. This application relates to video encoding and compression, and more particularly to video processing systems and methods for sign prediction in block-based video encoding. [Background technology]

[0002] Digital video is supported by a variety of electronic devices, including digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video gaming consoles, smartphones, video teleconferencing devices, and video streaming devices. Electronic devices transmit and receive or otherwise communicate digital video data over communication networks and / or store the digital video data on storage devices. Due to limited bandwidth capacity in communication networks and limited memory resources in storage devices, video coding may be used to compress video data according to one or more video coding standards before the video data is communicated or stored. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Expert Group (MPEG) coding, and others. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]

[0003] Implementations of the present disclosure provide a video decoding method for transform coefficient sign prediction at the video decoder side. [Means for solving the problem]

[0004] The video decoding method may include selecting, by one or more processors, a set of transform coefficient candidates for transform coefficient code prediction from dequantized transform coefficients. The dequantized transform coefficients are associated with transform blocks of a video frame from the video. The video decoding method may further include applying, by the one or more processors, a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of transform coefficient candidates. The video decoding method may further include determining, by the one or more processors, a combination of code candidates associated with the selected hypothesis as a set of predictive codes for the set of transform coefficient candidates. The video decoding method may further include estimating, by the one or more processors, original codes of the set of transform coefficient candidates based on the set of predictive codes and a sequence of code signaling bits received from the video encoder. The video decoding method may further include updating, by the one or more processors, the dequantized transform coefficients based on the estimated original codes of the set of transform coefficient candidates.

[0005] An implementation of the present disclosure also provides a video decoding device for transform coefficient code prediction at a video decoder side. The video decoding device may include a memory configured to store a video including a plurality of video frames and one or more processors coupled to the memory. The one or more processors may be configured to select a set of transform coefficient candidates from dequantized transform coefficients for transform coefficient code prediction. The dequantized transform coefficients are associated with transform blocks of video frames from the video. The one or more processors may be further configured to apply a template-based hypothesis generation scheme to select hypotheses from a plurality of candidate hypotheses for the set of transform coefficient candidates. The one or more processors may be further configured to determine a combination of code candidates associated with the selected hypotheses as a set of predictive codes for the set of transform coefficient candidates. The one or more processors may be further configured to estimate original codes of the set of transform coefficient candidates based on the set of predictive codes and a sequence of code signaling bits received from the video encoder. The one or more processors may be further configured to update the dequantized transform coefficients based on the estimated original codes of the set of transform coefficient candidates.

[0006] Implementations of the present disclosure also provide a non-transitory computer-readable storage medium having stored thereon instructions, which, when executed by one or more processors, cause the one or more processors to perform a video decoding method for transform coefficient code prediction at a video decoder. The video decoding method may include selecting a set of transform coefficient candidates from transform coefficients of dequantized code predictions. The dequantized transform coefficients are associated with transform blocks of a video frame from the video. The video decoding method may further include applying a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of transform coefficient candidates. The video decoding method may further include determining a combination of code candidates associated with the selected hypothesis as a set of predicted codes for the set of transform coefficient candidates. The video decoding method may further include estimating original codes of the set of transform coefficient candidates based on the set of predicted codes and a sequence of code signaling bits received through a bitstream from the video encoder. The video decoding method may further include updating the dequantized transform coefficients based on the original codes of the estimated set of transform coefficient candidates. The bitstream is stored in the non-transitory computer-readable storage medium.

[0007] Implementations of this disclosure also provide a non-transitory computer-readable storage medium storing a bitstream decodable by a video method. The video method includes selecting a set of transform coefficient candidates from dequantized transform coefficients for transform coefficient code prediction. The dequantized transform coefficients are associated with transform blocks of a video frame from the video. The video method includes applying a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of transform coefficient candidates. The video method includes determining a combination of code candidates associated with the selected hypothesis as a set of predictive codes for the set of transform coefficient candidates. The video method includes estimating original codes of the set of transform coefficient candidates based on the set of predictive codes and a sequence of code signaling bits received through a bitstream from a video encoder. The video method includes updating the dequantized transform coefficients based on the estimated original codes of the set of transform coefficient candidates.

[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present disclosure.

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a block diagram illustrating an example system for encoding and decoding video blocks according to some implementations of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an example video encoder according to some implementations of the present disclosure. [Figure 3] FIG. 2 is a block diagram illustrating an example video decoder according to some implementations of the present disclosure. [Figure 4A]1 is a graphical representation illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 4B] 1 is a graphical representation illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 4C] 1 is a graphical representation illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 4D] 1 is a graphical representation illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 4E] 1 is a graphical representation illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some implementations of the present disclosure. [Figure 5] 1 is a graphical representation illustrating a top-left scan order of transform coefficients within a coefficient group according to some examples. [Figure 6] 1 is a graphical representation illustrating a low-frequency non-separable transform (LFNST) process according to some examples. [Figure 7] 10 is a graphical representation showing the top left region of the linear transform coefficients input to a forward LFNST according to some examples. [Figure 8] 10 is a graphical representation illustrating search areas for intra-template matching according to some examples. [Figure 9] 1 is a graphical representation illustrating an exemplary process of code prediction according to some examples. [Figure 10] 10 is a graphical representation illustrating calculation of a cost function for symbol prediction according to some examples. [Figure 11]FIG. 2 is a block diagram illustrating an example code prediction process in block-based video coding according to some implementations of the present disclosure. [Figure 12] 1 is a graphical representation illustrating an example hypothesis generation based on a linear combination of templates according to some implementations of the present disclosure. [Figure 13A] 1 is a graphical representation illustrating an exemplary implementation of an existing code prediction scheme according to some examples. [Figure 13B] 1 is a graphical representation illustrating an example implementation of a vector-based code prediction scheme according to some implementations of the present disclosure. [Figure 14A] 10 is a graphical representation illustrating an example calculation of a left-diagonal cost function along the left diagonal direction according to some implementations of the present disclosure. [Figure 14B] 10 is a graphical representation illustrating an example calculation of a right-diagonal cost function along the right-diagonal direction according to some implementations of the present disclosure. [Figure 15] 1 is a flowchart of an example method for code prediction in block-based video coding according to some implementations of the present disclosure. [Figure 16] 10 is a flowchart of another example method for code prediction in block-based video coding according to some implementations of the present disclosure. [Figure 17] FIG. 1 is a block diagram illustrating a computing environment coupled with a user interface according to some implementations of the present disclosure. [Figure 18] 1 is a flowchart of an example video decoding method for transform coefficient sign prediction at the video decoder side according to some implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Reference will now be made in detail to particular implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0012] It should be noted that terms such as "first," "second," and the like used in the description of this disclosure, the claims, and the accompanying drawings are used to distinguish between objects and are not used to describe a particular order or sequence. It should be understood that such terms may be interchanged under appropriate conditions such that the embodiments of the disclosure described herein may be performed in an order other than that shown in the accompanying drawings or described in this disclosure.

[0013] 1 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0014] In some implementations, destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of transferring encoded video data from source device 12 to destination device 14. In one example, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.

[0015] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, storage device 32 may correspond to a file server or another intermediate storage device capable of storing encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, or local disk drives. Destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or any combination thereof suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0016] 1, source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may include a camera phone or a video phone. However, implementations described in this disclosure may be applicable to video encoding generally and may be applicable to wireless and / or wired applications.

[0017] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0018] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem and receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0019] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0020] Video encoder 20 and video decoder 30 may operate according to proprietary or industry standards, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the present disclosure is not limited to a particular video encoding / decoding standard and is applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data according to any of these current or future standards.

[0021] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If implemented partially in software, the electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.

[0022] 2 is a block diagram illustrating an example video encoder 20 according to some implementations described herein. Video encoder 20 may perform intra-predictive and inter-predictive coding of video blocks within video frames. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. Note that the term "frame" is sometimes used synonymously with the terms "image" or "picture" in the field of video coding.

[0023] 2, video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partition unit 45, an intra-prediction processing unit 46, and an intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes an inverse quantization unit 58 for video block reconstruction, an inverse transform processing unit 60, and an adder 62. An in-loop filter 63, such as a deblocking filter, may be disposed between adder 62 and DPB 64 to filter block boundaries and remove block artifacts from the reconstructed video data. In addition to the deblocking filter, another in-loop filter, such as an SAO filter and / or an adaptive in-loop filter (ALF), may also be used to filter the output of adder 62. In some examples, the in-loop filter may be omitted, and the decoded video block may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.

[0024] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from video source 18 shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or reference pictures) for use in encoding video data by video encoder 20 (e.g., in an intra-predictive coding mode or an inter-predictive coding mode). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with other components of video encoder 20 or may be off-chip relative to those components.

[0025] As shown in FIG. 2, after receiving video data, partitioning unit 45 in prediction processing unit 41 partitions the video data into video blocks. This partitioning may also include partitioning the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predefined division structure, such as a quad-tree (QT) structure associated with the video data. A video frame may be or be considered as a two-dimensional array or matrix of samples having sample values. The samples in the array may also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be divided into multiple video blocks, for example, by using QT partitioning. A video block may also be or be considered as a two-dimensional array or matrix of samples having sample values, although with smaller dimensions than a video frame. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further partitioned into one or more block partitions or sub-blocks (which may again form blocks), e.g., by repeatedly using QT partitioning, binary-tree (BT) partitioning, or triple-tree (TT) partitioning, or any combination thereof. Note that the term “block” or “video block” as used herein may be a portion of a frame or picture, particularly a rectangular (square or non-square) portion. For example, with reference to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or may be or correspond to a corresponding block, e.g., a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB).Alternatively or additionally, a block or video block may be or correspond to a sub-block such as CTB, CB, PB, TB, etc.

[0026] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., coding rate and distortion level). Prediction processing unit 41 may provide the resulting intra-predictively coded block or inter-predictively coded block (e.g., a predictive block) to summer 50 to generate a residual block and to summer 62 to reconstruct the coded block for later use as part of a reference frame. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy coding unit 56.

[0027] To select an appropriate intra-prediction coding mode for a current video block, intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0028] In some implementations, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating motion vectors that indicate the displacement of video blocks in the current video frame relative to predictive blocks in a reference frame according to a predetermined pattern in the sequence of video frames. Motion estimation performed by motion estimation unit 42 may be a process of generating motion vectors that may estimate the motion of video blocks. The motion vectors may indicate, for example, the displacement of video blocks in the current video frame or picture relative to predictive blocks in a reference frame. The predetermined pattern may designate video frames in the sequence as P frames or B frames. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to the determination of motion vectors by motion estimation unit 42 for inter prediction, or may utilize motion estimation unit 42 to determine block vectors.

[0029] A predictive block of a video block may be or correspond to a block or reference block of a reference frame that is deemed to closely match the video block to be encoded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of square differences (SSD), or other difference metrics. In some implementations, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional-pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches for whole pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.

[0030] Motion estimation unit 42 calculates a motion vector for a video block in an inter-predictively coded frame by comparing the position of the video block to the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy coding unit 56.

[0031] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may locate the predictive block to which the motion vector points in one of the reference frame lists, obtain the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual block of pixel difference values by subtracting pixel values of the predictive block provided by motion compensation unit 44 from pixel values of the current video block being coded. The pixel difference values forming the residual block may include luma difference components, chroma difference components, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be integrated and are shown separately in FIG. 2 for conceptual purposes.

[0032] In some implementations, the intra BC unit 48 may generate a vector and fetch a predictive block in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is in the same frame as the current block being coded, and the vector is referred to as a block vector rather than a motion vector. Specifically, the intra BC unit 48 may determine an intra prediction mode to use to code the current block. In some examples, the intra BC unit 48 may code the current block using various intra prediction modes, e.g., during separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 may then select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using the rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics from the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to create the coded block, as well as the bit rate (i.e., number of bits) used to create the coded block. Intra BC unit 48 may calculate ratios from the distortions and rates of the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0033] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction according to implementations described herein. In either case, for intra block copying, the predictive block may be a block that is deemed to closely match the block to be coded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identifying the predictive block may include calculating values for sub-integer pixel positions.

[0034] Regardless of whether the predictive block is a block from the same frame via intra prediction or a block from a different frame via inter prediction, video encoder 20 may form a residual block by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. The pixel difference values forming the residual block may include both luma and chroma component differences.

[0035] Intra-prediction processing unit 46 may intra-predict the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra-block copy prediction performed by intra BC unit 48, as described above. Specifically, intra-prediction processing unit 46 may determine an intra-prediction mode to use to encode the current block. For example, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.

[0036] After prediction processing unit 41 determines a predictive block for the current video block by inter-prediction or intra-prediction, adder 50 forms a residual block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0037] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0038] Following quantization, entropy coding unit 56 may use an entropy coding technique to encode the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding methodology or technique. The encoded bitstream may then be transmitted to video decoder 30 as shown in FIG. 1 or archived to storage device 32 as shown in FIG. 1 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also encode motion vectors and other syntax elements of the current video frame being coded using entropy coding techniques.

[0039] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain to generate reference blocks for prediction of other video blocks. A reconstructed residual block may be generated in this manner. As described above, motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0040] Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to create a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.

[0041] 3 is a block diagram illustrating an example video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.

[0042] In some examples, units of video decoder 30 may be tasked with performing implementations of the present application. Also, in some examples, implementations of the present disclosure may be divided among one or more of the units of video decoder 30. For example, intra BC unit 85 may perform implementations of the present application alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85, and the functions of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.

[0043] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. A data buffer (DPB) 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in an intra-prediction coding mode or an inter-prediction coding mode). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous dynamic random access memory (SDRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), or other types of memory devices. 3, video data memory 79 and DPB 92 are depicted as two separate components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30 or may be off-chip relative to those components.

[0044] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 may use entropy decoding techniques to decode the bitstream to obtain quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.

[0045] When a video frame is coded as an intra-predictive coded (e.g., I) frame or for intra-coded predictive blocks within other types of frames, intra prediction unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.

[0046] When a video frame is coded as an inter-predictive (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 creates one or more predictive blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the predictive blocks may be created from a reference frame in one of the reference frame lists. Video decoder 30 may construct the reference frame lists, e.g., List 0 and List 1, using a default construction technique based on the reference frames stored in DPB 92.

[0047] In some examples, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 creates a predictive block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be within the same reconstructed region of the picture as the current video block processed by video encoder 20.

[0048] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by analyzing motion vectors and other syntax elements, and then use the prediction information to create a predictive block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) used to encode the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), construction information regarding one or more of the frame's reference frame list, the motion vectors of each inter-predictively coded video block of the frame, the inter-prediction status of each inter-predictively coded video block of the frame, and other information for decoding video blocks in the current video frame.

[0049] Similarly, intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block was predicted using intra BC mode, construction information regarding which video blocks of the frame are within the reconstruction region and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.

[0050] Motion compensation unit 82 may also perform interpolation using an interpolation filter used by video encoder 20 during encoding of the video block to calculate sub-integer pixel interpolated values of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter used by video encoder 20 from the received syntax element and use that interpolation filter to create the predictive block.

[0051] Inverse quantization unit 86 uses the same quantization parameter calculated by video encoder 20 for each video block in a video frame to inverse quantize the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80 to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct residual blocks in the pixel domain.

[0052] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs a decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. The decoded video block is sometimes referred to as a reconstructed block of the current video block. To further process the decoded video block, an in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF filter, may be disposed between adder 90 and the DPB. In some examples, in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by adder 90. The decoded video block in a given frame is then stored in DPB 92, which stores reference frames used for subsequent motion compensation of the next video block. DPB 92 or a memory device separate from DPB 92 may store the decoded video for later presentation on a display device, such as display device 34 of FIG. 1.

[0053] In a typical video coding process (including, for example, a video encoding process and a video decoding process), a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In another example, a frame may be monochromatic and thus include only one two-dimensional array of luma samples.

[0054] As shown in FIG. 4A, video encoder 20 (or, more specifically, partitioning unit 45) generates a coded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs arranged consecutively from left to right and from top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, which may be one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that CTUs in this disclosure are not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to encode the samples in the coding tree blocks. The syntax elements describe how a video sequence may be reconstructed in video decoder 30, including characteristics of different types of units of coded blocks of pixels, as well as inter- or intra-prediction, intra-prediction mode, motion vectors, and other parameters. For a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements used to code samples of the coding tree block. A coding tree block may be an N×N block of samples.

[0055] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of the CTU to divide the CTU into smaller CUs. As depicted in FIG. 4C , 64×64 CTU 400 is first partitioned into four smaller CUs, each with a block size of 32×32. Of the four smaller CUs, CU 410 and CU 420 are each partitioned into four 16×16 CUs by block size. Two 16×16 CUs, CU 430 and CU 440, are each further partitioned into four 8×8 CUs by block size. FIG. 4D depicts a quad tree data structure showing the final result of the partitioning process of CTU 400 depicted in FIG. 4C , where each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32×32 to 8×8. Similar to the CTU depicted in FIG. 4B, each CU may include a CB for luma samples, two corresponding coding blocks for chroma samples of the same size frame, and syntax elements used to encode the samples of the coding block. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures used to encode the samples of the coding block. Note that the quadtree partitioning depicted in FIGS. 4C and 4D is for illustrative purposes only, and a CTU may be divided into CUs based on quadtree / ternary / binary tree partitioning to accommodate various local characteristics. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by a binary tree structure and a ternary tree structure. As shown in FIG. 4E, there are multiple possible partition types for a coding block with width W and height H: quad-partition, vertical 2-partition, horizontal 2-partition, vertical 3-partition, vertically extended 3-partition, horizontally extended 3-partition, and horizontally extended 3-partition.

[0056] In some implementations, video encoder 20 may further partition the coding blocks of a CU into one or more M×N PBs. A PB may include a rectangular (square or non-square) block of samples to which the same prediction, inter- or intra-prediction, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PB. In a monochrome picture or a picture with three separate color planes, a PU may include a single PB and syntax structures used to predict the PB. Video encoder 20 may generate predictive luma blocks, predictive Cb blocks, and predictive Cr blocks for the luma, Cb, and Cr PBs of each PU of a CU.

[0057] Video encoder 20 may generate the predictive blocks of a PU using intra prediction or inter prediction. If video encoder 20 generates the predictive blocks of a PU using intra prediction, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of a frame associated with the PU. If video encoder 20 generates the predictive blocks of a PU using inter prediction, video encoder 20 may generate the predictive blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0058] After video encoder 20 generates the predictive luma block, the predictive Cb block, and the predictive Cr block for one or more PUs of a CU, video encoder 20 may generate the luma residual block of the CU by subtracting the predictive luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate the Cb residual block and the Cr residual block of the CU, such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and such that each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0059] Further, as shown in FIG. 4C , video encoder 20 may use quadtree partitioning to decompose the luma residual block, Cb residual block, and Cr residual block of a CU into one or more luma transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block may include a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and syntax structures used to transform the samples of the transform block.

[0060] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of the TU to generate a Cr coefficient block of the TU.

[0061] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to a process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby achieving further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may apply an entropy coding technique to encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a sequence of bits forming a representation of the encoded frame and associated data, which may be stored on a storage device 32 or transmitted to a destination device 14.

[0062] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks of PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing coding blocks for each CU of the frame, video decoder 30 may reconstruct the frame.

[0063] As mentioned above, video coding achieves video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that intra-block copy (IBC) can be considered as intra-frame prediction or a third mode. Of the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict a current video block from a reference video block.

[0064] However, as video data capture technology is constantly improving and video block sizes become finer to preserve video data details, the amount of data required to represent the motion vectors of the current frame is also increasing significantly. One way to overcome this challenge is to benefit from the fact that a group of adjacent CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between these adjacent CUs. Therefore, the motion information of spatially adjacent CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by investigating their spatial and temporal correlations, which is also called the "motion vector predictor (MVP)" of the current CU.

[0065] Instead of encoding the actual motion vector of the current CU into the video bitstream (e.g., the actual motion vector is determined by motion estimation unit 42 as described above in connection with FIG. 2), the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to create a motion vector difference (MVD) for the current CU. By doing so, the motion vector determined for each CU of a frame by motion estimation unit 42 does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream may be significantly reduced.

[0066] Similar to the process of selecting a predictive block in a reference frame during inter-frame prediction of a codeblock, a set of rules may be adopted by both the video encoder 20 and the video decoder 30 to construct a motion vector candidate list (also called a "merge list") for the current CU using potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then select one element from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from the video encoder 20 to the video decoder 30; the index of the selected motion vector predictor in the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU. Therefore, only the index of the selected motion vector predictor needs to be transmitted from the video encoder 20 to the video decoder 30.

[0067] A brief description of transform coefficient coding in a block-based video coding process (e.g., in the Enhanced Compression Model (ECM)) is provided herein. Specifically, each transform block is first divided into multiple coefficient groups (CGs), each including transform coefficients of a 4x4 sub-block of luma components and a 2x2 sub-block of chroma components. The coding of transform coefficients within a transform block is performed on a coefficient group-by-coefficient group basis. For example, the coefficient groups within a transform block are scanned and coded based on a first predetermined scanning order. When coding each coefficient group, the transform coefficients of the coefficient group are scanned based on a second predetermined scanning order within each sub-block. In the ECM, the same top-left scanning order is applied to scan the coefficient groups within a transform block and the different transform coefficients within each coefficient group (e.g., both the first predetermined scanning order and the second predetermined scanning order are top-left scanning orders). Figure 5 is a graphical representation illustrating a top-left scanning order of transform coefficients within a coefficient group according to some examples. The numbers 0 through 15 in FIG. 5 indicate the corresponding scan order of each transform coefficient within a coefficient group.

[0068] According to the transform coefficient coding scheme in ECM, first, for each transform block, a flag is signaled to indicate whether the transform block contains a non-zero transform coefficient. If there are at least non-zero transform coefficients in the transform block, the location of the last non-zero transform coefficient scanned according to the top-left scan order is explicitly signaled from the video encoder 20 to the video decoder 30. Once the location of the last non-zero transform coefficient is signaled, flags are further signaled for all coefficient groups coded before the last coefficient group (i.e., the coefficient group containing the last non-zero coefficient). Similarly, the number of the flag indicates whether each coefficient group contains a non-zero transform coefficient. If the flag for a coefficient group is equal to zero (indicating that all transform coefficients in the coefficient group are zero), no further information needs to be sent regarding that coefficient group. Otherwise (e.g., if the flag for a coefficient group is equal to one), the absolute value and the sign of each transform coefficient in the coefficient group are signaled in the bitstream according to the scan order. However, in existing designs, the signs of the transform coefficients are bypass coded (e.g., the context model is not applied), leading to inefficient transform coding in the current design. According to this disclosure, an improved LFNST process involving sign prediction of transform coefficients is described in more detail below, such that transform coding efficiency may be improved.

[0069] FIG. 6 is a graphical representation illustrating an LFNST process, according to some examples. In VVC, a secondary transform tool (e.g., LFNST) is applied after the primary transform to compress the energy of transform coefficients of intra-coded blocks. As shown in FIG. 6, in video encoder 20, forward LFNST 604 is applied between forward primary transform 603 and quantization 605, and in video decoder 30, inverse LFNST 608 is applied between inverse quantization 607 and inverse primary transform 609. For example, an LFNST process may include both forward LFNST 604 and inverse LFNST 608. As some examples, for a 4×4 forward LFNST 604, there may be 16 input coefficients; for an 8×8 forward LFNST 604, there may be 64 input coefficients; for a 4×4 inverse LFNST 608, there may be 8 input coefficients; and for an 8×8 inverse LFNST 608, there may be 16 input coefficients.

[0070] In the forward LFNST 604, a non-separable transform of variable transform size is applied based on the size of the coding block, which may be expressed using a matrix multiplication process. For example, assume that the forward LFNST 604 is applied to a 4x4 block. The samples in the 4x4 block may be expressed using a matrix X as shown in the following equation (1):

number

[0071] TIFF0007719949000002.tif8170

number

[0072] In the above equation (1) or (2), X refers to the coefficient matrix obtained through the forward linear transform 603, and X ij refers to the linear transform coefficients in matrix X. Then, the forward LFNST 604 is applied according to equation (3) as follows:

number

[0073] TIFF0007719949000005.tif48170

[0074] In some implementation forms, in the LFNST process, a reduced non-separable transformation kernel can be applied. For example, based on the above formula (3), the forward LFNST 604 is based on direct matrix multiplication, which is costly in terms of the memory resources for storing calculation operations and transformation coefficients. Therefore, by using a reduced non-separable transformation kernel in the LFNST design and mapping an N-dimensional vector to an R-dimensional vector in another space when R < N, the implementation cost of the LFNST process can be reduced. For example, instead of using an N×N matrix for the transformation kernel, an R×N matrix as shown in formula (4) is used as the transformation kernel of the forward LFNST 604.

Number

[0075] In the above formula (4), T R×N 's R basis vectors are generated by selecting the first R bases of the original N-dimensional transformation kernel (i.e., N×N). Further, assuming that T R×N is orthogonal, the inverse transformation matrix of the inverse LFNST 608 is the transpose of the forward transformation matrix T R×N .

[0076] For 8×8 LFNST, when a factor N / R=4 is applied, the forward LFNST 604 reduces a 64×64 transform matrix to a 16×48 transform matrix, and the inverse LFNST 608 reduces a 64×64 inverse transform matrix to a 48×16 inverse transform matrix. This is achieved by applying the LFNST process to 8×8 sub-blocks within the top-left region of the primary transform coefficients. Specifically, when a 16×48 forward LFNST is applied, 48 transform coefficients are taken as input from three 4×4 sub-blocks within the top-left 8×8 sub-block (excluding the bottom-right 4×4 sub-block). In some examples, the LFNST process is restricted to be applicable only when all transform coefficients outside the top-left 4×4 sub-block are zero, indicating that all first-order-only transform coefficients must be zero when LFNST is applied. Furthermore, to control worst-case complexity (in terms of per-pixel multiplication), the LFNST matrices for 4x4 and 8x8 coded blocks are forced to be 8x16 and 8x48 transforms, respectively. For 4xM and Mx4 coded blocks (M>4), the non-separable transform matrix of the LFNST is 16x16.

[0077] In LFNST transform signaling, there are a total of four transform sets, and two non-separable transform kernels are enabled for each transform set in the LFNST design. A transform set is selected from the four transform sets according to the intra prediction mode of the intra block. The mapping from intra prediction mode to transform sets is predetermined as shown in Table 1 below. For the current block (81<=predModeIntra<=83), if one of the three Cross-Component Linear Model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used, transform set "0" is selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate is indicated by signaling an LFNST index in the bitstream.

[0078] [Table 1]

[0079] In some examples, because LFNST is restricted to be applied to an intra block when all transform coefficients outside the first 16x16 sub-block are zero, LFNST index signaling depends on the position of the last valid (i.e., non-zero) transform coefficient. For example, for 4x4 and 8x8 coded blocks, the LFNST index is signaled only if the position of the last valid transform coefficient is less than 8. For other coded block sizes, the LFNST index is signaled only if the position of the last valid transform coefficient is less than 16. Otherwise (i.e., when the LFNST index is not signaled), it is inferred that the LFNST index is zero, i.e., LFNST is disabled.

[0080] Furthermore, to reduce the size of the buffer for caching transform coefficients, LFNST is not allowed if the width or height of the current coding block is larger than the maximum transform size (i.e., 64) signaled in the sequence parameter set (SPS). On the other hand, LFNST is only applied when the primary transform is DCT2. Furthermore, LFNST is applied to intra-coded blocks in both intra-slices and inter-slices, and for both luma and chroma components. When a dual-tree or local-tree is enabled (i.e., when the partitions of the luma and chroma components are not aligned), LFNST indices are signaled separately for the luma and chroma components (i.e., different LFNST transforms can be applied to the luma and chroma components). Otherwise, when a single-tree is applied (when the partitions of the luma and chroma components are aligned), LFSNT is applied only to the luma component, with a single LFNST index being signaled.

[0081] The LFNST design in ECM is similar to that in VVC, except that an additional LFNST kernel is introduced to improve energy compaction of residual samples for large block sizes. Specifically, when the width or height of a transform block is 16 or greater, a new LFNST transform is introduced in the upper-left region of the low-frequency transform coefficients generated from the primary transform. In the current ECM, as shown in FIG. 7, the low-frequency region includes six 4×4 sub-blocks (e.g., the six 4×4 sub-blocks shown in gray in FIG. 7) in the upper-left corner of the primary transform coefficients. In this case, the number of coefficient inputs to the forward LFNST 604 is 96. Furthermore, to control worst-case computational complexity, the number of coefficient outputs of the forward LFNST 604 is set to 32. Specifically, for a W×H transform block with W>=16 and H>=16, a 32×96 forward LFNST is applied, which takes 96 transform coefficients from the six 4×4 sub-blocks in the upper-left region as input and outputs 32 transform coefficients. On the other hand, the 8x8 LFNST in ECM utilizes the transform coefficients of all four 4x4 sub-blocks as input and outputs 32 transform coefficients (i.e., a 32x64 matrix for the forward LFNST 604 and a 64x32 matrix for the inverse LFNST 608). This differs from VVC, in which the 8x8 LFNST is applied only to the three 4x4 sub-blocks in the upper-left region to generate only 16 transform coefficients (i.e., a 16x48 matrix for the forward LFNST 604 and a 48x16 matrix for the inverse LFNST 608). Furthermore, the total number of LFNST sets increases from four in VVC to 35 in ECM. As with VVC, the selection of an LFNST set depends on the intra-prediction mode of the current coding unit, and each LFNST set contains three different transform kernels.

[0082] In some examples, in addition to the DCT2 transform used in HEVC, a multiple transform selection (MTS) scheme is applied to transform the residuals of both inter-coded and intra-coded blocks. The MTS scheme uses multiple transforms selected from the DCT8 and DST7 transforms.

[0083] For example, two control flags are specified at the sequence level to enable the MTS scheme for intra mode and inter mode separately. When the MTS scheme is enabled at the sequence level, another CU-level flag is further signaled to indicate whether the MTS scheme is applied or not. In some implementations, the MTS scheme is only applied to the luma component. Furthermore, the MTS scheme is signaled only if the following conditions are met: (a) both the width and height are less than or equal to 32, and (b) the coded block flag (CBF) is equal to 1. If the CU flag in the MTS is equal to zero, the DCT2 is applied in both the horizontal and vertical directions. If the CU flag in the MTS is equal to 1, two other flags are additionally signaled to indicate the horizontal and vertical transform types, respectively. Table 2 below shows the mapping between the horizontal and vertical control flags in the MTS and the applied transforms.

[0084] [Table 2]

[0085] Regarding the precision of the transform matrix, all MTS transform coefficients have 6-bit precision, the same as the DCT2 core transform. Given that VVC supports all transform sizes used in HEVC, all transform cores used in HEVC, including 4-point, 8-point, 16-point, and 32-point DCT-2 transforms and 4-point DST-7 transforms, remain the same as VVC. Meanwhile, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, and 32-point DST-7 and DCT-8, are also supported in the VVC transform design. Furthermore, to reduce the complexity of large-sized DST-7 and DCT-8, when either the width or height is equal to 32, the high-frequency transform coefficients located outside the 16x16 low-frequency region in the DST-7 and DCT-8 transform blocks are set to zero (also known as zeroing out).

[0086] In VVC, only DST7 and DCT8 transform kernels are used, in addition to DCT2, for intra- and inter-coding. For intra-coding, the statistical properties of the residual signal usually depend on the intra-prediction mode. Additional linear transforms may be useful to handle the diversity of residual characteristics.

[0087] Additional primary transforms, including DCT5, DST4, DST1, and the identity transform (IDT), are used in ECM. An MTS set is also created depending on the TU size and intra-mode information. Sixteen different TU sizes are considered, and for each TU size, five different classes can be considered depending on the intra-mode information. Four different transform pairs are considered per class (same as in VVC). A total of 80 different classes can be considered, but in many cases, several of these different classes share the same transform set. Therefore, the resulting lookup table (LUT) has 58 (less than 80) unique entries.

[0088] For angular modes, joint symmetry across TU shapes and intra prediction is considered. Thus, mode i (i>34) with TU shape A×B may be mapped to the same class corresponding to mode j=(68-i) with TU shape B×A. However, for each transform pair, the order of the horizontal and vertical transform kernels is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same class with the vertical and horizontal transform kernels swapped. For wide-angle modes, the closest conventional angular mode is used to determine the transform set. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80.

[0089] Intra-template matching prediction is an example of an intra-prediction mode that copies a prediction block from the reconstructed portion of the current frame, and the L-shaped template of the prediction block matches the current template. For a given search range, video encoder 20 searches for a template most similar to the current template in the reconstructed portion of the current frame (e.g., based on SAD cost) and uses the corresponding block as the prediction block. Then, video encoder 20 signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in a given search area including (a) the current CTU (R1), (b) the upper-left CTU (R2), (c) the upper CTU (R3), and (d) the left CTU (R4), as shown in FIG. 8. Intra-template matching is valid for CUs with width and height sizes of 64 or less. Meanwhile, the intra-template matching prediction mode is indicated by signaling a flag at the CU level. When intra-template matching is applied to a coding block whose width or height is between 4 and 16 (inclusive), the primary transform applied to the corresponding dimension is set to DCT-VII. Otherwise (i.e., when the width or height is less than 4 or greater than 16), DCT-II is applied to that dimension.

[0090] 9 is a graphical representation illustrating an exemplary process of code prediction according to some examples. In some implementations, code prediction may be intended to estimate the codes of transform coefficients in a transform block from samples of its neighboring blocks, and to code the difference between each estimated code and the corresponding true code with a "0" (or a "1") to indicate that the estimated code is the same (or not the same) as the true code. If the codes can be accurately estimated at a high rate (e.g., 90% or 95% of the codes are correctly estimated), the difference between the estimated code and the true code tends to be 0, which can be efficiently entropy coded by CABAC when compared with the bypass-coded codes of transform coefficients in VVC.

[0091] Generally, there is a high correlation between samples at the boundary between a current block and its neighboring blocks, and this high correlation can be exploited by a sign prediction scheme to predict the signs of the transform coefficients of the current block. As shown in Figure 9, assume that there are M non-zero transform coefficients in the current block (each of the M signs is either + or -). In this case, the total number of possible combinations of signs is 2 M The code prediction scheme uses each combination of codes to generate a corresponding hypothesis (e.g., reconstructed samples at the top and left boundaries of the current block), compares the reconstructed samples in the corresponding hypothesis with extrapolated samples from adjacent blocks, and obtains the sample difference (e.g., SSD or SAD) between the reconstructed samples and the extrapolated samples. M The code combination that minimizes the sample difference (out of three possible combinations) is selected as the predicted code in the current block.

[0092] In some implementations, the M corresponding transform coefficients may be processed by an inverse quantization operation and an inverse transform to generate a corresponding hypothesis for each combination of the M codes to obtain residual samples, as shown in Figure 9. The residual samples may be summed with the prediction samples to obtain reconstructed samples, including the reconstructed samples at the top and left boundaries of the current block (as shown in the L-shaped gray area 902).

[0093] In some implementations, a cost function is used to select the code combination, which measures the spatial discontinuity between samples at the boundary between the current block and its neighboring blocks. Instead of using the L2 norm (SSD), the cost function can be based on the L1 norm (SAD), as shown in equation (5) below.

number

[0094] In the above formula (5), B i,n (i=-2,-1) represents the neighboring samples of the current block from the upper neighboring block. C m,j (j=-2,-1) represents the neighboring samples of the current block from the left neighboring block. P 0,n and P m,0 , and , respectively denote the corresponding reconstructed samples at the top and left boundaries of the current block. N and M denote the width and height of the current block, respectively. Figure 10 shows the corresponding samples P of the current block used to calculate the cost function for code prediction. 0,n and P m,0 and the corresponding sample B in the adjacent block i,n and C m,j Indicates that.

[0095] In some implementations, to avoid the complexity of performing multiple inverse transforms, a template-based hypothesis reconstruction method may be applied in the code prediction scheme. Each template may be a set of reconstructed samples at the upper and left boundaries of the current block, and may be obtained by applying an inverse transform to a coefficient matrix, with certain coefficients set to 1 and all other coefficients equal to 0. Given that the inverse transform (e.g., DCT, DST) is linear, the corresponding hypothesis may be generated by a linear combination of a set of pre-computed templates.

[0096] In some implementations, the predictive codes are grouped into two sets, each coded by a single CABAC context. For example, the first set contains predictive codes for transform coefficients in the upper left corner of the transform block, and the second set contains predictive codes for transform coefficients in all other positions of the transform block.

[0097] Several exemplary deficiencies present in current designs of sign prediction schemes are identified herein. In a first example, sign prediction in current ECMs is only applicable to predicting signs for transform coefficients in transform blocks to which only a linear transform (e.g., a DCT transform and a DST transform) is applied. As mentioned above, to achieve improved energy compaction of residual samples of intra-coded blocks, LFNST may be applied to transform coefficients from the linear transform. However, in transform blocks to which LFNST is applied in current ECM designs, sign prediction is bypassed.

[0098] In the second example, to control the complexity of the code prediction, a predetermined maximum number of predicted codes ("L") is used for the transform block. max In current ECMs, video encoders determine the maximum number of values (e.g., L) based on a tradeoff between complexity and coding efficiency. max = 8) and send that value to the video decoder. Furthermore, for each transform block, the video encoder or decoder may scan all transform coefficients in raster scan order, and the first L maxnon-zero transform coefficients are selected as candidate transform coefficients for sign prediction. Such uniform treatment of different transform coefficients in a transform block may not be optimal in terms of the accuracy of sign prediction. For example, for transform coefficients with relatively large magnitudes, predicting their signs may be more likely to achieve correct prediction. This is because using incorrect signs for these transform coefficients tends to have a larger impact on reconstructed samples on block boundaries than the impact caused by using transform coefficients with relatively small magnitudes.

[0099] In a third example, a video encoder or decoder can encode the correctness of a predicted code instead of directly encoding an explicit code value. For example, for a transform coefficient with a positive code, if its predicted code is also positive, only a bin "0" needs to be indicated in the bitstream from the video encoder to the video decoder. In this case, the predicted code is the same as the true (or original) code of the transform coefficient, indicating that the code prediction for this transform coefficient is correct. Otherwise (e.g., if the predicted code is negative while the true code is positive), a bin "1" may be included in the bitstream from the video encoder to the video decoder. If all codes are predicted correctly, the corresponding bins indicated in the bitstream are 0, which can be efficiently entropy coded by CABAC. If some of the codes are predicted incorrectly, the corresponding bins indicated in the bitstream are 1. While arithmetic coding and an appropriate context model can be efficient in coding bins according to their corresponding probabilities, there are still significant bits generated in the bitstream to indicate code values.

[0100] In the fourth example, the current design of code prediction in ECM uses the spatial discontinuity between samples at the boundary between the current block and its neighboring block to select the best code prediction combination. To capture the spatial discontinuity, the L1 norm of the gradient difference along the vertical and horizontal directions is utilized. However, since the distribution of image signals is usually non-uniform, using only the vertical and horizontal directions may not accurately capture the spatial discontinuity.

[0101] In accordance with the present disclosure, to address one or more of the above-described exemplary deficiencies, video processing methods and systems for symbol prediction in block-based video coding are provided herein. The methods and systems disclosed herein can improve coding efficiency of symbol prediction while allowing for ease of use in hardware codec implementations. The methods and systems disclosed herein can improve coding efficiency of transform blocks that apply symbol prediction techniques to transform coefficients of the blocks.

[0102] For example, as described above, code prediction may predict the signs of transform coefficients within a transform block based on the correlation between boundary samples (also called boundary samples) located at or near the boundary between the transform block and its spatially adjacent blocks. Given that the existence of correlation does not depend on which specific transform is applied, the two coding tools (i.e., LFNST and code prediction) do not interfere with each other and can be applied jointly. Furthermore, because LFNST further compacts the energy of transform coefficients from a linear transform, the code prediction of LFNST transform coefficients may be more accurate than the code prediction of a linear transform. This is because an incorrect code prediction of a transform coefficient from LFNST may cause even greater discrepancies in the smoothness of boundary samples. Therefore, according to the present disclosure, a harmonization scheme is disclosed herein that enables the combination of LFNST and code prediction to improve the coding efficiency of transform coefficient coding. Furthermore, a template-based hypothesis generation scheme is also disclosed herein that reconstructs boundary samples for different combinations of predicted codes to reduce the number of inverse transforms.

[0103] In another example, instead of giving equal treatment to different transform coefficients in a transform block for selecting transform coefficient candidates for sign prediction as described above, a higher weight may be given to transform coefficients whose signs may lead to inconsistencies between boundary samples of adjacent blocks, assuming that these transform coefficients are more easily predicted. According to the present disclosure, the method and system disclosed herein may select transform coefficient candidates for sign prediction (e.g., transform coefficients whose signs are predicted for a transform block) based on one or more selection criteria to improve the accuracy of sign prediction. For example, to improve the accuracy of sign prediction, transform coefficients that have a greater influence on reconstructed boundary samples (rather than transform coefficients that have a smaller influence on reconstructed boundary samples) are selected as transform coefficient candidates for sign prediction.

[0104] In yet another example, when the signs of transform coefficients in a transform block are predicted with high accuracy (e.g., when the accuracy of the predicted sign is higher than a threshold such as 80% or 90%), there is a strong correlation between boundary samples of the transform block and its neighboring blocks. In this case, a situation commonly arises in which there may be consecutive transform coefficients (e.g., particularly several non-zero transform coefficients at the beginning of the transform block) that can be correctly predicted in most scenarios. In such a scenario, a single bin (instead of multiple bins) can be used to indicate whether the signs of all consecutive transform coefficients are correctly predicted, thereby saving the signaling overhead of code prediction. According to the present disclosure, a vector-based code prediction scheme is disclosed herein to reduce the signaling overhead of code prediction. Unlike existing code predictions that individually predict the sign of each non-zero transform coefficient, the disclosed vector-based code prediction scheme groups a set of consecutive non-zero transform coefficient candidates and predicts their corresponding signs together, so that the average number of bins (or bits) used to indicate the accuracy of the predicted sign can be efficiently reduced.

[0105] In yet another example, using only the vertical and horizontal directions may not accurately capture spatial discontinuities between samples at the boundary between a current block and its neighboring blocks. Therefore, more directions may be introduced to more accurately capture spatial discontinuities. In accordance with the present disclosure, an improved cost function is disclosed herein that considers both vertical and horizontal gradients and diagonal gradients to more accurately capture spatial discontinuities.

[0106] 11 is a block diagram illustrating an example symbol prediction process 1100 in block-based video coding according to some implementations of the present disclosure. In some implementations, the symbol prediction process 1100 may be performed by transform processing unit 52. In some implementations, the symbol prediction process 1100 may be performed by one or more processors (e.g., one or more video processors) of video encoder 20 or decoder 30. Throughout this disclosure, LFNST is used as an example of a secondary transform without loss of generality. It is contemplated herein that other examples of secondary transforms may also be applied.

[0107] In existing designs of ECM, sign prediction is disabled for transform blocks to which LFNST is applied. However, the principle of sign prediction is to predict the sign of a transform coefficient based on the correlation between boundary samples of a transform block and its spatially adjacent blocks, which is independent of the specific transform type (e.g., whether the transform type is a linear transform or a secondary transform) or transform core (e.g., whether the transform core is a DCT or a DST) applied to the transform block. Therefore, sign prediction and LFNST can be jointly applied to further improve the efficiency of the transform coding herein. According to the present disclosure, a sign prediction process 1100 can be applied to predict the sign of a transform coefficient in a transform block, and the linear transform and the secondary transform are jointly applied.

[0108] An exemplary overview of the code prediction process 1100 is provided herein. First, the code prediction process 1100 may perform a coefficient generation operation 1102 by applying a primary transform and a secondary transform to a transform block of a video frame from a video to generate transform coefficients for the transform block. Next, the code prediction process 1100 may perform a coefficient selection operation 1104 by selecting a set of transform coefficient candidates from the transform coefficients of the code prediction. Subsequently, the code prediction process 1100 may perform a hypothesis generation operation 1106 by applying a template-based hypothesis generation scheme to select a hypothesis from multiple candidate hypotheses for the set of transform coefficient candidates. Furthermore, the code prediction process 1100 may perform a code generation operation 1108 by determining a combination of code candidates associated with the selected hypothesis as a set of predicted codes for the set of transform coefficient candidates. Operations 1102, 1104, 1106, and 1108 are each described in more detail below.

[0109] For example, transform processing unit 52 of video encoder 20 may transform the residual video data into transform coefficients of a transform block by jointly applying a primary transform and a secondary transform (e.g., as shown in FIG. 6 in which forward primary transform 603 and forward LFNST 604 are applied together). A predetermined number (e.g., L) of non-zero transform coefficients may be selected as candidate transform coefficients from the transform coefficients of the transform block based on one or more selection criteria described below, where 1≦L≦the maximum number of possible codes. Then, by applying a template-based hypothesis generation scheme, multiple candidate hypotheses may be generated using different combinations of candidate codes for each of the L candidate transform coefficients, resulting in a total of 2 LThis results in L candidate hypotheses. Each candidate hypothesis may include reconstructed samples at the top and left boundaries of the transform block. A cost for each candidate hypothesis reconstruction may then be calculated using a cost function incorporating combined gradients along the horizontal, vertical, and diagonal directions. A candidate hypothesis associated with the smallest cost may be determined from the multiple candidate hypotheses as a hypothesis for predicting the signs of the L candidate transform coefficients. For example, the combination of code candidates used to generate the candidate hypothesis associated with the smallest cost is used as the predicted code for the L candidate transform coefficients.

[0110] First, the symbol prediction process 1100 may perform a coefficient generation operation 1102, in which a primary transform (e.g., DCT, DST, etc.) and a secondary transform (e.g., LFNST) may be jointly applied to a transform block to generate transform coefficients of the transform block. For example, a primary transform may be applied to the transform block to generate primary transform coefficients of the transform block. Then, an LFNST may be applied to the transform block to generate LFNST transform coefficients based on the primary transform coefficients.

[0111] The symbol prediction process 1100 may then perform a coefficient selection operation 1104, in which a set of candidate transform coefficients for symbol prediction may be selected from the transform coefficients of the transform block based on one or more selection criteria. The selection of candidate transform coefficients may maximize the number of candidate transform coefficients that can be correctly predicted, thereby improving the accuracy of symbol prediction.

[0112] In some implementations, the set of transform coefficient candidates may be selected from the transform coefficients of the transform block based on the magnitudes of the transform coefficients. For example, the set of transform coefficient candidates may include one or more transform coefficients having a magnitude greater than the remaining transform coefficients in the transform block.

[0113] Generally, for transform coefficients with larger magnitudes, the predicted codes of these transform coefficients are more likely to be correct. This is because these transform coefficients with larger magnitudes tend to have a greater impact on the quality of reconstructed samples, and using an incorrect code for these transform coefficients may increase the likelihood of discontinuities between boundary samples between a transform block and its spatially adjacent blocks. Based on this rationale, a set of transform coefficient candidates for code prediction may be selected from the transform coefficients of a transform block based on the magnitudes of the non-zero transform coefficients within the transform block. For example, all non-zero transform coefficients within a transform block may be scanned and sorted to form a coefficient list in descending order of magnitude. The transform coefficient with the largest magnitude may be selected from the coefficient list and placed as the first transform coefficient candidate in the set of transform coefficient candidates, the transform coefficient with the second largest magnitude may be selected from the coefficient list and placed as the second transform coefficient candidate in the set of transform coefficient candidates, and so on until the number of selected transform coefficient candidates reaches a predetermined number L. In some implementations, when selecting a set of transform coefficient candidates, the quantization index of the transform coefficients may be used to represent the magnitudes of the transform coefficients.

[0114] In some implementations, a set of transform coefficient candidates may be selected from the transform coefficients of a transform block based on a coefficient scanning order for entropy coding applied to video coding. Because raw video content may contain abundant low-frequency information, the magnitudes of non-zero transform coefficients resulting from processing the video content tend to be large at low-frequency positions and small toward high-frequency positions. Therefore, a coefficient scanning order (such as a zigzag scan, a top-left scan, a horizontal scan, or a vertical scan) may be used in modern video codecs to scan the transform coefficients in a transform block for entropy coding. By using this coefficient scanning order, transform coefficients with larger magnitudes (usually corresponding to lower frequencies) are scanned before transform coefficients with smaller magnitudes (usually corresponding to higher frequencies). Based on this rationale, a set of transform coefficient candidates for code prediction disclosed herein may be selected from the transform coefficients of a transform block based on a coefficient scanning order for entropy coding. For example, a coefficient list may be obtained by scanning all transform coefficients in the transform block using the coefficient scanning order. Then, the first L non-zero transform coefficients in the coefficient list may be automatically selected as a set of transform coefficient candidates for code prediction.

[0115] In some implementations, for an intra-coded block, a set of candidate transform coefficients for code prediction may be selected from the transform coefficients of the block based on the intra-prediction direction of the block. For example, both the video encoder 20 and the video decoder 30 may determine and store multiple scan orders consistent with the intra-prediction direction (e.g., 67 intra-prediction directions in VVC and ECM) as a look-up table. When encoding the transform coefficients of the intra block, the video encoder 20 or the video decoder 30 may identify a scan order from among the scan orders that is closest to the intra-prediction of the intra block. The video encoder 20 or the video decoder 30 may use the identified scan order to scan all non-zero transform coefficients of the intra block to obtain a coefficient list and select the first L non-zero transform coefficients from the coefficient list as a set of candidate transform coefficients.

[0116] In some implementations, video encoder 20 may determine a scanning order for the transform coefficients of a transform block and signal the determined scanning order to video decoder 30. One or more new syntax elements indicating the determined scanning order may be signaled through the bitstream. For example, multiple fixed scanning orders (e.g., for different transform block sizes and coding modes) may be predetermined by video encoder 20 and pre-shared with video decoder 30. Then, after selecting a scanning order from the fixed scanning orders, video encoder 20 need only signal a single index indicating the selected scanning order to video decoder 30. In another example, one or more new syntax elements may be used to enable signaling of any selected scanning order of the transform coefficients. In some implementations, the one or more syntax elements may be signaled at various coding levels, e.g., sequence parameter set (SPS), picture parameter set (PPS), picture (or slice) level, CTU (or CU) level, etc.

[0117] In some implementations, a set of transform coefficient candidates may be selected from the transform coefficients of a transform block based on their influence scores on the reconstructed boundary samples of the transform block. Specifically, as shown in equation (5) above, the selection of a combination of signs (i.e., predicted codes, or code predictors) is based on a cost function for minimizing discontinuities in the gradients of samples between the current transform block and its spatially adjacent blocks. Therefore, the signs of transform coefficients that have a relatively large influence on the reconstructed samples at the top and left boundaries of the current transform block tend to be more likely to be accurately predicted, because a reversal of these signs may cause large variations in the smoothness between the boundary samples calculated in equation (5). To maximize the rate of accurate sign prediction, the signs of these transform coefficients (i.e., transform coefficients that have a larger influence on the reconstructed boundary samples) may be predicted before other transform coefficients (i.e., transform coefficients that have a smaller influence on the reconstructed boundary samples). Based on this rationale, a set of transform coefficient candidates for sign prediction disclosed herein may be selected based on their influence scores on the reconstructed samples at the top and left boundaries of the current transform block.

[0118] For example, video encoder 20 or decoder 30 may sort all transform coefficients based on a measurement of their corresponding influence scores on the reconstructed boundary samples of the transform block. If a transform coefficient has a larger influence score on the reconstructed boundary samples, the transform coefficient may be assigned a smaller index in the code prediction candidate list because it is more easily predicted accurately. The set of transform coefficient candidates disclosed herein may be the L transform coefficients with the L smallest indices in the code prediction candidate list.

[0119] In some implementations, different criteria may be applied to quantify the influence score of a transform coefficient on a reconstructed boundary sample. For example, a value measuring the energy of the variation of the reconstructed boundary sample caused by the transform coefficient may be used as the influence score, which may be obtained (in the L1 norm) as follows:

number

[0120] In the above formula (6), C i,j represents the transform coefficient at position (i,j) in the transform block. i,j (l,k) is the conversion coefficient C i,j where N and M represent the width and height of the transform block, respectively. V represents the influence score of the transform coefficient at position (i,j).

[0121] In another example, the L1 norm in equation (6) above can be replaced with an L2 norm, so that the influence score (e.g., a measure of the energy of the variation of the reconstructed boundary samples caused by the transform coefficients) can be calculated using the L2 norm as follows:

number

[0122] According to the present disclosure, in the above (6) and (7), (e.g., T i,j (0,n) and T i,j Although only the top and left boundary samples (denoted by (m,0)) are used in the calculation, the transform coefficient selection scheme disclosed in this specification can also be applied to any code prediction scheme by changing the reconstructed samples of the current transform block used in the corresponding cost function.

[0123] The code prediction process 1100 may then perform a hypothesis generation operation 1106, in which a template-based hypothesis generation scheme may be applied to select a hypothesis for the set of transform coefficient candidates from a plurality of candidate hypotheses. First, multiple combinations of code candidates may be determined for the set of transform coefficient candidates based on the total number of coefficients included in the set of transform coefficient candidates. For example, if there are a total of L transform coefficient candidates, the multiple combinations of code candidates for the set of transform coefficient candidates may be determined ... L transform coefficient candidates. L There can be multiple combinations of candidate codes. Each candidate code can be either a negative sign (-) or a positive sign (+). Each combination of candidate codes may contain a total of L negative or positive signs. For example, if L=2, then multiple combinations of candidate codes can be 2 of the candidate codes. 2 = 4 combinations, which are (+,+), (+,-), (-,-), and (-,-), respectively.

[0124] A template-based hypothesis generation scheme may then be applied to generate multiple candidate hypotheses for multiple combinations of code candidates, respectively. To reduce the complexity of the inverse primary and secondary transforms that need to be performed, the template-based hypothesis generation scheme disclosed herein can be used to optimize the generation of reconstructed boundary samples of the transform block. Two exemplary approaches for implementing the template-based hypothesis generation scheme are disclosed herein. It is contemplated that other exemplary approaches for implementing the template-based hypothesis generation scheme are also possible, and this is not limited herein.

[0125] In a first exemplary approach, a corresponding candidate hypothesis for each combination of code candidates may be generated based on a linear combination of templates, resulting in multiple candidate hypotheses for multiple combinations of code candidates. Each template may correspond to a candidate transform coefficient from a set of candidate transform coefficients. Each template may represent a group of reconstructed samples at the top and left boundaries of a transform block. Each template may be generated by applying an inverse secondary transform and an inverse linear transform to the transform block, and each of the set of candidate transform coefficients is set to zero except for the candidate transform coefficient corresponding to the template, which is set to one (e.g., the candidate transform coefficient corresponding to the template is set to one, and the remaining candidate transform coefficients are each set to zero).

[0126] For example, the corresponding candidate hypotheses for each combination of code candidates may be set to be a linear combination of templates. For the templates corresponding to each candidate transform coefficient, the weights of each template may be set to be the magnitudes of the dequantized transform coefficients corresponding to each candidate transform coefficient. An example of hypothesis generation based on a linear combination of templates is shown in FIG. 12, which will be described in more detail below.

[0127] To predict the signs of candidate transform coefficients, video encoder 20 or decoder 30 may examine all candidate hypotheses before identifying a hypothesis associated with a combination of candidate codes that can minimize a cost value calculated from a cost function. In the first exemplary approach described above, each candidate hypothesis may be generated based on a combination of multiple templates, which is relatively complex considering the computations (e.g., additions, multiplications, and shifts) for each sample included in such a combination. To reduce the computational complexity associated with identifying a hypothesis that minimizes a cost value calculated from a cost function, a second exemplary approach is introduced herein.

[0128] In a second exemplary approach, multiple combinations of code candidates associated with multiple candidate hypotheses may be treated as multiple hypothesis indices for the multiple candidate hypotheses. For example, digital 0 and digital 1 may be configured to represent a positive sign (+) and a negative sign (-), respectively. A combination of code candidates corresponding to a candidate hypothesis may be used as a unique representation (i.e., a hypothesis index) for the candidate hypothesis. For example, assume there are three predicted codes (e.g., L=3). Hypothesis index 000 may represent a candidate hypothesis generated by setting all three code candidates to be positive (e.g., the three code candidates are (+, +, +)). Similarly, hypothesis index 010 may represent a candidate hypothesis generated by setting the first and third code candidates to be positive and the second code candidate to be negative (e.g., the three code candidates are (+, -, +)).

[0129] Then, multiple hypothetical indexes Gray Based on the code order, multiple candidate hypotheses may be generated, such that a reconstructed sample of a previous candidate hypothesis having a previous hypothesis index may be used to generate a current candidate hypothesis having a current hypothesis index. The current hypothesis index of the current candidate hypothesis may be a value of one of the multiple hypothesis indexes. Gray In the code order, the current hypothesis index may immediately follow the previous hypothesis index of the previous hypothesis candidate. The current hypothesis index may be generated by changing the code candidate associated with the previous hypothesis index from positive (or negative) to negative (or positive). For example, the current hypothesis index may be obtained by changing a single "0" (or "1") in the previous hypothesis index to "1" (or "0").

[0130] For example, multiple hypothesis indexes can be GrayThe hypothesis indexes may be reordered based on the code order to generate a reordered sequence of hypothesis indexes. For a first hypothesis index in the reordered sequence of hypothesis indexes, a first candidate hypothesis corresponding to the first hypothesis index may be generated by applying an inverse secondary transform and an inverse linear transform to the transform block, with each of the set of candidate transform coefficients set to 1. For a second hypothesis index in the reordered sequence of hypothesis indexes that immediately follows the first hypothesis index, a second candidate hypothesis corresponding to the second hypothesis index may be generated based on (a) the first candidate hypothesis corresponding to the first hypothesis index and (b) an adjusting term for the second candidate hypothesis. Table 3 below shows an example process for generating all candidate hypotheses for an LFNST when the number of candidate transform coefficients is three (e.g., L=3).

[0131] [Table 3]

[0132] In the above Table 3, the first column is 3 The second column shows the hypothesis indexes corresponding to the combinations of the code candidates, using digital 0 and 1 to represent the positive sign (+) and negative sign (-), respectively. The hypothesis indexes in the second column are Gray The codes are ordered (e.g., 000, 001, 011, 010, 110, 111, 101, 100). The third column shows the candidate hypotheses corresponding to the combinations of the candidate codes and the hypothesis indexes. The fourth column shows the calculation of the candidate hypotheses.

[0133] In Table 3, TXYZ in the fourth column represent corresponding templates (i.e., reconstructed samples at the top and left boundaries of a transform block), which may be generated by applying an inverse transform to a coefficient matrix of the transform block, with certain transform coefficients set to 1 and all other transform coefficients equal to zero. For example, T100 represents a corresponding template generated by applying an inverse transform to a coefficient matrix, with only the transform coefficient corresponding to the first code candidate set to 1 and all transform coefficients in the coefficient matrix set to 0. C0, C1, and C2 represent the absolute values of the dequantized transform coefficients associated with the first, second, and third code candidates, respectively.

[0134] Referring to Table 3, for a first hypothesis index 000, a first candidate hypothesis H000 may be generated by applying an inverse secondary transform and an inverse linear transform to a coefficient matrix associated with the transform block, with each of the candidate transform coefficients set to 1. For a second hypothesis index 001, which immediately follows the first hypothesis index 000, the second candidate hypothesis H001 may be generated based on (a) the first candidate hypothesis H000 and (b) an adjustment term for the second candidate hypothesis (e.g., −C2*T001). Similarly, for a third hypothesis index 011, which immediately follows the second hypothesis index 001, the third candidate hypothesis H011 may be generated based on (a) the second candidate hypothesis H001 and (b) an adjustment term for the third candidate hypothesis (e.g., −C1*T010). For a fourth hypothesis index 010 that immediately follows the third hypothesis index 011, a fourth candidate hypothesis H010 may be generated based on (a) the third candidate hypothesis H011 and (b) an adjustment term for the fourth candidate hypothesis (e.g., C2*T001).

[0135] Subsequently, a hypothesis associated with the smallest cost may be determined from multiple candidate hypotheses based on a cost function incorporating combined gradients along the horizontal, vertical, and diagonal directions. As explained above, if the cost function utilizes only horizontal and vertical gradients (e.g., as shown above in Equation (5)), the cost function may not perform well for image signals with high heterogeneity. According to the present disclosure, gradients along one or more diagonal directions are also utilized to improve the accuracy of the cost function. For example, two diagonal directions, including a left diagonal direction and a right diagonal direction, may also be incorporated into the cost function. For example, the cost functions for the two diagonal directions may be written according to the following Equations (8) and (9):

number

number

[0136] In the above formula (8) or formula (9), B -1,n-1 , B -2,n-2 , B -1,n+1 , and B -2,n+2 represents the neighboring samples of the transformed block from the upper neighboring block. m-1,-1 , C m-2,-2 , C m+1,-1 , and C m+2,-2 represents the neighboring samples from the left neighboring block of a transform block. 0,n and P m,0 represent the reconstructed samples on the top and left boundaries of the transform block, respectively. N and M represent the width and height of the transform block, respectively. costD1 and costD1 represent the left and right diagonal cost functions for the left and right diagonal directions, respectively.

[0137] The two diagonal cost functions (e.g., costD1 and costD2) may be used in conjunction with a horizontal-vertical cost function (e.g., costHV shown in Equation (5) above). A cost function for sign prediction may then be determined based on the horizontal-vertical cost function incorporating gradients along the horizontal and vertical directions, the left diagonal cost function incorporating gradients along the left diagonal direction, and the right diagonal cost function incorporating gradients along the right diagonal direction. For example, the cost function may be a weighted sum of the horizontal-vertical cost function, the left diagonal cost function, and the right diagonal cost function, as described in Equation (10).

number

[0138] In the above equation (10), ω refers to the weights of the left and right diagonal cost functions.

[0139] In another example, the cost function may be the minimum of the horizontal-vertical cost function, the left diagonal cost function, and the right diagonal cost function, as described in equation (11).

number

[0140] Compared to equation (5) shown above, the cost functions of equations (10) or (11) disclosed herein may require more neighboring pixels to support the cost functions costD1, costD2 along the diagonal direction, which will be explained in more detail below with reference to Figures 14A-14B.

[0141] In some implementations, a cost corresponding to each candidate hypothesis may be determined using equation (10) or equation (11) above. Multiple costs may then be calculated for the multiple candidate hypotheses, respectively. A minimum cost may be determined from the multiple costs. The candidate hypothesis associated with the minimum cost may be determined from the multiple candidate hypotheses and selected as the hypothesis for code prediction.

[0142] The code prediction process 1100 may then perform a code generation operation 1108, in which the combination of code candidates associated with the selected hypothesis is determined to be a set of predictive codes for the set of candidate transform coefficients. For example, the combination of code candidates (e.g., L code candidates) used to generate the selected hypothesis may be used as the predictive codes for the L candidate transform coefficients.

[0143] In some implementations, the code generation operation 1108 may include applying a vector-based code prediction scheme to the set of predictive codes to generate a sequence of code signaling bits for the set of candidate transform coefficients. A bitstream including the sequence of code signaling bits may be generated by the video encoder 20 and stored in the storage device 32 of FIG. 1. Alternatively or additionally, the bitstream may be transmitted to the video decoder 30 through the link 16 of FIG. 1.

[0144] As explained above, if the sign of a transform coefficient in a transform block is predicted properly, it is highly likely that the signs of multiple consecutive transform coefficients can be correctly predicted. In this case, the signaling scheme of existing code prediction designs is obviously inefficient in terms of overhead for signaling the code value of a transform block, since it needs to signal multiple bins "0" to individually indicate that the corresponding sign of each transform coefficient can be correctly predicted. An exemplary implementation of an existing code prediction scheme is described in more detail below with reference to Figure 13A.

[0145] According to the present disclosure, the efficiency of code signaling may be improved by applying the vector-based code prediction scheme disclosed herein. Specifically, transform coefficient candidates of a transform block may be divided into multiple groups, and the codes of the transform coefficient candidates within each group may be jointly predicted. In this case, if the original codes (or true codes) of the transform coefficient candidates within a group are the same as their respective predicted codes, a bin with a value of “0” may be sent in the bitstream to indicate that all codes in the group are correctly predicted. Otherwise (i.e., if at least one transform coefficient candidate within the group has an original code different from the predicted code), a bin with a value of “1” may be signaled in the bitstream first to indicate that not all codes of the transform coefficient candidates within the group are correctly predicted. Then, additional bins may also be signaled in the bitstream from the video encoder 20 to the video decoder 30 to individually indicate the corresponding correctness of each predicted code within the group. An exemplary implementation of the vector-based code prediction scheme disclosed herein is described in more detail below with reference to FIG. 13B.

[0146] In some implementations, a set of transform coefficient candidates may be divided into multiple groups of transform coefficient candidates. For each group of transform coefficient candidates, one or more code signaling bits may be generated for the group of transform coefficient candidates based on whether the original code of the group of transform coefficient candidates is identical to the predicted code of the group of transform coefficient candidates. In response to the original code of the group of transform coefficient candidates being identical to the predicted code of the group of transform coefficient candidates, a bin having a zero value (“0”) may be generated and added to the bitstream as a code signaling bit. For example, the bitstream may include “0” indicating that the predicted code of the group of transform coefficient candidates is correctly predicted.

[0147] On the other hand, a bin having a value of one ("1") may be generated in response to the original code of the group of transform coefficient candidates not being identical to the predicted code of the group of transform coefficient candidates. A set of additional bins may also be generated to signal the corresponding correctness of the predicted code of the group of transform coefficient candidates. The bin having a value of one and the set of additional bins may then be added to the bitstream as code signaling bits. For example, the set of additional bins may be an XOR result of the original code and the predicted code of the group of transform coefficient candidates. An additional bin having a value of "0" may indicate that the predicted code of the transform coefficient candidate corresponding to the additional bin is correctly predicted, while an additional bin having a value of "1" may indicate that the predicted code of the transform coefficient candidate corresponding to the additional bin is incorrectly predicted. The bitstream may include (a) a "1" to indicate that the predicted code of the group of transform coefficient candidates is incorrectly predicted, and (b) a set of additional bins to indicate which predicted codes are correctly predicted and which predicted codes are incorrectly predicted.

[0148] In some implementations, the size of each group of candidate transform coefficients may be adaptively changed based on one or more predetermined criteria. The one or more predetermined criteria may include the width or height of the transform block, the coding mode of the transform block (e.g., intra-coding or inter-coding), the number of non-zero transform coefficients in the transform block, etc. In some implementations, the size of each group of candidate transform coefficients may be signaled in the bitstream at various coding levels, such as the SPS, PPS, slice or picture level, the CTU or CU level, or the transform block level.

[0149] In some implementations, one or more constraints may be applied to limit the application scenario of the vector-based sign prediction scheme disclosed herein. For example, the vector-based sign prediction scheme disclosed herein may be applied to process the signs of a first portion of the transform coefficients in a transform block, while the signs of a second portion of the transform coefficients in the transform block may be processed using an existing sign prediction scheme. In a further example, the vector-based sign prediction scheme disclosed herein may be applicable to the first N (e.g., N=2, 3, 4, 5, 6, 7, or 8, etc.) non-zero transform coefficient candidates from the transform block, while the signs of other transform coefficient candidates from the transform block may be processed using the existing sign prediction scheme shown in Figure 13A, which will be described later in this disclosure.

[0150] According to the present disclosure, the symbol prediction process 1100 disclosed herein may be disabled under some scenarios. For example, when an LFNST is applied to a coding block coded using intra-template matching mode, the primary transform may be DCT-VII. Given that the LFNST core transform in ECM is primarily trained when the primary transform is DCT-II, the corresponding LFNST transform coefficients of an intra-template matching block may exhibit different characteristics compared to the coefficients of other LFNST blocks. Based on this rationale, if the current coding block is an intra-template matching block and coded using LFNST, the symbol prediction process 1100 may be disabled.

[0151] According to the present disclosure, to control the computational complexity of code prediction, the maximum number of predictive codes for LFNST blocks may be different from the maximum number of predictive codes for non-LFNST blocks. For example, the maximum number of predictive codes for LFNST blocks may be set to 6 (or 4), while the maximum number of predictive codes for non-LFNST blocks may have a value other than 6 (or 4). Furthermore, different values for the maximum number of predictive codes may be applied to video blocks that employ LFNST and video blocks that do not employ LFNST. In some implementations, video encoder 20 may determine the maximum number of predictive codes for LFNST blocks based on the encoder's corresponding complexity or performance priority and may signal the maximum number to video decoder 30. When the maximum number of predictive codes for LFNST blocks is signaled to video decoder 30, the maximum number may be signaled, for example, at various coding levels, such as the sequence parameter set (SPS), picture parameter set (PPS), picture or slice level, or CTU or CU level. In some implementations, video encoder 20 may determine different values for the maximum number of predictive codes for video blocks that apply LFNST and video blocks that do not apply LFNST, and signal the values of the maximum number from video encoder 20 to video decoder 30.

[0152] According to this disclosure, assuming that the transform coefficients of both the primary transform and the secondary transform are fixed, the video encoder 20 or the video decoder 30 may pre-calculate templates (e.g., template samples) for different combinations of different transform block sizes and primary and secondary transform combinations. The video encoder 20 or the video decoder 30 may store the templates (e.g., template samples) in an internal or external memory to avoid the complexity of creating template samples on the fly for optimized implementations. The template samples may be stored with different decimal precisions to achieve different trade-offs between storage size and sample precision. For example, the video encoder 20 or the video decoder 30 may scale the floating samples of the template by a fixed factor (e.g., 64, 128, or 256) and round the scaled samples to their nearest integer. The rounded samples may be stored in memory. Then, when the template is used to reconstruct a candidate hypothesis, the corresponding samples may first be unscaled to their original precision to ensure that the generated samples in the candidate hypothesis are within the correct dynamic range.

[0153] FIG. 12 is a graphical representation illustrating exemplary hypothesis generation based on a linear combination of templates according to some implementations of the present disclosure. In FIG. 12, four patterned blocks 0-3 represent transform coefficient candidates whose signs are predicted. Factors C0, C1, C2, and C3 represent corresponding values of dequantized transform coefficients of the four transform coefficient candidates. Templates 0-3 may correspond to the four transform coefficient candidates 0-3, respectively. For example, template 0 corresponding to transform coefficient candidate 0 may be generated by applying an inverse secondary transform and an inverse linear transform to the transform block, with transform coefficient candidate 0 set to one and the remaining transform coefficient candidates in the transform block set to zero. Templates 1-3 may be generated in a similar manner, respectively. Hypothesis candidates may be generated by adding templates 0-1 and weights C0-C3, respectively.

[0154] 13A is a graphical representation illustrating an exemplary implementation of an existing code prediction scheme according to some examples. FIG. 13B is a graphical representation illustrating an exemplary implementation of a vector-based code prediction scheme according to some implementations of the present disclosure. An exemplary comparison between the existing code prediction scheme and the vector-based code prediction scheme disclosed herein is presented herein with reference to FIGS. 13A-13B.

[0155] In Figures 13A and 13B, a transform block has six non-zero transform coefficients selected as transform coefficient candidates for sign prediction. The transform coefficient candidates are scanned from the coefficient matrix of the transform block using a raster scan order. Figures 13A-13B also show the original and predicted codes of the transform coefficient candidates. For example, the original and predicted codes of a first transform coefficient candidate having a value of "-2" are both "-" (represented as "1" in Figures 13A-13B). The original and predicted codes of a second transform coefficient candidate having a value of "3" are both "+" (represented as "0" in Figures 13A-13B). The original and predicted codes of a third transform coefficient candidate having a value of "1" are "+" and "-", respectively (represented as "0" and "1" in Figures 13A-13B). The original code of the third transform coefficient candidate is incorrectly predicted. As shown in FIGS. 13A and 13B, except for the third transform coefficient, the original codes of all other transform coefficient candidates are the same as their corresponding predicted codes (ie, correctly predicted).

[0156] Referring to FIG. 13A, a total of six bins (i.e., 0, 0, 1, 0, 0, and 0) are generated, each corresponding to a candidate transform coefficient. The six bins may be generated by performing an XOR operation between the original codes and predicted codes of the six candidate transform coefficients. The six bins may be used to indicate the corresponding correctness of the six predicted codes. For example, the first and second bins, each having a value of "0," indicate that the predicted codes of the first and second candidate transform coefficients are correct. The third bin, having a value of "1," indicates that the predicted code of the third candidate transform coefficient is incorrect. The six bins may be sent to CABAC for entropy coding.

[0157] Referring to FIG. 13B, the vector-based code prediction scheme disclosed herein divides six transform coefficient candidates into three groups, each containing two consecutive transform coefficient candidates. Because the codes of the transform coefficient candidates in Groups 0 and 2 can be correctly predicted, only two bins each containing a value of "0" are generated for the two groups. In the case of Group 1, because it contains a third transform coefficient candidate whose code cannot be correctly predicted, a bin containing a value of "1" (underlined in FIG. 13B) is generated and signaled in the bitstream to indicate that the group contains at least a transform coefficient candidate whose original code differs from its predicted code. Subsequently, for the third and fourth coefficients in Group 1, two additional bins containing a value of "1" and a value of "0" are generated to indicate whether their codes can be correctly predicted. Correspondingly, when the vector-based code prediction scheme disclosed herein is applied, a total of five bins are generated for CABAC, which has fewer bits than the bins generated by the existing code prediction scheme shown in FIG. 13A. Therefore, by applying the vector-based code prediction scheme disclosed herein, the signaling overhead may be reduced and the coding efficiency of the transform blocks may be improved.

[0158] According to the present disclosure, a raster scan order is used to obtain transform coefficient candidates from the coefficient matrix of the transform block as shown in Figure 13B, but any other scan order may be used to select transform coefficient candidates for sign prediction. For example, the transform coefficient candidates may be selected based on one or more selection criteria described above. Similar descriptions will not be repeated herein.

[0159] 14A is a graphical representation illustrating an exemplary calculation of a left diagonal cost function along the left diagonal direction according to some implementations of the present disclosure. FIG. 14B is a graphical representation illustrating an exemplary calculation of a right diagonal cost function along the right diagonal direction according to some implementations of the present disclosure. Compared with the above equation (6) for calculating costHV, the left diagonal cost function costD1 or the right diagonal cost function costD2 shown in the above equation (10) or (11) may require more neighboring pixels (shown as marked pixels in areas 1402, 1404, and 1406 in FIGS. 14A-14B) to support the calculation of the cost functions costD1, costD2 along the diagonal direction. If these pixels in areas 1402, 1404, and 1406 are unavailable, a nearest-neighbor padding method may be adopted to fill these unavailable positions. For example, B in area 1406 -1,4 If is not available, B -1,4 To fill the position of B -1,4 B is the closest available pixel to -1,3 is used (e.g., B -1,4 =B -1,3 (B in Area 1402 -1,-1 (C -1,-1 (also written as "B") -1,-2 (C -1,-1 ), B -2,-1 (C -1,-1 ), and B -2,-2 (C -1,-1 ) is unavailable, two exemplary methods are disclosed herein for filling in the unavailable positions.

[0160] In a first exemplary method, each unavailable location may be filled by weighting its nearest available pixel, as shown in equations (12)-(15) below.

number

number

number

number

[0161] In a second exemplary method, some of the unavailable locations may each be filled with its nearest available pixel. -1,-2 If is not available, this is C 0,-2 Filled with B -2,-1 If is not available, B -2,0 However, B -2,-2 and B -1,-1 If σ is not available, they may be filled with the average of their two nearest neighbors calculated according to equations (12) and (15) above.

[0162] In accordance with the present disclosure, while the calculation of the cost function shown in equation (10) or equation (11) above uses the left and right diagonals (i.e., 135° and 45° as shown in FIGS. 14A-14B) for illustrative purposes, it is contemplated that any other measurement components (e.g., continuity measurements along any one or more directions) may be incorporated into the calculation of the cost function for code prediction.

[0163] 15 is a flowchart of an example method 1500 for code prediction in block-based video coding according to some implementations of the present disclosure. The method 1500 may be implemented by a video processor associated with the video encoder 20 or the video decoder 30 and may include steps 1502-1508 described below. Some of the steps may be optional for implementing the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 15.

[0164] In step 1502, the video processor may apply a primary transform and a secondary transform to a transform block of a video frame from the video to generate transform coefficients for the transform block.

[0165] In step 1504, the video processor may select a set of candidate transform coefficients for sign prediction from the transform coefficients.

[0166] In step 1506, the video processor may apply a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of candidate transform coefficients.

[0167] In step 1508, the video processor may determine the combination of code candidates associated with the selected hypothesis to be a set of predicted codes for the set of candidate transform coefficients.

[0168] 16 is a flowchart of another exemplary method 1600 for code prediction in block-based video coding according to some implementations of the present disclosure. Method 1600 may be implemented by a video processor associated with video encoder 20 or video decoder 30 and may include steps 1602-1616 described below. Specifically, steps 1606-1610 of method 1600 may be performed as an exemplary implementation of step 1506 of method 1500. Some of the steps may be optional for implementing the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 16.

[0169] In step 1602, the video processor may apply a primary transform and a secondary transform to a transform block of a video frame from the video to generate transform coefficients for the transform block.

[0170] In step 1604, the video processor may select a set of candidate transform coefficients for sign prediction from the transform coefficients.

[0171] In step 1606, the video processor may determine multiple combinations of code candidates for the set of candidate transform coefficients based on the total number of candidate transform coefficients in the set of candidate transform coefficients.

[0172] In step 1608, the video processor may apply a template-based hypothesis generation scheme to generate multiple candidate hypotheses for multiple combinations of code candidates, respectively.

[0173] In step 1610, the video processor may select a hypothesis associated with the smallest cost from a plurality of candidate hypotheses based on a cost function that incorporates combined gradients along the horizontal, vertical, and diagonal directions.

[0174] In step 1612, the video processor may determine the combination of code candidates associated with the selected hypothesis to be a set of predicted codes for the set of candidate transform coefficients.

[0175] In step 1614, the video processor may apply a vector-based code prediction scheme to the set of predicted codes to generate a sequence of code signaling bits for the set of candidate transform coefficients.

[0176] In step 1616, the video processor may generate a bitstream including a sequence of code signaling bits.

[0177] According to this disclosure, method 1500 of Figure 15 and method 1600 of Figure 16 may be performed on a video encoder side or a video decoder side. When method 1500 of Figure 15 and method 1600 of Figure 16 are performed on a video encoder side, these methods may be considered as encoding methods for sign prediction of transform coefficients at the video encoder side. When method 1500 of Figure 15 and method 1600 of Figure 16 are performed on a video decoder side, these methods may be considered as decoding methods for sign prediction of transform coefficients at the video decoder side. An exemplary decoding method for transform coefficient sign prediction at the video decoder side is provided below with reference to Figure 18.

[0178] 17 illustrates a computing environment 1710 coupled with a user interface 1750, according to some implementations of the present disclosure. The computing environment 1710 may be part of a data processing server. The computing environment 1710 includes a processor 1720, a memory 1730, and an input / output (I / O) interface 1740.

[0179] The processor 1720 typically controls the overall operation of the computing environment 1710, such as operations related to display, data acquisition, data communication, and image processing. The processor 1720 may include one or more processors for executing instructions called for performing all or some of the steps in the methods described above. Additionally, the processor 1720 may include one or more modules that facilitate interaction between the processor 1720 and other components. The processor 1720 may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.

[0180] Memory 1730 is configured to store various types of data to support the operation of computing environment 1710. Memory 1730 may include predefined software 1732. Examples of such data include instructions for any applications or methods run on computing environment 1710, video data sets, image data, etc. Memory 1730 may be implemented using any type of volatile or non-volatile memory device, or combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0181] The I / O interface 1740 provides an interface between the processor 1720 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1740 may be coupled to an encoder and a decoder.

[0182] In some implementations, a non-transitory computer-readable storage medium is also provided, for example in memory 1730, that includes a plurality of programs executable by processor 1720 in computing environment 1710 for performing the methods described above. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or datastream that includes encoded video information (e.g., video information including one or more syntax elements) generated by an encoder (e.g., video encoder 20 of FIG. 2) using, for example, the encoding method described above, for use by a decoder (e.g., video decoder 30 of FIG. 3) in decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0183] In some implementations, a computing device is also provided that includes one or more processors (e.g., processor 1720) and a non-transitory computer-readable storage medium or memory 1730 storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above upon execution of the plurality of programs.

[0184] In some implementations, a computer program product is also provided that includes a plurality of programs, e.g., in memory 1730, executable by processor 1720 in computing environment 1710 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0185] In some implementations, the computing environment 1710 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0186] 18 is a flowchart of a video decoding method 1800 for transform coefficient sign prediction at the video decoder side according to some implementations of the present disclosure. The method 1800 may be performed by a video processor associated with the video decoder 30 and may include steps 1802-1810 described below. Some of the steps may be optional for implementing the disclosure provided herein. Furthermore, some of the steps may be performed simultaneously or in a different order than that shown in FIG. 18.

[0187] In step 1802, the video processor may select a set of candidate transform coefficients from the dequantized transform coefficients for transform coefficient sign prediction. The dequantized transform coefficients are associated with transform blocks of video frames from the video. The dequantized transform coefficients of a transform block in video decoder 30 may be equivalent to the transform coefficients of the transform block in video encoder 20.

[0188] In some implementations, a video processor may receive a bitstream including a sequence of code signaling bits and quantized transform coefficients associated with a transform block. The video processor may generate dequantized transform coefficients from the quantized transform coefficients via dequantization unit 86 of FIG. 3.

[0189] In some implementations, the video processor may select a set of transform coefficient candidates from the dequantized transform coefficients based on the magnitudes of the dequantized transform coefficients. In some implementations, the video processor may select a set of transform coefficient candidates from the dequantized transform coefficients based on the magnitudes of the quantization indexes of the dequantized transform coefficients. In some implementations, the video processor may select a set of transform coefficient candidates from the dequantized transform coefficients based on a coefficient scanning order of entropy coding applied to the video coding.

[0190] In some implementations, the video processor may select a set of candidate transform coefficients from the dequantized transform coefficients based on their influence scores on the reconstructed boundary samples of the transform block. For example, the influence scores of the dequantized transform coefficients on the reconstructed boundary samples are measured as the L1 norm of the variation of each dequantized transform coefficient on the reconstructed boundary samples. In another example, the influence scores of the dequantized transform coefficients on the reconstructed boundary samples are measured as the L2 norm of the variation of each dequantized transform coefficient on the reconstructed boundary samples.

[0191] In some implementations, the video processor may perform operations such as those described above with reference to coefficient selection operation 1104 of Figure 11 to select a set of candidate transform coefficients from the dequantized transform coefficients. The video processor may also perform operations such as those described above with reference to step 1504 of Figure 15 to select a set of candidate transform coefficients from the dequantized transform coefficients. Similar descriptions will not be repeated herein.

[0192] In step 1804, the video processor may apply a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for a set of candidate transform coefficients.

[0193] In some implementations, the video processor may perform operations as described above with reference to hypothesis generation operation 1106 of Figure 11. The video processor may also perform operations as described above with respect to step 1506 of Figure 15 to select a hypothesis from multiple candidate hypotheses, and similar descriptions will not be repeated herein.

[0194] In step 1806, the video processor may determine the combination of code candidates associated with the selected hypothesis to be a set of predicted codes for the set of candidate transform coefficients.

[0195] In some implementations, the video processor may perform operations as described above with reference to code generation operation 1108 of Figure 11. The video processor may also perform operations as described above with respect to step 1508 of Figure 15 to determine a set of predictive codes for a set of candidate transform coefficients, a description of which will not be repeated herein.

[0196] At step 1808, the video processor may estimate original signs of the set of candidate transform coefficients based on the set of predictive signs and the sequence of sign signaling bits received from the video encoder.

[0197] For example, referring to Figure 13B, a set of predicted codes may include group #0 with values (1,0), group #2 with values (1,0), and group #3 with values (1,0), where 1 indicates a negative sign and 0 indicates a positive sign. A sequence of code signaling bits may include bit "0" in group #0, bits "1,1,0" in group #2, and bit "0" in group #3. Because the bit in group #0 has value "0", indicating that the predicted code for this group with value (1,0) is the same as the original code, the estimated original code for group #0 is determined to be (1,0). Because the first bit of the bits in group #1 has a value of "1," indicating that the predicted code (1,0) of this group is not the same as the original code, the estimated original code (1,0) of group #1 is determined to be the XOR result of the predicted code (1,0) of this group and the second and third bits of group #1, which are "1,0" (e.g., estimated original code = XOR((1,0),(1,0)) = (0,0)). Because the bits in group #2 have a value of "0," indicating that the predicted code of this group, which has a value of (1,0), is the same as the original code, the estimated original code of group #2 is determined to be (1,0). The estimated original codes of the set of candidate transform coefficients are then formed by concatenating the estimated original codes of groups #0, #1, and #2, respectively, and include (1,0,0,0,1,0).

[0198] In step 1810, the video processor may update the dequantized transform coefficients based on the estimated original signs of the set of candidate transform coefficients. For example, the video processor may use the estimated original signs as the true signs for the dequantized transform coefficients in the transform block corresponding to the set of candidate transform coefficients.

[0199] In some implementations, after the dequantized transform coefficients are updated, the video processor may further apply an inverse primary transform and an inverse secondary transform to the dequantized transform coefficients to generate residual samples in a residual block corresponding to the transform block. The inverse secondary transform corresponds to a secondary transform including an LFNST. The inverse primary transform corresponds to a primary transform including a DCT-II, DCT-V, DCT-VIII, DST-I, DST-IV, DST-VII, or identity transform.

[0200] In some implementations, the sequence of code signaling bits for the set of transform coefficient candidates is generated by the video encoder by applying a vector-based code prediction scheme to another set of predictive codes for another set of transform coefficient candidates selected at the video encoder side, the other set of transform coefficient candidates being transform coefficients at the video encoder side that correspond to the set of transform coefficient candidates at the video decoder side.

[0201] In some implementations, applying the vector-based code prediction scheme to another set of predictive codes for another set of transform coefficient candidates further includes dividing the other set of transform coefficient candidates into multiple groups of transform coefficient candidates, and, for each group of transform coefficient candidates, generating one or more code signaling bits for the group of transform coefficient candidates based on whether the original code of the group of transform coefficient candidates is identical to the predictive code of the group of transform coefficient candidates.

[0202] In some implementations, generating one or more code signaling bits for the group of transform coefficient candidates includes generating a bin having a value of zero in response to an original code of the group of transform coefficient candidates being identical to a predicted code of the group of transform coefficient candidates, and adding the bin to the bitstream as a code signaling bit.

[0203] In some implementations, generating one or more code signaling bits for the group of transform coefficient candidates includes generating a bin having a value of 1 in response to an original code of the group of transform coefficient candidates not being identical to a predicted code of the group of transform coefficient candidates, generating a set of additional bins to signal the corresponding correctness of the predicted code of the group of transform coefficient candidates, and adding the bin and the set of additional bins to the bitstream as code signaling bits.

[0204] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosure. Many modifications, variations and alternative implementations will be apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0205] Unless otherwise specified, the order of steps of the method according to the present disclosure is intended to be illustrative only, and the steps of the method according to the present disclosure are not limited to the order specifically described above, and may be changed according to actual conditions. Furthermore, at least one of the steps of the method according to the present disclosure may be adjusted, combined, or deleted according to actual requirements.

[0206] The examples have been chosen and described to explain the principles of the disclosure, to enable those skilled in the art to understand the disclosure in various implementations, and to make best use of the underlying principles and various implementations with various modifications as suited to the particular use contemplated. Therefore, it should be understood that the scope of the disclosure is not limited to the particular examples of implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the disclosure.

Claims

1. 1. A video decoding method for transform coefficient sign prediction, comprising: selecting a set of candidate transform coefficients for the transform coefficient sign prediction from dequantized transform coefficients, the dequantized transform coefficients being associated with transform blocks of a video frame; applying a template-based hypothesis generation scheme to select a hypothesis from a plurality of candidate hypotheses for the set of candidate transform coefficients; determining a combination of code candidates associated with the selected hypothesis as a set of predicted codes for the set of candidate transform coefficients; estimating original codes of the set of candidate transform coefficients based on a set of prediction codes and a sequence of code signaling bits received from a video encoder; updating the dequantized transform coefficients based on the estimated original signs for the set of candidate transform coefficients; 1. A video decoding method comprising:

2. receiving a bitstream including the sequence of code signaling bits and quantized transform coefficients associated with the transform block; generating the dequantized transform coefficients from the quantized transform coefficients; The video decoding method of claim 1 further comprising:

3. 2. The video decoding method of claim 1, further comprising applying an inverse primary transform and an inverse secondary transform to the dequantized transform coefficients to generate residual samples in a residual block corresponding to the transform block.

4. the inverse secondary transform corresponds to a secondary transform including a low frequency non-separable transform (LFNST); and / or 4. The video decoding method of claim 3, wherein the inverse linear transform corresponds to a linear transform comprising DCT-II, DCT-V, DCT-VIII, DST-I, DST-IV, DST-VII, or an identity transform.

5. selecting the set of candidate transform coefficients selecting the set of candidate transform coefficients from the dequantized transform coefficients based on magnitudes of the dequantized transform coefficients; selecting the set of candidate transform coefficients from the dequantized transform coefficients based on magnitudes of quantization indexes of the dequantized transform coefficients; selecting the set of candidate transform coefficients from the dequantized transform coefficients based on a coefficient scanning order of an entropy coding applied to video coding; or selecting the set of candidate transform coefficients from the dequantized transform coefficients based on influence scores of the dequantized transform coefficients on reconstructed boundary samples of the transform block; 10. The video decoding method of claim 1, comprising at least one of:

6. 6. The video decoding method of claim 5, wherein the influence score of the dequantized transform coefficients relative to the reconstructed boundary sample is measured as an L1 norm or an L2 norm of the variation of each dequantized transform coefficient relative to the reconstructed boundary sample.

7. applying the template-based hypothesis generation scheme to select the hypotheses; determining a plurality of combinations of code candidates for the set of transform coefficient candidates based on a total number of transform coefficient candidates in the set of transform coefficient candidates; applying the template-based hypothesis generation scheme to generate the plurality of candidate hypotheses for the plurality of combinations of code candidates, respectively; determining a hypothesis associated with a minimum cost from the plurality of candidate hypotheses based on a cost function incorporating combined gradients along horizontal, vertical, and diagonal directions; The video decoding method of claim 1 , comprising:

8. applying the template-based hypothesis generation scheme to generate the plurality of candidate hypotheses for the plurality of combinations of candidate codes, respectively; 8. The video decoding method of claim 7, comprising generating corresponding candidate hypotheses for each combination of candidate codes based on a linear combination of templates.

9. each template corresponds to a candidate transform coefficient in the set of candidate transform coefficients; each template is generated by applying an inverse secondary transform and an inverse linear transform to the transform block, and each of the set of transform coefficient candidates is set to zero except for the transform coefficient candidate corresponding to the template, which is set to one; 9. The video decoding method of claim 8.

10. applying the template-based hypothesis generation scheme to generate the plurality of candidate hypotheses for the plurality of combinations of candidate codes, respectively; determining the plurality of combinations of candidate codes to be a plurality of hypothesis indices for the plurality of candidate hypotheses, respectively; ordering the plurality of hypothesis indexes based on a Gray code ordering of the plurality of hypothesis indexes to generate a permuted sequence of hypothesis indexes; for a first hypothesis index in the sorted sequence of hypothesis indexes, generating a first candidate hypothesis corresponding to the first hypothesis index by applying an inverse secondary transform and an inverse linear transform to the transform block, wherein each of the set of candidate transform coefficients is set to 1; for a second hypothesis index in the sorted sequence of hypothesis indexes immediately following the first hypothesis index, generating the second candidate hypothesis corresponding to the second hypothesis index based on the first candidate hypothesis corresponding to the first hypothesis index and an adjustment term for the second candidate hypothesis; 8. The video decoding method of claim 7, comprising:

11. 8. The video decoding method of claim 7, wherein the cost functions are determined based on a horizontal-vertical cost function that incorporates gradients along the vertical and horizontal directions, a left-diagonal cost function that incorporates gradients along the left diagonal direction, and a right-diagonal cost function that incorporates gradients along the right diagonal direction.

12. The sequence of code signaling bits for the set of candidate transform coefficients is encoded by the video encoder as: by applying a vector-based code prediction scheme to a different set of predictive codes on a different set of candidate transform coefficients selected at the video encoder; The video decoding method of claim 1 , wherein the other set of candidate transform coefficients are transform coefficients at the video encoder side that correspond to the set of candidate transform coefficients at a video decoder side.

13. applying the vector-based code prediction scheme to the different set of predictive codes for the different set of candidate transform coefficients to generate the sequence of code signaling bits; dividing the other set of candidate transform coefficients into a plurality of groups of candidate transform coefficients; generating, for each group of candidate transform coefficients, one or more code signaling bits for the group of candidate transform coefficients based on whether an original code of the group of candidate transform coefficients is identical to a predicted code of the group of candidate transform coefficients; The video decoding method of claim 12 further comprising:

14. generating the one or more code signaling bits for the group of candidate transform coefficients; in response to determining that the original codes of the group of candidate transform coefficients are the same as the predicted codes of the group of candidate transform coefficients; generating bins with a value of zero; adding said bins to a bitstream as code signaling bits; 14. The video decoding method of claim 13, comprising:

15. generating the one or more code signaling bits for the group of candidate transform coefficients; In response to determining that the original codes of the group of candidate transform coefficients are not identical to the predicted codes of the group of candidate transform coefficients, generating bins with a value of 1; generating a set of additional bins for signaling corresponding correctness of the predictive codes of the group of candidate transform coefficients; adding said bin and said set of additional bins to a bitstream as code signaling bits; 14. The video decoding method of claim 13, comprising:

16. 1. A video decoding apparatus for transform coefficient sign prediction, comprising: a memory configured to store a bitstream to be decoded; one or more processors coupled to the memory and configured to perform the video decoding method of any of claims 1 to 15 on the bitstream; A video decoding device comprising:

17. 16. A non-transitory computer-readable storage medium storing instructions and a bitstream to be decoded, the instructions, when executed by one or more processors, causing the one or more processors to perform a video decoding method according to any of claims 1 to 15 on the bitstream. A non-transitory computer-readable storage medium.

18. 16. A computer program comprising instructions that, when executed by a computing device having one or more processors, cause the computing device to store a bitstream and to perform the video decoding method of any of claims 1 to 15 on the bitstream.

19. A method for storing a bitstream in a non-transitory computer-readable storage medium, the method comprising decoding the stored bitstream by a video decoding method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and apparatus for residual code prediction in the transform domain

    JP2021516016A

  • Low-complexity sign prediction for video coding

    US20180176556A1