Code Prediction for Block-Based Video Coding

The video decoding method and device enhance video coding efficiency by predicting transform coefficients through hypothesis generation and cost function-based selection, addressing compression challenges in limited bandwidth and memory scenarios.

JP7821879B2Active Publication Date: 2026-02-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024527601
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-15
Filing Date
2022-11-08
Publication Date
2026-02-27
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing video data while maintaining video quality due to limited bandwidth and memory resources, particularly in the prediction of transform coefficients.

Method used

Implementing a video decoding method and device that utilize sign prediction and code prediction of transform coefficients through hypothesis generation and cost function-based selection, along with a non-transitory computer-readable storage medium for executing these processes.

Benefits of technology

Enhances video coding efficiency by improving the prediction of transform coefficients, thereby optimizing compression while maintaining video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007821879000020
    Figure 0007821879000020
  • Figure 0007821879000021
    Figure 0007821879000021
  • Figure 0007821879000022
    Figure 0007821879000022
Patent Text Reader

Abstract

An implementation of the present disclosure provides a video decoding apparatus and method for sign prediction of transform coefficients at a video decoder side. The method may include receiving a bitstream including a sequence of sign signaling bits. The method may further include determining a sign prediction region in a transform block of a video frame from a video performing the sign prediction of the transform coefficients of the transform block, and generating, by the one or more processors, a plurality of candidate hypotheses of a set of transform coefficient candidates associated with the sign prediction region of the transform block. The method may also include selecting, by the one or more processors, a hypothesis from the plurality of candidate hypotheses as a set of predicted signs of the set of transform coefficient candidates based on a cost function, and estimating, by the one or more processors, an original sign of the set of transform coefficient candidates based on the set of predicted signs and the sequence of sign signaling bits.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based on and claims priority to U.S. Provisional Application No. 63 / 277,705, filed November 10, 2021. This application further claims priority to PCT Application No. PCT / US22 / 43607, filed September 15, 2022, which in turn claims priority to U.S. Provisional Application No. 63 / 244,317, filed September 15, 2021, and U.S. Provisional Application No. 63 / 250,797, filed September 30, 2021. This application further claims priority to PCT Application No. PCT / US22 / 40442, filed August 16, 2022, which in turn claims priority to U.S. Provisional Application No. 63 / 233,940, filed August 17, 2021. The contents of all of the above applications are incorporated herein by reference in their entirety. [Technical Field]

[0002] This application relates to video coding and compression, and more particularly to video processing systems and methods for code prediction in block-based video coding. [Background technology]

[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, videoconferencing devices, and video streaming devices. Electronic devices transmit, receive, or communicate digital video data over communication networks and / or store the digital video data in storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding may be used to compress video data according to one or more video coding standards before communicating or storing it. For example, video coding standards include Versatile Video Coding (VVC), Joint Search and Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Experts Committee (MPEG) coding, and the like. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention

[0004] Implementations of this disclosure provide a video decoding method for sign prediction of transform coefficients, the video decoding method may include receiving a bitstream including a sequence of sign signaling bits. video The decoding method may further include determining, by one or more processors, a sign prediction region in a transform block of a video frame from the video performing the sign prediction of the transform coefficients of the transform block, and generating, by the one or more processors, a plurality of candidate hypotheses for a set of candidate transform coefficients associated with the sign prediction region of the transform block. videoThe decoding method may include selecting, by the one or more processors, a hypothesis from the plurality of candidate hypotheses as a set of predicted codes for the set of candidate transform coefficients based on a cost function. video The decoding method may include estimating, by the one or more processors, original codes of the set of candidate transform coefficients based on the set of predicted codes and the sequence of code signaling bits.

[0005] An implementation of the present disclosure also provides a video decoding device for code prediction of transform coefficients. The video decoding device may include a memory configured to store a bitstream including a sequence of code signaling bits and one or more processors coupled to the memory. The one or more processors may be configured to determine a code prediction region in a transform block of a video frame from a video to perform the code prediction of the transform coefficients of the transform block and generate a plurality of candidate hypotheses for a set of transform coefficient candidates associated with the code prediction region of the transform block. The one or more processors may be further configured to select one hypothesis from the plurality of candidate hypotheses as a set of predicted codes for the set of transform coefficient candidates based on a cost function. The one or more processors may also be configured to estimate original codes for the set of transform coefficient candidates based on the set of predicted codes and the sequence of code signaling bits.

[0006] Implementations of the present disclosure also provide a non-transitory computer-readable storage medium having stored thereon a bitstream including a sequence of sign signaling bits and instructions that, when executed by one or more processors, cause the one or more processors to perform a video decoding method for sign prediction of transform coefficients. The video decoding method may include determining a sign prediction region in a transform block of a video frame from a video performing the sign prediction of the transform coefficients of the transform block, and generating a plurality of candidate hypotheses for a set of candidate transform coefficients associated with the sign prediction region of the transform block. video The decoding method may further include selecting one hypothesis from the plurality of candidate hypotheses as a set of predicted codes for the set of candidate transform coefficients based on a cost function. video The decoding method may include estimating original codes of the set of candidate transform coefficients based on the set of predicted codes and the sequence of code signaling bits. 。

[0007] Implementations of the present disclosure include video Decryption Also provided is a non-transitory computer-readable storage medium having stored thereon a bitstream decodable by the method. Decryption The method includes, by one or more processors, decoding a sign prediction region in a transform block of a video frame from a video source that performs sign prediction of transform coefficients of the transform block. Decision and generating, by the one or more processors, a plurality of candidate hypotheses for a set of candidate transform coefficients associated with the code prediction region of the transform block. The video decoding method may further include selecting, by the one or more processors, a hypothesis from the plurality of candidate hypotheses as a set of predicted codes for the set of candidate transform coefficients based on a cost function. The video decoding method may also include estimating, by the one or more processors, original codes for the set of candidate transform coefficients based on the set of predicted codes and a sequence of code signaling bits in the bitstream.

[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. [Brief explanation of the drawings]

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.

[0010] [Figure 1] FIG. 1 is a block diagram illustrating an example system for encoding and decoding video blocks, according to some implementations of the present disclosure. [Figure 2] 1 is a block diagram illustrating an example video encoder according to some implementations of the present disclosure. [Figure 3] 1 is a block diagram illustrating an example video decoder according to some implementations of the present disclosure. [Figure 4A] 1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some implementations of the present disclosure. [Figure 4B] 1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some implementations of the present disclosure. [Figure 4C] 1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some implementations of the present disclosure. [Figure 4D] 1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some implementations of the present disclosure. [Figure 4E] 1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some implementations of the present disclosure. [Figure 5] 1 is a graphical representation illustrating a top-left scan order of transform coefficients within a set of coefficients, according to some embodiments. [Figure 6] 1 is a graphical representation illustrating a Low Frequency Non-Separable Transform (LFNST) process, according to some embodiments. [Figure 7] 10 is a graphical representation showing the top left region of the primary transform coefficients input to a forward LFNST, according to some embodiments. [Figure 8] 1 is a graphical representation illustrating a search area for intra-template matching, according to some embodiments. [Figure 9] 1 is a graphical representation illustrating an exemplary process of code prediction, according to some embodiments. [Figure 10] 1 is a graphical representation illustrating the calculation of a cost function for symbol prediction, according to some embodiments. [Figure 11] 1 is a graphical representation illustrating two exemplary scalar quantizers used for dependent scalar quantization, according to some embodiments. [Figure 12A] 1 is a graphical representation illustrating state transitions using a four-state state machine used for dependent scalar quantization, according to some embodiments. [Figure 12B] 12B is a table illustrating exemplary quantizer selection in response to the state transitions of FIG. 12A in accordance with some embodiments. [Figure 13] FIG. 2 is a block diagram illustrating an example code prediction process in block-based video coding, according to some implementations of the present disclosure. [Figure 14] 1 is a graphical representation illustrating an example hypothesis generation based on a linear combination of templates, according to some implementations of the present disclosure. [Figure 15A] 1 is a graphical representation illustrating an exemplary implementation of an existing code prediction scheme, according to some embodiments. [Figure 15B] 1 is a graphical representation illustrating an example implementation of a vector-based code prediction scheme, according to some implementations of the present disclosure. [Figure 16A] 10 is a graphical representation illustrating an exemplary calculation of a left-diagonal cost function along the left diagonal direction, according to some implementations of the present disclosure. [Figure 16B] 10 is a graphical representation illustrating an exemplary calculation of a right-diagonal cost function along the right-diagonal direction, according to some implementations of the present disclosure. [Figure 17] 10 is a flowchart of a method for capturing dominant gradient directions in neighboring reconstructed samples of a current block, according to some implementations of the present disclosure. [Figure 18A]10 is a graphical representation illustrating example template samples and gradient filter windows in gradient-based selection of sample extrapolation direction of a cost function, according to some implementations of the present disclosure. [Figure 18B] 1 is a graphical representation illustrating an example Histogram of Gradients (HoG) for gradient-based selection of sample extrapolation directions for a cost function, according to some implementations of the present disclosure. [Figure 19] 1 is a graphical representation illustrating a sign prediction region for predicting signs of transform coefficients, according to some embodiments. [Figure 20] 1 is a flowchart of an example code prediction method in block-based video coding according to some implementations of the present disclosure. [Figure 21] 1 is a flowchart of an example video encoding method for sign prediction of transform coefficients performed by a video encoder, according to some implementations of the present disclosure. [Figure 22] 1 is a flowchart of an example video decoding method for sign prediction of transform coefficients performed by a video decoder, according to some implementations of the present disclosure. [Figure 23] FIG. 1 is a block diagram illustrating a computing environment coupled to a user interface, according to some implementations of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in an understanding of the subject matter described herein. However, it will be apparent to those skilled in the art that various alternative embodiments can be used and the subject matter can be practiced without these specific details without departing from the scope of the claims. For example, it will be apparent to those skilled in the art that the subject matter described herein can be implemented in multiple types of electronic devices having digital video capabilities.

[0012] It should be understood that terms such as "first," "second," and the like, used in the specification and claims of this disclosure and in the accompanying drawings, are used to distinguish between objects and are not used to describe any particular order or sequence. It should be understood that the data so used may be substituted under appropriate conditions, thereby enabling the implementation of the embodiments of the disclosure described herein to be carried out in an order other than that shown in the accompanying drawings or described in this disclosure.

[0013] 1 is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel, according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 include wireless communication capabilities.

[0014] In some implementations, destination device 14 may receive the encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of transferring encoded video data from source device 12 to destination device 14. In one embodiment, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to destination device 14.

[0015] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or another intermediate storage device that may store the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or download. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network-attached storage (NAS) device, or a local disk drive. Destination device 14 may access the encoded video data over any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or any combination thereof, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination thereof.

[0016] 1 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface receiving video data from a video content provider, and / or a computer graphics system generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may include a camera phone or video phone. However, implementations described in this disclosure may be applicable to video encoding generally and may be applicable to wireless and / or wired applications.

[0017] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access, decoding, and / or playback by destination device 14 or another device. Output interface 22 may further include a modem and / or a transmitter.

[0018] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem to receive encoded video data over link 16. The encoded video data communicated over link 16 or provided to storage device 32 may include various syntax elements generated by video encoder 20 that are used to decode the video data by video decoder 30. Such syntax elements may be included within the encoded video data that is transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0019] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0020] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as, for example, VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of those standards. It should be understood that the present disclosure is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or future standards.

[0021] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or combinations thereof. If implemented partially in software, the electronic device may store software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated within the respective devices as part of a combined encoder / decoder (CODEC).

[0022] 2 is a block diagram illustrating an example video encoder 20 according to some implementations described herein. The video encoder 20 can perform intra-predictive and inter-predictive coding of video blocks within video frames. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. Note that the term "frame" may be used synonymously with the terms "image" or "picture" in the field of video coding.

[0023] As shown in FIG. 2, video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a division unit 45, an intra-prediction processing unit 46, and an intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, may be disposed between adder 62 and DPB 64 to filter block boundaries and remove block artifacts from the reconstructed video data. In addition to the deblocking filter, another in-loop filter, such as an SAO filter and / or an adaptive in-loop filter (ALF), may be used to filter the output of adder 62. In some embodiments, the in-loop filter may be omitted, and the decoded video block may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.

[0024] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained from video source 18, for example, as shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used in encoding video data by video encoder 20 (e.g., in intra-predictive coding mode or inter-predictive coding mode). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various embodiments, video data memory 40 may be on-chip with other components of video encoder 20 or off-chip relative to those components.

[0025] As shown in FIG. 2, after receiving video data, partitioning unit 45 in prediction processing unit 41 partitions the video data into video blocks. This partitioning may include dividing the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predetermined partitioning structure, such as a quadtree (QT) structure associated with the video data. A video frame may be, or may be considered as, a two-dimensional array or matrix of samples having sample values. The samples in the array may be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame may be partitioned into multiple video blocks, for example, using QT partitioning. A video block has smaller dimensions than a video frame, but may also be, or may be considered as, a two-dimensional array or matrix of samples having sample values. The number of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further divided into one or more block partitions or sub-block partitions (which may again form blocks), e.g., by iteratively using QT partitioning, binary tree (BT) partitioning, ternary tree (TT) partitioning, or any combination thereof. Note that, as used herein, the term "block" or "video block" may refer to a portion of a frame or picture, particularly a rectangular (square or non-square) portion. For example, with reference to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or a corresponding block such as a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB). Alternatively or additionally, a block or video block may be or correspond to a sub-block such as a CTB, CB, PB, TB, etc.

[0026] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., code rate and distortion level). Prediction processing unit 41 provides the resulting intra- or inter-predictively coded block (e.g., a predictive block) to summer 50 to generate a residual block and to summer 62 to reconstruct a coded block for later use as part of a reference frame. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partition information, and other such syntax information, to entropy coding unit 56.

[0027] To select an appropriate intra-prediction coding mode for a current video block, intra-prediction processing unit 46 within prediction processing unit 41 performs intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block being coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0028] In some implementations, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating motion vectors that indicate the displacement of video blocks in the current video frame relative to predictive blocks in a reference frame according to a predetermined pattern in the sequence of video frames. Motion estimation performed by motion estimation unit 42 may be the process of generating motion vectors that estimate the motion of video blocks. The motion vectors may, for example, indicate the displacement of video blocks in the current video frame or picture relative to predictive blocks in a reference frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to how motion estimation unit 42 determines motion vectors for inter prediction, and may utilize motion estimation unit 42 to determine the block vectors.

[0029] A prediction block for a video block may be or may correspond to a block or reference block of a reference frame that is deemed to closely match the video block to be encoded, with respect to pixel differences, which may be determined by sum of absolute differences (SAD) or sum of squared differences (SSD), or other difference metrics. In some implementations, video encoder 20 may calculate values ​​for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, and other fractional-pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion searches for whole-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.

[0030] Motion estimation unit 42 calculates a motion vector for a video block by comparing the position of the video block in the inter-predictively coded frame with the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy coding unit 56.

[0031] The motion compensation performed by motion compensation unit 44 may include obtaining or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may determine which predictive block in a reference frame list the motion vector points to, obtain the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then subtracts pixel values ​​of the predictive block provided by motion compensation unit 44 from pixel values ​​of the current video block being coded to form a residual block of pixel difference values. The pixel difference values ​​forming the residual block may include luma difference components, chroma difference components, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be integrated together, but are shown separately in FIG. 2 for conceptual purposes.

[0032] In some implementations, the intra BC unit 48 may generate a vector to obtain a predictive block in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is within the same frame as the current block being coded, and the vector is referred to as a block vector rather than a motion vector. In particular, the intra BC unit 48 may determine an intra prediction mode to be used to code the current block. In some embodiments, the intra BC unit 48 may encode the current block using various intra prediction modes, e.g., during separate coding passes, and test their performance using rate-distortion analysis. The intra BC unit 48 may then select and use an appropriate intra prediction mode from among the tested intra prediction modes and generate a corresponding intra mode indicator. For example, the intra BC unit 48 may calculate rate-distortion values ​​for the various tested intra prediction modes using rate-distortion analysis, and select and use the intra prediction mode with the best rate-distortion characteristics from among the tested modes as the appropriate intra prediction mode. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to produce the coded block, as well as the bitrate (i.e., number of bits) used to produce the coded block. Intra BC unit 48 can calculate ratios from the distortions and rates of the various coded blocks to determine which intra-prediction mode exhibits the optimal rate-distortion value for the block.

[0033] In other embodiments, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform the functions for intra BC prediction according to the implementations described herein. In either case, for intra block copying, the predictive block may be a block that is deemed to closely match the block to be coded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identification of the predictive block may include calculating values ​​for sub-integer pixel positions.

[0034] Regardless of whether the predictive block is from the same frame via intra prediction or a different frame via inter prediction, video encoder 20 may form a residual block by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being encoded to form pixel difference values. The pixel difference values ​​that form the residual block may include both luma and chroma component differences.

[0035] Intra-prediction processing unit 46 may intra-predict the current video block, instead of the inter-prediction performed by motion estimation unit 42 and motion compensation unit 44 or the intra-block copy prediction performed by intra BC unit 48, as described above. In particular, intra-prediction processing unit 46 may determine the intra-prediction mode to be used to encode the current block. For example, intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, e.g., during separate-pass encoding, and intra-prediction processing unit 46 (or, in some embodiments, a mode selection unit) may select and use an appropriate intra-prediction mode from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.

[0036] After prediction processing unit 41 determines a predictive block for the current video block via inter- or intra-prediction, adder 50 forms a residual block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 converts the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0037] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be varied by adjusting a quantization parameter. In some embodiments, quantization unit 54 may perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0038] Following quantization, entropy coding unit 56 encodes the quantized transform coefficients into a video bitstream using an entropy coding technique, such as context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be transmitted to video decoder 30, as shown in FIG. 1, or archived to storage device 32, as shown in FIG. 1, for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also use entropy coding techniques to encode motion vectors and other syntax elements for the current video frame being coded.

[0039] Inverse quantization unit 58 and inverse transform processing unit 60 may apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for generating reference blocks for predicting other video blocks. Reconstructed residual blocks may be generated. As described above, motion compensation unit 44 may generate motion-compensated prediction blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction blocks to calculate sub-integer pixel values ​​used for motion estimation.

[0040] Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.

[0041] 3 is a block diagram illustrating an example video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.

[0042] In some embodiments, a unit of video decoder 30 may be responsible for performing implementations of the present disclosure. Also, in some embodiments, implementations of the present disclosure may be divided into one or more units of video decoder 30. For example, intra BC unit 85 may perform implementations of the present disclosure alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some embodiments, video decoder 30 may not include intra BC unit 85, and the functionality of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.

[0043] Video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, obtained via wired or wireless network communication of video data from a local video source such as a camera, or obtained by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. DPB 92 of video decoder 30 stores reference video data used in decoding video data by video decoder 30 (e.g., in intra- or inter-prediction coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For ease of explanation, video data memory 79 and DPB 92 are shown in Figure 3 as two separate components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some embodiments, video data memory 79 may be on-chip with other components of video decoder 30 or may be off-chip with respect to those components.

[0044] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 may use entropy decoding techniques to decode the bitstream to obtain quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.

[0045] If the video frame is coded as an intra-prediction coded (e.g., I) frame, or relative to an intra-coded predictive block of another type of frame, intra prediction unit 84 of prediction processing unit 81 may generate predictive data for the video block of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.

[0046] If a video frame is encoded as an inter-predictively coded (e.g., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more predictive blocks of video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each predictive block may be generated from a reference frame in one of the reference frame lists. Video decoder 30 may construct the reference frame lists, e.g., List 0 and List 1, using a default construction technique based on the reference frames stored in DPB 92.

[0047] In some embodiments, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a predictive block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be within the same reconstructed region of the picture as the current video block processed by video encoder 20.

[0048] Motion compensation unit 82 and / or intra BC unit 85 determine prediction information for video blocks of the current video frame by analyzing the motion vectors and other syntax elements, and then use this prediction information to generate a predictive block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra- or inter-prediction) to use in encoding the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), configuration information for one or more reference frame lists of the frame, the motion vector for each inter-prediction coded video block of the frame, the inter-prediction state for each inter-prediction coded video block of the frame, and other information for decoding the video blocks of the current video frame.

[0049] Similarly, intra BC unit 85 uses some of the received syntax elements, such as flags, to determine that the current video block was predicted using intra BC mode, configuration information regarding which video blocks of the frame are within the reconstructed region and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction states for each intra BC predicted video block of the frame, and other information for decoding the video blocks of the current video frame.

[0050] Motion compensation unit 82 may also perform interpolation using an interpolation filter used by video encoder 20 during encoding of the video block to calculate interpolated values ​​of sub-integer pixels of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter used by video encoder 20 from the received syntax element and use this interpolation filter to generate the predictive block.

[0051] Inverse quantization unit 86 uses the same quantization parameter calculated by video encoder 20 for each video block in a video frame to inverse quantize and determine the degree of quantization of the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct residual blocks in the pixel domain.

[0052] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, summer 90 reconstructs a decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. The decoded video block may be referred to as a reconstructed block of the current video block. An in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF, may be disposed between summer 90 and DPB 92 to further process the decoded video block. In some embodiments, in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by summer 90. The decoded video block in a given frame is then stored in DPB 92, which stores reference frames used for subsequent motion compensation of the next video block. DPB 92, or a memory device separate from DPB 92, may also store the decoded video for later display on a display device, such as display device 34 of FIG.

[0053] In a typical video coding process (e.g., including a video encoding process and a video decoding process), a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In another example, a frame may be monochrome and therefore include only one two-dimensional array of luma samples.

[0054] As shown in FIG. 4A, video encoder 20 (or more specifically, division unit 45) generates a coded representation of a frame by first dividing the frame into a set of CTUs. A video frame may include an integer number of CTUs arranged consecutively in raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, which may be one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that CTUs in this disclosure are not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the nature of different types of units of coded blocks of pixels and how a video sequence may be reconstructed at video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters. In monochrome pictures or pictures with three distinct color planes, a CTU may include a single coding tree block and syntax elements for encoding samples of the coding tree block. A coding tree block may be an NxN block of samples.

[0055] To achieve superior performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree block of a CTU to partition the CTU into smaller CUs. As shown in FIG. 4C , a 64×64 CTU 400 is first partitioned into four smaller CUs, each with a block size of 32×32. Of the four smaller CUs, CU 410 and CU 420 are each partitioned into four CUs with a block size of 16×16. Two 16×16 CUs, 430 and 440, are each further partitioned into four CUs with a block size of 8×8. FIG. 4D shows a quad tree data structure illustrating the final result of the partitioning process of CTU 400 as shown in FIG. 4C , where each leaf node of the quad tree corresponds to one CU with a respective size ranging from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU may include a CB for luma samples, two corresponding coding blocks for chroma samples of the same size frame, and syntax elements for encoding the samples of the coding block. In monochrome pictures or pictures with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. Note that the quadtree partitioning shown in FIGS. 4C and 4D is for illustrative purposes only; a CTU can be divided into CUs based on quadtree / ternary / binary tree partitioning to adapt to various local characteristics. In a multi-type tree structure, a CTU is divided by a quadtree structure, and each quadtree leaf CU can be further divided by binary and ternary tree structures. As shown in FIG. 4E, there are multiple possible division types for a coding block with width W and height H: 4-way partitioning, vertical 2-way partitioning, horizontal 2-way partitioning, vertical 3-way partitioning, vertically extended 3-way partitioning, horizontally extended 3-way partitioning, and horizontally extended 3-way partitioning.

[0056] In some implementations, video encoder 20 may further divide the coding blocks of a CU into one or more M×N PBs. A PB may include rectangular (square or non-square) blocks of samples to which the same prediction, inter or intra prediction, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In monochrome pictures or pictures with three separate color planes, a PU may include a single PB and syntax structures for predicting the PB. Video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma PB, Cb PB, and Cr PB of each PU of a CU.

[0057] Video encoder 20 may use intra prediction or inter prediction to generate the predictive blocks of a PU. If video encoder 20 uses intra prediction to generate the predictive blocks of a PU, it may generate the predictive blocks of the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate the predictive blocks of the PU, it may generate the predictive blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0058] After generating the predicted luma block, Cb block, and Cr block of one or more PUs of a CU, video encoder 20 may generate a luma residual block of the CU by subtracting the predicted luma block of the CU from the original luma coding block such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block of the CU such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0059] Further, as shown in FIG. 4C , video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks, respectively. A transform block may include a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some embodiments, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In monochrome pictures or pictures with three separate color planes, a TU may include a single transform block and syntax structures for transforming the transform block samples.

[0060] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of the TU to generate a Cr coefficient block of the TU.

[0061] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to possibly reduce the amount of data for representing the transform coefficients and provide further compression. After quantizing the coefficient block, video encoder 20 may apply an entropy coding technique to encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a sequence of bits forming a representation of the encoded frame and associated data, which may be stored on storage device 32 or transmitted to destination device 14.

[0062] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks of PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing the coding blocks of each CU of the frame, video decoder 30 may reconstruct the frame.

[0063] As mentioned above, video coding mainly uses two modes, i.e., intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction), to achieve video compression. Note that intra-block copy (IBC) can be considered as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding efficiency more than intra-frame prediction because it uses motion vectors to predict a current video block from a reference video block.

[0064] However, with the advancement of video data collection technology, the video block size that preserves the details of the video data becomes finer, and the amount of data required to represent the motion vectors of the current frame also increases substantially. One way to solve this problem is to benefit from the fact that a group of neighboring CUs in both the spatial and temporal domains not only have similar video data for prediction, but also have similar motion vectors between these neighboring CUs. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by exploring their spatial and temporal correlations, also called the "motion vector predictor (MVP)" of the current CU.

[0065] Instead of encoding the actual motion vector of the current CU into the video bitstream (e.g., determining the actual motion vector by motion estimation unit 42 as described above in connection with FIG. 2), the motion vector predictor for the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. This eliminates the need to encode the motion vector determined by motion estimation unit 42 for each CU of a frame into the video bitstream, significantly reducing the amount of data for representing motion information in the video bitstream.

[0066] Similar to the process of selecting a predictive block in a reference frame during inter-frame prediction of a coding block, both video encoder 20 and video decoder 30 can utilize a set of rules to construct a motion vector candidate list (also called a "merge list") for the current CU using potential candidate motion vectors associated with CUs that are spatially adjacent to and / or temporally co-located with the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. This eliminates the need to transmit the motion vector candidate list itself from video encoder 20 to video decoder 30; the index of a selected motion vector predictor from the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to encode and decode the current CU using the same motion vector predictor in the motion vector candidate list. Therefore, only the index of the selected motion vector predictor needs to be transmitted from video encoder 20 to video decoder 30.

[0067] A brief description of transform coefficient coding in a block-based video coding process (e.g., Enhanced Compression Model (ECM)) is provided herein. Specifically, each transform block is first divided into multiple coefficient groups (CGs), each of which includes a 4x4 sub-block of transform coefficients for luma components and a 2x2 sub-block of transform coefficients for chroma components. The coding of transform coefficients within a transform block is performed on a coefficient group basis. For example, the coefficient groups within a transform block are scanned and coded based on a first predetermined scanning order. When coding each coefficient group, the transform coefficients within the coefficient group are scanned based on a second predetermined scanning order within each sub-block. In ECM, the same top-left scanning order is applied to scan the coefficient groups within a transform block and the different transform coefficients within each coefficient group (e.g., the first predetermined scanning order and the second predetermined scanning order are both top-left scanning orders). Figure 5 is a graphical representation illustrating the top-left scanning order of transform coefficients within a coefficient group according to some embodiments. The numbers 0 through 15 in Figure 5 indicate the corresponding scanning order of each transform coefficient within the coefficient group.

[0068] According to the transform coefficient coding scheme in ECM, a flag is first signaled for each transform block to indicate whether the transform block contains a non-zero transform coefficient. If at least one non-zero transform coefficient exists in the transform block, the position of the last non-zero transform coefficient scanned according to the top-left scanning order is explicitly signaled from video encoder 20 to video decoder 30. After the position of the last non-zero transform coefficient is signaled, flags are further signaled for all coefficient groups coded before the last coefficient group (i.e., the coefficient group including the last non-zero coefficient). Similarly, the number of flags indicates whether each coefficient group contains a non-zero transform coefficient. If the flag for a coefficient group is equal to 0 (indicating that all transform coefficients in the coefficient group are zero), no further information needs to be transmitted for that coefficient group. Otherwise (e.g., if the flag for a coefficient group is equal to 1), the absolute value and sign of each transform coefficient in the coefficient group are signaled in the bitstream according to the scanning order. However, in existing designs, the signs of the transform coefficients are bypass coded (e.g., the context model is not applied), resulting in poor transform coding efficiency. Consistent with the present disclosure, an improved LFNST process involving sign prediction of the transform coefficients is described in more detail below so that the efficiency of transform coding can be improved.

[0069] FIG. 6 is a graphical representation illustrating an LFNST process according to some embodiments. In VVC, a secondary transform tool (e.g., LFNST) is applied to compress the energy of transform coefficients of intra-coded blocks after the primary transform. As shown in FIG. 6, a forward LFNST 604 is applied between a forward primary transform 603 and quantization 605 in video encoder 20, and an inverse LFNST 608 is applied between inverse quantization 607 and an inverse primary transform 609 in video decoder 30. For example, an LFNST process may include both a forward LFNST 604 and an inverse LFNST 608. As some examples, for a 4×4 forward LFNST 604, there may be 16 input coefficients; for an 8×8 forward LFNST 604, there may be 64 input coefficients; for a 4×4 inverse LFNST 608, there may be 8 input coefficients; and for an 8×8 inverse LFNST 608, there may be 16 input coefficients.

[0070] In forward LFNST 604, a non-separable transform with a variable transform size is applied based on the size of the coding block, which can be represented using a matrix multiplication process. For example, assume that forward LFNST 604 is applied to a 4x4 block. The samples in the 4x4 block can be represented using a matrix X as shown in equation (1) below.

number

[0071] The matrix X is expressed as the vector

number

number

[0072] In the above formula (1) or (2), X represents the coefficient matrix obtained by the forward linear transformation 603, and X ijrepresents the linear transformation coefficients in matrix X. Next, the forward LFNST 604 is applied as follows according to Equation (3).

Number

[0073] In the above Equation (3),

Number

Number

Number

[0074] In some implementations, a reduced non-separable transformation kernel can be applied to the LFNST process. For example, based on the above Equation (3), the forward LFNST 604 is based on direct matrix multiplication, which is expensive in terms of computational operations and memory resources for storing transformation coefficients. Therefore, using a reduced non-separable transformation kernel in the LFNST design can reduce the implementation cost of the LFNST process by mapping an N-dimensional vector to an R-dimensional vector in another space (R < N). For example, instead of using an N×N matrix for the transformation kernel, as shown in Equation (4), an R×N matrix is used as the transformation kernel for the forward LFNST 604.

Number

[0075] In the above formula (4), T R×N The R basis vectors in T are generated by selecting the first R basis of the original N-dimensional transformation kernel (i.e., N×N). R×N is orthogonal, the inverse transformation matrix of the inverse LFNST608 is the forward transformation matrix T R×N is the transpose of

[0076] For an 8x8 LFNST, when coefficients N / R=4 are applied, a 64x64 transform matrix is ​​reduced to a 16x48 transform matrix for the forward LFNST 604, and a 64x64 inverse transform matrix is ​​reduced to a 48x16 inverse transform matrix for the inverse LFNST 608. This is achieved by applying the LFNST process to an 8x8 sub-block in the top-left region of primary transform coefficients. Specifically, when a 16x48 forward LFNST is applied, the input is 48 transform coefficients from three 4x4 sub-blocks in the top-left 8x8 sub-block (excluding the bottom-right 4x4 sub-block). In some embodiments, the LFNST process is restricted to be applicable only when all transform coefficients outside the top-left 4x4 sub-block are zero, implying that when LFNST is applied, all primary-only transform coefficients must be zero. Furthermore, to control worst-case complexity (in terms of per-pixel multiplication), the LFNST matrices for 4x4 and 8x8 coding blocks are constrained to 8x16 and 8x48 transforms, respectively. For 4xM and Mx4 coding blocks (M>4), the LFNST non-separable transform matrix is ​​16x16.

[0077] LFNST transform signaling has a total of four transform sets, and the LFNST design enables two non-separable transform kernels per transform set. One transform set is selected from the four transform sets depending on the intra prediction mode of the intra block. The mapping from intra prediction mode to transform sets is predetermined as shown in Table 1 below. If one of the three cross-component linear model (CCLM) modes (e.g., INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81<=intra prediction mode<=83), transform set "0" is selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate is indicated by signaling an LFNST index in the bitstream. Table 1 [Table 1]

[0078] In some embodiments, because LFNST is restricted to be applied to intra-blocks when all transform coefficients outside the first 16x16 sub-block are zero, signaling of the LFNST index depends on the position of the last significant (i.e., non-zero) transform coefficient. For example, for 4x4 and 8x8 coding blocks, the LFNST index is signaled only if the position of the last significant transform coefficient is less than 8. For other coding block sizes, the LFNST index is signaled only if the position of the last significant transform coefficient is less than 16. Otherwise (i.e., if the LFNST index is not signaled), it is inferred that the LFNST index is 0, i.e., LFNST is disabled.

[0079] To reduce the size of the buffer that caches transform coefficients, LFNST is disallowed if the width or height of the current coding block is larger than the maximum transform size (i.e., 64) signaled in the sequence parameter set (SPS). LFNST is only applied when the primary transform is DCT2. LFNST is also applied to intra-coded blocks, both intra- and inter-slice, and to both luma and chroma components. When a binary tree or local tree is enabled (when the division of luma and chroma components is misaligned), LFNST indices are signaled separately for luma and chroma components (i.e., different LFNST transforms can be applied to luma and chroma components). Otherwise, when a single tree is applied (when the division of luma and chroma components is aligned), LFNST is applied only to the luma component, for which a single LFNST index is signaled.

[0080] The LFNST design in ECM is similar to that in VVC, except that an additional LFNST kernel is introduced to improve energy compaction for residual samples with large block sizes. Specifically, when the width or height of a transform block is 16 or more, a new LFNST transform is introduced in the upper-left region of the low-frequency transform coefficients generated from the primary transform. In the current ECM, as shown in Figure 7, the low-frequency region is divided into six 4x4 sub-blocks (e.g., Figure 7) in the upper-left corner of the primary transform coefficients. in 7 contains six 4x4 sub-blocks (shown in grey). In this case, the number of coefficient inputs to the forward LFNST 604 is 96. Furthermore, to control worst-case computational complexity, the number of coefficient outputs of the forward LFNST 604 is set to 32.

[0081] Specifically, for W×H transform blocks (W>=16 and H>=16), a 32×96 forward LFNST is applied, taking 96 transform coefficients from the six 4×4 sub-blocks in the upper-left region as input and outputting 32 transform coefficients. In contrast, the 8×8 LFNST in ECM uses all transform coefficients from the four 4×4 sub-blocks as input and outputs 32 transform coefficients (i.e., a 32×64 matrix for the forward LFNST 604 and a 64×32 matrix for the inverse LFNST 608). This differs from VVC, where the 8×8 LFNST is applied to only the three 4×4 sub-blocks in the upper-left region, producing only 16 transform coefficients (a 16×48 matrix for the forward LFNST 604 and a 48×16 matrix for the inverse LFNST 608). Furthermore, the total number of LFNST sets increases from four in VVC to 35 in ECM. Similar to VVC, the selection of an LFNST set depends on the intra prediction mode of the current coding unit, and each LFNST set contains three different transform kernels.

[0082] In some embodiments, in addition to the DCT2 transform used in HEVC, a Multiple Transform Selection (MTS) scheme is applied to transform the residuals of both inter-coded and intra-coded blocks. The MTS scheme uses multiple transforms selected from the DCT8 and DST7 transforms.

[0083] For example, two control flags are specified at the sequence level to enable the MTS scheme in intra mode and inter mode, respectively. When the MTS scheme is enabled at the sequence level, another CU-level flag is further signaled to indicate whether the MTS scheme is applied. In some implementations, the MTS scheme is applied only to the luma component. Furthermore, the MTS scheme is signaled only if the following conditions are met: (a) both the width and height are less than or equal to 32, and (b) the coded block flag (CBF) is equal to 1. When the MTS CU flag is equal to 0, the DCT2 is applied in both the horizontal and vertical directions. When the MTS CU flag is equal to 1, two other flags are further signaled to indicate the transform type in the horizontal and vertical directions, respectively. The mapping between the MTS horizontal and vertical control flags and the applied transform is shown in Table 2 below. Table 2 [Table 2]

[0084] Regarding the precision of the transform matrix, all MTS transform coefficients have 6-bit precision, the same as the DCT2 core transform. Given that VVC supports all transform sizes used in HEVC, all transform cores used in HEVC remain the same as VVC, including 4-point, 8-point, 16-point, and 32-point DCT-2 transforms and 4-point DST-7 transforms. Other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, and 32-point DST-7 and DCT-8, are also supported in the VVC transform design. Furthermore, to reduce the complexity of large-sized DST-7 and DCT-8, high-frequency transform coefficients located outside the 16x16 low-frequency region are set to 0 (also called zero-out) when either the width or height of the DST-7 or DCT-8 transform block is equal to 32.

[0085] In VVC, in addition to DCT2, only DST7 and DCT8 transform kernels are used for intra-coding and inter-coding. For intra-coding, the statistical properties of the residual signal usually depend on the intra-prediction mode. Additional linear transforms may be useful to handle the diversity of residual properties.

[0086] Additional linear transforms, such as DCT5, DST4, DST1, and identity transform (IDT), are used in ECM. The MTS set is also created depending on the TU size and intra-mode information. Sixteen different TU sizes are considered, with five different classes for each TU size depending on the intra-mode information. For each class, four different transform pairs are considered (same as in VVC). A total of 80 different classes are considered, but some of these different classes often share the same transform set. Therefore, the resulting look-up table (LUT) has 58 unique entries (less than 80).

[0087] For angular modes, the joint symmetry across TU shapes and intra prediction is considered. Thus, mode i (i>34) with TU shape A×B can be mapped to the same class corresponding to mode j=(68-i) with TU shape B×A. However, for each transform pair, the order of the horizontal and vertical transform kernels is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same class, with the vertical and horizontal transform kernels swapped. For wide-angle modes, the closest conventional angular mode is used to determine the transform set. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80.

[0088] Intra-template matching prediction is an example of an intra-prediction mode in which a predicted block whose L-shaped template matches the current template is copied from the reconstructed portion of the current frame. For a given search range, video encoder 20 searches for a template most similar to the current template in the reconstructed portion of the current frame (e.g., based on SAD cost) and uses the corresponding block as the predicted block. Then, video encoder 20 signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is generated by matching the L-shaped causal neighboring blocks of the current block with other blocks within a given search area, as shown in FIG. 8, which includes (a) the current CTU (R1), (b) the top-left CTU (R2), (c) the topmost CTU (R3), and (d) the leftmost CTU (R4). Intra-template matching is valid for CUs whose width and height are 64 or less. Meanwhile, the intra-template matching prediction mode is indicated by signaling a CU-level flag. If intra template matching is applied to a coding block whose width or height is between 4 and 16 (inclusive), the linear transform applied to the corresponding dimension is set to DCT-VII. Otherwise (i.e., if the width or height is less than 4 or greater than 16), DCT-II is applied to the dimension.

[0089] 9 is a graphical representation illustrating an example process of sign prediction, according to some embodiments. In some implementations, sign prediction may aim to estimate the sign of a transform coefficient in a transform block from samples of its neighboring blocks. The accuracy of the estimated sign may be coded according to a context model to indicate whether the sign prediction is correct or not. For example, one context model may be ,eachThe difference between an estimated code and the corresponding true code can be coded with a "0" (or a "1") to indicate that the estimated code is the same (or not the same) as the true code. If the codes can be accurately estimated a high percentage of the time (e.g., 90% or 95% of the codes are correctly estimated), the difference between the estimated code and the true code will tend to be zero, which allows for efficient entropy coding by CABAC compared to bypass-coded codes of transform coefficients in VVC.

[0090] Consistent with some implementations of the present disclosure, to further improve coding efficiency, other context models may be used to entropy code the accuracy of the code prediction. For example, in a first method, the accuracy of the code prediction may be entropy coded using the context of the level values ​​(i.e., intensity) of the transform coefficients. This is because the greater the intensity, the higher the probability of making a correct prediction. For example, when such a method is applied, the range of possible intensity values ​​may be divided into several segments, and different contexts may be assigned to the code prediction entropy coding of coefficients in different segments.

[0091] In a second method, the sign prediction can be entropy coded using the context of the scan position of the transform coefficient. For example, one single context may be applied to coding the transform coefficients located in the first L positions of a transform block, while another context may be applied to coding the signs of coefficients at other positions. Alternatively, all sign prediction positions in a transform block may be classified into different groups according to their respective importance (or probability of correct prediction), and a different context may be assigned to each group separately.

[0092] In a third method, the symbol prediction can be coded using coding-related information such as coding mode context, block size, and component channel information. For example, different contexts can be used for inter and intra modes. In another embodiment, different context models can be used for different blocks. In yet another embodiment, different contexts can be used for luma and chroma components.

[0093] Although the above context modeling schemes are described individually, it is contemplated that any of them can be used together. In fact, each method can be freely combined with the others to achieve different context designs.

[0094] Generally, there is a high correlation between samples at the boundary between a current block and its neighboring blocks, which can be utilized in a sign prediction scheme to predict the signs of the transform coefficients of the current block. Assuming that there are M non-zero transform coefficients in the current block (the M signs are + or -, respectively), as shown in Figure 9, then the total number of possible code combinations is 2 M The code prediction scheme uses each code combination to generate a corresponding hypothesis (e.g., the reconstructed samples at the top and left boundaries of the current block), compares the reconstructed samples in the corresponding hypothesis with extrapolated samples from neighboring blocks, and obtains the sample difference (e.g., SSD or SAD) between the reconstructed and extrapolated samples. The sample difference is minimized (2 M The code combination (among possible code combinations) is selected as the predicted code in the current block.

[0095] In some implementations, for each combination of M codes, M corresponding transform coefficients may be processed by an inverse quantization operation and an inverse transform to generate a corresponding hypothesis to obtain residual samples, which may be added to prediction samples to obtain reconstructed samples, including reconstructed samples at the top and left boundaries of the current block (shown in the L-shaped gray region 902).

[0096] In some implementations, a cost function that measures the spatial discontinuity between samples at the boundary of the current block and its neighboring blocks is used to select the code combination. Instead of using the L2 norm (SSD), the cost function may be based on the L1 norm (SAD), as shown in equation (5) below.

number

[0097] In the above formula (5), B i,n (i=-2,-1) represents the adjacent sample from the adjacent block on top of the current block. C m,j (j=-2,-1) represents the adjacent sample from the leftmost adjacent block of the current block. 0,n and P m,n , and , respectively denote the corresponding reconstructed samples at the top and left edges of the current block. N and M denote the width and height of the current block, respectively. Figure 10 shows the corresponding samples P of the current block used to calculate the cost function for symbol prediction. 0,n and P m,0 , and the corresponding sample B of the adjacent block i,n and C m,j Shows.

[0098] In some implementations, to avoid the complexity of performing multiple inverse transforms, a template-based hypothesis reconstruction method can be applied to the code prediction scheme. Each template may be a set of reconstructed samples at the top and left boundaries of the current block, and can be obtained by applying the inverse transform to a coefficient matrix with certain coefficients set to 1 and all other coefficients equal to 0. Assuming that the inverse transform (e.g., DCT, DST) is linear, the corresponding hypothesis can be generated by a linear combination of a set of pre-computed templates.

[0099] In some implementations, the predicted codes are grouped into two sets, each coded by a single CABAC context. For example, the first set includes predicted codes for transform coefficients in the upper left corner of the transform block, and the second set includes predicted codes for transform coefficients in all other positions of the transform block.

[0100] Like HEVC, VVC uses scalar quantization. In some implementations, scalar quantization in VVC may be implemented as dependent scalar quantization. Dependent scalar quantization refers to an approach in which the set of allowable reconstructed values ​​of a transform coefficient depends on the value of the transform coefficient level preceding the current transform coefficient level in reconstruction order. The main effect of this approach is that allowable reconstructed vectors are more densely packed in an N-dimensional vector space (N represents the number of transform coefficients in a transform block) compared to traditional independent scalar quantization used in HEVC. That is, for a constant average number of allowable reconstructed vectors per unit volume in N dimensions, the average distortion between an input vector and its nearest reconstructed vector is reduced.

[0101] Dependent scalar quantization may be implemented by (a) defining two scalar quantizers with different reconstruction levels and (b) defining a process for switching between the two scalar quantizers. FIG. 11 shows two example scalar quantizers used for VVC dependent scalar quantization according to some implementations of this disclosure. As shown in FIG. 11, the VVC quantization design applies two scalar quantizers, denoted by Q0 and Q1. The positions of the available reconstruction levels are uniquely specified by the quantization step size Δ. In this implementation, the selection between the two scalar quantizers Q0 and Q1 is not explicitly signaled in the bitstream. Instead, the quantizer used for a current transform coefficient is determined by the parity of the transform coefficient level preceding the current transform coefficient in encoding order by video encoder 20 or reconstruction order by video decoder 30.

[0102] In some implementations, switching between the two scalar quantizers is performed via a state machine. For example, Figure 12A is a graphical representation illustrating state transitions using a four-state state machine used for dependent scalar quantization, according to some embodiments. As shown in Figure 12, the state can take on four different values: 0, 1, 2, and 3, which are uniquely determined by the parity of the transform coefficient level that precedes the current transform coefficient in encoding / reconstruction order.

[0103] In some implementations, at the start of dequantization of a transform block, the state is set to 0. The transform coefficients are reconstructed in scan order (i.e., in the same order as entropy decoding).

[0104] After the current transform coefficient is reconstructed, the state is updated according to the state machine. For example, in FIG. 12A, k indicates the value of the transform coefficient level. At each state, the next state is determined based on the parity of the transform coefficient level k, i.e., (k&1). The next state when (k&1)==1 is different from the next state when (k&1)==0. As shown in FIG. 12A, the state machine includes two arrows pointing from each of the four states to two different states. FIG. 12B is a table illustrating exemplary quantizer selection according to the state transitions of FIG. 12A. For example, according to FIGS. 12A and 12B, at state 1, when (k&1)==0, the next state is 2, and when (k&1)==1, the next state is 0.

[0105] Similarly, at the decoder, the reconstructed quantization index of one transform coefficient can be calculated according to equation (6). quantIdx=(abs(k)<<1)-(state&1) (6) where abs() is a function that calculates the absolute value of the input, and state is the current state of the state machine when analyzing the level of the current transform coefficient. At the decoder side, the reconstructed transform coefficients after dequantization can be obtained according to Equation (7). transCoeff=quantIdx·Δ (7)

[0106] Several exemplary deficiencies present in current designs of sign prediction schemes are identified herein. In a first example, sign prediction in current ECMs is only applicable to predicting the signs of transform coefficients in transform blocks to which only a linear transform (e.g., a DCT and a DST transform) is applied. As mentioned above, applying LFNST to transform coefficients from a linear transform can improve energy compaction of residual samples of intra-coded blocks. However, in current ECM designs, sign prediction is bypassed for transform blocks to which LFNST is applied.

[0107] In the second example, a predetermined maximum number of predicted symbols ("L") is used for the transform block. max”) to control the complexity of the code prediction. In current ECMs, video encoders determine the maximum number (e.g., L max = 8) and transmits the value to the video decoder. Furthermore, the video encoder or decoder may scan all transform coefficients for each transform block in raster scan order, and the first L max The non-zero transform coefficients are selected as candidate transform coefficients for sign prediction. Treating different transform coefficients in a transform block equally in this way may not be optimal in terms of the accuracy of sign prediction. For example, for transform coefficients with relatively large magnitudes, predicting their signs is more likely to result in a correct prediction. This is because using the wrong sign for these transform coefficients tends to have a larger impact on reconstructed samples at block boundaries than using transform coefficients with relatively small magnitudes.

[0108] In a third example, rather than directly encoding explicit code values, a video encoder or decoder can encode the accuracy of the predicted code. For example, for a transform coefficient with a positive sign, if its predicted code is also positive, only bin "0" needs to be indicated in the bitstream from the video encoder to the video decoder. In this case, the predicted code is the same as the true code (or original code) of the transform coefficient, indicating that the code prediction for this transform coefficient is correct. Otherwise (e.g., when the predicted code is negative but the true code is positive), the bitstream from the video encoder to the video decoder may include bin "1." If all codes are predicted correctly, the corresponding bins indicated in the bitstream are 0, which can be efficiently entropy coded by CABAC. If some codes are predicted incorrectly, the corresponding bins indicated in the bitstream are 1. Although arithmetic coding plus an appropriate context model is efficient for coding bins according to their corresponding probabilities, there are still non-negligible bits generated in the bitstream to indicate code values.

[0109] In the fourth example, the spatial discontinuity between samples at the boundary between the current block and its neighboring blocks is used to select the best code prediction combination in the current design of code prediction in ECM. To capture the spatial discontinuity, the L1 norm of gradient differences along the vertical and horizontal directions is utilized. However, since the distribution of image signals is usually non-uniform, using only the vertical and horizontal directions may not accurately capture the spatial discontinuity.

[0110] Consistent with this disclosure, video processing methods and systems for symbol prediction in block-based video coding are provided herein to address one or more of the above-mentioned exemplary deficiencies. The methods and systems disclosed herein can improve coding efficiency of symbol prediction while allowing for ease of use in hardware codec implementations. The methods and systems disclosed herein can improve coding efficiency of transform blocks that apply symbol prediction techniques to transform coefficients of the transform blocks.

[0111] For example, as described above, sign prediction can predict the signs of transform coefficients within a transform block based on the correlation between boundary samples (also called edge samples) located at or near the boundary between the transform block and its spatially adjacent blocks. Given that the existence of correlation does not depend on which specific transform is applied, the two coding tools (i.e., LFNST and sign prediction) do not interfere with each other and can be applied jointly. Furthermore, because LFNST further compresses the energy of the transform coefficients of a linear transform, sign prediction of LFNST transform coefficients is more accurate than that of a linear transform. This is because an incorrect sign prediction of a transform coefficient from LFNST can cause greater discrepancies in the smoothness of boundary samples. Therefore, consistent with the present disclosure, a harmonization scheme is disclosed herein that enables the combination of LFNST and sign prediction to improve the coding efficiency of transform coefficient coding. Furthermore, a template-based hypothesis generation scheme is also disclosed herein that reconstructs edge samples for different combinations of predicted signs to reduce the number of inverse transforms.

[0112] In another example, for selecting transform coefficient candidates for sign prediction as described above, instead of treating different transform coefficients in a transform block equally, a higher weight may be given to transform coefficients if their signs are more easily predicted, which may result in inconsistencies between boundary samples of adjacent blocks. Consistent with the present disclosure, the methods and systems disclosed herein may select transform coefficient candidates for sign prediction (e.g., transform coefficients whose signs are predicted for a transform block) based on one or more selection criteria to improve the accuracy of sign prediction. For example, transform coefficients that have a greater influence on reconstructed border samples (rather than transform coefficients that have a smaller influence on reconstructed border samples) are selected as transform coefficient candidates for sign prediction, thereby improving the accuracy of sign prediction.

[0113] In yet another example, when the signs of transform coefficients in a transform block are predicted with high accuracy (e.g., when the accuracy of the predicted signs is higher than a threshold, such as 80% or 90%), there is a strong correlation between the boundary samples of the transform block and its neighboring blocks. In this case, it commonly occurs that there may be consecutive transform coefficients (e.g., particularly several non-zero transform coefficients at the beginning of the transform block) that can be correctly predicted in many scenarios. In such scenarios, a single bin (instead of multiple bins) can be used to indicate whether the signs of all consecutive transform coefficients are correctly predicted, thereby saving the signaling overhead of code prediction. Consistent with the present disclosure, a vector-based code prediction scheme that reduces the signaling overhead of code prediction is disclosed herein. Unlike existing code predictions that individually predict the signs of each non-zero transform coefficient, the disclosed vector-based code prediction scheme groups a set of consecutive non-zero transform coefficient candidates and predicts their corresponding signs jointly, thereby efficiently reducing the average number of bins (or bits) used to indicate the accuracy of the predicted signs.

[0114] In yet another example, using only the vertical and horizontal directions may not accurately capture spatial discontinuities between samples at the boundary between a current block and its adjacent blocks. This allows for more directions to be introduced to more accurately capture spatial discontinuities. Consistent with this disclosure, an improved cost function is disclosed herein that considers both vertical and horizontal gradients and diagonal gradients to more accurately capture spatial discontinuities.

[0115] 13 is a block diagram illustrating an example symbol prediction process 1100 in block-based video coding according to some implementations of the present disclosure. In some implementations, the symbol prediction process 1300 may be performed by transform processing unit 52. In some implementations, the symbol prediction process 1300 may be performed by one or more processors (e.g., one or more video processors) of video encoder 20 or decoder 30. Throughout this disclosure, LFNST is used as an example of a secondary transform without loss of generality. It is contemplated that other examples of secondary transforms may also be applicable herein.

[0116] In existing designs of ECM, sign prediction is disabled for transform blocks to which LFNST is applied. However, the principle of sign prediction is to predict the sign of transform coefficients based on the correlation between border samples of a transform block and its spatially neighboring blocks, and is independent of the specific transform type (e.g., linear or secondary) or transform core (e.g., DCT or DST) applied to the transform block. Therefore, in this specification, sign prediction and LFNST can be applied together to further improve the efficiency of transform coding. Consistent with the present disclosure, the sign prediction process 1300 can be applied to predict the sign of transform coefficients in a transform block to which a linear transform and a secondary transform are jointly applied.

[0117] An exemplary overview of the code prediction process 1300 is provided herein. First, the code prediction process 1300 may perform a coefficient generation operation 1302 by applying a primary transform and a secondary transform to a transform block of a video frame from a video to generate transform coefficients for the transform block. Next, the code prediction process 1300 may perform a coefficient selection operation 1304 by selecting a set of transform coefficient candidates from the transform coefficients for code prediction. Subsequently, the code prediction process 1300 may perform a hypothesis generation operation 1306 by applying a template-based hypothesis generation scheme to select one hypothesis from multiple candidate hypotheses for the set of transform coefficient candidates. Furthermore, the code prediction process 1300 may perform the code generation operation 1108 by determining a combination of code candidates associated with the selected hypothesis that results in a set of predicted codes for the set of transform coefficient candidates. Operations 1302, 1304, 1306, and 1308 are each described in more detail below.

[0118] For example, transform processing unit 52 of video encoder 20 may transform the residual video data into transform coefficients of a transform block by jointly applying a primary transform and a secondary transform (e.g., jointly applying forward primary transform 603 and forward LFNST 604 as shown in FIG. 6 ). A predetermined number (e.g., L) of non-zero transform coefficients from the transform coefficients of the transform block may be selected as candidate transform coefficients based on one or more selection criteria described below, where 1≦L≦the maximum number of predictable codes. A template-based hypothesis generation scheme may then be applied to generate multiple candidate hypotheses using combinations of different code candidates for each of the L candidate transform coefficients, thus generating a total of 2 LL candidate hypotheses can be generated. Each candidate hypothesis may include reconstructed samples at the top and left boundaries of the transform block. A cost function incorporating combined gradients along the horizontal, vertical, and diagonal directions can then be used to calculate the cost of each candidate hypothesis reconstruction. A candidate hypothesis associated with the smallest cost can be determined from the multiple candidate hypotheses as a hypothesis for predicting signs of the L candidate transform coefficients. For example, the combination of sign candidates used to generate the candidate hypothesis associated with the smallest cost is used as the predicted sign of the L candidate transform coefficients.

[0119] First, the symbol prediction process 1300 may perform a coefficient generation operation 1302, in which a primary transform (e.g., DCT, DST, etc.) and a secondary transform (e.g., LFNST) may be jointly applied to a transform block to generate transform coefficients of the transform block. For example, a primary transform may be applied to the transform block to generate primary transform coefficients of the transform block. Then, an LFNST may be applied to the transform block to generate LFNST transform coefficients based on the primary transform coefficients.

[0120] Subsequently, the symbol prediction process 1300 may perform a coefficient selection operation 1304, which may select a set of candidate transform coefficients for symbol prediction from the transform coefficients of the transform block based on one or more selection criteria. The selection of candidate transform coefficients may maximize the number of candidate transform coefficients that can be correctly predicted, thereby improving the accuracy of symbol prediction.

[0121] In some implementations, the set of candidate transform coefficients may be selected from the transform coefficients of the transform block based on the strength of the transform coefficients. For example, the set of candidate transform coefficients may include one or more transform coefficients that have a strength greater than the remaining transform coefficients in the transform block.

[0122] In general, the stronger the magnitude of transform coefficients, the more likely the predicted signs of these transform coefficients are to be correct. This is because the stronger the magnitude of these transform coefficients, the more likely they are to have an impact on the quality of reconstructed samples, and using incorrect signs for these transform coefficients may easily cause discontinuities between boundary samples between the transform block and its spatially adjacent blocks. Based on this rationale, a set of candidate transform coefficients for sign prediction can be selected from the transform coefficients of a transform block based on the strengths of the non-zero transform coefficients within the transform block.

[0123] There are several possible ways to implement an intensity-based sorting scheme of transform coefficients for code prediction. In a first implementation, this scheme directly uses the reconstructed transform coefficients after inverse quantization (i.e., the inverse quantized transform coefficients) for sorting, so that inverse quantized transform coefficients with greater intensities are placed before transform coefficients with smaller intensities for code prediction. For example, all non-zero transform coefficients in a transform block can be scanned and sorted to create a coefficient list according to their descending order of intensity. The transform coefficient with the greatest intensity can be selected from the coefficient list and placed as the first transform coefficient candidate in a set of transform coefficient candidates, and the transform coefficient with the second greatest intensity can be selected from the coefficient list and placed as the second transform coefficient candidate in the set of transform coefficient candidates, and this is repeated until the number of selected transform coefficient candidates reaches a predetermined number L.

[0124] In the second implementation, instead of directly using the dequantized transform coefficients, the quantization index of each transform coefficient (e.g., quantIdx obtained according to equation (6)) is used to represent the intensity of the transform coefficient and can be used for such sorting. As shown in equation (7), the value of one dequantized transform coefficient is equal to the product of its quantization index quantIdx and the corresponding step size Δ, and the step size is the same for the dequantization of all transform coefficients in one transform block. Therefore, the two implementations are actually mathematically identical. However, if the quantization index quantIdx can be obtained in the analysis stage (earlier than the acquisition of the dequantized transform coefficients), the second implementation may offer certain advantages when implemented by some specific hardware.

[0125] In some implementations, a set of candidate transform coefficients may be selected from the transform coefficients of a transform block based on a coefficient scanning order for entropy coding applied to video coding. Because natural video content contains abundant low-frequency information, the magnitude of non-zero transform coefficients obtained from processing the video content tends to be large at low-frequency locations and small toward high-frequency locations. Therefore, modern video codecs can scan and entropy code the transform coefficients in a transform block using a coefficient scanning order (such as a zigzag scan, a top-left scan, a horizontal scan, or a vertical scan). By using this coefficient scanning order, transform coefficients with larger magnitudes (usually corresponding to lower frequencies) are scanned before transform coefficients with smaller magnitudes (usually corresponding to higher frequencies). Based on this rationale, a set of candidate transform coefficients for code prediction disclosed herein can be selected from the transform coefficients of a transform block based on a coefficient scanning order for entropy coding. For example, a coefficient list can be obtained by scanning all transform coefficients in the transform block using the coefficient scanning order. The first L non-zero transform coefficients in the coefficient list can then be automatically selected as a set of candidate transform coefficients for sign prediction.

[0126] In some implementations, for an intra-coded block, a set of candidate transform coefficients for code prediction can be selected from the transform coefficients of the block based on the intra-prediction direction of the block. For example, multiple scan orders coherent with the intra-prediction direction (e.g., 67 intra-prediction directions in VVC and ECM) can be determined and stored as lookup tables in both the video encoder 20 and the video decoder 30. When encoding the transform coefficients of the intra block, the video encoder 20 or the video decoder 30 can identify, from the scan order, a scan order that is closest to the intra-prediction of the intra block. The video encoder 20 or the video decoder 30 can scan all non-zero transform coefficients of the intra block using the identified scan order to obtain a coefficient list, and select the first L non-zero transform coefficients from the coefficient list as a set of candidate transform coefficients.

[0127] In some implementations, video encoder 20 may determine a scanning order for transform coefficients of a transform block and signal the determined scanning order to video decoder 30. One or more new syntax elements indicating the determined scanning order may be signaled through the bitstream. For example, multiple fixed scanning orders (e.g., for different transform block sizes and coding modes) may be predetermined by video encoder 20 and pre-shared with video decoder 30. Then, after selecting a scanning order from the fixed scanning orders, video encoder 20 need only signal a single index to indicate the selected scanning order to video decoder 30. In another embodiment, one or more new syntax elements may be used to enable signaling of any selected scanning order of transform coefficients. In some implementations, one or more syntax elements may be signaled at various coding levels, such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture (or slice) level, or a CTU (or CU) level.

[0128] In some implementations, a set of candidate transform coefficients may be selected from the transform coefficients of a transform block based on their influence scores on the reconstructed edge samples of the transform block. Specifically, as shown in equation (5) above, the selection of the code combination (i.e., predicted code or code predictor) is based on a cost function for minimizing the discontinuity of the gradient of samples between the current transform block and its spatially adjacent blocks. Therefore, the signs of transform coefficients that have a relatively large influence on the reconstructed samples at the top and left edges of the current transform block tend to be more likely to be accurately predicted because inverting these signs can result in a large variation in the smoothness between the boundary samples, as calculated in equation (5). To maximize the rate of accurate code prediction, the signs of these transform coefficients (i.e., those that have a large influence on the reconstructed edge samples) may be predicted before other transform coefficients (i.e., those that have a small influence on the reconstructed edge samples). Based on this rationale, the set of candidate transform coefficients for code prediction disclosed herein may be selected based on their influence scores on the reconstructed samples at the top and left edges of the current transform block.

[0129] For example, video encoder 20 or decoder 30 may sort all transform coefficients based on a measure of their corresponding influence scores on the reconstructed border samples of the transform block. If a transform coefficient has a large influence score on the reconstructed border samples, it may be assigned a small index in the code prediction candidate list because it is easier to predict correctly. The set of transform coefficient candidates disclosed herein may be the L transform coefficients with the L smallest indices in the code prediction candidate list.

[0130] In some implementations, different criteria can be applied to quantify the influence score of a transform coefficient on the reconstructed edge samples. For example, a value obtained by measuring the energy of the fluctuation of the reconstructed edge samples caused by the transform coefficient can be used as the influence score (in the L1 norm), which can be obtained as follows:

number

[0131] The above formula ( 8 ) then, C i,j represents the transform coefficient at position (i,j) in the transform block. i,j (l,k) is the conversion coefficient C i,j where N and M represent the width and height of the transform block, respectively. V represents the influence score of the transform coefficient at position (i,j).

[0132] In another example, the L1 norm in equation (8) above can be replaced with the L2 norm, so that the influence score (e.g., a measure of the energy of the variation in the reconstructed border samples caused by the transform coefficients) can be calculated using the L2 norm as follows:

number

[0133] Consistent with the present disclosure, in the above formulas (8) and (9), (e.g., T i,j (0,n) and T i,j Although only the top and left edge samples (denoted by (m,0)) are used in the calculation, the transform coefficient selection scheme disclosed herein can be applied to any code prediction scheme by changing the reconstructed samples of the current transform block used in the corresponding cost function.

[0134] Similar to the transform coefficient intensity-based scheme, there are various possible ways to implement the influence score-based scheme. In a first implementation, the scheme is implemented by multiplying the dequantized transform coefficients C i,j In a second implementation, the quantization index value quantIdx is applied instead to obtain the intensities of the dequantized transform coefficients C i,j Specifically, when the quantization index value quantIdx is applied, equations (8) and (9) become equations (10) and (11) below.

number

number

[0135] The two methods are in fact mathematically identical, since the value of one dequantized transform coefficient is equal to the product of its quantization index quantIdx and the corresponding step size Δ, and the step size is the same for the dequantization of all transform coefficients in one transform block.

[0136] Subsequently, the code prediction process 1300 may perform a hypothesis generation operation 1106, in which a template-based hypothesis generation scheme may be applied to select one hypothesis from multiple candidate hypotheses for the set of candidate transform coefficients. Initially, a combination of multiple candidate code codes may be determined for the set of candidate transform coefficients based on the total number of coefficients included in the set of candidate transform coefficients. For example, if there are a total of L candidate transform coefficients, the combination of multiple candidate code codes for the set of candidate transform coefficients may be determined ... LEach code candidate may be a combination of L code candidates. Each code candidate may be a negative sign (-) or a positive sign (+). Each code candidate combination may include a total of L negative signs or positive signs. For example, when L=2, the combination of multiple code candidates may be 2, which are (+,+), (+,-), (-,-), and (-,-), respectively. 2 = 4 code candidate combinations.

[0137] A template-based hypothesis generation scheme can then be applied to generate multiple candidate hypotheses for each combination of multiple code candidates. To reduce the complexity of the inverse primary and secondary transforms that need to be performed, the template-based hypothesis generation scheme disclosed herein can be used to optimize the generation of reconstructed edge samples for the transform block. Two exemplary approaches for implementing the template-based hypothesis generation scheme are disclosed herein. Other exemplary approaches for implementing the template-based hypothesis generation scheme are contemplated as possible, and this is not a limitation of this specification.

[0138] In the first exemplary approach, a candidate hypothesis corresponding to each combination of code candidates can be generated based on a linear combination of templates, thereby generating multiple candidate hypotheses for each combination of code candidates. Each template may correspond to a candidate transform coefficient from a set of candidate transform coefficients. Each template may represent reconstructed samples at the top and left boundaries of a transform block. Each template may be generated by applying an inverse secondary transform and an inverse linear transform to the transform block, where each candidate transform coefficient in the set of candidate transform coefficients is set to 0 except for the candidate transform coefficient corresponding to the template, which is set to 1 (e.g., the candidate transform coefficient corresponding to the template is set to 1, while the remaining candidate transform coefficients are set to 0).

[0139] For example, the candidate hypotheses corresponding to each code candidate combination can be set to be a linear combination of templates. For the templates corresponding to each candidate transform coefficient, the weight of each template can be set to the magnitude of the dequantized transform coefficient corresponding to each candidate transform coefficient. An example of hypothesis generation based on a linear combination of templates is shown in FIG. 14 and described in more detail below.

[0140] To predict the signs of candidate transform coefficients, video encoder 20 or decoder 30 may traverse all candidate hypotheses before identifying a hypothesis associated with a combination of candidate codes that can minimize a cost value calculated from a cost function. In the first exemplary approach described above, each candidate hypothesis can be generated based on a combination of multiple templates, but the sample-by-sample computations (e.g., additions, multiplications, and shifts) associated with such combinations are relatively complex. To reduce the computational complexity associated with identifying a hypothesis that minimizes a cost value calculated from a cost function, a second exemplary approach is introduced herein.

[0141] In the second exemplary approach, combinations of multiple code candidates associated with multiple candidate hypotheses can be treated as multiple hypothesis indices for the multiple candidate hypotheses. For example, digital 0 and 1 can be configured to represent a positive sign (+) and a negative sign (-), respectively. A combination of code candidates corresponding to a candidate hypothesis can be used as a unique representation (i.e., a hypothesis index) for the candidate hypothesis. For example, assume there are three predicted codes (e.g., L=3). Hypothesis index 000 can represent a candidate hypothesis generated by setting all three code candidates to positive (e.g., the three code candidates are (+,+,+)). Similarly, hypothesis index 010 can represent a candidate hypothesis generated by setting the first and third code candidates to positive and the second code candidate to negative (e.g., the three code candidates are (+,-,+)).

[0142] Next, multiple candidate hypotheses can be generated based on the Gary code order of the multiple hypothesis indexes, whereby a reconstructed sample of a previous candidate hypothesis having a previous hypothesis index can be used to generate a current candidate hypothesis having a current hypothesis index. The current hypothesis index of the current candidate hypothesis can be immediately after the previous candidate hypothesis index of the previous candidate hypothesis according to the Gary code order of the multiple hypothesis indexes. The current hypothesis index can be generated by changing the sign candidate associated with the previous hypothesis index from positive (or negative) to negative (or positive). For example, the current hypothesis index can be obtained by changing a single "0" (or "1") in the previous hypothesis index to "1" (or "0").

[0143] For example, the multiple hypothesis indexes may be sorted based on the Gary code order of the multiple hypothesis indexes to generate a sequence of sorted hypothesis indexes. For a first hypothesis index in the sequence of sorted hypothesis indexes, a first candidate hypothesis corresponding to the first hypothesis index may be generated by applying an inverse secondary transform and an inverse linear transform to a transform block in which each transform coefficient candidate of the set of transform coefficient candidates is set to 1. For a second hypothesis index immediately following the first hypothesis index in the sequence of sorted hypothesis indexes, a second candidate hypothesis corresponding to the second hypothesis index may be generated based on (a) the first candidate hypothesis corresponding to the first hypothesis index and (b) an adjustment term of the second candidate hypothesis. Table 3 below shows an example process for generating all candidate hypotheses for an LFNST when the number of transform coefficient candidates is three (e.g., L=3). Table 3 [Table 3]

[0144] In Table 3 above, the first column is 3The first column shows the combinations of 000, 001, 011, 010, 110, 111, 101, 100, respectively. The second column shows the hypothesis indexes corresponding to the combinations of the code candidates, using digital 0 and 1 to represent positive (+) and negative (-) signs. The hypothesis indexes in the second column are ordered according to Gary code order (e.g., 000, 001, 011, 010, 110, 111, 101, 100). The third column shows the candidate hypotheses corresponding to the combinations of the code candidates and the hypothesis indexes. The fourth column shows the calculation of the candidate hypotheses.

[0145] In Table 3, TXYZ in the fourth column represent corresponding templates (i.e., reconstructed samples at the top and left boundaries of a transform block), which can be generated by applying an inverse transform to a coefficient matrix of a transform block in which certain transform coefficients are set to 1 and all other transform coefficients are equal to 0. For example, T100 represents a corresponding template generated by applying an inverse transform to a coefficient matrix in which only the transform coefficient corresponding to the first code candidate is set to 1 and all other transform coefficients in the coefficient matrix are set to 0. C0, C1, and C2 represent the absolute values ​​of the dequantized transform coefficients associated with the first, second, and third code candidates, respectively.

[0146] Referring to Table 3, for the first hypothesis index 000, a first candidate hypothesis H000 is generated by applying an inverse secondary transform and an inverse linear transform to a coefficient matrix associated with a transform block with each candidate transform coefficient set to 1. For the second hypothesis index 001, which immediately follows the first hypothesis index 000, the second candidate hypothesis H001 may be generated based on (a) the first candidate hypothesis H000 and (b) an adjustment term for the second candidate hypothesis (e.g., -C2*T001). Similarly, for the third hypothesis index 011, which immediately follows the second hypothesis index 001, the third candidate hypothesis H011 may be generated based on (a) the second candidate hypothesis H001 and (b) an adjustment term for the third candidate hypothesis (e.g., -C1*T010). For the fourth hypothesis index 010, which immediately follows the third hypothesis index 011, the fourth candidate hypothesis H010 can be generated based on (a) the third candidate hypothesis H011 and (b) the adjustment term of the fourth candidate hypothesis (e.g., C2*T001).

[0147] The hypothesis associated with the minimum cost can then be determined from multiple candidate hypotheses based on a cost function incorporating joint gradients along the horizontal, vertical, and diagonal directions. As mentioned above, if the cost function only utilizes gradients along the horizontal and vertical directions (e.g., as shown in Equation (5) above), it may not work well for image signals with high heterogeneity. Consistent with the present disclosure, gradients along one or more diagonal directions are also utilized to improve the accuracy of the cost function. For example, two diagonal directions, including a left diagonal direction and a right diagonal direction, may be incorporated into the cost function. For example, the cost function for the two diagonal directions can be described according to Equations (12) and (13) below.

number

number

[0148] The above formula ( 12 ) or ( 13 ) So, B -1,n-1 , B-2,n-2 , B -1,n+1 , and B -2,n+2 represents the adjacent sample from the neighboring block at the top edge of the transform block. m-1,-1 , C m-2,-2 , C m+1,-1 , and C m+2,-2 represents the adjacent sample from the leftmost adjacent block of the transform block. 0,n and P m,0 represent the reconstructed samples at the top and left boundaries of the transform block, respectively. N and M represent the width and height of the transform block, respectively. costD1 and costD 2 represent the left diagonal cost function and the right diagonal cost function for the left diagonal direction and the right diagonal direction, respectively.

[0149] Two diagonal cost functions (e.g., costD1 and costD2) can be used jointly with horizontal and vertical cost functions (e.g., costHV shown in Equation (5) above). The cost function for sign prediction can then be determined based on the horizontal and vertical cost functions incorporating gradients along the horizontal and vertical directions, the left diagonal cost function incorporating gradients along the left diagonal direction, and the right diagonal cost function incorporating gradients along the right diagonal direction. For example, the cost function may be a weighted sum of the horizontal and vertical cost functions, the left diagonal cost function, and the right diagonal cost function, as described in Equation (14). cost=costHV+ω(costD1+costD2) (14)

[0150] In the above equation (14), ω represents the weights of the left and right diagonal cost functions.

[0151] In another example, the cost function may be the minimum of the horizontal and vertical cost functions, the left diagonal cost function, and the right diagonal cost function, as described in equation (15). cost=min{costHV,costD1,costD2} (15)

[0152] Compared to equation (5) above, the cost functions of equations (14) or (15) disclosed herein may require more neighboring pixels to support the cost functions costD1 and costD2 along the diagonal direction, which will be explained in more detail below with reference to Figures 16A and 16B.

[0153] In some implementations, a cost corresponding to each candidate hypothesis can be determined using equation (14) or (15) above. Then, multiple costs can be calculated for each of the multiple candidate hypotheses. A minimum cost can be determined from the multiple costs. The candidate hypothesis associated with the minimum cost can be determined from the multiple candidate hypotheses and selected as the hypothesis for code prediction.

[0154] Subsequently, the code prediction process 1300 may perform a code generation operation 1108, in which a combination of code candidates associated with the selected hypothesis is determined to result in a set of predicted codes for the set of candidate transform coefficients. For example, the combination of code candidates (e.g., L code candidates) used to generate the selected hypothesis can be used as the predicted codes for the L candidate transform coefficients.

[0155] In some implementations, the code generation operation 1308 may include applying a vector-based code prediction scheme to the set of predicted codes to generate a sequence of code signaling bits for the set of candidate transform coefficients. A bitstream including the sequence of code signaling bits may be generated by the video encoder 20 and stored in the storage device 32 of FIG. 1. Alternatively, or in addition, the bitstream may be transmitted to the video decoder 30 via the link 16 of FIG. 1.

[0156] As mentioned above, if the sign of a transform coefficient in a transform block is predicted properly, it is highly likely that the signs of multiple consecutive transform coefficients can be correctly predicted. In this case, the signaling scheme of existing code prediction designs is obviously inefficient in terms of overhead for signaling the code value of a transform block, because it needs to signal multiple bins "0" to individually indicate that the corresponding sign of each transform coefficient can be correctly predicted. An exemplary implementation of the existing code prediction scheme will be described in more detail below with reference to Figure 15A.

[0157] Consistent with the present disclosure, the efficiency of code signaling can be improved by applying the vector-based code prediction scheme disclosed herein. Specifically, transform coefficient candidates of a transform block can be divided into multiple groups, and the codes of the transform coefficient candidates within each group can be jointly predicted. In this case, if the original codes (or true codes) of the transform coefficient candidates within a group are the same as their respective predicted codes, only bins with a value of “0” need to be transmitted in the bitstream to indicate that all codes within the group are correctly predicted. Otherwise (i.e., if there is at least one transform coefficient candidate within the group whose original code is different from the predicted code), bins with a value of “1” can be signaled first in the bitstream to indicate that not all codes of the transform coefficient candidates within the group are correctly predicted. Then, additional bins can also be signaled in the bitstream from the video encoder 20 to the video decoder 30 to indicate the corresponding accuracy of each predicted code within the group. An exemplary implementation of the vector-based code prediction scheme disclosed herein is described in more detail below with reference to FIG. 15B.

[0158] In some implementations, the set of transform coefficient candidates can be divided into multiple groups of transform coefficient candidates. For each group of transform coefficient candidates, one or more code signaling bits can be generated for the group of transform coefficient candidates to indicate the accuracy of the predicted code. As described above, various context models can be used to encode the accuracy of the code prediction, and the code signaling bits can be generated according to the context model used for the coding block.

[0159] In one embodiment, a sign signaling bit may be generated based on whether the original sign of the candidate set of transform coefficients is the same as the predicted sign of the candidate set of transform coefficients. In response to the original sign of the candidate set of transform coefficients being the same as the predicted sign of the candidate set of transform coefficients, a bin with a value of zero ("0") may be generated and added to the bitstream as a sign signaling bit. For example, the bitstream may include a "0" indicating that the predicted sign of the candidate set of transform coefficients was correctly predicted.

[0160] Alternatively, a bin with a value of one ("1") may be generated in response to the original codes of the candidate transform coefficient set not being identical to the predicted codes of the candidate transform coefficient set. An additional set of bins may also be generated to signal the corresponding accuracy of the predicted codes of the candidate transform coefficient set. The bin with a value of one and the set of additional bins may then be added to the bitstream as code signaling bits. For example, the set of additional bins may be an XOR result of the original codes and the predicted codes of the candidate transform coefficient set. An additional bin with a value of "0" may indicate that the predicted codes of the candidate transform coefficient set corresponding to the additional bin were correctly predicted, while an additional bin with a value of "1" may indicate that the predicted codes of the candidate transform coefficient set corresponding to the additional bin were incorrectly predicted. The bitstream may include (a) a "1" indicating that the predicted codes of the candidate transform coefficient set were incorrectly predicted, and (b) a set of additional bins indicating which predicted codes were correctly predicted and which predicted codes were incorrectly predicted.

[0161] The sign signaling bit can be generated similarly using other context models. For example, the sign signaling bit indicates that the predicted sign of the candidate transform coefficient set was correctly predicted. The value is "0" Bin and bins with a value of one ("1") indicating that the predicted sign is incorrect. An additional set of bins can also be generated to signal the corresponding accuracy of the predicted signs of the candidate transform coefficient sets that have incorrect sign predictions.

[0162] In some implementations, the size of each candidate set of transform coefficients may be adaptively changed based on one or more predetermined criteria, which may include the width or height of the transform block, the coding mode of the transform block (e.g., intra-coding or inter-coding), the number of non-zero transform coefficients in the transform block, etc. In some implementations, the size of each candidate set of transform coefficients may be signaled in the bitstream at various coding levels, such as the SPS, PPS, slice or picture level, the CTU or CU level, or the transform block level.

[0163] In some implementations, one or more constraints may be applied to limit application scenarios of the vector-based sign prediction scheme disclosed herein. For example, the vector-based sign prediction scheme disclosed herein may be applied to process signs of a first portion of transform coefficients in a transform block, while signs of a second portion of transform coefficients in the transform block may be processed using an existing sign prediction scheme. In a further example, the vector-based sign prediction scheme disclosed herein may be applied to the first N (e.g., N=2, 3, 4, 5, 6, 7, or 8) non-zero candidate transform coefficients from a transform block, while signs of a second portion of transform coefficients in the transform block may be processed using an existing sign prediction scheme. from The signs of the other candidate transform coefficients can be processed using an existing sign prediction scheme as shown in FIG. 15A, which will be described later in this disclosure.

[0164] Consistent with this disclosure, the code prediction process 1100 disclosed herein may be disabled in some scenarios. For example, when an LFNST is applied to a coding block coded using intra-template matching, the primary transform may be a DCT-VII. Given that the core transform of the LFNST in an ECM is primarily trained when the primary transform is a DCT-II, the corresponding LFNST transform coefficients of an intra-template matching block may exhibit different characteristics compared to the coefficients of other LFNST blocks. Based on this rationale, the code prediction process 1300 may be disabled when the current coding block is an intra-template matching block and coded using an LFNST.

[0165] Consistent with this disclosure, the maximum number of predicted codes for an LFNST block and the maximum number of predicted codes for a non-LFNST block may be different to control the computational complexity of code prediction. For example, the maximum number of predicted codes for an LFNST block may be set to 6 (or 4), while the maximum number of predicted codes for a non-LFNST block may have a value different from 6 (or 4). Furthermore, different values ​​for the maximum number of predicted codes may be applied to video blocks that apply LFNST and video blocks that do not apply LFNST. In some implementations, video encoder 20 may determine the maximum number of predicted codes for an LFNST block based on the encoder's corresponding complexity or performance priority and may signal the maximum number to video decoder 30. If the maximum number of predicted codes for an LFNST block is signaled to video decoder 30, it may be signaled at various coding levels, such as the sequence parameter set (SPS), picture parameter set (PPS), picture or slice level, or CTU or CU level. In some implementations, video encoder 20 may determine different values ​​for the maximum number of predicted codes for video blocks that apply LFNST and video blocks that do not apply LFNST, and signal the values ​​of the maximum number from video encoder 20 to video decoder 30.

[0166] Consistent with this disclosure, assuming that the transform coefficients of both the primary and secondary transforms are fixed, the video encoder 20 or video decoder 30 can pre-compute templates (e.g., template samples) for different transform block sizes and different combinations of primary and secondary transforms. To avoid the complexity of generating template samples on-the-fly for optimized implementations, the video encoder 20 or video decoder 30 can store the templates (e.g., template samples) in internal or external memory. The template samples may be stored with different decimal precisions to achieve different tradeoffs between storage size and sample precision. For example, the video encoder 20 or video decoder 30 can scale the floating samples of the template by a fixed factor (e.g., 64, 128, or 256) and then round the scaled samples to the nearest integer. The rounded samples may be stored in memory. Then, when reconstructing a candidate hypothesis using the template, the corresponding samples are first descaled to their original precision to ensure that the samples generated by the candidate hypothesis are within the correct dynamic range.

[0167] FIG. 14 is a graphical representation illustrating exemplary hypothesis generation based on a linear combination of templates according to some implementations of the present disclosure. In FIG. 14, four patterned blocks 0-3 represent transform coefficient candidates whose signs are predicted. Coefficients C0, C1, C2, and C3 represent corresponding values ​​of dequantized transform coefficients of the four transform coefficient candidates. Templates 0-3 may correspond to the four transform coefficient candidates 0-3, respectively. For example, template 0 corresponding to transform coefficient candidate 0 can be generated by applying an inverse secondary transform and an inverse linear transform to the transform block, where transform coefficient candidate 0 is set to 1 and the remaining transform coefficient candidates in the transform block are set to 0. Similarly, templates 1-3 can be generated, respectively. Hypothesis candidates can be generated by adding templates 0-1 and weights C0-C3, respectively.

[0168] 15A is a graphical representation illustrating an example implementation of an existing code prediction scheme according to some embodiments. FIG. 15B is a graphical representation illustrating an example implementation of a vector-based code prediction scheme according to some implementations of the present disclosure. An example comparison between the existing code prediction scheme and the vector-based code prediction scheme disclosed herein is described herein with reference to FIG. 15A and FIG. 15B.

[0169] In Figures 15A and 15B, there are six non-zero transform coefficients in the transform block that are selected as candidate transform coefficients for sign prediction. The candidate transform coefficients are scanned from the coefficient matrix of the transform block using a raster scan order. The original and predicted signs of the candidate transform coefficients are shown in Figures 15A and 15B. 5B For example, the original code and predicted code of the first transform coefficient candidate with a value of "-2" are both "-" (see FIGS. 15A and 15B). 5B The original and predicted signs of the second transform coefficient candidate, which has a value of 3, are both "+" (see Figure 1). 15 A and Fig. 15 B is represented as "0"). The original sign of the third transform coefficient candidate, whose value is "1", is "+" and the predicted sign is "-" (Figure 15 A and Fig. 15 15A and 15B, the original signs of the third transform coefficient candidate are the same as their corresponding predicted signs (i.e., are correctly predicted).

[0170] As shown in FIG. 15A, a total of six bins (i.e., 0, 0, 1, 0, 0, and 0) are generated, each corresponding to a candidate transform coefficient. The six bins can be generated by performing an XOR operation between the original codes and predicted codes of the six candidate transform coefficients. The six bins can be used to indicate the corresponding accuracy of the six predicted codes. For example, values ​​of "0" in the first and second bins indicate that the predicted codes of the first and second candidate transform coefficients are correct. A value of "1" in the third bin indicates that the predicted code of the third transform coefficient is incorrect. The six bins can be sent to CABAC for entropy coding.

[0171] As shown in FIG. 15B, the vector-based code prediction scheme disclosed herein divides six transform coefficient candidates into three groups, each containing two consecutive transform coefficient candidates. Because the codes of the transform coefficient candidates in groups #0 and #2 can be correctly predicted, only two bins with a value of “0” are generated for these two groups. Because group #1 contains a third transform coefficient candidate whose code cannot be correctly predicted, a bin with a value of “1” (underlined in FIG. 15B) is generated and signaled in the bitstream to indicate that this group contains at least one transform coefficient candidate whose original code is different from its predicted code. Subsequently, two additional bins with values ​​of “1” and “0” are generated for the third and fourth transform coefficient candidates in group #1 to indicate whether their codes can be correctly predicted. Correspondingly, when the vector-based code prediction scheme disclosed herein is applied, a total of five bins are generated for CABAC, which have fewer bits than the bins generated by the existing code prediction scheme shown in FIG. 15A. Therefore, by applying the vector-based code prediction scheme disclosed herein, the signaling overhead can be reduced and the coding efficiency of transform blocks can be improved.

[0172] Consistent with the present disclosure, as shown in Figure 15B, a raster scan order is used to obtain transform coefficient candidates from the coefficient matrix of the transform block, but other scan orders can also be used to select transform coefficient candidates for sign prediction. For example, the transform coefficient candidates can be selected based on one or more selection criteria described above. Similar descriptions will not be repeated here.

[0173] 16A is a graphical representation illustrating an exemplary calculation of a left diagonal cost function along the left diagonal direction according to some implementations of the present disclosure. FIG. 16B is a graphical representation illustrating an exemplary calculation of a right diagonal cost function along the right diagonal direction according to some implementations of the present disclosure. 5 ), the left diagonal cost function costD1 or the right diagonal cost function costD2 shown in equation (14) or (15) above may require more neighboring pixels to support the calculation of the cost functions costD1 and costD2 along the diagonal direction (areas 1602, 1604, and 1606 in FIGS. 16A and 16B). 60 6). If these pixels in regions 1602, 1604, and 1606 are unavailable, the nearest padding method can be used to fill these unavailable positions. For example, B -1,4 If is not available, B -1、4 B is the closest available pixel to -1、3 Using B -1,4 Fill in the position of (e.g., B -1,4 =B -1,3 ). B in area 1602 -1,-1 (C -1,-1 (also indicated by B) -1,-2 (C -1,-1 ), B -2,-1 (C -1,-1 ), and B -2、-2 (C -1,-1 ) is unavailable, several exemplary methods are disclosed herein for filling in the unavailable positions.

[0174] In a first exemplary method, each unavailable location can be filled by weighting its nearest available pixel, as shown in equations (16)-(19) below. B -2,-2 =(B -2,0 +C 0,-2 )*0.5 (16) B -2,-1 =(2*B -2,0 +C 0,-1 ) / 3 (17) B -1,-2 =(B -1,0 +2*C 0,-2 ) / 3 (18) B -1,-1 =(B -1,0 +C 0,-1 )*0.5 (19)

[0175] In a second exemplary method, some of the unavailable locations can be filled with their nearest available pixels. For example, B -1,-2 If is not available, it is C 0,-2 Filled with B -2,-1 If B is not available, -2,0 However, B -2,-2 and B -1,-1 If σ is not available, it can be filled in with the average value of their nearest neighboring pixels calculated according to equations (16) and (17) above.

[0176] In the third exemplary method, only available neighboring reconstructed samples are used to calculate the cost function. If unavailable reconstructed samples are included in the cost calculation for one boundary sample along the top / left boundary of the current block, those samples are not used in the cost calculation for the corresponding direction. For example, in Figure 16B, boundary sample P 0,0 , P 0,1 , P 1,0 and P 1,0 Only P is used to calculate the cost function value in the upper right direction. 0,2 , P 0,3 , P 2,0 , and P3,0 is not used because it refers to at least one reference sample for which cost calculations are not available.

[0177] Consistent with this disclosure, the left and right diagonals (i.e., 135° and 45° as shown in FIGS. 16A and 16B ) are used for illustrative purposes in calculating the cost functions shown in equations (14) or (15) above, although it is contemplated that other measurement elements (e.g., measures of continuity along one or more arbitrary directions) can be incorporated into the calculation of the cost function for code prediction.

[0178] In a fourth implementation, a sample extrapolation method based on gradient analysis of texture information between adjacent samples of one current block can be implemented to improve the accuracy of the cost function for code prediction. Instead of always using a fixed extrapolation direction (e.g., vertical extrapolation for the above adjacent samples and horizontal extrapolation for the horizontally adjacent samples), texture analysis of reconstructed samples adjacent to the top and left edges of the current block can be performed at both the encoder and decoder, and the most dominant direction of the gradient of the adjacent samples can be selected to extrapolate boundary samples along the top and left edges of the current block.

[0179] For example, Figure 17 is a flowchart of a method 1700 for capturing dominant gradient directions in neighboring reconstructed samples of a current block according to some implementations of the present disclosure. Method 1700 may be implemented by a video processor associated with video encoder 20 or video decoder 30 and may include steps 1702-1712 described below. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in Figure 17.

[0180] In step 1702, reference samples of the current block are selected for gradient derivation. In some implementations, the reference samples form a template. For example, FIG. 18A is a graphical representation illustrating example template samples and a gradient filter window in gradient-based selection of a sample extrapolation direction of a cost function according to some implementations of the present disclosure. As shown in FIG. 18A, a template of N rows and columns of reconstructed samples adjacent to the top and left edges of the current block is used as reference samples for gradient derivation. In the example of FIG. 18A, the size of the template is equal to 3.

[0181] In step 1704, a histogram of gradients (HoG) is initialized. For example, the HoG can be generated using a number of entries, with each entry of the histogram of gradients corresponding to the cumulative strength of gradients in a predetermined angular direction. Each entry can be initialized as 0. For example, FIG. 18B is a graphical representation illustrating an exemplary histogram of gradients (HoG) for gradient-based selection of sample extrapolation directions of a cost function according to some implementations of the present disclosure. In practice, different predetermined directions may be used in the disclosed gradient analysis scheme. In one implementation, the same directions as those defined for the 65 angular directions of normal intra prediction in VVC / ECM are used.

[0182] In step 1706, a gradient filter window is applied to the reference samples to calculate their respective gradients. As shown in Figure 18A, one NxN gradient filter window is applied to each template sample in the middle row / column of the template (i.e., the filter window is centered at the sample position) to calculate its corresponding horizontal gradient G h and the vertical gradient G v Calculate each.

[0183] The angle (Angle) and amplitude (Amp) of the gradient of each reference sample are calculated in step 1708. For example, the gradient of the sample can be calculated by equations (20) and (21).

number

[0184] In step 1710, the gradient angle can be converted to one predetermined direction and the corresponding entry in the HoG can be updated. For example, see FIG. B As shown in Figure 1, the amplitude of the HoG at each angle is updated by adding the magnitude (Amp) of the sample gradient at that angle. The resulting amplitude is the cumulative magnitude.

[0185] In step 1712, the HoG entry with the greatest cumulative strength is selected as the direction used to extrapolate neighboring samples in the cost function for the current block. B As shown, the largest entry is circled.

[0186] In the above method, the direction with the greatest strength is selected as the sample extrapolation direction. Such a method is not necessarily reliable in the presence of some noise (e.g., quantization error and / or coding noise caused by other coding modules). To solve this problem, certain conditions can be applied before using the dominant gradient direction as the extrapolation direction.

[0187] For example, in one implementation, a selected dominant gradient direction may be enabled for sample extrapolation in the cost function calculation only if there are enough template samples belonging to the selected gradient direction (e.g., if the proportion of samples belonging to the selected direction is sufficiently large, such as exceeding a predetermined threshold). Otherwise (e.g., if the number of template samples belonging to the selected direction is not sufficient), the default extrapolation (e.g., vertical extrapolation of the top-adjacent samples and horizontal extrapolation of the left-adjacent samples) still applies.

[0188] In another implementation, a selected dominant gradient direction may be enabled for sample extrapolation in the calculation of the cost function only if the magnitude of the gradient associated with that dominant gradient direction is sufficiently large (e.g., if the ratio of the magnitude of the selected gradient direction to the sum of the magnitudes of all gradient directions is greater than another predetermined threshold). Otherwise (e.g., if the magnitude of the gradient in the selected direction is not significant), the default extrapolation still applies.

[0189] In yet another implementation, the above constraints are applied together: the selected direction is valid for the current block's sample extrapolation only if the number of template samples associated with the selected direction is large enough and the gradient strength is large enough; otherwise, the default extrapolation is still applied.

[0190] In yet another implementation, when predicting a code within a block, if the selected dominant gradient direction is valid, a direction orthogonal to the selected direction is used for sample extrapolation in calculating the cost function. For example, if the selected direction is 45 degrees, 135 degrees is used as the direction for extrapolating samples along the top and bottom boundaries of the current block. Similarly, if the selected direction is 135 degrees, 45 degrees is used as the direction for extrapolating samples along the top and left boundaries of the current block.

[0191] In the implementation disclosed above, sign prediction is applicable only to predict the signs of coefficients located in the top-left 4x4 sub-block of a transform block (e.g., as shown in FIG. 7). Generally, this method is reasonable because most of the energy of a common transform block is usually concentrated in a few low-frequency transform coefficients. Correspondingly, the signs of transform coefficients in the top-left corner are statistically easier to predict than those in other locations. However, this assumption is not always correct. For example, for inter-blocks with complex motion fields (e.g., sub-block inter-mode, in which an inter-block is divided into multiple sub-blocks, each with its own motion vector), numerous edges may be generated along the boundaries between different motions in the predicted signal. In such cases, applying a DCT / DST transform may generate non-negligible high-frequency transform coefficients that may extend beyond the top-left 4x4 corner. The signs of these high-frequency coefficients cannot be predicted according to the sign prediction design disclosed above.

[0192] In some alternative implementations, the region of transform coefficients input to the forward LFNST and subsequently used for sign prediction can be expanded from some top-left 4x4 sub-blocks (e.g., as shown in Figure 7) to further improve coding performance. For example, Figure 19 is a graphical representation illustrating a sign prediction region for predicting the signs of transform coefficients, according to some embodiments.

[0193] Specifically, the sign of the top-left A × B region of the current transform block can be selected for prediction, A and B can be calculated according to equation (22). A=min(width,TH) B=min(height,TH) (22) where width and height are the width and height of the transform block, and TH is the maximum region size (also called the region size threshold) for code prediction. Manipulating the value of TH may provide different tradeoffs in performance and complexity of the code prediction technique. Increasing the region for code prediction increases the number of transform coefficients for which codes are predicted, but may increase computational complexity due to the increased number of code combinations that need to be tested at both the encoder and decoder.

[0194] In some implementations, different methods can be used to determine the value of TH. In one implementation, a fixed value (e.g., 8, 16, and 32) can be used for TH in all sequences and coding scenarios. Given that the region size threshold TH is fixed, the encoder and decoder use the same value (i.e., without signaling) to determine corresponding regions when searching for transform coefficients for code prediction. In another implementation, the encoder can be given the flexibility to determine an optimal region size threshold according to the specific characteristics of the video sequence and its preferred performance / complexity tradeoff, and signal the corresponding value from the encoder to the decoder.

[0195] In one exemplary method, the region size threshold may be selected from a fixed set so that only a fixed number of bits need to be signaled to the decoder in the bitstream to indicate which size is selected. For example, assuming there are four allowable region size thresholds {4, 8, 16, 32}, only two bits need to be signaled to indicate the particular value selected from the set. In another exemplary method, the region size threshold may be adaptively determined for a transform block at the encoder side. Thus, the value of the region size threshold may be any number. In such a case, a fixed-length codeword cannot be used because its value is unknown; instead, several variable-length codewords (e.g., Exponential-Golomb, Unary Code, etc.) can be applied to indicate the determined region size threshold value in the bitstream.

[0196] Furthermore, when the value of TH is transmitted in a bitstream, it can be signaled at different levels, such as the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, or coded block level. When the value of TH is signaled at the sequence level (e.g., VPS or SPS), it means that one and the same region size threshold is shared by all transform blocks in a video sequence. When the value of TH is signaled at the picture / slice level (i.e., PPS or slice header), the region size threshold can vary from picture to picture or slice to slice, respectively. Similarly, when the value of TH is signaled at the coded block, it provides maximum adaptability of the coded prediction region, but it would consume more coding bits if this value had to be signaled separately for each coded block.

[0197] On the decoder side, the decoder can determine the code prediction region based on the size of the transform block and the region size threshold TH. If TH is a fixed value, it may be preprogrammed in the decoder without being signaled in the bitstream. The decoder may use this preprogrammed fixed value for all sequences and scenarios. On the other hand, if TH is not a fixed value but a value determined by the encoder, it is signaled to the decoder in the bitstream. Therefore, the decoder determines the code prediction region based on the size of the transform block and the signaled region size threshold TH.

[0198] In some implementations, the decoder can first determine at which level the region size threshold is signaled. As described above, the TH can be signaled at different levels, such as the VPS, SPS, PPS, slice header, or coded block level. The decoder can determine the scope to which the signaled TH should be applied based on the level at which the TH is signaled. For example, if the TH is signaled at the VPS or SPS level, the decoder applies the TH to all transform blocks in the video sequence. If the TH is signaled at the PPS or slice header level, the decoder applies the TH to all transform blocks in the current picture / slice and reads different THs for different pictures / slices. If the TH is signaled at the coded block level, the decoder applies different THs read from the bitstream to different coded blocks.

[0199] If the region size threshold is a value adaptively determined for the transform block, the decoder can determine the value of TH based on the codeword signaled in the bitstream. As mentioned above, TH can be signaled using a fixed-length codeword if it is selected from a predetermined group of values, or can be signaled using a variable-length codeword if it is an arbitrary number adaptively determined for the transform block.

[0200] In some implementations, after determining which value of TH applies to the current transform block, the decoder may determine a code prediction region based on the size of the transform block and the region size threshold TH, for example, according to equation (22). For example, as shown in Figure 19, a width A of the code prediction region is determined to be the smaller of the width of the transform block and the region size threshold TH, and a height B of the code prediction region is determined to be the smaller of the height of the transform block and the region size threshold TH.

[0201] The extended code prediction space scheme does not interfere with any of the code prediction techniques described above, and it is believed that code prediction techniques can be easily adapted and combined with the extended code prediction space to improve coding performance. As just one specific example, the extended code prediction space can be combined with code prediction reordering and LFNST code prediction.

[0202] 20 is a flowchart of an exemplary code prediction method 2000 in block-based video coding according to some implementations of the present disclosure. The method 2000 may be implemented by a video processor associated with the video encoder 20 or the video decoder 30 and may include steps 2002-2008 described below. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 20.

[0203] In step 2002, the video processor may apply a primary transform and a secondary transform to a transform block of a video frame from the video to generate transform coefficients for the transform block.

[0204] In step 2004, the video processor may select a set of candidate transform coefficients from the transform coefficients for sign prediction.

[0205] In step 2006, the video processor may apply a template-based hypothesis generation scheme to select one hypothesis from a plurality of candidate hypotheses for a set of candidate transform coefficients.

[0206] In step 2008, the video processor may determine a combination of code candidates associated with the selected hypothesis that results in a set of predicted codes for the set of candidate transform coefficients.

[0207] Consistent with this disclosure, method 2000 and FIG. 20 may be performed on a video encoder side or a video decoder side. When performed on a video encoder side, method 2000 may be considered as an encoding method for sign prediction of transform coefficients at a video encoder side. When performed on a video decoder side, method 2000 may be considered as a decoding method for sign prediction of transform coefficients at a video decoder side. On the video encoder side An exemplary encoding method for sign prediction of transform coefficients and an exemplary decoding method for sign prediction of transform coefficients at the video decoder side are provided below with reference to Figures 21 and 22, respectively.

[0208] FIG. 21 is a flowchart of an exemplary video encoding method 2100 for sign prediction of transform coefficients performed by a video encoder according to some implementations of the present disclosure. Method 2100 may be implemented by a video processor associated with video encoder 20 and may include steps 2102-2122 described below. Specifically, steps 2102-2108 of method 2100 may be performed as an exemplary implementation of step 2002 of method 2000. Steps 2112-2116 of method 2100 may be performed as an exemplary implementation of step 2006 of method 2000. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 21.

[0209] In step 2102, the video processor may apply a linear transform to a transform block of a video frame from the video to generate coefficients of the transform block.

[0210] In step 2104, the video processor may select a region size threshold and generate a codeword signaling the region size threshold. As described above, the region size threshold TH may be a fixed value or a value adaptively determined by the video processor. The region size threshold may be signaled to the decoder using the codeword. For example, if the region size threshold is selected from a group of predetermined values, the codeword may be a fixed length, e.g., 2 bits. If the region size threshold is an arbitrary value determined to be optimal for the transform block, the codeword may be variable length. If the region size threshold is a fixed value, it may not be signaled.

[0211] In step 2106, the video processor may determine a code prediction region of the transform block based on the region size threshold. For example, the code prediction region may be determined according to equation (22), as described above.

[0212] In step 2108, the video processor may apply a secondary transform to the sub-block in the code prediction domain to generate transform coefficients of the transform block. In some implementations, the secondary transform may be a forward LFNST.

[0213] In step 2110, the video processor may select a set of candidate transform coefficients from the transform coefficients for sign prediction.

[0214] In step 2112, the video processor may determine a combination of multiple code candidates for the set of candidate transform coefficients based on the total number of candidate transform coefficients in the set of candidate transform coefficients.

[0215] In step 2114, the video processor may apply a template-based hypothesis generation scheme to generate multiple candidate hypotheses for each combination of multiple code candidates.

[0216] In step 2116, the video processor may select a hypothesis associated with the smallest cost from a plurality of candidate hypotheses based on a cost function. In some implementations, the cost function may be calculated by extrapolating neighboring samples based on a dominant gradient direction. In some implementations, as described above, the dominant gradient direction must satisfy certain conditions before being used as an extrapolation direction for purposes of calculating the cost function. The video processor may determine whether these conditions are satisfied. If so, extrapolation of neighboring samples may be performed along the dominant gradient direction when calculating the cost function. Otherwise, a default extrapolation direction (e.g., vertical extrapolation for the top-neighboring sample and horizontal extrapolation for the left-neighboring sample) may be used.

[0217] In step 2118, the video processor may determine a combination of code candidates associated with the selected hypothesis that results in a set of predicted codes for the set of candidate transform coefficients.

[0218] In step 2120, the video processor may generate a sequence of code signaling bits for the set of candidate transform coefficients according to the context model. As described above, various context models and combinations thereof can be used to encode the accuracy of the predicted codes. In some embodiments, the sequence of code signaling bits may be generated by applying a vector-based code prediction scheme.

[0219] In step 2122, the video processor may generate a bitstream including a sequence of code signaling bits and a codeword signaling a size region threshold (if signaling is required).

[0220] 22 is a flowchart of an example video decoding method 2200 for sign prediction of transform coefficients performed by a video decoder according to some implementations of the present disclosure. Method 2200 may be implemented by a video processor associated with video decoder 30 and may include steps 2202-2216 described below. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 22.

[0221] In step 2202, the video processor may determine a code prediction region of a transform block of a video frame from the video. In some embodiments, the code prediction region may be determined based on the size of the transform block and a region size threshold. As described above, the region size threshold TH may be a fixed value or a value adaptively determined by the video encoder. If the value of the region size threshold is adaptively determined by the video encoder, the video processor may determine the value of the region size threshold from a codeword signaled in the bitstream. If the value of the region size threshold is a fixed value and not signaled, the video processor may obtain the value of the region size threshold from a decoder memory. In some implementations, the video processor may determine the code prediction region according to equation (22).

[0222] In step 2204, the video processor may select a set of candidate transform coefficients from the dequantized transform coefficients associated with the symbol prediction domain. The set of candidate transform coefficients is used for symbol prediction of the transform coefficients. The dequantized transform coefficients are associated with a transform block. The dequantized transform coefficients of the transform block in video decoder 30 may be equal to the transform coefficients of the transform block in video encoder 20.

[0223] In some implementations, a video processor may receive a bitstream including a sequence of code signaling bits and quantized transform coefficients associated with a transform block. The video processor may generate dequantized transform coefficients from the quantized transform coefficients via inverse quantization unit 86 of FIG. 3.

[0224] In some implementations, the video processor can select a set of transform coefficient candidates from the dequantized transform coefficients associated with the code prediction domain based on the strength of the dequantized transform coefficients. In some implementations, the video processor can select a set of transform coefficient candidates from the dequantized transform coefficients based on the strength of the quantization indexes of the dequantized transform coefficients. In some implementations, the video processor can select a set of transform coefficient candidates from the dequantized transform coefficients based on a coefficient scanning order of entropy coding applied to the video coding.

[0225] In some implementations, the video processor may select a set of candidate transform coefficients from the dequantized transform coefficients associated with the coded prediction region based on influence scores of the dequantized transform coefficients on the reconstructed border samples of the transform block. For example, the influence scores of the dequantized transform coefficients on the reconstructed border samples are measured as the L1 norm of the variation of each dequantized transform coefficient on the reconstructed border samples. In another implementation, the influence scores of the dequantized transform coefficients on the reconstructed border samples are measured as the L2 norm of the variation of each dequantized transform coefficient on the reconstructed border samples.

[0226] In step 2206, the video processor may determine a combination of multiple code candidates for the set of candidate transform coefficients based on the total number of candidate transform coefficients in the set of candidate transform coefficients.

[0227] In step 2208, the video processor may apply a template-based hypothesis generation scheme to generate multiple candidate hypotheses for each combination of multiple code candidates.

[0228] In step 2210, the video processor may select a hypothesis associated with the smallest cost from a plurality of candidate hypotheses based on a cost function, which, in some implementations, may be calculated by extrapolating neighboring samples based on a dominant gradient direction, as described above.

[0229] In step 2212, the video processor may determine a combination of code candidates associated with the selected hypothesis that results in a set of predicted codes for the set of candidate transform coefficients.

[0230] In step 2214, the video processor may estimate original signs of the set of candidate transform coefficients based on the set of predicted signs and the sequence of sign signaling bits received from the video encoder.

[0231] For example, referring to FIG. 15B, the set of predicted codes may include group #0 with values ​​(1,0), group #2 with values ​​(1,0), and group #3 with values ​​(1,0), where 1 indicates a negative sign and 0 indicates a positive sign. The sequence of code signaling bits may include bit "0" of group #0, bits "1,1,0" of group #2, and bit "0" of group #3. Since the bit of group #0 has value "0", indicating that the predicted code having value (1,0) in this group is correct (e.g., the predicted code is the same as the original code), the estimated original code of group #0 is determined to be (1,0). The first bit of the bits in group #1 has a value of "1," indicating that the predicted code (1,0) of this group is incorrect (e.g., the predicted code is not the same as the original code), so the estimated original code (1,0) of group #1 is determined to be the XOR result of the predicted code (1,0) of this group and the second and third bits of group #1, "1,0" (e.g., estimated original code = XOR((1,0),(1,0)) = (0,0)). The bits in group #2 have a value of "0," indicating that the predicted code having the value (1,0) of this group is correct (e.g., the predicted code is the same as the original code), so the estimated original code of group #2 is determined to be (1,0). Thus, the estimated original codes of the set of candidate transform coefficients are formed by concatenating the estimated original codes of groups #0, #1, and #2, respectively, which include (1,0,0,0,1,0). As mentioned above, a video encoder can use different context models to encode the accuracy of the predicted code.

[0232] In step 2216, the video processor may update the dequantized transform coefficients based on the estimated original signs of the set of candidate transform coefficients. For example, the video processor may use the estimated original signs as the true signs of the dequantized transform coefficients in the transform block corresponding to the set of candidate transform coefficients.

[0233] In some implementations, after the dequantized transform coefficients are updated, the video processor may further apply an inverse primary transform and an inverse secondary transform to the dequantized transform coefficients to generate residual samples in a residual block corresponding to the transform block. The inverse secondary transform corresponds to a secondary transform including an LFNST. The inverse primary transform corresponds to a linear transform including a DCT-II, DCT-V, DCT-VIII, DST-I, DST-IV, DST-VII, or identity transform.

[0234] In some implementations, the sequence of sign signaling bits for the set of transform coefficient candidates is generated by the video encoder by applying a vector-based sign prediction scheme to another set of predicted signs for another set of transform coefficient candidates selected by the video encoder side, the other set of transform coefficient candidates being transform coefficients at the video encoder side that correspond to the set of transform coefficient candidates at the video decoder side.

[0235] In some implementations, applying the vector-based code prediction scheme to the set of other predicted codes of the set of other transform coefficient candidates further includes dividing the set of other transform coefficient candidates into a plurality of transform coefficient candidate groups, and for each transform coefficient candidate group, generating one or more code signaling bits for the transform coefficient candidate group based on whether the original code of the transform coefficient candidate group is the same as the predicted code of the transform coefficient candidate group.

[0236] In some implementations, generating one or more sign signaling bits for the set of candidate transform coefficients includes generating a bin with a value of 0 in response to an original code of the set of candidate transform coefficients being identical to a predicted code of the set of candidate transform coefficients, and adding the bin to the bitstream as a sign signaling bit. In some implementations, generating one or more sign signaling bits for the set of candidate transform coefficients includes generating a bin with a value of 1 in response to an original code of the set of candidate transform coefficients not being identical to a predicted code of the set of candidate transform coefficients, generating a set of additional bins to signal a corresponding accuracy of the predicted code of the set of candidate transform coefficients, and adding the bin and the set of additional bins to the bitstream as sign signaling bits.

[0237] 23 illustrates a computing environment 2310 coupled to a user interface 2350 according to some implementations of the present disclosure. The computing environment 2310 may be part of a data processing server. For example, the video processor in the video encoder 20 or the video decoder 30 described above may be implemented using the computing environment 2310. The computing environment 2310 includes a processor 2320, a memory 2330, and an input / output (I / O) interface 2340.

[0238] The processor 2320 typically controls the overall operation of the computing environment 2310, such as operations related to display, data acquisition, data communication, and image processing. The processor 2320 may include one or more processors for executing instructions to perform all or a portion of the steps of the methods described above. The processor 2320 may also include one or more modules that facilitate interaction between the processor 2320 and other components. The processor 2320 may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.

[0239] The memory 2330 is configured to store various types of data to support the operation of the computing environment 2310. The memory 2330 may include predetermined software 2332. Examples of such data include instructions for any applications or methods operating in the computing environment 2310, video data sets, image data, etc. The memory 2330 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.

[0240] The I / O interface 2340 provides an interface between the processor 2320 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 2340 may be coupled to an encoder and a decoder.

[0241] In some implementations, a non-transitory computer-readable storage medium including a plurality of programs, e.g., in memory 2330, executable by processor 2320 in computing environment 2310 for performing the above-described methods is also provided. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or datastream including encoded video information (e.g., video information including one or more syntax elements) generated by an encoder (e.g., video encoder 20 in FIG. 2) using, for example, the above-described encoding method used by a decoder (e.g., video decoder 30 in FIG. 3) in decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0242] In some implementations, a computing device is also provided that includes one or more processors (e.g., processor 2320) and a non-transitory computer-readable storage medium or memory 2330 having stored thereon a plurality of programs executable by the one or more processors, the one or more processors being configured to perform the methods described above when executing the plurality of programs.

[0243] In some implementations, a computer program product is also provided that includes a plurality of programs, e.g., in memory 2330, executable by the processor 2320 in the computing environment 2310, for performing the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0244] In some embodiments, the computing environment 2310 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0245] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosure. Many modifications, variations and alternative implementations will become apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0246] Unless otherwise specified, the order of steps in the method according to the present disclosure is intended as an example only, and the steps in the method according to the present disclosure are not limited to the order specifically described above, and may be changed according to actual conditions. In addition, at least one step of the steps in the method according to the present disclosure may be adjusted, combined, or deleted according to actual needs.

[0247] The examples have been chosen and described to explain the principles of the present disclosure and to enable others skilled in the art to understand the present disclosure in terms of various implementations and to optimally utilize the underlying principles and various implementations with various modifications suited to the particular use intended. Therefore, it should be understood that the scope of the present disclosure should not be limited to the specific implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the present disclosure.

Claims

1. receiving a bitstream including a sequence of code signaling bits; determining a sign prediction region in a transform block of a video frame for performing sign prediction of transform coefficients of the transform block; generating a plurality of candidate hypotheses for a set of candidate transform coefficients associated with the sign prediction region of the transform block; selecting one hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients based on a cost function, the cost function being calculated by extrapolating neighboring samples of the transform block in an extrapolation direction determined based on a dominant gradient direction, the dominant gradient direction being the direction with the greatest intensity in a histogram of gradients (HoG) obtained based on cumulative intensities of gradients in each of a plurality of angular directions; and estimating original codes of the set of candidate transform coefficients based on the set of predicted codes and the sequence of code signaling bits.

2. The video decoding method of claim 1 , wherein the code prediction region is determined based on a size of the transform block and a region size threshold.

3. 3. The video decoding method of claim 2, wherein a width of the code prediction region is determined to be the smaller of a width of the transform block and the region size threshold, and a height of the code prediction region is determined to be the smaller of a height of the transform block and the region size threshold.

4. the region size threshold is a fixed value for all transform blocks of the video frame and is not signaled in the bitstream; or the region size threshold is an adaptively determined value for the transform block, and the value of the region size threshold is signaled in the bitstream; or the value of the region size threshold is selected from a predetermined group of values ​​and is signaled using a fixed length codeword; or the value of the region size threshold is an arbitrary value determined for the transform block and signaled using a variable length codeword, or the value of the region size threshold is signaled in a video parameter set, a sequence parameter set, a picture parameter set, or at slice header and coded block level. The video decoding method of claim 2.

5. The bitstream further comprises quantized transform coefficients associated with the transform blocks, and the video decoding method further comprises: generating dequantized transform coefficients from the quantized transform coefficients; 2. The video decoding method of claim 1, further comprising: updating the dequantized transform coefficients based on the estimated original signs of the set of candidate transform coefficients.

6. 2. The video decoding method of claim 1, further comprising applying an inverse linear transform and an inverse low-frequency non-separable transform (LFNST) to the dequantized transform coefficients associated with the code prediction domain to generate residual samples in a residual block corresponding to the transform block.

7. The step of generating the plurality of candidate hypotheses for the set of candidate transform coefficients comprises: determining a combination of a plurality of code candidates for the set of transform coefficient candidates based on a total number of transform coefficient candidates in the set of transform coefficient candidates; 2. The video decoding method of claim 1, further comprising: applying a template-based hypothesis generation scheme to generate each of the plurality of candidate hypotheses for the plurality of candidate code combinations.

8. The video decoding method of claim 1 , wherein the sequence of sign signaling bits for the set of candidate transform coefficients indicates accuracy of the predicted sign of each candidate transform coefficient.

9. estimating the original codes of the set of candidate transform coefficients based on the set of predicted codes and the sequence of code signaling bits, 9. The video decoding method of claim 8, comprising determining the original signs of the candidate transform coefficients using or correcting the predicted signs of the candidate transform coefficients depending on the accuracy indicated by each of the sign signaling bits.

10. the sequence of code signaling bits is entropy coded based on the strength context of each of the candidate transform coefficients; or the sequence of code signaling bits is entropy coded based on the scan position context of each of the transform coefficients; or The video decoding method of claim 8 , wherein the sequence of code signaling bits is entropy coded based on a coding mode context of the transform block, a block size, or component channel information.

11. a memory configured to store a bit stream including a sequence of code signaling bits; coupled to the memory; and and one or more processors configured to perform the method of any one of claims 1 to 10 on a bitstream.

12. determining a sequence of code signaling bits; determining a sign prediction region in a transform block of a video frame for performing sign prediction of transform coefficients of the transform block; generating a plurality of candidate hypotheses for a set of candidate transform coefficients associated with the sign prediction region of the transform block; selecting one hypothesis from the plurality of candidate hypotheses as a set of predicted signs for the set of candidate transform coefficients based on a cost function, the cost function being calculated by extrapolating neighboring samples of the transform block in an extrapolation direction determined based on a dominant gradient direction, the dominant gradient direction being the direction with the greatest intensity in a histogram of gradients (HoG) obtained based on cumulative intensities of gradients in each of a plurality of angular directions; and estimating original codes of the set of candidate transform coefficients based on the set of predicted codes and the sequence of code signaling bits.

13. generating a bitstream according to the video coding method of claim 12; storing said bitstream.

14. A computer program product containing instructions for storing a bitstream containing encoded video data decoded by a method according to any one of claims 1 to 10.

15. generating a bitstream according to the video coding method of claim 12; transmitting said bitstream.

Citation Information

Patent Citations

  • Low-complexity sign prediction for video coding

    US20180176556A1

  • Sign prediction in video coding

    US20190208225A1

  • Method for encoding a digital image and associated decoding method, devices, user terminal and computer programs

    WO2017115028A1

  • Method and apparatus for harmonizing multiple sign bit hiding and residual sign prediction

    WO2019172797A1

  • Method and apparatus for residual sign prediction in transform domain

    WO2019172798A1