Zero residual flag coding

Zero residual flagging techniques optimize video encoding and decoding by deriving contexts from neighboring block prediction modes, addressing inefficiencies in representing intra-prediction directions and motion vectors, resulting in improved compression ratios and reduced data requirements.

JP2026012244APending Publication Date: 2026-01-23TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025179939
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-11
Filing Date
2025-10-24
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing video coding techniques face inefficiencies in representing intra-prediction directions and motion vectors, leading to suboptimal compression ratios and increased bit usage for less probable directions, which affects the overall efficiency of video encoding and decoding processes.

Method used

The implementation of zero residual or zero coefficient flagging methods, which involve deriving contexts for decoding zero-skip flags based on neighboring block prediction modes and modes, to optimize the encoding and decoding of video blocks, thereby reducing redundant bit usage.

Benefits of technology

Enhances video encoding and decoding efficiency by reducing bit usage for less probable intra-prediction directions and motion vectors, leading to improved compression ratios and reduced data requirements for video transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012244000001_ABST
    Figure 2026012244000001_ABST
Patent Text Reader

Abstract

This disclosure relates to an improved video coding scheme for zero residual or zero coefficient flags.SOLUTION: For example, a method for decoding a current block in a video stream is disclosed. The method may include determining a zero skip flag of at least one neighboring block of a current block as a reference zero skip flag, determining a prediction mode of the current block to be either an intra prediction mode or an inter prediction mode, deriving at least one context for decoding the zero skip flag of the current block based on the reference zero skip flag and the current prediction mode, and decoding the zero skip flag of the current block according to the at least one context.SELECTED DRAWING: Figure 16
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application is based on and claims the benefit of U.S. Non-Provisional Patent Application No. 17 / 573,299, filed January 11, 2022, which claims priority to U.S. Provisional Patent Application No. 63 / 209,261, entitled "ZERO RESIDUAL FLAG CODING," filed June 10, 2021. Both applications are incorporated herein by reference in their entireties.

[0002] This disclosure relates generally to a set of advanced video coding / decoding techniques, and more particularly to an improved coding scheme for zero residual or zero coefficient flagging. [Background technology]

[0003] The discussion of the background art provided herein is intended to generally present the context for the present disclosure. The inventors' work is not admitted expressly or implicitly as prior art to the present disclosure to the extent that that work is described in this background section, along with aspects of the description that may not otherwise be admitted as prior art at the time of filing of this application.

[0004] Video coding and decoding may be performed using inter-picture prediction with motion compensation. Uncompressed digital video may include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luma samples and associated fully sampled or subsampled chroma samples. The series of pictures may have a fixed or variable picture rate (also called frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920 x 1080, a frame rate of 60 frames per second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.

[0005] One goal of video coding and video decoding is to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, sometimes by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques in which an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to a coding / decoding process in which the original video information is not fully preserved during coding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for its intended purpose, even with some information loss. For video, lossy compression is widely adopted in many applications. The amount of tolerable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of film or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect different distortion tolerances. That is, generally, higher distortion tolerance allows for coding algorithms that result in higher losses and higher compression ratios.

[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.

[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture can be called an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session, or as still images. The samples of the intra-predicted block can then be transformed to the frequency domain, and the transform coefficients thus generated can be quantized before entropy coding. Intra-prediction refers to a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are required for a given quantization step size to represent the block after entropy coding.

[0008] For example, traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to code / decode blocks based on surrounding sample data and / or metadata that precede the block of data being intra-coded or intra-decoded in decoding order, e.g., obtained during the encoding and / or decoding of spatial neighbors. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses reference data only from the current picture being reconstructed, and not from other reference pictures.

[0009] Intra-prediction may take many different forms. When two or more such techniques are available in a given video coding technology, the technique used may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In certain cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and the intra-coding parameters of a block of video may be coded separately or collectively included in the mode's codeword. The codeword used for a given mode, sub-mode, and / or parameter combination may affect coding efficiency gains via intra-prediction, and therefore may also affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, refined in H.265, and further refined in newer coding techniques such as joint search model (JEM), versatile video coding (VVC), and benchmark set (BMS). In general, intra prediction can use available neighboring sample values ​​to form a predictor block. For example, available values ​​of a particular neighboring sample set along a particular direction and / or line can be copied into the predictor block. A reference to the direction used can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions specified by the 33 possible intra-predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra-modes specified in H.265). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction from which neighboring samples are used to predict sample 101. For example, arrow (102) indicates that sample (101) is predicted from one or more neighboring samples to the upper right, at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more neighboring samples to the lower left of sample (101), at an angle of 22.5 degrees from horizontal.

[0012] 1A, a square block (104) of 4x4 samples (indicated by a thick dashed line) is depicted in the upper left. The square block (104) contains 16 samples, each labeled with "S," its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Because the size of the block is 4x4 samples, S44 is located in the lower right. Also shown are examples of reference samples that follow a similar numbering scheme. The reference sample is labeled R and has its Y position (e.g., row index) and X position (column index) relative to the block (104). Both H.264 and H.265 use predicted samples that neighbor the block being reconstructed.

[0013] Intra-picture prediction of block 104 may begin by copying reference sample values ​​from neighboring samples according to a signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block 104 indicating the prediction direction of the arrow (102), i.e., the sample is predicted from one or more prediction samples to the upper right, at a 45-degree angle from the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate a reference sample, especially when the direction is not evenly divisible by 45 degrees.

[0015] The number of possible directions has increased as video coding technology continues to develop. In H.264 (2003), for example, nine different directions are available for intra prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions as of the time of this disclosure. Experimental studies have been conducted to help identify the most appropriate intra prediction directions, and these most appropriate directions can be coded with fewer bits using specific techniques of entropy coding, accepting a specific bit penalty for the direction. Furthermore, the direction itself may be predictable from neighboring directions used in the intra prediction of decoded neighboring blocks.

[0016] FIG. 1B shows a schematic diagram (180) showing 65 intra-prediction directions according to JEM to illustrate the increasing number of prediction directions in various coding techniques that have evolved over time.

[0017] Methods for mapping bits representing intra-prediction directions to prediction directions in a coded video bitstream can vary between video coding techniques and can range, for example, from simple direct mappings of prediction directions to intra-prediction modes to complex adaptive schemes including codewords, most-probable modes, and similar techniques. However, in all cases, there may be certain directions of intra-prediction that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, in a well-designed video coding technique, these less-probable directions may be represented with more bits than more-probable directions.

[0018] Inter-picture prediction, or inter-prediction, may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used to predict a newly reconstructed picture or picture part (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture to be used (similar to the temporal dimension).

[0019] In some video compression techniques, the current MV applicable to a particular area of ​​sample data can be predicted from other MVs, e.g., from other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and precede the current MV in decoding order. Doing so can significantly reduce the overall amount of data required to code the MV by relying on the removal of redundancy in correlated MVs, thereby increasing compression efficiency. MV prediction can work effectively because, for example, when coding an input video signal derived from a camera (known as natural video), areas larger than the area to which a single MV is applicable have a statistical likelihood of moving in a similar direction in the video sequence and therefore, in some cases, can be predicted using similar motion vectors derived from MVs in neighboring areas. As a result, the actual MV of a given area is similar or identical to the MV predicted from the surrounding MVs. Such MVs can further be represented, after entropy coding, with fewer bits than would be used if the MV were coded directly rather than predicted from one or more neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example due to rounding errors when computing the predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms specified by H.265, the one described below is a technique that we will refer to as "spatial merging" from here on.

[0021] Specifically, referring to Figure 2, the current block (201) contains samples that the encoder found during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., the last reference picture (in decoding order), using the MV associated with any one of five surrounding samples represented by A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, MV prediction can use predictors from the same reference picture as neighboring blocks. Summary of the Invention [Means for solving the problem]

[0022] Aspects of this disclosure provide methods and apparatus for video encoding and decoding to improve coding schemes for zero residual or zero coefficient flags.

[0023] In some example implementations, a method for decoding a current block in a video stream is disclosed. The method may include: determining a zero-skip flag of at least one neighboring block of the current block as a reference zero-skip flag; determining a prediction mode of the current block as either an intra prediction mode or an inter prediction mode; deriving at least one context for decoding the zero-skip flag of the current block based on the reference zero-skip flag and the current prediction mode; and decoding the zero-skip flag of the current block according to the at least one context.

[0024] In the above implementations, the at least one neighboring block includes a first neighboring block and a second neighboring block of the current block, and the reference zero skip flag includes a first reference zero skip flag and a second reference zero skip flag corresponding to the first neighboring block and the second neighboring block, respectively. In some implementations, deriving at least one context based on the reference zero skip flag and the current prediction mode may include, in response to the current prediction mode being an inter prediction mode, deriving a first mode-dependent reference zero skip flag based on the first prediction mode and the first reference zero skip flag of the first neighboring block, deriving a second mode-dependent reference zero skip flag based on the second prediction mode and the second reference zero skip flag of the second neighboring block, and deriving at least one context based on the first mode-dependent reference zero skip flag and the second mode-dependent reference zero skip flag. In some implementations, deriving the first mode-dependent reference zero skip flag may include: setting the first mode-dependent reference zero skip flag to a value of the first reference zero skip flag in response to the first prediction mode being an inter prediction mode; and setting the first mode-dependent reference zero skip flag to a value indicating no skipping in response to the first prediction mode not being an inter prediction mode. Deriving the second mode-dependent reference zero skip flag may include: setting the second mode-dependent reference zero skip flag to a value of the second reference zero skip flag in response to the second prediction mode being an inter prediction mode; and setting the second mode-dependent reference zero skip flag to a value indicating no skipping in response to the second prediction mode not being an inter prediction mode.

[0025] In some of the above implementations, deriving at least one context based on the first mode-dependent reference zero-skip flag and the second mode-dependent reference zero-skip flag may include deriving the context of the at least one context as a sum of the first mode-dependent reference zero-skip flag and the second mode-dependent reference zero-skip flag.

[0026] In some of the above implementations, the first neighboring block and the second neighboring block may include the above neighboring block and the left neighboring block of the current block, respectively.

[0027] In some of the above implementations, when the prediction mode of the first neighboring block is an intra-copy mode, the first prediction mode is considered to be an inter-prediction mode, and when the prediction mode of the second neighboring block is an intra-copy mode, the second prediction mode is considered to be an inter-prediction mode.

[0028] In some of the above implementations, a prediction mode flag to indicate whether the current block is an inter-coded block is signaled before signaling the zero-skip flag for the current block.

[0029] In any of the above implementations, when the prediction mode associated with the current block is an intra block copy mode, the current prediction mode is considered to be an inter prediction mode.

[0030] In some of the above implementations, the at least one context may include a first context and a second context used to code the zero skip flag of the current block when the current block is intra-coded and inter-coded, respectively.

[0031] In some other example implementations, a method for decoding a current block in a video stream is disclosed. The method may include: determining a zero-skip flag of the current block to indicate whether the current block is an all-zero block; deriving at least one set of contexts for decoding a prediction mode flag based on a value of the zero-skip flag, where the prediction mode flag indicates whether the current block should be intra-decoded or inter-decoded; and decoding the prediction mode flag of the current block according to the at least one set of contexts.

[0032] In some of the above implementations, when the zero-skip flag of the current block indicates that the current block is not an all-zero block, the at least one set of contexts may include a first set of contexts, and when the zero-skip flag of the current block indicates that the current block is an all-zero block, the at least one set of contexts includes a second set of contexts different from the first set of contexts.

[0033] In some of the above implementations, the set of at least one context further depends on the value of at least one reference zero skip flag of at least one neighboring block of the current block. Further, the at least one neighboring block may include an upper neighboring block and a left neighboring block of the current block of the video.

[0034] In some other example implementations, a method for decoding a current block in a video bitstream is disclosed. The method may include: parsing the video bitstream to identify a zero-skip flag of the current block; determining whether the current block is an all-zero block according to the zero-skip flag; and, in response to the current block being an all-zero block, setting a prediction mode flag of the current block to indicate an inter-prediction mode rather than parsing it from the video bitstream.

[0035] In some of the above implementations, the method may further include using the prediction mode flag of the current block when decoding neighboring blocks of the current block.

[0036] Aspects of the present disclosure also provide a device or apparatus having a memory for storing computer instructions and a processor configured to execute the computer instructions to perform any of the above method implementations.

[0037] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform the above-described method implementations for video decoding and / or encoding.

[0038] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0039] [Figure 1A] FIG. 10 is a schematic diagram of an example subset of intra-prediction directional modes. [Figure 1B] FIG. 2 illustrates exemplary intra-prediction directions. [Figure 2]FIG. 1 is a schematic diagram illustrating a current block and its surrounding spatial merge candidates for motion vector prediction in one example. [Figure 3] FIG. 3 is a schematic diagram illustrating a simplified block diagram of a communication system (300) according to an exemplary embodiment. [Figure 4] FIG. 4 is a schematic diagram illustrating a simplified block diagram of a communication system (400) according to an exemplary embodiment. [Figure 5] FIG. 2 is a schematic diagram illustrating a simplified block diagram of a video decoder according to an example embodiment. [Figure 6] FIG. 1 is a schematic diagram illustrating a simplified block diagram of a video encoder according to an example embodiment. [Figure 7] FIG. 2 is a block diagram illustrating a video encoder according to another example embodiment. [Figure 8] FIG. 2 is a block diagram illustrating a video decoder according to another example embodiment. [Figure 9] FIG. 1 illustrates a coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 11] FIG. 10 illustrates another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 12] FIG. 10 illustrates another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 13] 1A and 1B illustrate a scheme for dividing a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present disclosure. [Figure 14] 10A and 10B illustrate another scheme for dividing a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present disclosure. [Figure 15] FIG. 10 illustrates another scheme for dividing a coding block into multiple transform blocks, according to an exemplary embodiment of the present disclosure. [Figure 16]1 is a flowchart according to an exemplary embodiment of the present disclosure. [Figure 17] FIG. 1 is a schematic diagram illustrating a computer system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0040] Figure 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes, for example, multiple terminal devices that can communicate with each other via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) may perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be implemented, for example, in media serving applications.

[0041] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of coded video data, which may be implemented, for example, during video conferencing applications. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., of a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device according to the reconstructed video data.

[0042] In the example of FIG. 3 , the terminal devices 310, 320, 330, and 340 may be embodied as a server, a personal computer, and a smartphone, although the applicability of the principles underlying the present disclosure is not so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated videoconferencing equipment, and the like. The network 350 represents any number or type of network that conveys coded video data between the terminal devices 310, 320, 330, and 340, including, for example, wired (wired) and / or wireless communication networks. The communication network 350 may exchange data over circuit-switched channels, packet-switched channels, and / or other types of channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network 350 may not be important to the operation of the present disclosure unless explicitly described herein.

[0043] 4 illustrates the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applied to other video-enabled applications including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0044] A video streaming system may include a video source (401), such as a video capture subsystem (413), which may include a digital camera, for creating a stream of uncompressed video pictures or images (402). In one example, the stream of video pictures (402) includes samples recorded by the digital camera of the video source 401. The stream of video pictures (402), shown in bold to emphasize its high data volume compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) including a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), shown with thin lines to emphasize its low data volume compared to the stream of uncompressed video pictures (402), may be stored directly on the streaming server (405) or on a downstream video device (not shown) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410), for example, within the electronic device (430). The video decoder (410) decodes the input copy of the encoded video data (407) and creates an output stream of video pictures (411) that is uncompressed and can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) may be encoded according to a particular video coding / compression standard. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC, as well as other video coding standards.

[0045] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0046] 5 shows a block diagram of a video decoder (510) according to any of the following embodiments of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0047] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one coded video sequence may be decoded at a time, with the decoding of each coded video sequence being independent of other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data or a streaming source that transmits the coded video data. The receiver (531) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective processing circuits (not shown). The receiver (531) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (515) may be located between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be separate and external to the video decoder (510) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (510), for example, to combat network jitter, and there may be another buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may be unnecessary or may be small. For use with best-effort packet networks such as the Internet, a sufficiently sized buffer memory (515) may be required, and its size may be relatively large.Such buffer memory may be implemented with an adaptive size and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0048] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device, such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a supplemental enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The entropy coding of the coded video sequence may be in accordance with a video coding technique or standard and may be in accordance with various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract from the coded video sequence a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, motion vectors, etc.

[0049] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0050] The reconstruction of the symbols (521) may involve several different processing or functional units, depending on the type of video picture or portion thereof being coded (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. The units that are included and how they are included may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following several processing or functional units is not shown for the sake of simplicity.

[0051] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these functional units will interact closely with each other and may be at least partially integrated with each other. However, to clearly explain the various functions of the disclosed subject matter, the following disclosure will adopt a conceptual subdivision into functional units.

[0052] The first unit may include a scalar / inverse transform unit (551), which may receive quantized transform coefficients as well as control information from the parser (520) including information indicating which type of inverse transform to use, block size, quantization coefficients / parameters, quantization scaling matrices, etc. as symbol(s) (521). The scalar / inverse transform unit (551) may output blocks containing sample values ​​that can be input to an aggregator (555).

[0053] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate blocks of the same size and shape as the block being reconstructed using information from surrounding blocks that have already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558), for example, buffers the partially reconstructed and / or fully reconstructed current picture. In some implementations, the aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0054] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (553) may access a reference picture memory (557) to fetch samples used for inter-picture prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added to the output of the scalar / inverse transform unit (551) by an aggregator (555) to generate output sample information (the output of unit 551 may be referred to as a residual sample or residual signal). The address in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches the prediction samples may be controlled by a motion vector, available to the motion-compensated prediction unit (553) in the form of a symbol (521), which may have, for example, an X component, a Y component (shift), and a reference picture component (time). Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, and may be associated with a motion vector prediction mechanism, etc.

[0055] The output samples of the aggregator (555) may undergo various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to previously reconstructed and loop-filtered sample values ​​as well as meta-information obtained during decoding of a coded picture or previous portion (in decoding order) of the coded video sequence. As described in more detail below, several types of loop filters may be included as part of the loop filter unit 556, in various orders.

[0056] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) and also stored in a reference picture memory (557) for use in future inter-picture prediction.

[0057] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and any unused current picture buffer can be reallocated before beginning reconstruction of the next coded picture.

[0058] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, a profile may select specific tools from among all tools available in the video compression technique or standard as tools intended for use only under that profile. To comply with a standard, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further limited by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled in the coded video sequence.

[0059] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0060] 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG.

[0061] The video encoder (603) may receive video samples from a video source (601) (which, in the example of FIG. 6, is not part of the electronic device (620)), which may capture video image(s) to be coded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0062] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCb, RGB, XYZ, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that can store previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures or images that impart motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., used. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.

[0063] According to some example embodiments, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units, as described below. For simplicity, coupling is not shown. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions associated with the video encoder (603) optimized for a particular system design.

[0064] In some example embodiments, the video encoder (603) may be configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and one or more reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that which a (remote) decoder would create, even if the embedded decoder 633 processes the video stream coded by the source coder 630 without entropy coding (because in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the coded video bitstream may be lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding the symbol stream leads to bit-accurate results regardless of the decoder's location (local or remote), the contents in the reference picture memory (634) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​for reference picture samples that the decoder will "see" when using prediction during decoding. This fundamental principle of reference picture synchrony (and the resulting drift if synchrony cannot be maintained, for example, due to channel error) is used to improve coding quality.

[0065] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, such as the video decoder (510), already described in detail above in conjunction with Figure 5. Referring also briefly to Figure 5, however, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633) within the encoder.

[0066] At this point, it can be said that any decoder technology, except for parsing / entropy decoding, which may only exist in the decoder, may also necessarily need to exist in the corresponding encoder in substantially the same functional form. For this reason, the subject matter of the disclosure may focus on the decoder operation, which is similar to the decoding part of the encoder. Therefore, a description of the encoder technology can be omitted, since it is the reverse of the decoder technology described comprehensively. Only in certain areas or aspects will a more detailed description of the encoder be provided below.

[0067] In operation, in some example implementations, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes color channel differences (or residuals) between pixel blocks of the input picture and pixel blocks of reference picture(s) that may be selected as prediction reference(s) for the input picture. The terms “residue” and its adjective form “residual” may be used interchangeably.

[0068] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may preferably be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content with reconstructed reference pictures obtained by a far-end (remote) video decoder (without transmission errors).

[0069] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new picture. The predictor (635) may operate on sample blocks, pixel blocks at a time, to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).

[0070] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0071] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by lossless compression of the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0072] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) and prepare them for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0073] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures are often assigned as one of the following picture types:

[0074] An intra picture (I-picture) may be one that can be coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0075] A predicted picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0076] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, many predicted pictures may use more than two reference pictures and associated metadata in the reconstruction of a single block.

[0077] A source picture is generally spatially subdivided into multiple sample coding blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures. Source pictures or intermediate processed pictures may also be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same method, as described in more detail below.

[0078] The video encoder (603) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0079] In some exemplary embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, SEI messages, VUI parameter set fragments, etc.

[0080] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits temporal or other correlation between pictures. For example, a particular picture being encoded / decoded, called the current picture, may be divided into blocks. If a block in the current picture resembles a reference block in a previously coded, yet-buffered reference picture in the video, it may be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0081] In some exemplary embodiments, bi-prediction techniques can be used for inter-picture prediction. Such bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which advance the current picture in the video in decoding order (but may be in the past or future, respectively, in display order). A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block can be jointly predicted by a combination of the first reference block and the second reference block.

[0082] Additionally, merge mode techniques may be used to improve coding efficiency in inter-picture prediction.

[0083] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs within a picture may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU may include three parallel coding tree blocks (CTBs), i.e., one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one CU of 64x64 pixels or four CUs of 32x32 pixels. One or more of the 32x32 blocks may each be further divided into four CUs of 16x16 pixels. In some exemplary embodiments, each CU may be analyzed during encoding to determine the prediction type of that CU from various prediction types, such as inter-prediction type and intra-prediction type. A CU may be divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed on a prediction block basis. The division of a CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. A luma PB or a chroma PB may include a matrix of sample values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0084] 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of this disclosure. The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. The exemplary video encoder (703) may be used in place of the example video encoder (403) of FIG. 4.

[0085] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) then determines, for example, using rate-distortion optimization (RDO), whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode. If it is determined that the processing block is coded in intra-mode, the video encoder (703) may encode the processing block into a coded picture using intra-prediction techniques, and if it is determined that the processing block is coded in inter-mode or bi-predictive mode, the video encoder (703) may encode the processing block into a coded picture using inter-prediction techniques or bi-prediction techniques, respectively. In some exemplary embodiments, a merge mode may be used as a sub-mode of inter-picture prediction, in which motion vectors are derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other exemplary embodiments, there may be motion vector components applicable to the current block. Thus, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode decision module, to determine the prediction mode of a processing block.

[0086] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725), which are coupled to each other as shown in the exemplary configuration of Figure 7.

[0087] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures in display order), generate inter-prediction information (e.g., a description of redundancy information, motion vectors, merge mode information according to the inter-coding technique), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on video information encoded using a decoding unit 633 incorporated in the example encoder 620 of FIG. 6 (shown as residual decoder 728 of FIG. 7, as described in further detail below).

[0088] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with previously coded blocks in the same picture, generate transformed quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). The intra encoder (722) may calculate intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0089] The general-purpose controller (721) may be configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines a prediction mode for a block and provides a control signal to the switch (726) based on the prediction mode. For example, if the prediction mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select an intra-mode result for use by the residual calculator (723) and control the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. If the prediction mode for the block is inter-mode, the general-purpose controller (721) controls the switch (726) to select an inter-prediction result for use by the residual calculator (723) and control the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.

[0090] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and a prediction result for the block selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate the transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and intra-prediction information. The decoded blocks are processed appropriately to generate decoded pictures, which can be buffered in a memory circuit (not shown) and used as reference pictures.

[0091] The entropy encoder (725) may be configured to format a bitstream to include the encoded blocks and to perform entropy coding. The entropy encoder (725) may be configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Residual information may not be present when coding blocks in a merged sub-mode of either an inter mode or a bi-prediction mode.

[0092] 8 shows a diagram of an exemplary video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) may be used in place of the video decoder (410) of the example of FIG. 4.

[0093] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), all coupled together, as shown in the exemplary configuration of Figure 8.

[0094] The entropy decoder (871) may be configured to reconstruct, from a coded picture, certain symbols that represent the syntax elements of which the coded picture is composed. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information) that can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, merged submode, or another submode), certain samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), residual information in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter-mode or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information may undergo inverse quantization and be provided to the residual decoder (873).

[0095] The inter decoder (880) may be configured to receive the inter prediction information and generate inter prediction results based on the inter prediction information.

[0096] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0097] The residual decoder (873) may be configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (871) (datapath not shown as this may only be a small amount of control information).

[0098] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual as output by the residual decoder (873) and the prediction result (possibly as output by an inter-prediction module or an intra-prediction module) to form a reconstructed block that forms part of a reconstructed picture as part of the reconstructed video. Note that other appropriate operations, such as deblocking operations, may also be performed to improve visual quality.

[0099] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In some exemplary embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0100] When looking at coding block partitioning, in some example implementations, a predetermined pattern may be applied. As shown in FIG. 9, an example four-way partitioning tree may be used, starting from a first predetermined level (e.g., the 64x64 block level) to a second predetermined level (e.g., the 4x4 level). For example, the base block may follow four partitioning options, indicated by 902, 904, 906, and 908, and the partitions represented by R may be recursively partitioned in that the same partitioning tree shown in FIG. 9 may be repeated at lower scales down to the lowest level (e.g., the 4x4 level). In some implementations, additional restrictions may be applied to the partitioning scheme of FIG. 9. In the implementation of FIG. 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be used but not repeatedly, while square partitions may be used repeatedly. Subsequent partitioning of FIG. 9 by recursion, as needed, generates a final set of coding blocks. Such a scheme may be applied to one or more of the color channels.

[0101] FIG. 10 illustrates another exemplary predetermined partitioning pattern that enables forming a partitioning tree through recursive partitioning. As shown in FIG. 10, an exemplary 10-way partitioning structure or pattern may be predefined. The root block may start from a predetermined level (e.g., from the 128×128 level or the 64×64 level). The exemplary partitioning structure of FIG. 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. A partition type having three subpartitions, indicated by 1002, 1004, 1006, and 1008 in the second column of FIG. 10, may be referred to as a "T" partition. The "T" partitions 1002, 1004, 1006, and 1008 may also be referred to as left T, upper T, right T, and lower T. In some implementations, none of the rectangular partitions in FIG. 10 may be further subdivided. A coding tree depth may be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth of the root node or root black of a 128x128 block may be set to 0, and after the root block is split one more time following FIG. 10, the coding tree depth increases by 1. In some implementations, only the all-square partitions of 1010 may allow recursive splitting to the next level of the split tree following the pattern of FIG. 10. In other words, recursive splitting is not possible for the square partitions of patterns 1002, 1004, 1006, and 1006. Subsequent splitting following FIG. 10 by recursion generates a final set of coding blocks, as needed. Such a scheme may be applied to one or more of the color channels.

[0102] After partitioning or dividing the base block according to any of the above division procedures or other procedures, a final set of partitions or coding blocks may still be obtained. Each of these partitions may be at one of various division levels. Each partition may be referred to as a coding block (CB). In the various exemplary division implementations described above, each resulting CB may be of any allowed size and division level. They are called coding blocks because they may form a unit for which some basic coding / decoding decisions may be made and coding / decoding parameters may be optimized, determined, and signaled in an encoded video bitstream. The highest level in the final partition represents the depth of the coding block division tree. A coding block may be a luma coding block or a chroma coding block.

[0103] In some other example implementations, a quadtree structure may be used to recursively divide the base luma blocks and base chroma blocks into coding units. Such a division structure may be referred to as a coding tree unit (CTU), and the CTU is divided into coding units (CUs) by using the quadtree structure to adapt the division to various local characteristics of the base CTU. In such implementations, implicit quadtree division may be performed at picture boundaries, such that blocks continue quadtree division until their size fits within the picture boundaries. The term CU is used to collectively refer to units of luma coding blocks (CBs) and chroma coding blocks (CBs).

[0104] In some implementations, the CB may be further divided. For example, the CB may be further divided into multiple prediction blocks (PBs) for the purpose of intra-frame or inter-frame prediction during the coding and decoding processes. In other words, the CB may be further partitioned into different subpartitions, where individual prediction decisions / configurations may be made. In parallel, the CB may be further divided into multiple transform blocks (TBs) for the purpose of describing the level at which a transform or inverse transform of video data is performed. The division scheme of the CB into PBs and TBs may or may not be the same. For example, each division scheme may be performed using a unique procedure based on, for example, various characteristics of the video data. The division schemes of the PBs and TBs may be independent in some exemplary implementations. The division schemes and boundaries of the PBs and TBs may be correlated in some other exemplary implementations. In some implementations, for example, the TBs may be divided after PB division, and in particular, each PB may be determined following the division of the coding block and then further divided into one or more TBs. For example, in some implementations, the PB may be divided into one, two, four, or other number of TBs.

[0105] In some implementations, the luma channel and the chroma channels may be processed differently to divide the base block into coding blocks and further into prediction blocks and / or transform blocks. For example, in some implementations, division of the coding block into prediction blocks and / or transform blocks may be allowed for the luma channel, but such division of the coding block into prediction blocks and / or transform blocks may not be allowed for the chroma channel(s). In such implementations, therefore, transform and / or prediction of the luma block may be performed only at the coding block level. In another example, the minimum transform block size of the luma channel and the chroma channel(s) may be different, e.g., the coding block of the luma channel may be allowed to be divided into smaller transform blocks and / or predictive blocks than the chroma channels. In yet another example, the maximum depth of division of the coding block into transform blocks and / or predictive blocks may be different between the luma channel and the chroma channels, e.g., the coding block of the luma channel may be allowed to be divided into deeper transform blocks and / or predictive blocks than the chroma channel(s). As a specific example, a luma coding block may be divided into transform blocks of multiple sizes that can be represented by a recursive division down to a maximum of two levels, allowing transform block shapes such as square, 2:1 / 1:2, 4:1 / 1:4, etc., and transform block sizes from 4 x 4 to 64 x 64. However, for chroma blocks, only the largest possible transform block designated for the luma block may be allowed.

[0106] In some example implementations for dividing a coding block into PBs, the depth, shape, and / or other characteristics of the PB division may depend on whether the PB is intra-coded or inter-coded.

[0107] The division of a coding block (or a prediction block) into transform blocks may be performed recursively or non-recursively, further considering the transform blocks at the boundaries of the coding block or the prediction block, in various exemplary manners, including but not limited to quadtree division and predetermined pattern division. In general, the resulting transform blocks may be at different division levels, may not be the same size, and may not be square in shape (e.g., they may be rectangular with several allowed sizes and aspect ratios).

[0108] In some implementations, a coding partition tree scheme or structure may be used. The coding partition tree schemes used for the luma channel and the chroma channel may not be the same. In other words, the luma channel and the chroma channel may have separate coding tree structures. Furthermore, whether the luma channel and the chroma channel use the same coding partition tree structure or different coding partition tree structures, and the actual coding partition tree structure to be used, may depend on whether the slice being coded is a P slice, a B slice, or an I slice. For example, for an I slice, the chroma channel and the luma channel may have separate coding partition tree structures or coding partition tree structure modes, while for a P slice or a B slice, the luma channel and the chroma channel may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one coding partition tree structure, and the chroma channel may be partitioned into chroma CBs by another coding partition tree structure.

[0109] Specific exemplary implementations of the division of coding blocks and transform blocks are described below. In one such exemplary implementation, a base coding block may be divided into coding blocks using the recursive quadtree division described above. At each level, local video data characteristics may determine whether to continue further quadtree division of a particular partition. The resulting CBs may be at various quadtree division levels with different sizes. The decision of whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CB level (or at the CU level in the case of three color channels). Each CB may be further divided into one, two, four, or other number of PBs according to a PB division type. The same prediction process may be applied within one PB, and related information is sent to the decoder on a PB-by-PB basis. After obtaining residual blocks by applying a prediction process based on the PB division type, the CB may be divided into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific implementation, the CB or TB may not be limited to a square shape. Furthermore, in this particular example, the PBs may be square or rectangular in shape for inter prediction, and only square in intra prediction. A coding block may be further divided into, for example, four square-shaped TBs. Each TB may be further divided recursively (using quad-tree partitioning) into smaller TBs called residual quad-trees (RQTs).

[0110] Another specific example for partitioning a base coding block into a CB and other PBs and / or TBs is described below. For example, instead of using multiple partition unit types as shown in FIG. 10, a quadtree with nested multi-type trees using bipartite and tripartite segmentation structures may be used. The separation of the concepts of CB, PB, and TB (i.e., the division of the CB into PBs and / or TBs, and the division of the PB into TBs) may be abandoned except in cases where the CB has a size too large for the maximum transform length, which may require further division. This exemplary division scheme may be designed to support greater flexibility in the CB division shape so that both prediction and transformation can be performed at the CB level without further division. In such a coding tree structure, the CB may have either a square or rectangular shape. Specifically, a coding tree block (CTB) may first be divided by a quadtree structure. Then, the leaf nodes of the quadtree may be further divided by a multi-type tree structure. FIG. 11 shows an example of a multi-type tree structure. Specifically, the exemplary multitype tree structure of FIG. 11 includes four split types: vertical bisection (SPLIT_BT_VER) (1102), horizontal bisection (SPLIT_BT_HOR) (1104), vertical trisection (SPLIT_TT_VER) (1106), and horizontal trisection (SPLIT_TT_HOR) (1108). CB then corresponds to a leaf of the multitype tree. In this exemplary implementation, as long as CB is not too large relative to the maximum transform length, this segmentation is used for both prediction and transform processing without further splitting. This means that in most cases, CB, PB, and TB have the same block size in a quadtree with a nested multitype tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of CB.

[0111] An example of a quadtree with a nested multitype tree coding block structure for block division of one CTB is shown in FIG. 12. More specifically, FIG. 12 shows that a CTB 1200 is quadtree-divided into four square partitions 1202, 1204, 1206, and 1208. A decision to further use the multitype tree structure of FIG. 11 for division is made for each quadtree-divided partition. In the example of FIG. 12, partition 1204 is not further divided. Partitions 1202 and 1208 each adopt a different quadtree division. In partition 1202, the second-level quadtree-divided upper-left partition, upper-right partition, lower-left partition, and lower-right partition adopt the third-level division of quadtree, 1104 of FIG. 11, undivided, and 1108 of FIG. 11, respectively. Partition 1208 employs another quadtree partitioning, and the second-level quadtree-partitioned upper-left, upper-right, lower-left, and lower-right partitions employ the third-level partitioning of 1106 in FIG. 11 , unsplit, unsplit, and 1104 in FIG. 11 , respectively. Two of the subpartitions of the third-level upper-left partition of 1208 are further split according to 1104 and 1108. Partition 1206 employs the second-level partitioning pattern according to 1102 in FIG. 11 into two partitions, and the two partitions are further split at the third level according to 1108 and 1102 in FIG. 11 . A fourth-level partitioning is further applied to one of them according to 1104 in FIG. 11 .

[0112] In the above specific example, the maximum luma transform size may be 64 × 64, and the maximum supported chroma transform size may also be different from the luma, for example, 32 × 32. If the width or height of a luma coding block or a chroma coding block is larger than the maximum transform width or maximum transform height, the luma coding block or the chroma coding block may be automatically split horizontally and / or vertically to meet the horizontal and / or vertical transform size constraints.

[0113] In the specific example for dividing a base coding block into CBs described above, the coding tree scheme may support the ability for luma and chroma to have separate block tree structures. For example, in the case of P slices and B slices, the luma CTB and chroma CTB in one CTU may share the same coding tree structure. In the case of an I slice, for example, the luma and chroma may have separate coding block tree structures. When the separate block tree mode is applied, the luma CTB may be divided into luma CBs by one coding tree structure, and the chroma CTB is divided into chroma CBs by another coding tree structure. This means that a CU in an I slice can consist of a coding block for the luma component or a coding block for two chroma components, and a CU in a P slice or B slice always consists of coding blocks for all three color components unless the video is monochrome.

[0114] Exemplary implementations for dividing coding or prediction blocks into transform blocks and the coding order of the transform blocks are described in further detail below. In some exemplary implementations, transform division may support transform blocks of multiple shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with transform block sizes ranging from, for example, 4×4 to 64×64. In some implementations, if the coding block is 64×64 or smaller, transform block division may be applied only to the luma component, such that for chroma blocks, the transform block size is identical to the coding block size. Otherwise, if the width or height of the coding block is greater than 64, both the luma coding block and the chroma coding block may be implicitly divided into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.

[0115] In some example implementations, for both intra-coded and inter-coded blocks, a coding block may be further divided into multiple transform blocks with a division depth up to a predetermined number of levels (e.g., two levels). The division depth and size of the transform blocks may be related. An example mapping from the transform size of the current depth to the transform size of the next depth is shown below in Table 1.

[0116] [Table 1]

[0117] According to the example mapping in Table 1, for a 1:1 square block, the next level transform division may create four 1:1 square sub-transform blocks. The transform division may stop at, for example, 4x4. Thus, a transform size of 4x4 at the current depth corresponds to the same size of 4x4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level transform division creates two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next level transform division creates two 1:2 / 2:1 sub-transform blocks.

[0118] In some example implementations, further restrictions may be applied to the luma component of an intra-coded block. For example, at each level of transform partitioning, all sub-transform blocks may be constrained to have equal sizes. For example, for a 32x16 coding block, level 1 transform partitioning creates two 16x16 sub-transform blocks, and level 2 transform partitioning creates eight 8x8 sub-transform blocks. In other words, to keep the transform units equal in size, second-level partitioning must be applied to all first-level sub-blocks. An example of transform block partitioning for an intra-coded square block according to Table 1 is shown in FIG. 13, with the coding order indicated by the arrows. Specifically, 1302 denotes a square coding block. The first-level partitioning into four equal-sized transform blocks according to Table 1 is shown in 1304, with the coding order indicated by the arrows. The second-level partitioning of all first-level equal-sized blocks according to Table 1 into 16 equal-sized transform blocks is shown in 1306, with the coding order indicated by the arrows.

[0119] In some example implementations, the above restrictions on intra-coding may not apply to the luma component of an inter-coded block. For example, after the first level of transform partitioning, any one of the sub-transform blocks may be further partitioned independently at another level. Thus, the resulting transform blocks may or may not be of the same size. An example partitioning of an inter-coded block into transform blocks with a coding order is shown in FIG. 14. In the example of FIG. 14, an inter-coded block 1402 is partitioned into transform blocks at two levels according to Table 1. At the first level, the inter-coded block is partitioned into four transform blocks of equal size. Then, only one of the four transform blocks (but not all of them) is further partitioned into four sub-transform blocks, resulting in a total of seven transform blocks with two different sizes, as shown at 1404. An example coding order of these seven transform blocks is indicated by an arrow at 1404 in FIG. 14.

[0120] In some example implementations, some additional restrictions on the transform blocks may be applied to the chroma component(s). For example, for the chroma component(s), the transform block size may be as large as the coding block size, but may not be smaller than a predetermined size, e.g., 8x8.

[0121] In some other example implementations, for coding blocks whose width (W) or height (H) is greater than 64, both the luma coding block and the chroma coding block may be implicitly divided into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively.

[0122] Figure 15 further illustrates another alternative exemplary scheme for dividing a coding block or a predictive block into transform blocks. As shown in Figure 15, instead of using recursive transform division, a set of predetermined division types may be applied to a coding block according to the transform type of the coding block. In the specific example shown in Figure 15, one of six exemplary division types may be applied to divide the coding block into various numbers of transform blocks. Such a scheme may be applied to either a coding block or a predictive block.

[0123] More specifically, the partitioning scheme of FIG. 15 provides up to six partition types for any given transform type, as shown in FIG. 15. In this scheme, every coding block or predictive block may be assigned a transform type based, for example, on rate-distortion cost. In one example, the partition type assigned to a coding block or predictive block may be determined based on the transform partition type of the coding block or predictive block. As shown by the four partition types illustrated in FIG. 15, a particular partition type may correspond to the partition size and pattern (or partition type) of the transform block. Correspondence between various transform types and various partition types may be predefined. Exemplary correspondences are shown below, with capitalized labels indicating transform types that may be assigned to coding blocks or predictive blocks based on rate-distortion cost.

[0124] · PARTITION_NONE: Allocate a transformation size equal to the block size.

[0125] ·PARTITION_SPLIT: Allocates a transformation size that is 1 / 2 the width and 1 / 2 the height of the block size.

[0126] ·PARTITION_HORZ: Allocates a transformation size with the same width as the block size and half the height of the block size.

[0127] ·PARTITION_VERT: Allocates a transformation size with a width half the block size and a height equal to the block size.

[0128] PARTITION_HORZ4: Allocates a transformation size with the same width as the block size and 1 / 4 of the height of the block size.

[0129] PARTITION_VERT4: Allocates a transformation size that is 1 / 4 the width of the block size and the same height as the block size.

[0130] In the above example, all of the partition types shown in Figure 15 include uniform transform sizes for the partitioned transform blocks. This is not a limitation but merely an example. In some other implementations, mixed transform block sizes may be used for the partitioned transform blocks in a particular partition type (or pattern).

[0131] With reference to several example implementations of signaling and entropy coding of specific types of various blocks / units, for each intra- and inter-coding block / unit, a flag, i.e., a skip_txfm flag, may be signaled within the coded bitstream, as shown in the example syntax of Table 2 below and represented by the read_skip() function for retrieving these flags from the bitstream. This flag may indicate whether the transform coefficients are all zero in the current coding unit. This skip_txfm flag may alternatively be referred to as a zero-skip flag. In some example implementations, if this flag is signaled with, for example, a value of 1, another transform coefficient-related syntax, such as an EOB (End of Block), need not be signaled for any of the coding blocks within the coding unit but can be derived as a value or data structure predefined for and associated with the zero transform coefficient block. For inter-coding blocks, as shown by the example of Table 2, this flag may be signaled after a skip_mode flag, indicating that the coding unit may be skipped for various reasons. If the skip_mode flag value is true, the coding unit should be skipped and there is no need to signal the skip_txfm flag, and the skip_txfm flag is inferred to be 1. Otherwise, if the skip_mode flag is false, more information about the coding block / unit is included in the bitstream and the skip_txfm flag is additionally signaled to indicate whether the coding block / unit is all zeros or not.

[0132] [Table 2]

[0133] In this manner, the skip_txfm flag may be used to indicate whether the block associated with this flag is all zero and can be skipped. For example, if this flag is signaled as 1, the block may contain all-zero coefficients. Otherwise, if this flag is signaled as 0, at least some coefficients in the block are non-zero. In some implementations, a block may be partitioned or divided into transform blocks in various ways, as described above. Each of the t transform blocks may be further associated with an EOB skip flag, which may be signaled in the coded bitstream. If the EOB skip flag for a particular transform block is signaled as 1, the EOB of the transform block is not included in the bitstream, and the transform block may be skipped because it has no non-zero coefficients. However, if the EOB skip flag for a transform block is signaled as 0, it indicates that the transform block has at least one non-zero coefficient and the transform block cannot be skipped. The following example implementations focus on signaling and coding of the skip_txfm flag. For simplicity, the skip_txfm flag described above may be referred to as a skip flag.

[0134] A set of contexts may be designed for entropy coding of the skip_txfm flag (or zero skip flag) in the bitstream. A context for coding the skip_txfm flag of a particular current coding block may be selected from the set of contexts. The selection of a context may be indicated by an index. The selection from the set of contexts may depend on various factors. For example, the selection of an encoding context for the skip_txfm of the current block may depend on the skip_txfm flag value of at least one of the neighboring blocks of the current block. For example, the at least one neighboring block may include an upper and / or left neighboring block of the current block. For example, a total of three different contexts may be selected. In a particular example implementation, if none of the upper or left neighboring blocks are coded with a non-zero skip_txfm flag, a context index value of 0 is used to code the skip_txfm flag of the current block. If one of the neighboring blocks above and to the left is coded with a non-zero skip_txfm flag, a context index value of 1 may be used to code the skip_txfm flag of the current block. If both the neighboring blocks above and to the left are coded with a non-zero skip_txfm flag, a context value of 2 is used to code the skip_txfm flag of the current block. Such an implementation is based on the statistical observation that the more neighboring blocks have all-zero coefficients, the more likely the current block also has all-zero coefficients. An exemplary selection process from the set of contexts for coding the current skip_txfm flag is shown in Table 3 below.

[0135] [Table 3]

[0136] The following additional example implementations focus on schemes for further improving the entropy coding efficiency of the skip_txfm flag and several other flags in the bitstream, utilizing several statistical correlations between these flags and other flags within or across the current coding block (e.g., between neighboring blocks). In particular, the context for encoding these flags for a coding block is designed to depend on the prediction mode (either intra- or inter-prediction mode) of the coding block and / or its neighboring coding blocks.

[0137] For example, some of these example implementations may be based on the observation that, in addition to the correlation of skip_txfm flag values ​​between neighboring blocks, the probability that the skip_txfm value is equal to 1 for the current block may also depend significantly on whether the current block is an intra-coded or inter-coded block. By taking the prediction mode into consideration when determining the context for encoding the skip_txfm flag, more coding efficiency can be obtained as a result of the possible correlation between the skip_txfm flag value and the prediction mode.

[0138] The various exemplary implementations below may be used separately or combined in any order. Furthermore, each of these implementations may be embodied as part of an encoder and / or decoder and may be implemented in either hardware or software. For example, they may be hard-coded in dedicated processing circuitry (e.g., one or more integrated circuits). In another example, they may be implemented by one or more processors executing a program stored on a non-transitory computer-readable medium. Furthermore, the term "block" may refer to any unit of video information, and the term "block size" may refer to either the width or height of a block, or the maximum width and height, or the minimum width and height, or the region size (width x height), or the aspect ratio (width:height or height:width).

[0139] In some typical example implementations, multiple contexts may be designed to encode the skip_txfm flag of a current block of video in a bitstream. The context used to signal and encode the skip_txfm flag of a particular current block may depend on the prediction mode of the current block, for example, whether the current block is intra-coded or inter-coded.

[0140] In some example embodiments, if the current block is an inter-coded block, the context for signaling and coding the skip_txfm flag may depend on the following two conditions: a) whether at least one neighboring block of the current block (e.g., a block above and / or to the left) is intra-coded; and b) the skip_txfm flag of at least one neighboring block of the current block (e.g., the neighboring block above or to the left).

[0141] In one specific example implementation that uses both the upper and left neighboring blocks, if the current block is an inter-coded block, the context selection (context index value) implementation for encoding the current skip_txfm flag described above based on the neighboring skip_txfm flags may be further modified according to the following equation: above_skip_txfm=is_inter(above_mi)?above_mi→skip_txfm:0 (1) left_skip_txfm=is_inter(left_mi)?left_mi→skip_txfm:0 (2) context=above_skip_txfm+left_skip_txfm (3) Here, is_inter(left_mi) or is_inter(above_mi) returns a Boolean value indicating whether the left block or the above block is an inter-coded block, above_mi→skip_txfm indicates the transform skip flag (skip_txfm) of the above block, and left_mi→skip_txfm indicates the transform skip flag of the left block. These neighboring block skip_txfm flags may be referred to as reference skip_txfm flags as references for generating "above_skip_txfm" and "above_skip_txfm", and they are referred to as prediction mode-dependent neighboring skip_txfm flags. The context selection index may be the sum of the prediction mode-dependent neighboring skip_txfm flags.

[0142] Following the above example, there may be three exemplary contexts indicated by contexts 0, 1, and 2. If the current block is an intra-predicted block, determining the context for coding the current skip_txfm flag may follow the exemplary procedure in Table 3 or other procedures. However, if the current block is an inter-predicted block, the context selection is modified by the above equation so that the skip_txfm flag of the neighboring block considered in Table 3 is replaced with the prediction mode-dependent neighboring skip_txfm flag in the above equation. The prediction mode-dependent neighboring skip_txfm flag basically represents the actual neighboring skip_txfm flag that is modified according to the prediction mode of the neighboring block. Specifically, if the neighboring block is an inter-coded block, the actual skip_txfm flag of the neighboring block is retained as the prediction mode-dependent skip_txfm flag; otherwise, if the neighboring block is an intra-coded block, the prediction mode-dependent skip_txfm flag of the neighboring block is set to 0 (as if the neighboring block is not an all-zero coefficient, regardless of whether it is actually true). Essentially, the above implementation exploits the observation that if one or more neighboring blocks are intra-predicted, then the current block is likely not all non-zero.

[0143] The above formula is only used to generate an example of three different context indices that refer to three different contexts. Other mathematics can be formulated to make such distinctions, and there may be more than three different contexts and corresponding context indices.

[0144] In some other or further example implementations, the context for signaling and coding the skip_txfm flag may depend on the current block prediction mode, e.g., the following conditions: a) whether the current block is an intra-coded or inter-coded block; b) the skip_txfm value of at least one neighboring block, e.g., the neighboring block above and / or to the left.

[0145] In some of the above example implementations, if the current block (or neighboring block) is in intra block copy mode, it is considered an inter-coded block rather than an intra-predicted block when signaling the skip_txfm flag.

[0146] In some example implementations described above, a flag called an is_inter flag may be used to indicate whether the current block is an inter-coded block. Such a flag may be signaled before the skip_txfm flag in the bitstream so that a decoder may parse its value first and its use is to determine the decoding context of the skip_txfm flag that follows in the bitstream.

[0147] In some example implementations, if the current block is an intra-coded block, one set of contexts may be used to signal the skip_txfm flag. Otherwise, another set of contexts may be used to signal the skip_txfm flag. In other words, two context sets may be designed, and the selection of the context set for signaling may be first made by the prediction mode of the current block, and then, as described above, further selection may be made in each set of contexts based on other information such as neighboring skip_txfm flags and neighboring prediction modes.

[0148] In the above implementations, the context for signaling and coding of the skip_txfm flag of the current block may depend on its prediction mode and / or the prediction mode of its neighboring blocks and / or the skip_txfm flag of its neighboring blocks. In some other example implementations, the signaling and coding of the is_inter flag to indicate whether the current block is an inter-coded block may depend on the value of the skip_txfm flag of the current block.

[0149] For example, if the skip_txfm flag of the current block is equal to 0 (indicating that at least one coefficient in the current block is non-zero), one set of contexts may be used to signal the is_inter flag. Otherwise, another set of contexts may be used to signal the is_inter flag. This may be based on the observation that blocks with non-zero coefficients are statistically more likely to have similar prediction modes than otherwise, and similarly, blocks with all-zero coefficients are also statistically more likely to have similar prediction modes than otherwise.

[0150] In another further example, the context for signaling the is_inter flag may alternatively or additionally depend on the neighboring skip_txfm. For example, such a context may depend on the following exemplary conditions: a) whether the value of the skip_txfm flag of the current block is 0; b) whether the value of the skip_txfm flag of at least one neighboring block, for example, the block above / left, is 0. This scheme essentially exploits the statistical correlation between prediction modes and coefficient characteristics (all-zero or otherwise).

[0151] In these implementations, due to the interdependencies discussed above, the skip_txfm flag may be signaled before the is_inter flag in the bitstream.

[0152] In some further implementations, the is_inter flag may not be included in the bitstream, but may be implied as true (or 1) on the decoder side when the skip_txfm flag is equal to 1. This may be useful because even if there are no coefficients sent in the bitstream because the skip_txfm flag is 1 (all-zero coefficients), inter prediction is implied or imputed for this block and is used to decode neighboring blocks when those neighboring blocks are associated with a flag coded based on the prediction mode of this current block.

[0153] FIG. 16 shows a flowchart 1600 of an example method according to the principles underlying the above-described implementation for zero-skip flag coding. The example method flow starts at 1601. At S1610, the zero-skip flag of at least one neighboring block of the current block is determined as a reference zero-skip flag. At S1620, it is determined that the prediction mode for the current block is either an intra-prediction mode or an inter-prediction mode. At S1630, at least one context for decoding the zero-skip flag of the current block is derived based on the reference zero-skip flag and the current prediction mode. At S1640, the zero-skip flag of the current block is decoded according to the at least one context. The example method flow ends at S1699. The above-described method flow also applies to encoding.

[0154] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to luma blocks or chroma blocks.

[0155] The techniques described above may be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 17 illustrates a computer system (1700) suitable for implementing certain embodiments of the disclosed subject matter.

[0156] Computer software may be coded using any suitable machine code or computer language that can be subjected to assembly, compilation, linking, or similar mechanisms to produce code containing instructions that can be executed by one or more computer central processing units (CPUs) and graphics processing units (GPUs), etc., directly, or through interpretation and execution of microcode, etc.

[0157] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0158] 17 for computer system (1700) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (1700).

[0159] The computer system (1700) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).

[0160] The input human interface devices may include one or more (only one of each) of a keyboard (1701), a mouse (1702), a trackpad (1703), a touch screen (1710), a data glove (not shown), a joystick (1705), a microphone (1706), a scanner (1707), and a camera (1708).

[0161] The computer system (1700) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1710), data gloves (not shown), or joystick (1705), although some haptic feedback devices may not function as input devices), audio output devices (such as speakers (1709), headphones (not shown)), visual output devices (such as screens (1710), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of two-dimensional visual output or three-dimensional or higher-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0162] The computer system (1700) may also include human-accessible storage devices and associated media such as optical media including CD / DVD ROM / RW (1720) with media such as CD / DVD (1721), thumb drives (1722), removable hard drives or solid state drives (1723), legacy magnetic media such as tape or floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0163] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0164] The computer system (1700) also includes an interface (1754) to one or more communication networks (1755). The network may be, for example, wireless, wired, or optical. The network may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, or the like. Examples of networks include local area networks such as Ethernet and wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; television wired or wireless wide-area digital networks including cable television, satellite television, and terrestrial broadcast television; and vehicular and industrial networks including CANbus. Certain networks typically require an external network interface adapter attached to a specific general-purpose data port (e.g., a USB port on the computer system (1700)) or peripheral bus (1749), while other networks are typically integrated into the core of the computer system (1700) by attaching to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system (1700) can communicate with other entities. Such communication may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., from the CANbus to a particular CANbus device), or bidirectional, e.g., communication with other computer systems using local area digital networks or wide area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0165] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1740) of the computer system (1700).

[0166] The core (1740) may include one or more central processing units (CPUs) (1741), graphics processing units (GPUs) (1742), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1743), task-specific hardware accelerators (1744), graphics adapters (1750), etc. These devices may be connected via a system bus (1748), along with read-only memory (ROM) (1745), random access memory (1746), and internal mass storage (1747) such as internal hard drives or SSDs that are not user-accessible. In some computer systems, the system bus (1748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1748) or via a peripheral bus (1749). In one example, a screen (1710) may be connected to the graphics adapter (1750). Peripheral bus architectures include PCI, USB, and the like.

[0167] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) can execute certain instructions that may combine to form the above-described computer code. The computer code can be stored in ROM (1745) or RAM (1746). Transient data can also be stored in RAM (1746), and persistent data can be stored, for example, in internal mass storage (1747). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (1741), GPU (1742), mass storage (1747), ROM (1745), RAM (1746), etc.

[0168] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0169] As a non-limiting example, a computer system (1700) having the architecture, and in particular a core (1740), can provide functionality as a result of processor(s) (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage devices, as introduced above, as well as media associated with specific storage of the core (1740) that is non-transitory in nature, such as the core's internal mass storage (1747) or ROM (1745). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1740). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1740), and in particular the processor (including a CPU, GPU, FPGA, etc.) within the core (1740), to perform particular processes or particular portions of particular processes, including determining data structures stored in RAM (1746) and modifying such data structures according to processes defined by the software, as described herein. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1744)) that can operate in place of or together with software to perform particular processes or portions of particular processes described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0170] While this disclosure has described several exemplary embodiments, there are modifications, permutations, and various substitute equivalents that fall within the scope of this disclosure. In the implementations and embodiments described above, any of the process operations may be combined or configured in any quantity or order as desired. Also, two or more of the process operations described above may be performed in parallel. Thus, it will be appreciated that those skilled in the art can devise numerous systems and methods not explicitly shown or described herein that embody the principles of the present disclosure and thus are within its spirit and scope.

[0171] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI:Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted Block HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide-angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Conversion unit CTU: Coding Tree Unit PDPC: Position-dependent prediction combination ISP: Intra-subpartition SPS: Sequence parameter settings PPS: Picture Parameter Set APS: Adaptive Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-component adaptive loop filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-Component Sample Offset LSO: Local Sample Offset LR: Loop Recovery Filter AV1:AOMedia Video 1 AV2:AOMedia Video 2 [Explanation of symbols]

[0172] 300 Communication Systems 310,320,330,340 Terminal Devices 350 Network 400 Communication Systems 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 encoded video data 405 Streaming Server 406 Client Subsystem 407 Copy Video Data 408 Client Subsystem 409 Copy Video Data 410 Video Decoder 411 Video Picture Output Stream 412 Display 413 Video Capture Subsystem 420 Electronic Devices 430 Electronic Devices 501 Channel 510 Video Decoder 512 displays, rendering devices 515 buffer memory 520 Parser 521 Symbol 530 Electronic Devices 531 Receiver 551 Scaler / Descaler Unit 552 Intra-picture prediction unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Sources 603 Video Encoder 620 Electronic Devices 630 Source Coder 632 Coding Engine 633 decoder 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Video Sequences 645 Entropy Coder 650 Controller 660 Communication Channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Interencoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 1700 Computer Systems 1701 Keyboard 1702 Mouse 1703 Trackpad 1705 Joystick 1706 Microphone 1707 Scanner 1708 Camera 1709 Speaker 1710 Touch Screen 1720 CD / DVD ROM / RW 1721 Medium 1722 thumb drive 1723 Removable Hard Drive or Solid State Drive 1740 cores 1741 Central Processing Unit (CPU) 1742 Graphics Processing Unit (GPU) 1743 Field Programmable Gate Area (FPGA) 1744 Accelerator 1745 Read-Only Memory (ROM) 1746 Random Access Memory (RAM) 1747 Mass Storage 1748 System Bus 1749 Peripheral bus 1750 graphics adapter 1754 Interface 1755 Communication Network

Claims

1. 1. A method for decoding a current block in a video, comprising: determining a zero-skip flag of at least one neighboring block of the current block as a reference zero-skip flag; determining a prediction mode of the current block to be either an intra prediction mode or an inter prediction mode; deriving at least one context for decoding the zero skip flag of the current block based on the reference zero skip flag and the prediction mode of the current block; decoding the zero-skip flag of the current block according to the at least one context; A method comprising:

2. the at least one neighboring block includes a first neighboring block and a second neighboring block of the current block; The method of claim 1 , wherein the reference zero skip flags include a first reference zero skip flag and a second reference zero skip flag corresponding to the first adjacent block and the second adjacent block, respectively.

3. The step of deriving at least one context based on the reference zero skip flag and the prediction mode of the current block includes, in response to the prediction mode of the current block being an inter prediction mode, deriving a first mode-dependent reference zero skip flag based on a first prediction mode of the first neighboring block and the first reference zero skip flag; deriving a second mode-dependent reference zero skip flag based on a second prediction mode of the second neighboring block and the second reference zero skip flag; deriving the at least one context based on the first mode-dependent reference zero-skip flag and the second mode-dependent reference zero-skip flag; 3. The method of claim 2, comprising:

4. deriving a first mode dependent reference zero skip flag; setting the first mode-dependent reference zero skip flag to a value of the first reference zero skip flag in response to the first prediction mode being an inter prediction mode; in response to the first prediction mode not being an inter prediction mode, setting the first mode-dependent reference zero skip flag to a value indicating no skipping; Including, deriving a second mode dependent reference zero skip flag; setting the second mode-dependent reference zero skip flag to a value of the second reference zero skip flag in response to the second prediction mode being an inter prediction mode; in response to the second prediction mode not being an inter prediction mode, setting the second mode-dependent reference zero skip flag to a value indicating no skipping; 4. The method of claim 1, comprising:

5. deriving the at least one context based on the first mode-dependent reference zero-skip flag and the second mode-dependent reference zero-skip flag, deriving a context of the at least one context as a sum of the first mode-dependent reference zero-skip flag and the second mode-dependent reference zero-skip flag.

5. The method of claim 4, comprising:

6. The method of claim 1 , wherein the first neighboring block and the second neighboring block comprise an upper neighboring block and a left neighboring block of the current block, respectively.

7. When the prediction mode of the first neighboring block is an intra copy mode, the first prediction mode is regarded as an inter prediction mode; The method according to claim 1 , wherein when the prediction mode of the second neighboring block is an intra-copy mode, the second prediction mode is regarded as an inter-prediction mode.

8. 4. The method of claim 1, wherein before signaling the zero skip flag of the current block, a prediction mode flag is signaled to indicate whether the current block is an inter-coded block.

9. The method of claim 1 , wherein when a prediction mode associated with the current block is an intra block copy mode, the prediction mode of the current block is considered to be an inter prediction mode.

10. 4. The method of claim 1, wherein the at least one context includes a first context and a second context used to code the zero skip flag of the current block when the current block is intra-coded and inter-coded, respectively.

11. 1. A method for decoding a current block in a video, comprising: determining a zero-skip flag for the current block to indicate whether the current block is an all-zero block; deriving at least one set of contexts for decoding a prediction mode flag based on a value of the zero-skip flag, the prediction mode flag indicating whether the current block should be intra-decoded or inter-decoded; decoding the prediction mode flag of the current block according to the at least one set of contexts; A method comprising:

12. When the zero skip flag of the current block indicates that the current block is not an all-zero block, the at least one set of contexts includes a first set of contexts; 12. The method of claim 11, wherein when the zero-skip flag of the current block indicates that the current block is an all-zero block, the at least one set of contexts includes a second set of contexts that is different from the first set of contexts.

13. The method of claim 11 , wherein the set of at least one context further depends on a value of at least one reference zero-skip flag of at least one neighboring block of the current block.

14. 14. The method of claim 11, wherein the at least one neighboring block comprises an above neighboring block and a left neighboring block of the current block of the video.

15. 1. A method for decoding a current block in a video bitstream, comprising: parsing the video bitstream to identify a zero-skip flag for the current block; determining whether the current block is an all-zero block according to the zero skip flag; in response to the current block being an all-zero block, setting a prediction mode flag of the current block to indicate an inter prediction mode rather than parsing a prediction mode from the video bitstream; A method comprising:

16. The method of claim 15 , further comprising using the prediction mode flag of the current block when decoding neighboring blocks of the current block.

17. A device comprising circuitry configured to perform the method of any one of claims 1 to 3.

18. A device comprising circuitry configured to perform the method of any one of claims 11 to 13.

19. 17. A device comprising circuitry configured to perform the method of claim 15 or 16.

20. 4. A non-transitory computer-readable medium for storing instructions that, when executed by a process, are configured to encode a current block in a video by performing the method of any one of claims 1 to 3.

21. 14. A non-transitory computer-readable medium for storing instructions that, when executed by a process, are configured to encode a current block in a video by performing the method of any one of claims 11 to 13.

22. 17. A non-transitory computer-readable medium for storing instructions that, when executed by a process, are configured to encode a current block in a video by performing the method of claim 15 or 16.