Signaling for zero residual flag and prediction mode
By deriving skip transform flags and optimizing entropy coding contexts, the problem of large overhead of zero-residual flag signaling in the prior art is solved, and the efficiency and performance of video encoding are improved.
Patent Information
- Application Number
- CN202480005564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2024-04-25
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing video encoding technology, signaling notification of zero residual flag or zero transformation coefficient flag increases additional overhead without bringing coding gain, resulting in inefficient entropy coding.
Deriving skip transform flags is used to determine whether intra-block replication and inter-prediction modes are applied, entropy encoding context is optimized, signaling overhead is reduced, and signaling is explicitly notified of zero-residual flags only if necessary.
It improves entropy coding efficiency, reduces signaling overhead, and improves the overall performance of video encoding.
Smart Images

Figure CN120359752A_ABST
Abstract
Description
[0001] Incorporation by reference
[0002] This application claims the benefit of priority to U.S. Non - Provisional Application No. 18 / 632,916, filed on April 11, 2024, which claims the benefit of priority to U.S. Provisional Application No. 63 / 470,759, filed on June 2, 2023. Each of the above - mentioned applications is incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure describes a set of advanced video / stream encoding / decoding techniques. More specifically, the disclosed techniques relate to enhancements for signaling and entropy coding of zero - residual flags or zero - transform - coefficient flags. Background Art
[0004] Uncompressed digital video can include a sequence of pictures and may have specific bit - rate requirements for storage, data processing, and transmission bandwidth in streaming applications. One purpose of video encoding and decoding is to reduce the signaling overhead in the video bitstream through various compression and encoding techniques. Summary of the Invention
[0005] The present disclosure describes various embodiments of methods, devices, and computer - readable storage media for enhancing the signaling and entropy coding of zero - residual flags or zero - transform - coefficient flags.
[0006] According to one aspect, embodiments of the present disclosure provide a method for decoding a video bitstream performed by a decoder. The method includes: receiving a video bitstream including a current picture, the current picture including a current block, and the current block including a current transform block; determining a skip - transform flag indicating whether the current transform block has all - zero coefficients via one of the following: receiving the skip - transform flag from the video bitstream; or deriving the skip - transform flag; deriving at least one of the following flags based on the skip - transform flag: an intra - block copy flag indicating whether IntraBC (Intra - Block Copy) is applied to the current block; an inter - prediction flag indicating whether the current block is encoded in an inter - prediction mode; reconstructing the current block based on at least one of the following: the IntraBC flag, the inter - prediction flag.
[0007] According to another aspect, embodiments of the present disclosure provide a device or decoder for decoding a video bitstream. The device / decoder includes: a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the device / decoder to perform the above - mentioned method for video decoding and / or encoding.
[0008] In another aspect, embodiments of the present disclosure provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above methods for video decoding and / or encoding.
[0009] The above and other aspects and their implementations are described in more detail in the accompanying drawings, the specification, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0011] Figure 1 A schematic illustration showing a simplified block diagram of a communication system (100) according to an example embodiment;
[0012] Figure 2 A schematic illustration showing a simplified block diagram of a communication system (200) according to an example embodiment;
[0013] Figure 3 A schematic illustration showing a simplified block diagram of a video decoder according to an example embodiment;
[0014] Figure 4 A schematic illustration showing a simplified block diagram of a video encoder according to an example embodiment;
[0015] Figure 5 A block diagram showing a video encoder according to another example embodiment;
[0016] Figure 6 A block diagram showing a video decoder according to another example embodiment;
[0017] Figure 7 A scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0018] Figure 8 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0019] Figure 9 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0020] Figure 10 An example of dividing a basic block into coding blocks according to an example partitioning scheme;
[0021] Figure 11 A scheme for dividing a coding block into a plurality of transform blocks and the coding order of the transform blocks according to an example embodiment of the present disclosure;
[0022] Figure 12 Shows another scheme for dividing a coding block into a plurality of transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present disclosure;
[0023] Figure 13 Shows an example logic flow of the method in the present disclosure.
[0024] Figure 14 Shows a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. Detailed Description of the Invention
[0025] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, note that the present invention can be implemented in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any one of the embodiments to be described below. Also note that the present invention can be implemented as a method, apparatus, component, or system. Thus, embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.
[0026] Throughout the specification and claims, terms may have nuanced meanings that are presented or implied in contexts beyond the explicitly stated meanings. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" as used herein does not necessarily refer to different embodiments. Similarly, the phrase "in one implementation" or "in some implementations" as used herein does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" as used herein does not necessarily refer to different implementations. For example, it is meant that the claimed subject matter includes combinations of all or parts of the exemplary embodiments / implementations.
[0027] Typically, terms can be understood, at least in part, from their usage in context. For example, terms such as "and," "or," or "and / or" as used herein can include a variety of meanings, at least in part, depending on the context in which such terms are used. Generally, "or" (if used in an associative list, e.g., A, B, or C) is intended to mean: A, B, and C, used herein in an inclusive sense; and A, B, or C, used herein in an exclusive sense. Additionally, at least in part depending on the context, the terms "one or more" or "at least one" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Similarly, terms such as "a," "an," or "the" can likewise be understood to convey a singular usage or convey a plural usage, at least in part depending on the context. Further, the terms "based on" or "determined by" can be understood to not necessarily convey an exclusive set of factors, and can alternatively allow for additional factors that are not necessarily explicitly described, again at least in part depending on the context.
[0028] As Figure 1 shown, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present disclosure can not be so limited. Embodiments of the present disclosure can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, etc. The network (150) represents any number or type of network for transmitting encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switching, packet-switching, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0029] As an example of an application for the disclosed subject matter, Figure 2 illustrates the setup of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter can equally apply to other video applications, including, for example, video conferencing, digital TV (Television, TV) broadcasting, gaming, virtual reality, storing compressed video on digital media including CD (Compact Disc, CD), DVD (Digital Video Disc, DVD), memory sticks, etc.
[0030] As Figure 2As shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201), such as a digital camera device, for creating an uncompressed video picture or image stream (202). In an example, the video picture stream (202) includes samples recorded by the digital camera device of the video source (201). The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (204) (or encoded video bitstream). The video picture stream (202) may be processed by an electronic device (220) coupled to the video source (201) and including a video encoder (203). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume when compared to the uncompressed video picture stream (202). The encoded video data (204) may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems such as Figure 2 the client subsystems (206) and (208) in
[0031] Figure 3 can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes an incoming copy (207) of the encoded video data and creates an outgoing video picture stream (211) that is uncompressed and can be presented on a display (212) (e.g., a display screen) or other presentation device (not depicted). Figure 2 The video decoder (310) of an electronic device (330) is shown in block diagram form according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may be used in place of
[0032] the video decoder (210) in the example of Figure 3As shown, a receiver (331) may receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle playback timing, a buffer memory (315) may be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include: information for managing the operation of a video decoder (310), and potentially information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the encoded video sequences. The parser (320) may extract from the encoded video sequences a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder. The subgroups may include Groups of Picture (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (320) may also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequences. The reconstruction of the symbols (321) may involve multiple different processing or functional units. The units involved and how they are involved may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.
[0033] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive from the parser (320) the quantized transform coefficients as symbols (321) and control information, including information indicating which inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values that may be input into an aggregator (355).
[0034] In some cases, the output samples of the scaler / inverse transform (351) can belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by an intra picture prediction unit (352). In some cases, the intra picture prediction unit (352) can use surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (355) can add, on a per-sample basis, the prediction information that has been generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).
[0035] In other cases, the output samples of the scaler / inverse transform unit (351) can belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (353) can access the reference picture memory (357) based on a motion vector to obtain samples for inter picture prediction. After motion compensating the obtained reference samples according to the sign (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 can be referred to as residual samples or a residual signal) to generate output sample information.
[0036] The output samples of the aggregator (355) can undergo various loop filtering techniques in a loop filter unit (356) that includes several types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to a rendering device (312) and stored in the reference picture memory (357) for future inter picture prediction.
[0037] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) can be included in an electronic device (420). The electronic device (420) can also include a transmitter (440) (e.g., transmission circuitry). The video encoder (403) can be used in place of Figure 4 the video encoder (403) in the example of.
[0038] A video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a rate control component (450). In some embodiments, the component (450) may be functionally coupled to and control other functional units as described below. Parameters set by the component (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.
[0039] In some example embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and an (in - loop) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols to create sample data in a manner similar to the way a (remote) decoder would create sample data, although the in - loop decoder 433 processes the encoded video stream of the source encoder 430 without performing entropy coding (since in the video compression techniques contemplated in the disclosed subject matter, any compression between symbols and the encoded video bit stream in entropy coding may be lossless). At this point, it can be observed that any decoder technique other than parsing / entropy decoding that may be present only in the decoder may also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations related to the decoding part of the encoder. Thus, the description of encoder techniques can be simplified because encoder techniques are the reverse of the fully described decoder techniques. A more detailed description of the encoder is provided only in certain areas or aspects below.
[0040] During operation, in some example implementations, the source encoder (430) may perform motion - compensated predictive coding that predictively encodes an input picture by referring to one or more previously - encoded pictures of a video sequence designated as "reference pictures".
[0041] The in - loop video decoder (433) may decode the encoded video data of pictures that may be designated as reference pictures. The in - loop video decoder (433) replicates the decoding process that a video decoder may perform on a reference picture and may store the reconstructed reference picture in a reference picture cache (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that would be obtained by a distal (remote) video decoder.
[0042] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search in the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata that can be used as a suitable prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc.
[0043] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, the setting of parameters and subgroup parameters for encoding video data.
[0044] The outputs of all the above functional units can be subjected to entropy encoding in the entropy encoder (445). The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission via the communication channel (460), which can be a hardware / software link to a storage device storing the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0045] The controller (450) can manage the operations of the video encoder (403). During encoding, the controller (450) can assign a specific encoding picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be designated as one of the following picture types: intra pictures (I pictures), predictive pictures (P pictures), bi-predictive pictures (B pictures), multi-predictive pictures. Source pictures can generally be spatially subdivided into multiple sample encoding blocks, as described in further detail below.
[0046] Figure 5 A diagram of a video encoder (503) according to another example embodiment of the present disclosure is shown. The video encoder (503) is configured to receive sample values in a processing block (e.g., a prediction block) within a current video picture in a sequence of video pictures and encode the processing block into an encoded picture that is part of an encoded video sequence. An example video encoder (503) can be used in place of Figure 4 the video encoder (403) in the example of
[0047] For example, the video encoder (503) receives a matrix of sample values of the processing block. The video encoder (503) then uses, for example, Rate-Distortion Optimization (RDO) to determine whether to best encode the processing block using an intra mode, an inter mode, or a bi-predictive mode.
[0048] In Figure 5 the example of, the video encoder (503) includes an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the example arrangement of Figure 5 .
[0049] The inter-frame encoder (530) is configured to: receive samples of a current block (e.g., a processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order); generate inter-frame prediction information (e.g., a motion vector, merge mode information, a description of redundant information according to an inter-frame coding technique); and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique.
[0050] The intra-frame encoder (522) is configured to: receive samples of a current block (e.g., a processing block); compare the block with blocks that have been encoded in the same picture; and generate quantized coefficients after transformation; and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques).
[0051] The general controller (521) may be configured to determine general control data and control other components of the video encoder (503) based on the general control data to, for example, determine a prediction mode of a block and provide a control signal to the switch (526) based on the prediction mode.
[0052] The residual calculator (523) may be configured to calculate a difference (residual data) between a received block and a prediction result of a block selected from the intra-frame encoder (522) or the inter-frame encoder (530). The residual encoder (524) may be configured to encode the residual data to generate transform coefficients. Then, the transform coefficients are quantized to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to format a bitstream to include encoded blocks and perform entropy coding.
[0053] Figure 6 A diagram showing an example video decoder (610) according to another embodiment of the present disclosure is shown. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (610) may be used instead of Figure 4A video decoder (410) in an example of FIG.
[0054] exist Figure 6 In the example of FIG. 6 , the video decoder ( 610 ) includes Figure 6 An entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together are shown in the example arrangement of.
[0055] The entropy decoder (671) may be configured to reconstruct certain symbols according to the coded picture, which represent the syntax elements constituting the coded picture. The inter-frame decoder (680) may be configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) may be configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) may be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) may be configured to combine the residual output by the residual decoder (673) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) in the spatial domain to form a reconstructed block, which forms a part of the reconstructed picture as a part of the reconstructed video.
[0056] Note that the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using any suitable technology. In some example embodiments, the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using one or more processors executing software instructions.
[0057] Turning to block partitioning for encoding and decoding, the general partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After chunking or partitioning the basic block following any of the example partitioning processes or other processes or combinations thereof described below, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of the various partitioning levels in the partitioning hierarchy and can have various shapes. Each of the partitions can be referred to as a Coding Block (CB). For the various example partitioning implementations described further below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because they can form the units for which some basic encoding / decoding decisions can be made and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. The coding blocks can be luma coding blocks or chroma coding blocks. The CB tree structure for each color can be called a Coding Block Tree (CBT). The coding blocks for all color channels can be collectively referred to as Coding Units (CUs). The hierarchical structures for all color channels can be collectively referred to as Coding Tree Units (CTUs). The partitioning patterns or structures for the various color channels in a CTU can be the same or different.
[0058] In some implementations, the partitioning tree scheme or structure for the luma channel and the chroma channel may not need to be the same. In other words, the luma channel and the chroma channel can have separate coding tree structures or patterns. Additionally, whether the luma channel and the chroma channel use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being encoded is a P-slice, a B-slice, or an I-slice. For example, for an I-slice, the chroma channel and the luma channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P-slice or a B-slice, the luma channel and the chroma channel can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luma channel can be partitioned into CBs by one coding partitioning tree structure, and the chroma channel can be partitioned into chroma CBs by another coding partitioning tree structure.
[0059] In some example implementations, a predefined partitioning pattern can be applied to the basic block. As Figure 7As shown, the exemplary 4-way partitioning tree may start from a first predefined level (e.g., the 64×64 block level or other size, as the basic block size), and the basic block may be hierarchically partitioned down to a predefined lowest level (e.g., the 4×4 level). For example, the basic block may be subject to four predefined partitioning options or patterns indicated by 702, 704, 706, and 708, where the partition designated as R is allowed for recursive partitioning because the same partitioning option indicated by Figure 7 may be repeated to a lesser extent until the lowest level (e.g., the 4×4 level). In some implementations, additional restrictions may be applied to the Figure 7 partitioning scheme. In the Figure 7 implementation, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they are not allowed to be recursive, while square partitions are allowed to be recursive. If needed, following the Figure 7 recursive case, the partitioning generates a final set of coded blocks. The coding tree depth may be further limited to indicate the split depth starting from the root node or root block. For example, the coding tree depth of the root node or root block, such as a 64×64 block, may be set to 0, and after the root block is further split once following Figure 7 , the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from the 64×64 basic block to the 4×4 minimum partition will be 4 (starting from level 0). Such a partitioning scheme may be applied to one or more of the color channels. The scheme of Figure 7 may be followed to independently partition each color channel (e.g., the partitioning pattern or option in the predefined pattern may be independently determined for each of the color channels at each hierarchical level). Alternatively, two or more of the color channels may share the Figure 7 same hierarchical pattern tree (e.g., the same partitioning pattern or option in the predefined pattern may be selected for two or more color channels at each hierarchical level).
[0060] Figure 7 shows an extended partitioning tree that provides up to 4 partitioning types for any given transform type. In this scheme, the transform partitioning type is assigned to each prediction block based on the RD (Rate-Distortion, RD) advantage. The transform size assigned to the prediction block is determined in the following manner based on the transform partitioning type and the prediction block size:
[0061] · PARTITION_NONE: Assign a transform size equal to the block size.
[0062] · PARTITION_SPLIT: Assign a transform size that is 1 / 2 the width of the block size and 1 / 2 the height of the block size.
[0063] ·PARTITION_HORZ: Allocate a transform size that has the same width as the block size and a height that is 1 / 2 of the block size.
[0064] ·PARTITION_VERT: Allocate a transform size that has a width that is 1 / 2 of the block size and the same height as the block size.
[0065] In some example implementations, Figure 7 the partitioning types in
[0066] Figure 8 include a uniform transform size, and no recursion is used in partitioning type 708 (i.e., recursion level = 0). Figure 8 The example partitioning structures in Figure 8 include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example implementations, Figure 8 none of the rectangular partitions in Figure 8 are allowed to be further subdivided. The coding tree depth can be further limited to indicate the split depth starting from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after the root block
[0067] is further split once, the coding tree depth is incremented by 1. In some implementations, only the all-square partitions in 810 (represented by "R") are allowed to Figure 8 follow the Figure 8 pattern and be recursively partitioned to the next level of the partitioning tree.
[0068] In some example implementations, a coding tree unit (CTU) can be split into coding units (CUs) by using a quadtree structure represented as a coding tree to adapt to various local features. A decision on whether to use inter-picture (temporal) prediction or intra-picture (spatial) prediction to code a picture region is made at the CU level. Each CU can be further split into one, two, or four prediction units (PUs) according to the PU split type. Inside a PU, the same prediction process is applied, and relevant information is transmitted to the decoder based on the PU. After obtaining a residual block by applying the prediction process based on the PU split type, the CU can be partitioned into transform units (TUs) according to another quadtree structure such as the coding tree of the CU. In some example implementations, a CU or a TU can be only square-shaped, while a PU can be square-shaped or rectangular-shaped for an inter-picture prediction block. In some example implementations, a coding block can be further split into four square sub-blocks, and a transform is performed on each sub-block, i.e., TU. Each TU can be further recursively (using quadtree splitting) split into smaller TUs, which is referred to as a Residual Quad-Tree (RQT).
[0069] In some example implementations, at the picture boundary, an implicit quadtree split can be adopted such that the block will maintain the quadtree split until the size fits the picture boundary.
[0070] Such quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether a basic block or an intermediate block or partition is further quadtree split can be adapted to various local characteristics of the basic block or intermediate block / partition.
[0071] Another example implementation for partitioning a basic block into CB, PB (Prediction Block, PB), and / or TB (Transform Block, TB) is further described below. For example, instead of using a multi-partition unit type such as Figure 7 or Figure 8Rather than the multi-partition unit type shown, a quadtree with a nested multi-type tree using a binary and / or ternary split partitioning structure can be used. The separation of CB, PB, and TB can be dispensed with (i.e., CB is partitioned into PB and / or TB, and PB is partitioned into TB), unless when binning is required for a CB that is too large in size for the maximum transform length, in which case such a CB may need to be further split. This example partitioning scheme can be designed to support greater flexibility for the CB partitioning shape, such that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, the CB can have a square or rectangular shape. Specifically, the coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. Figure 9 An example of a nested multi-type tree structure using a binary or ternary split is shown. Specifically, Figure 9 The example multi-type tree structure includes four split types, which are referred to as vertical binary split (SPLIT_BT_VER), horizontal binary split (SPLIT_BT_HOR), vertical ternary split (SPLIT_TT_VER), and horizontal ternary split (SPLIT_TT_HOR). Then, the CB corresponds to the leaf of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, this split is used for both prediction and transform processing without any further partitioning. This means that, in most cases, the CB, PB, and TB have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CB. In some implementations, in addition to the binary or ternary split, Figure 9 the nested pattern can also include a quadtree split.
[0072] Figure 10 A specific example of a quadtree with a nested multi-type tree coding block structure with block partitioning for one basic block is shown. The basic block 1000 is quadtree split into four square partitions 1002, 1004, 1006, and 1008. A decision is made for each of the quadtree-split partitions to further split using Figure 9 the multi-type tree structure and the quadtree. In Figure 10In the example, partition 1004 is not further split. Partitions 1002 and 1008 each undergo another quadtree split. For partition 1002, the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree split respectively undergo a third-level split of the quadtree, a horizontal binary split, no split, and a horizontal ternary split. Partition 1208 undergoes another quadtree split, and the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree split respectively undergo a third-level split of a vertical ternary split, no split, no split, and a horizontal binary split. Partition 1006 is split into two partitions following a second-level split pattern that follows a vertical binary split, and these two partitions are further split in the third level according to a horizontal ternary split and a vertical binary split. According to the horizontal binary split, a fourth-level split is further applied to one of the third-level partitions.
[0073] For the above specific example, the maximum luminance transform size can be 64×64, and the maximum supported chrominance transform size can be different from the luminance at, for example, 32×32. Even though the example CBs above are generally not further split into smaller PBs and / or TBs, when the width or height of a luminance coding block or a chrominance coding block is greater than the maximum transform width or height, the luminance coding block or the chrominance coding block can also be automatically split in the horizontal direction and / or the vertical direction to meet the transform size limit in that direction. Figure 10 In the above specific example for dividing a basic block into CBs, and as described above, the coding tree scheme can support the ability for luminance and chrominance to have separate block tree structures. For example, for P slices and B slices, the luminance CTB and the chrominance CTB in a CTU can share the same coding tree structure. For example, for I slices, luminance and chrominance can have separate coded block tree structures. When applying separate block tree structures, the luminance CTB can be divided into luminance CBs through one coding tree structure, and the chrominance CTB can be divided into chrominance CBs through another coding tree structure. This means that a CU in an I slice can include coded blocks of the luminance component or coded blocks of two chrominance components, and unless the video is monochromatic, a CU in a P slice or a B slice always includes coded blocks of all three color components.
[0074]
[0075] When a coding block is further divided into multiple transform blocks, the transform blocks therein can be sorted in the bitstream according to various orders or scan patterns. Example implementations for dividing a coding block or a prediction block into transform blocks and the coding order of the transform blocks are further described in detail below. In some example implementations, as described above, the transform partitioning can support multiple shapes such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1 transform blocks, where the transform block size ranges from, for example, 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, the transform block partitioning can be applied only to the luminance component, such that for the chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, both the luminance coding block and the chrominance coding block can be implicitly split into multiple min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.
[0076] In some example implementations of the transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coding block can be further divided into multiple transform blocks with a partitioning depth of up to a predefined number of levels (e.g., 2 levels). The transform block partitioning depth and size can be related. For some example implementations, the following (Error! Reference source not found) shows the mapping from the transform size at the current depth to the transform size at the next depth.
[0077] Table 1: Transform Partitioning Size Settings
[0078] Transformation size at the current depth Transformation size at the next depth TX_4×4 TX_4×4 TX_8×8 TX_4×4 TX_16×16 TX_8×8 TX_32×32 TX_16×16 TX_64×64 TX_32×32 TX_4×8 TX_4×4 TX_8×4 TX_4×4 TX_8×16 TX_8×8 TX_16×8 TX_8×8 TX_16×32 TX_16×16 TX_32×16 TX_16×16 TX_32×64 TX_32×32 TX_64×32 TX_32×32 TX_4×16 TX_4×8 TX_16×4 TX_8×4 TX_8×32 TX_8×16 TX_32×8 TX_16×8 TX_16×64 TX_16×32 TX_64×16 TX_32×16
[0079] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform split can create four 1:1 square sub-transform blocks. The transform partitioning can stop, for example, at 4×4. Thus, the transform size of 4×4 at the current depth corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform split can create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next-level transform split can create two 1:2 / 2:1 sub-transform blocks.
[0080] In some example implementations, for the luminance component of intra-coded blocks, additional restrictions can be applied to the transform block partitioning. For example, for each level of the transform partitioning, all sub-transform blocks can be restricted to have equal sizes. For example, for a 32×16 coding block, the level 1 transform split creates two 16×16 sub-transform blocks, and the level 2 transform split creates eight 8×8 sub-transform blocks. In other words, the second-level split must be applied to all first-level sub-blocks to keep the transform units of equal size. Figure 11An example of transform block partitioning of an intra-coded square block performed according to Table 1 and the coding order indicated by the arrows is shown. Specifically, 1102 shows a square coding block. A first-level split into 4 equally sized transform blocks according to Table 1 and the coding order indicated by the arrows are shown in 1104. A second-level split of all first-level equal-sized blocks into 16 equal-sized transform blocks according to Table 1 and the coding order indicated by the arrows are shown in 1106.
[0081] In some example implementations, for the luminance component of an inter-coded block, the above restrictions for intra-coding may not be applied. For example, after the first-level transform split, any one of the sub-transform blocks can be further independently split at a higher level. Thus, the resulting transform blocks may or may not have the same size. Figure 12 An example split of an inter-coded block into transform blocks and its coding order are shown. In Figure 12 the example of, the inter-coded block 1202 is split into transform blocks in two levels according to Table 1. At the first level, the inter-coded block is split into 4 equally sized transform blocks. Then, only one (not all) of the four transform blocks is further split into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes, as shown in 1204. The example coding order of these 7 transform blocks is indicated by the Figure 12 arrows in 1204 of.
[0082] In some example implementations, for the chrominance component, some additional restrictions for transform blocks may be applied. For example, for the chrominance component, the transform block size can be as large as the coding block size, but not less than a predefined size such as 8×8.
[0083] In some other example implementations, for coding blocks whose width (W) or height (H) is greater than 64, both the luminance coding block and the chrominance coding block can be implicitly split into a plurality of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively. Here, in the present disclosure, "min(a,b)" can return the smaller value between a and b.
[0084] Zero transform coefficient flag coding
[0085] In video coding technologies such as AV1 (AOMedia Video 1), for each intra-coded block and inter-coded block, a flag, namely the skip_txfm flag, is signaled as shown in the following table indicated by the read_skip() function. This flag indicates whether all the transform coefficients in the current coded block are zero. If this flag is signaled with a value of 1, the syntax related to the transform coefficients such as EOB is not signaled and is derived as the value associated with the zero transform coefficient block. For inter-coded blocks, this flag is signaled after the skip mode flag. When skip_mode is true, the skip_txfm flag is not signaled and is inferred as 1. Otherwise, the skip_txfm flag is signaled. Table 2 below shows an example of intra-mode information syntax.
[0086] Table 2: Intra-mode Information Syntax
[0087]
[0088]
[0089] Table 3 below shows an example of inter-mode information syntax.
[0090] Table 3: Inter-mode Information Syntax
[0091]
[0092]
[0093] Table 4 below shows an example of skip syntax.
[0094] Table 4: Skip Syntax
[0095]
[0096] Skip flag semantics
[0097] In some example implementations, for a block such as a transform block, a skip flag can be used to indicate whether there can be transform coefficients to be read for that block. When the skip flag is equal to 0, it indicates that there are transform coefficients to be read (or the block has at least one non-zero transform coefficient). While when the skip flag is equal to 1, it indicates that there are no transform coefficients to be read (or the block has all non-zero transform coefficients).
[0098] In some example implementations, contexts for entropy encoding / decoding the above skip flags can be derived. For example, the derivation of the contexts can depend on the skip flag values of the upper and / or left neighboring blocks. Exemplarily, there can be a total of 3 candidate contexts, and they can be stored in an array. If neither the upper neighboring block nor the left neighboring block is encoded with a non-zero skip flag, context value 0 (i.e., array index 0) is used. If one of the upper neighboring block or the left neighboring block is encoded with a non-zero skip flag, context value 1 (i.e., array index 1) is used. If both the upper neighboring block and the left neighboring block are encoded with non-zero skip flags, context value 2 (i.e., array index 2) is used.
[0099] In some example implementations, the above context array can include, for example, TileSkipCdf[ctx]. Here, ctx is an index that can be calculated by a function as shown in Table 5 below.
[0100] Table 5: ctx Derivation
[0101]
[0102] In current video coding technologies such as AV1, when the current block is encoded as an intra block, the residual block can rarely be zero. However, a flag indicating whether the residual block is all zero is still signaled, which is additional overhead and not beneficial to the coding gain.
[0103] In the present disclosure, various implementations for improving video coding / decoding technologies including AV1, HEVC (High Efficiency Video Coding), VVC (Versatile Video Coding), VP9, etc. are disclosed. These implementations aim to improve entropy coding efficiency, optimize entropy coding contexts, and reduce signaling overhead.
[0104] Based on investigations and statistical observations of video coding technologies such as AV1, it is found that the values of the residual blocks are closely related to the prediction modes associated with the residual blocks. One observation shows that when the current block (i.e., the block currently being encoded / decoded) is encoded as an intra block, the residual block can rarely be zero. That is, the probability that all the residuals in the residual block are zero is very low (close to or equal to 0). The same observation can also apply to the transform coefficients and the quantized transform coefficients in the transform blocks of the current block. For example, refer to Figure 11。Without further partitioning, the current block 1102 can be a transform block. In the case of 1-level partitioning, the current block 1104 can be split into 4 transform blocks. In the case of 2-level partitioning, the current block 1106 can be split into 16 transform blocks. Thus, based on statistical observations, if the current block is encoded as an intra block (i.e., intra prediction is applied to the current blocks 1102, 1104, and 1106), then for Figure 11 any of the transform blocks shown, the probability that the transform block has all-zero transform coefficients (or quantized transform coefficients) is very low.
[0105] However, in current video coding techniques, a flag indicating whether a residual block (or transform block, quantized transform block) is all-zero is still signaled, which adds additional overhead and provides no benefit to coding gain. In the present disclosure, when a block (e.g., a residual block, transform block, or quantized transform block) is all-zero, this means that the parameters (residuals, coefficients, or quantized coefficients) in the block are all zero.
[0106] In the present disclosure, various embodiments for improving video coding / decoding techniques including AV1, HEVC, VVC, VP9, etc. are disclosed. These embodiments are aimed at at least improving entropy coding efficiency, optimizing entropy coding context, and reducing signaling overhead. These embodiments can be implemented in a decoder and / or an encoder.
[0107] In the present disclosure, if the prediction mode is not a smooth mode, or the prediction mode generates prediction samples according to a given prediction direction, then the mode is referred to as an angular mode or a directional mode. The angular mode can also be referred to as the directional mode.
[0108] In the present disclosure, the term "block" can refer to a transform block, an encoded block, a prediction block, a coding block, a coding unit (CU), etc. In the present disclosure, when referring to the block size, it can refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width × height), or the aspect ratio of the block (width:height, or height:width). The term chroma block can refer to a block in any chroma (color) channel. The direction of a reference frame is determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.
[0109] In the present disclosure, a sample can be interpreted as the pixel value of a pixel. It can generally refer to any component (luminance or chroma).
[0110] In the present disclosure, unless otherwise specified, signaling can include one or more sub-signalings. One or more sub-signalings can be transmitted together, or can be transmitted separately.
[0111] In the present disclosure, a flag, i.e., skip_txfm_flag, is used to indicate whether there are some transform coefficients to be read for a transform block. When skip_txfm_flag is equal to 1, it indicates that no transform coefficients need to be read for the transform block. That is to say, the read operation can be skipped, and all transform coefficients in the transform block are considered to be zero. By utilizing the correlation between the prediction mode and skip_txfm_flag, optimized signaling efficiency can be achieved and coding gain can be improved. Since a coding block can be partitioned into one or more transform blocks according to a partitioning scheme, there can be different implementations of skip_txfm_flag. In one example, skip_txfm_flag can be used at the transform block level such that each transform block has a dedicated skip_txfm_flag. In another example, skip_txfm_flag can be used at the coding block level such that a single skip_txfm_flag applies to all transform blocks in the current block.
[0112] In the present disclosure, the "is_inter" flag is used to indicate whether a block is predicted in an inter prediction mode; and "is_intrabc" is used to indicate whether a block is predicted in an IntraBC (Intra Block Copy) mode.
[0113] In the present disclosure, information such as a flag can be explicitly signaled in a video bitstream via, for example, a syntax element. When a flag is explicitly signaled, the decoder directly extracts the flag from the bitstream. The flag can also be implicitly derived by the decoder based on other information. In this case, the flag is not carried in the bitstream.
[0114] In one embodiment, the Intra Block Copy (IntraBC) mode can be allowed and used / signaled in the current block. The term "allowed" means that the IntraBC mode is an allowed option for the current block such that the IntraBC mode is allowed to be enabled or applied to the current block (i.e., the current block is encoded in the IntraBC mode). Conversely, when the IntraBC mode is not allowed or prohibited, this means that the IntraBC mode is not an option (i.e., the current block cannot be encoded in the IntraBC mode). Note that "allowed" does not necessarily mean that the IntraBC mode must be used currently, but rather means that the IntraBC mode is an option and is subject to further enabling.
[0115] In some example implementations, when IntraBC mode is allowed for the current block, if skip_txfm_flag is equal to 1 and is_inter is parsed as false (or 0), then the is_intrabc flag is not parsed but is derived as true (or 1). That is, in this case, there is no need to encode and signal the is_intrabc flag in the video bitstream, and there is no syntax element for explicitly carrying the is_intrabc flag. The decoder will derive this flag.
[0116] In some example implementations, when IntraBC mode is not allowed (i.e., cannot be used / signaled) in the current block and skip_txfm_flag is equal to 1, then the is_inter flag is not signaled / parsed but is derived as true. That is, when all transform blocks are zero but the current block is not intraBC - encoded, then the current block must be inter - predicted and no explicit signaling for the is_inter flag is required.
[0117] In one implementation, when IntraBC mode is allowed in the current block and skip_txfm_flag is equal to 1, a flag is signaled to indicate whether the current block is inter - encoded. If this flag indicates that the current block is not an inter - encoded block, then the current block is inferred to be an intra - block copy block. That is, if the current block is not inter - encoded and there is at least one all - zero transform block (in the current block), then the current block must be encoded in IntraBC mode.
[0118] In one implementation, when IntraBC mode is allowed in the current block and skip_txfm_flag is equal to 1, a flag is signaled to indicate whether the current block is an IntraBC block. If this flag indicates that the current block is not IntraBC - predicted, then the current block is inferred to be an inter - block. That is, if the current block is not an IntraBC block and there is at least one all - zero transform block (in the current block), then the current block must be encoded in inter - prediction mode.
[0119] In one implementation, when the current picture is one of the following: a key (intra) picture; an intra - slice; an intra - tile only; or an intra - sub - picture only, and if IntraBC mode is not allowed (i.e., cannot be used / signaled) in the current block, then skip_txfm_flag is not signaled but is derived as 0, indicating that the corresponding transform blocks in the current block can have non - zero coefficients. Note that, as mentioned before, skip_txfm_flag can be applied to one transform block or all transform blocks in the current block.
[0120] In one embodiment, there are more than one transform blocks in the current block, such as Figure 11 and Figure 12 shown, blocks 1104, 1106, and 1204. The value of skip_txfm_flag can be used to adjust or optimize the encoding order of the transform blocks. For example, when skip_txfm_flag is equal to 0, the encoding order of the transform blocks can be adjusted based on the number of transform blocks and / or the block size of the current block. Exemplarily, in Figure 11 the initial encoding order is shown by the arrows in block 1106. When skip_txfm_flag is equal to 0, the initial encoding order can be further adjusted. For example, the encoding order can be reversed. For another example, the encoding order can be sorted by the number of non-zero coefficients in each transform block - the transform block with more non-zero coefficients has a higher priority (encoded first) in the encoding order.
[0121] When skip_txfm_flag is equal to 0 and when encoding multiple transform blocks using the transform partitioning scheme (i.e., PARTITION_NONE, PARTITION_VERT, PARTITION_HORZ, PARTITION_SPLIT) described with reference to Figure 7 if, following the encoding order, all of the previous transform blocks except the last transform block are associated with zero residuals, the transform block level flag indicating whether the last transform block has all-zero residuals (EOB equal to 0) is not signaled but is inferred accordingly such that the flag indicates that the last transform block does not have all-zero residuals.
[0122] In one embodiment, the probability distribution of observing whether the transform blocks are all non-zero is related to the prediction mode of the current block. Thus, when entropy coding skip_txfm_flag, the prediction mode of the current block can be considered when selecting the entropy coding context. Specifically, the entropy coding context is different when the current block is encoded in palette mode compared to when it is not. In an example implementation, the context value depends on whether the current block is encoded in palette mode. For example, if the current block is encoded in palette mode, context value 1 can be selected, and if the current block is not encoded in palette mode, another context value can be selected.
[0123] In one embodiment, the probability distribution of observing whether the transform blocks are all non-zero is related to the block size of the current block. Thus, context derivation depends on the block size. For example, when the block size is greater than a threshold N, context 0 is used. When the block size is less than or equal to N, another context, such as context 1, is used. The threshold N can be pre-configured or can be signaled in the video bitstream.
[0124] In the present disclosure, embodiments are described for exemplary purposes. The various embodiments and / or implementations described in the present disclosure may be performed individually or in any order in combination. The features, advantages, and characteristics of the present solution may be combined in any suitable manner in one or more embodiments. According to the description herein, those of ordinary skill in the relevant art will recognize that the present solution may be practiced without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present solution. In addition, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored in a non-transitory computer-readable medium. In the present disclosure, the term block may be interpreted as a prediction block, a coding block, or a coding unit (CU).
[0125] Figure 13 FIG. 1300 is a flowchart showing an exemplary method for decoding a video bitstream that follows the principles described in the above embodiments. The exemplary method flow may include some or all of the following steps: S1310, receiving a video bitstream including a current picture, the current picture including a current block, and the current block including a current transform block; S1320, determining a skip transform flag indicating whether the current transform block has all-zero coefficients via one of the following: receiving the skip transform flag from the video bitstream; or deriving the skip transform flag; S1330, deriving at least one of the following flags based on the skip transform flag: an intra block copy flag indicating whether IntraBC (intra block copy) is applied to the current block; an inter prediction flag indicating whether the current block is encoded in an inter prediction mode; and S1340, reconstructing the current block based on at least one of the following: the IntraBC flag, the inter prediction flag.
[0126] In the present disclosure, the direction of a reference frame may be determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.
[0127] The above operations can be combined or arranged in any number or order as needed. Two or more of the steps and / or operations can be performed in parallel. The embodiments and implementations in the present disclosure can be used alone or in any order combination. The steps in one embodiment / method can be divided to form multiple sub-methods, and each of the sub-methods can be independent of the other steps in the embodiment and can form an independent solution. In addition, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term block can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block herein can also be used to refer to a transform block.
[0128] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 FIG. shows a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.
[0129] The computer software can be encoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.
[0130] The instructions can be executed on various types of computers or their components (including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.).
[0131] Figure 14 The components shown for the computer system (1800) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (1800).
[0132] A computer system (1800) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera device (1808).
[0133] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (1809), headphones (not depicted)), visual output devices (e.g., screen (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which may output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic graphics output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).
[0134] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1820) with media such as CD / DVD (1821), thumb drives (1822), removable hard disk drives or solid-state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0135] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0136] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, urban, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet, wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks including CAN bus, etc.
[0137] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1840) of the computer system (1800).
[0138] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1843), a hardware accelerator for specific tasks (1844), a graphics adapter (1850), etc. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. In an example, the screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.
[0139] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.
[0140] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art can envision various systems and methods that, although not explicitly shown or described herein, embody the principles of the disclosure and thus fall within the spirit and scope of the disclosure.
Claims
1. A method for decoding a video bitstream performed by a decoder, the method comprising: Receiving the video bitstream including a current transform block in a current block of a current picture; Determining a skip transform flag based on the received video bitstream; Deriving at least one of the following flags based on the skip transform flag: An intra block copy flag indicating whether intra block copy is applied to the current block; Or An inter prediction flag indicating whether the current block is encoded in an inter prediction mode; And Reconstructing the current block based on at least one of the intra block copy flag and the inter prediction flag derived according to the skip transform flag.
2. The method according to claim 1, wherein, When intra block copy is allowed for the current block, and deriving at least one of the following flags includes: Deriving the intra block copy flag as true in response to the skip transform flag being true and the inter prediction flag being false.
3. The method according to claim 1, wherein, When intra block copy is not allowed for the current block, and deriving at least one of the following flags includes: Deriving the inter prediction flag as true in response to the skip transform flag being true.
4. The method according to any one of claims 1 to 3, wherein: Intra block copy is allowed for the current block; The method further comprises: Receiving the inter prediction flag from the video bitstream; and Deriving at least one of the following flags includes: Deriving the intra block copy flag as true in response to the inter prediction flag being false.
5. The method according to any one of claims 1 to 3, wherein: Intra block copy is allowed for the current block; The skip transform flag is true; The method further comprises: Receiving the intra block copy flag from the video bitstream; and Deriving at least one of the following flags includes: Deriving the inter prediction flag as true in response to the intra block copy flag being false.
6. The method according to any one of claims 1 to 3, wherein The selection of the context for entropy decoding the skip transform flag depends on whether the current block is encoded in a palette mode, and the selection is different when the current block is encoded in the palette mode compared to when the current block is not encoded in the palette mode.
7. The method according to any one of claims 1 to 3, wherein, The selection of the context for entropy decoding the skip transform flag depends on the size of the current block.
8. The method according to any one of claims 1 to 3, further comprising: Selecting a first context to be used for entropy decoding the skip transform flag in response to the size of the current block being greater than a predetermined threshold; And Selecting a second context to be used for entropy decoding the skip transform flag in response to the size of the current block being equal to or less than the predetermined threshold.
9. The method according to any one of claims 1 to 3, wherein The current block includes a plurality of transform blocks, and the method further comprises: Adjusting the decoding order of the plurality of transform blocks based on the size of the current block and the number of the plurality of transform blocks in response to the skip transform flag being false.
10. The method according to claim 9, wherein, Adjusting the decoding order includes reversing the decoding order used when the skip transform flag is true.
11. The method according to any one of claims 1 to 3, wherein: The current block includes a plurality of transform blocks; Following the coding order, the plurality of transform blocks includes a last transform block and N transform blocks before the last transform block, and the method further includes: In response to: 1) the skip transform flag being false; And 2) all of the N transform blocks having been determined to have all-zero coefficients, determining that the last transform block has at least one non-zero coefficient, or determining that the flag indicating whether the last transform block has all non-zero coefficients is false.
12. A device for decoding a video bitstream, the device comprising a memory for storing computer instructions and a processor in communication with the memory, wherein, When the processor executes the computer instructions, the processor is configured to cause the device to perform the following operations: Receive the video bitstream including the current picture, the current picture including a current block, and the current block including a current transform block; Determine a skip transform flag indicating whether the current transform block has all-zero coefficients via one of the following: Receive the skip transform flag from the video bitstream; or Derive the skip transform flag; Derive at least one of the following flags based on the skip transform flag: An intra block copy flag indicating whether intra block copy is applied to the current block; Or An inter prediction flag indicating whether the current block is encoded in an inter prediction mode; And Reconstruct the current block based on at least one of the following: the intra block copy flag, the inter prediction flag.
13. The device according to claim 12, wherein, Intra block copy is allowed for the current block, and wherein, when the processor is configured to cause the device to derive at least one of the following flags, the processor is configured to cause the device to perform the following operations: In response to the skip transform flag being true and the inter prediction flag being false, derive the intra block copy flag as true.
14. The apparatus according to claim 12, wherein, Intra block copy is not allowed for the current block, and wherein, when the processor is configured to cause the device to derive at least one of the following flags, the processor is configured to cause the device to perform the following operations: In response to the skip transform flag being true, derive the inter prediction flag as true.
15. The device according to any one of claims 12 to 14, wherein: Intra block copy is allowed for the current block; When the processor executes the computer instructions, the processor is further configured to cause the device to perform the following operations: Receive the inter prediction flag from the video bitstream; and Wherein, when the processor is configured to cause the device to derive at least one of the following flags, the processor is configured to cause the device to perform the following operations: In response to the inter prediction flag being false, derive the intra block copy flag as true.
16. The device according to any one of claims 12 to 14, wherein: Intra block copy is allowed for the current block; The skip transform flag is true; When the processor executes the computer instructions, the processor is further configured to cause the device to perform the following operations: Receive the intra block copy flag from the video bitstream; and When the processor is configured to cause the device to derive at least one of the following flags, the processor is configured to cause the device to perform the following operations: In response to the intra block copy flag being false, derive that the inter prediction flag is true.
17. The apparatus according to any one of claims 12 to 14, wherein The selection of the context for entropy decoding the skip transform flag depends on whether the current block is encoded in palette mode, and is different when the current block is encoded in palette mode compared to when the current block is not encoded in palette mode.
18. The apparatus according to any one of claims 12 to 14, wherein The selection of the context for entropy decoding the skip transform flag depends on the size of the current block.
19. The device according to any one of claims 12 to 14, wherein, When the processor executes the computer instructions, the processor is further configured to cause the device to perform the following operations: In response to the size of the current block being greater than a predetermined threshold, select a first context to be used for entropy decoding the skip transform flag; and In response to the size of the current block being equal to or less than the predetermined threshold, select a second context to be used for entropy decoding the skip transform flag.
20. A non-transitory storage medium for storing computer-readable instructions, the computer-readable instructions causing the processor in a decoder for decoding a video bitstream to perform the following operations when executed: Receive the video bitstream including a current picture, the current picture including a current block, and the current block including a current transform block; Determine a skip transform flag indicating whether the current transform block has all-zero coefficients via one of the following: Receive the skip transform flag from the video bitstream; or Derive the skip transform flag; Derive at least one of the following flags based on the skip transform flag: An intra block copy flag indicating whether intra block copy is applied to the current block; or An inter prediction flag indicating whether the current block is encoded in inter prediction mode; and Reconstruct the current block based on at least one of: the intra block copy flag, the inter prediction flag.