Video coding and decoding method, video code stream storage method, equipment and storage medium
By decoding the zero-skip flag in video encoding technology, the problem of inefficient processing of zero residual or zero coefficient flags in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202510508745.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-11
- Filing Date
- 2022-01-28
- Publication Date
- 2025-06-13
AI Technical Summary
Existing video encoding technology is inefficient when dealing with zero residual or zero coefficient flags, resulting in redundancy and low efficiency in the encoding and decoding process.
The encoding and decoding process is optimized by determining the zero skip flag for the current block and based on the zero skip flag and prediction mode of the adjacent block, the encoding and decoding process is derived.
Improve the efficiency of video encoding and decoding, reduce redundant information, and improve the performance of the encoding and decoding process.
Smart Images

Figure CN120151518A_ABST
Abstract
Description
Incorporation by Reference
[0001] This application claims priority to U.S. Non - Provisional Patent Application No. 17 / 573,299, filed on January 11, 2022, and U.S. Provisional Application No. 63 / 209,261, titled "Zero - Residual Flag Coding and Decoding", filed on June 10, 2021. The entire contents of both applications are incorporated herein by reference. Technical Field
[0002] This disclosure generally relates to a set of advanced video encoding / decoding techniques, and more particularly, to an improved coding and decoding scheme for zero - residual flags or zero - coefficient flags. Background Art
[0003] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of filing of the present application, and it has never been expressly or implicitly admitted as prior art to this application.
[0004] Inter - picture prediction with motion compensation can be used for video encoding and decoding. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated full - sampled or subsampled chrominance samples. The series of pictures has a fixed or variable picture rate (or frame rate), such as 60 pictures per second or 60 frames per second. Uncompressed video has specific bit - rate requirements. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a chrominance subsampling of 4:2:0, with 8 bits per pixel per color channel, requires a bandwidth of nearly 1.5 Gbit / s. Such a one - hour video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding is to reduce the redundant information of an uncompressed input video signal through compression. Video compression can help reduce the requirements for the above bandwidth and / or storage space, and in some cases, can reduce it by two or more orders of magnitude. Lossless compression, lossy compression, and combinations of both can be used. Lossless compression refers to a technique in which an exact copy of the original signal is reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not completely retained during encoding and cannot be completely recovered during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application despite some information loss. In the case of video, lossy compression is widely used in many applications. The amount of allowable distortion depends on the application. For example, users of certain consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable through a specific coding algorithm can be selected or adjusted to reflect various distortion tolerances: higher allowable distortion generally allows coding algorithms that result in higher losses and higher compression ratios.
[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0007] Video codec technology can include known intra-frame coding techniques. In intra-frame coding, sample values are represented without reference to samples or other data of previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in the intra-frame mode, the picture can be referred to as an intra-frame picture. Intra-frame pictures and their derivatives (such as independent decoder refresh pictures) can be used to reset the decoder state and can therefore be used as the first picture in an encoded video bitstream and a video session, or as a still image. Then, the samples of the blocks after intra-frame prediction can be transformed into the frequency domain, and the so-generated transform coefficients can be quantized before entropy coding. Intra-frame prediction represents a technique that minimizes the sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.
[0008] As is known from, for example, MPEG-2 generation coding techniques, traditional intra coding does not use intra prediction. However, some newer video compression techniques include: attempts to encode / decode blocks based on, for example, surrounding sample data and / or metadata that are obtained during spatially adjacent encoding and / or decoding and that precede, in decoding order, the data block being intra encoded or decoded. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that, in at least some cases, intra prediction uses only reference data from the current picture in reconstruction and not reference data from other reference pictures.
[0009] There can be many different forms of intra prediction. When more than one such technique is available in a given video coding technique, the technique used can be referred to as an intra prediction mode. One or more intra prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes and / or can be associated with various parameters, and the mode / sub-mode information and intra coding parameters for a video block can be included in a mode codeword, which can be encoded separately or jointly. For a given mode, sub-mode, and / or parameter combination, which codeword is used can affect the coding efficiency gain through intra prediction, and the same is true for the entropy coding technique used to convert the codeword into a bitstream.
[0010] A certain mode of intra prediction was introduced with H.264, modified in H.265, and further modified in newer coding techniques such as Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). Generally, for intra prediction, available adjacent sample values that have become available can be used to form a predictor block. For example, the available values of a particular set of adjacent samples along a particular direction and / or row can be copied into the predictor block. The reference to the direction used can be encoded in the bitstream or can itself be predicted.
[0011] Reference Figure 1A , depicted in the lower right, is a subset of 9 predictor directions specified among the 33 possible intra predictor directions in H.265 (corresponding to the 33 angular modes of the 35 intra modes specified in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the directions according to which the sample at 101 is predicted using adjacent samples. For example, arrow (102) indicates that sample (101) is predicted based on one or more adjacent samples in the upper right at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more adjacent samples in the lower left of sample (101) at a 22.5-degree angle to the horizontal direction.
[0012] Still referring to Figure 1A, a square block (104) including 4×4 samples is shown in the upper left (represented by a thick dashed line). The square block (104) consists of 16 samples, and each sample is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension within the block (104). Since the block is of 4×4 sample size, S44 is located in the lower right corner. Example reference samples following a similar numbering scheme are also shown. The reference samples are labeled with "R", its Y position (e.g., row index) relative to the block (104), and its X position (e.g., column index). In H.264 and H.265, neighboring prediction samples adjacent to the block in reconstruction are used.
[0013] Intra prediction within the picture of block 104 can start by copying reference sample values from neighboring samples according to a signalized prediction direction. For example, assume that the encoded video bitstream includes signaling that, for this block 104, indicates the prediction direction of arrow (102) - that is, samples are predicted based on one or more prediction samples in the upper right at a 45-degree angle to the horizontal direction. In such cases, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Then sample S44 is predicted based on reference sample R08.
[0014] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a reference sample, especially when the direction is not divisible by 45 degrees.
[0015] As video coding techniques continue to develop, the number of possible directions increases. For example, in H.264 (in 2003), 9 different directions can be used for intra prediction. This increased to 33 in H.265 (in 2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, accepting certain bit costs for the directions. Additionally, the direction itself can sometimes be predicted based on neighboring directions for intra prediction of already decoded neighboring blocks.
[0016] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to JEM is shown to illustrate the increase in the number of prediction directions in various coding techniques over time.
[0017] The manner in which bits representing intra prediction directions are mapped to prediction directions in an encoded video bitstream can vary with different video coding techniques; and can range, for example, from a simple direct mapping from prediction directions to intra prediction modes, to codewords, to complex adaptive schemes involving most probable modes and the like. However, in all cases, there may be certain directions for intra prediction in a video content that are statistically less likely to occur than some other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, those less likely directions will be representable by a larger number of bits than the more likely directions.
[0018] Inter-picture prediction or inter-frame prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used for prediction of a newly reconstructed picture or picture portion (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture in the current reconstruction. The MV can have two dimensions X and Y, or three dimensions, where the third dimension is an indication of the reference picture in use (approximate temporal dimension).
[0019] In some video compression techniques, the current MV applicable to a certain region of sample data can be predicted from other MVs, for example, from those other MVs that are related to other regions of sample data that are spatially adjacent to the region in the reconstruction and that are prior to the current MV in the decoding order. Doing so can significantly reduce the total amount of data required to encode the MVs by relying on removing redundancy in the related MVs, thereby increasing the compression efficiency. The MV prediction can be performed efficiently, for example, because when encoding an input video signal (referred to as natural video) derived from a camera, there is a statistical likelihood that regions larger than the region to which a single MV applies move in a similar direction in the video sequence. Thus, in some cases, similar motion vectors derived from MVs of adjacent regions can be used for prediction. This results in the actual MV of a given region being similar or identical to the MV predicted from the surrounding MVs. After entropy coding, such MVs can in turn be represented by a smaller number of bits than would be used if the MVs were directly encoded rather than predicted from one or more adjacent MVs. In some cases, the MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, the MV prediction itself may be lossy, for example, due to rounding errors when calculating the predicted value from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (Recommendation ITU-T H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified in H.265, the technique described in this application is what is hereinafter referred to as "spatial merge".
[0021] Please refer to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on a previous block of the same size that has generated a spatial offset. Additionally, the MV can be derived from metadata associated with one or more reference pictures instead of directly encoding the MV. For example, using the MV associated with any one of five surrounding samples of A 0 , A 1 and B 0 , B 1 , B 2 (corresponding to 202 to 206 respectively), the MV is derived from the metadata of the nearest reference picture (in decoding order). In H.265, MV prediction can use the prediction values of the same reference picture also used by adjacent blocks. Summary of the Invention
[0022] Aspects of the present disclosure provide methods, devices, and storage media for video coding and decoding, for improving the coding and decoding schemes of zero residual or zero coefficient flags.
[0023] In some example embodiments, a method for decoding a current block in a video is disclosed. The method may include determining a zero skip flag of at least one adjacent block of the current block as a reference zero skip flag; determining a prediction mode for the current block as an intra prediction mode or an inter prediction mode; deriving at least one context for decoding the zero skip flag of the current block based on the reference zero skip flag and the prediction mode of the current block; and decoding the zero skip flag of the current block according to the at least one context.
[0024] In the above-described embodiment, the at least one neighboring block includes a first neighboring block and a second neighboring block of the current block; and the reference zero skip flag includes a first reference zero skip flag and a second reference zero skip flag corresponding to the first neighboring block and the second neighboring block, respectively. In some embodiments, deriving the at least one context based on the reference zero skip flag and the prediction mode of the current block includes: in response to the prediction mode of the current block being an inter prediction mode: deriving a first mode-related reference zero skip flag based on the first prediction mode of the first neighboring block and the first reference zero skip flag; deriving a second mode-related reference zero skip flag based on the second prediction mode of the second neighboring block and the second reference zero skip flag; and deriving the at least one context based on the first mode-related reference zero skip flag and the second mode-related reference zero skip flag. In some embodiments, deriving the first mode-related reference zero skip flag includes: in response to the first prediction mode being an inter prediction mode, setting the first mode-related reference zero skip flag to the value of the first reference zero skip flag; and in response to the first prediction mode not being an inter prediction mode, setting the first mode-related reference zero skip flag to a value indicating non-skip. Deriving the second mode-related reference zero skip flag includes: in response to the second prediction mode being an inter prediction mode, setting the second mode-related reference zero skip flag to the value of the second reference zero skip flag; and in response to the second prediction mode not being an inter prediction mode, setting the second mode-related reference zero skip flag to a value indicating non-skip.
[0025] In some of the above embodiments, deriving the at least one context based on the first mode-related reference zero skip flag and the second mode-related reference zero skip flag includes: deriving one context of the at least one context as the sum of the first mode-related reference zero skip flag and the second mode-related reference zero skip flag.
[0026] In some of the above embodiments, the first neighboring block and the second neighboring block may include a top neighboring block and a left neighboring block of the current block, respectively.
[0027] In some of the above embodiments, when the prediction mode of the first neighboring block is an intra copy mode, the first prediction mode is regarded as an inter prediction mode; and when the prediction mode of the second neighboring block is an intra copy mode, the second prediction mode is regarded as an inter prediction mode.
[0028] In some of the above embodiments, before signaling the zero skip flag of the current block, a prediction mode flag for indicating whether the current block is an inter-coded block is signaled.
[0029] In some of the above embodiments, when the prediction mode associated with the current block is the intra block copy mode, the prediction mode of the current block is regarded as an inter prediction mode.
[0030] In some of the above embodiments, the at least one context includes a first context and a second context, and the first context and the second context are used to encode the zero skip flag of the current block when the current block is intra-coded and inter-coded, respectively.
[0031] In some other example embodiments, a method for encoding and decoding a current block in a video is disclosed. The method may include determining a zero skip flag of the current block to indicate whether the current block is an all-zero block; based on the value of the zero skip flag, deriving a set of at least one context for decoding a prediction mode flag, the prediction mode flag indicating whether to perform intra decoding or inter decoding on the current block; and decoding the prediction mode flag of the current block according to the set of at least one context.
[0032] In some of the above embodiments, when the zero skip flag of the current block indicates that the current block is not an all-zero block, the set of at least one context includes a set of first contexts; and when the zero skip flag of the current block indicates that the current block is an all-zero block, the set of at least one context includes a set of second contexts different from the set of first contexts.
[0033] In some of the above embodiments, the set of at least one context further depends on the value of at least one reference zero skip flag of at least one neighboring block of the current block.
[0034] In some other example embodiments, a method for decoding a current block in a video is disclosed. The method may include parsing a video bitstream to identify a zero skip flag of the current block; determining whether the current block is an all-zero block according to the zero skip flag; and in response to the current block being an all-zero block, setting a prediction mode flag of the current block to indicate an inter prediction mode, rather than parsing the prediction mode flag from the video bitstream.
[0035] In some of the above embodiments, the method further includes using the prediction mode flag of the current block in decoding adjacent blocks of the current block.
[0036] Aspects of the present disclosure also provide a device or apparatus having a memory for storing computer instructions and a processor configured to execute the computer instructions to implement any of the above methods.
[0037] Aspects of the present disclosure also provide a non - transitory computer - readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform any of the above - described method embodiments for video decoding and / or encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0039] Figure 1A A schematic diagram showing an exemplary subset of intra - prediction direction modes.
[0040] Figure 1B An illustration showing exemplary intra - prediction directions.
[0041] Figure 2 A schematic diagram showing spatial merge candidates for motion - vector prediction for a current block and its surroundings in one example.
[0042] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment.
[0043] Figure 4 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment.
[0044] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment.
[0045] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment.
[0046] Figure 7 A block diagram of a video encoder according to another example embodiment.
[0047] Figure 8 A block diagram of a video decoder according to another example embodiment.
[0048] Figure 9 A scheme of coding block partitioning according to an example embodiment of the present disclosure.
[0049] Figure 10 Another scheme of coding block partitioning according to an example embodiment of the present disclosure.
[0050] Figure 11 Another scheme of coding block partitioning according to an example embodiment of the present disclosure.
[0051] Figure 12 Another scheme of coding block partitioning according to an example embodiment of the present disclosure.
[0052] Figure 13 Shows a scheme for partitioning an encoded block into a plurality of transform blocks and the encoding order of the transform blocks according to an exemplary embodiment of the present disclosure.
[0053] Figure 14 Shows another scheme for partitioning an encoded block into a plurality of transform blocks and the encoding order of the transform blocks according to an exemplary embodiment of the present disclosure.
[0054] Figure 15 Shows another scheme for partitioning an encoded block into a plurality of transform blocks according to an exemplary embodiment of the present disclosure.
[0055] Figure 16 Shows a flowchart according to an exemplary embodiment of the present disclosure.
[0056] Figure 17 Shows a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0057] Figure 3 Is a simplified block diagram of a communication system (300) according to an embodiment disclosed in the present application. The communication system (300) includes a plurality of terminal devices, and the terminal devices can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first terminal device (310) and a second terminal device (320) interconnected through a network (350). In Figure 3 In the embodiment, the first terminal device (310) and the second terminal device (320) perform unidirectional data transmission. For example, the first terminal device (310) can encode video data (such as a video picture stream collected by the first terminal device (310)) for transmission to the second terminal device (320) through the network (350). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The second terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video data, and display video pictures according to the recovered video data. Unidirectional data transmission is relatively common in applications such as media services.
[0058] In another embodiment, a communication system (300) includes a third terminal device (330) and a fourth terminal device (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conference. For bidirectional data transmission, each of the third terminal device (330) and the fourth terminal device (340) can encode video data (such as a video picture stream captured by the terminal device) for transmission over a network (350) to the other of the third terminal device (330) and the fourth terminal device (340). Each of the third terminal device (330) and the fourth terminal device (340) can also receive the encoded video data transmitted by the other of the third terminal device (330) and the fourth terminal device (340), can decode the encoded video data to recover the video data, and can display video pictures on an accessible display device based on the recovered video data.
[0059] In Figure 3 an embodiment, the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340) can be servers, personal computers, and smart phones, but the scope of application of the basic principles disclosed in this application is not limited thereto. The embodiments disclosed in this application are applicable to desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, etc. The network (350) represents any number or type of network that conveys encoded video data between the first terminal device (310), the second terminal device (320), the third terminal device (330), and the fourth terminal device (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched, packet-switched, and / or other types of channels. The network can include a telecommunications network, a local area network, a wide area network, and / or the Internet. For the purposes of this application, unless explicitly explained herein, the architecture and topology of the network (350) may be irrelevant to the operations disclosed in this application.
[0060] As an example, Figure 4 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The subject matter disclosed in this application is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, compressed video storage on digital media including CDs, DVDs, memory sticks, etc.
[0061] A video streaming system may include an acquisition subsystem (413), which may include a video source (401) such as a digital camera to create an uncompressed video picture or image stream (402). In an embodiment, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. Compared with the encoded video data (404) (or encoded video bitstream), the uncompressed video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (402) may be processed by an electronic device (420), which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (402), the encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (404) (or encoded video bitstream (404)), which may be stored on a streaming server (405) for future use or directly used for downstream video devices (not shown). One or more streaming client subsystems, such as Figure 4 the client subsystem (406) and the client subsystem (408) in
[0062]
[0063] may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an uncompressed output video picture stream (411) that can be presented on a display (412) (such as a display screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in the present disclosure. In some streaming systems, the encoded video data (404), video data (407), and video data (409) (such as video bitstreams) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application may be used in the context of the VVC standard and other video coding standards.
[0062] It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0063] Figure 5 is a block diagram of a video decoder (510) according to an embodiment disclosed below of the present application. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 the video decoder (410) in the embodiment.
[0064] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), and the channel may be a hardware / software link leading to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective processing circuits (not labeled). The receiver (531) may separate the encoded video sequence from other data. To prevent network jitter, a buffer memory (515) may be configured between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided outside the video decoder (510) and separated from the video decoder (510) (not labeled). In other applications, a buffer memory (not labeled) is provided outside the video decoder (510) to, for example, prevent network jitter, and another buffer memory (515) may exist inside the video decoder (510) to, for example, handle the playback timing. And when the receiver (531) receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the buffer memory may be relatively large. Such a buffer memory may have an adaptive size and may be at least partially implemented in an operating system or a similar element (not labeled) outside the video decoder (510).
[0065] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device (such as a display screen), such as the display device (512). The display device may or may not be a component of the electronic device (530), but may be coupled to the electronic device (530), as Figure 5 shown. The control information for the display device may be a Supplemental Enhancement Information (SEI message) or a parameter set segment (not labeled) of Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., transform coefficients), quantizer parameter values, motion vectors, and so on.
[0066] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0067] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and the multiple processing or functional units below are not described.
[0068] In addition to the functional blocks already mentioned, the video decoder (510) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these functional units interact closely with each other and can be integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, a conceptual subdivision of the functional units is adopted in the following disclosure.
[0069] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives the quantized transform coefficients as symbols (521) and control information from the parser (520), including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).
[0070] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; for example, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed surrounding block information stored in the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) adds the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.
[0071] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for inter-picture prediction. After motion-compensating the extracted samples according to the symbol (521), these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 is referred to as residual samples or a residual signal), thereby generating output sample information. The motion compensation prediction unit (553) obtaining prediction samples from an address within the reference picture memory (557) may be controlled by a motion vector, and the motion vector is in the form of the symbol (521) for use by the motion compensation prediction unit (553), where the symbol (521) includes, for example, X, Y components (displacements) and a reference picture component (time). Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and motion compensation may also be associated with a motion vector prediction mechanism, and so on.
[0072] The output samples of the aggregator (555) may be employed by various loop filtering techniques in the loop filter unit (554). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters may be used for the loop filter unit (556) as a symbol (521) from the parser (520). However, in other embodiments, the video compression techniques may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. Some types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in further detail below.
[0073] The output of the loop filter unit (556) may be a sample stream that may be output to the display device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.
[0074] Once fully reconstructed, certain encoded pictures may be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.
[0075] A video decoder (510) may perform decoding operations according to a predetermined video compression technique employed, for example, in the ITU-T H.265 standard. An encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, the complexity of the encoded video sequence is within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0076] In one embodiment, a receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0077] Figure 6 is a block diagram of a video encoder (603) according to an embodiment disclosed in the present application. The video encoder (603) is provided in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the embodiment.
[0078] The video encoder (603) may receive video samples from a video source (601) (not Figure 6 part of the electronic device (620) in the embodiment), and the video source may capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) may be implemented as part of the electronic device (620).
[0079] A video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCb, RGB, XYZ, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures or images, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. being used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0080] According to an embodiment, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Performing an appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to these units. For simplicity, the couplings are not labeled in the figure. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum allowed motion vector search range, etc. The controller (650) can be used for other suitable functions that involve optimizing the video encoder (503) for a particular system design.
[0081] In some embodiments, the video encoder (603) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data, even though the embedded decoder 633 processes the encoded video stream through the source encoder 630 without entropy encoding (since in the video compression techniques contemplated in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is used to improve the encoding quality.
[0082] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, for example, which has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5 , when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) of the encoder.
[0083] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, this application sometimes focuses on the decoder operation, which is related to the decoding part of the encoder. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. Only some regions or aspects of the encoder will be described in more detail below.
[0084] During operation, in some embodiments, the source encoder (630) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes an input picture with reference to one or more previously encoded pictures designated as "reference pictures" in a video sequence. In this way, the coding engine (632) encodes the difference (residue) between a pixel block in a color channel of the input picture and a pixel block of a reference picture, which may be selected as a prediction reference for the input picture. The terms "residue" and its adjectival form "residual" may be used interchangeably.
[0085] The local video decoder (633) may decode the encoded video data of a picture that may be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the coding engine (632) may be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 6 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has the same content (no transmission errors) as the reconstructed reference picture that will be obtained by a remote video decoder.
[0086] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new picture. The predictor (635) may operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0087] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.
[0088] The outputs of all the above functional units may be entropy encoded in the entropy encoder (645). The entropy encoder (645) losslessly compresses the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0089] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over the communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0090] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures may typically be assigned to any of the following picture types:
[0091] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0092] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and reference index to predict the sample values of each block.
[0093] A bi - predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0094] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures. For other purposes, the source picture or the intermediate processed picture can be subdivided into other types of blocks. The partitioning of the encoded blocks and other types of blocks can follow or can not follow the same way, as described in further detail below.
[0095] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0096] In an embodiment, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0097] The captured video can be a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the specific picture being encoded / decoded is segmented into blocks, and the specific picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension that identifies the reference picture.
[0098] In some embodiments, bidirectional prediction techniques can be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past or future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be jointly predicted by a combination of the first reference block and the second reference block.
[0099] In addition, the merge mode technique can be used in inter-picture prediction to improve coding efficiency.
[0100] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a video picture sequence is segmented into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three parallel coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) by a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs. Each of one or more 32×32 blocks can be further divided into 4 16×16 pixel CUs. In one embodiment, each CU can be analyzed during encoding to determine the prediction type for the CU among various prediction types, such as inter-frame prediction type or intra-frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. The splitting of the CU into PUs (or PBs of different color channels) can be performed in various spatial modes. For example, a luminance or chrominance PB can include a matrix of sample values (e.g., luminance values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, and 16x8 samples, etc.
[0101] Figure 7FIG. is a diagram of a video encoder (703) according to another illustrative embodiment disclosed in the present application. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and encode the processing block into an encoded picture that is part of an encoded video sequence. In the present embodiment, the video encoder (703) is configured to replace Figure 4 the video encoder (303) in the embodiment.
[0102] In one embodiment, the video encoder (703) receives a matrix of sample values for a processing block, which may be, for example, a prediction block of 8×8 samples. Then, the video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-predictive mode to encode the processing block. When it is determined to encode the processing block in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into the encoded picture; and when it is determined to encode the processing block in the inter mode or the bi-predictive mode, the video encoder (703) may use inter prediction or bi-predictive techniques, respectively, to encode the processing block into the encoded picture. In some illustrative embodiments, the merge mode may be a sub-mode of inter-picture prediction, where a motion vector is derived from one or more motion vector prediction values without the aid of encoded motion vector components external to the prediction value. In certain other illustrative embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) includes components explicitly shown in Figure 7 , such as a mode decision module (not shown) for determining the prediction mode of the processing block.
[0103] In Figure 7 the embodiment, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the exemplary arrangement in Figure 7 .
[0104] The inter encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundancy information according to inter coding techniques, a motion vector, merge mode information), and calculate an inter prediction result (e.g., a prediction block) based on the inter prediction information using any suitable technique. In some embodiments, the reference picture is decoded using a decoding unit 633 (as shown in Figure 6 the example encoder 620 embedded inFigure 7 As shown by the residual decoder 728 (see details below), a decoded reference picture decoded based on the encoded video information.
[0105] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture in some cases, generate quantization coefficients after transformation, and also generate intra prediction information in some cases (e.g., based on intra prediction direction information of one or more intra coding techniques). The intra encoder (722) calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.
[0106] The general controller (721) is used to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the prediction mode for the block is the inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and add the inter prediction information to the bitstream.
[0107] The residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is used to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is used to convert the residual data from the time domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various illustrative embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and used as a reference picture.
[0108] An entropy encoder (725) is used to format a bitstream to produce encoded blocks and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (such as intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. It should be noted that when encoding a block in the merge submode of the inter mode or the bi - directional prediction mode, there is no residual information.
[0109] Figure 8 is a diagram of a video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (810) is used to replace Figure 4 the video decoder (410) in the embodiment.
[0110] In Figure 8 an embodiment, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as schematically arranged in Figure 8 .
[0111] The entropy decoder (871) can be used to reconstruct certain symbols according to the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode used to encode the block (such as the intra mode, the inter mode, the bi - directional prediction mode, the merge submode, or another submode), prediction information (such as intra prediction information or inter prediction information) that can identify certain samples or metadata for prediction by the intra decoder (872) or the inter decoder (880), residual information in the form of, for example, quantized transform coefficients, and so on. In an embodiment, when the prediction mode is the inter or bi - directional prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is the intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be de - quantized and provided to the residual decoder (873).
[0112] The inter decoder (880) is used to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0113] The intra decoder (872) is used to receive the intra prediction information and generate a prediction result based on the intra prediction information.
[0114] The residual decoder (873) is used to perform inverse quantization to extract the dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (873) may also use certain control information (for obtaining the quantizer parameter QP), which may be provided by the entropy decoder (871) (the data path is not marked as it is only low-data-volume control information).
[0115] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture, and the reconstructed picture may be part of the reconstructed video. It should be noted that other suitable operations such as deblocking operations can be performed to improve the visual quality.
[0116] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810). In an embodiment, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).
[0117] Turning to encoding block partitioning, and in some example embodiments, a predetermined pattern can be applied. As Figure 9 shown, an example 4-way partitioning tree can be adopted starting from a first predetermined level (e.g., 64×64 block level) and going down to a second predefined level (e.g., 4×4 level). For example, the basic block can be limited to four partitioning options indicated by 902, 904, 906, and 908, where the partition designated as R allows recursive partitioning, Figure 9 the same partitioning tree shown can be repeated at a lower scale until the lowest level (e.g., 4×4 level). In some embodiments, additional restrictions can be applied to Figure 9 the partitioning scheme. In Figure 9 the embodiments, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) can be allowed, but they may not be allowed to be recursive, while square partitions are allowed to be recursive. If needed, the final set of encoded blocks is generated according to Figure 9 the recursive partitioning. Such a scheme can be applied to one or more color channels.
[0118] Figure 10 shows another example of a predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. As Figure 10 shown, an example 10-way partitioning structure or pattern can be predefined. The root block can start at a predefined level (e.g., from the 128×128 level, or the 64×64 level). Figure 10 The example partitioning structure of includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Figure 10 1002, 1004, 1006, and 1008 in the second row of indicate partition types with 3 sub-partitions and can be referred to as "T-shaped" partitions. The "T-shaped" partitions 1002, 1004, 1006, and 1008 can be referred to as left T-shaped, top T-shaped, right T-shaped, and bottom T-shaped. In some embodiments, Figure 10 the rectangular partitions of do not allow further subdivision. The coding tree depth can be further defined to indicate the depth of splitting from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., a 128×128 block) can be set to 0, and after the root block is further split once according to the Figure 10 pattern, the coding tree depth is increased by 1. In some embodiments, only all the square partitions in 1010 can be recursively partitioned to the next level of the partitioning tree according to the Figure 10 pattern. In other words, for the square partitions with patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If needed, the final set of coded blocks is generated according to the Figure 10 recursive partitioning. Such a scheme can be applied to one or more color channels.
[0119] After dividing or partitioning the basic block according to any of the above partitioning processes or other processes, similarly, a final set of partitions or coded blocks can be obtained. Each of these partitions can be at one of various partitioning levels. Each partition can be referred to as a coded block (CB). For the above various example partitioning embodiments, each resulting CB can be of any allowed size and partitioning level. They are called coded blocks because they can form units for which some basic encoding / decoding decisions can be made, and the encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest level in the final partition represents the depth of the coded block partitioning tree. The coded block can be a luminance coded block or a chrominance coded block.
[0120] In some other example embodiments, a quadtree structure can be used to recursively divide a basic luminance and chrominance block into coding units. Such a division structure can be referred to as a coding tree unit (CTU), and the coding tree unit is divided into coding units (CUs) by using the quadtree structure so that the partitioning adapts to various local characteristics of the underlying CTU. In such embodiments, an implicit quadtree division is performed at the picture boundary such that the blocks will maintain the quadtree division until the size fits the picture boundary. The term CU is used to collectively refer to the units of luminance and chrominance coding blocks (CBs).
[0121] In some embodiments, a CB can be further partitioned. For example, for the purpose of intra-frame or inter-frame prediction during the encoding and decoding processes, a CB can be further partitioned into multiple prediction blocks (PBs). In other words, a CB can be further divided into different sub-partitions where separate prediction decisions / configurations can be made. In parallel, for the purpose of depicting the level at which the transform or inverse transform of the video data is performed, a CB can be further partitioned into multiple transform blocks (TBs). The partitioning schemes of a CB into PBs and TBs can be the same or can be different. For example, each partitioning scheme can be performed using its own process based on various characteristics of the video data, for example. In some example embodiments, the PB and TB partitioning schemes can be independent. In some other example embodiments, the PB and TB partitioning schemes and boundaries can be related. In some embodiments, for example, a TB can be partitioned after the PB partitioning, specifically, each PB is further partitioned into one or more TBs after the partitioning of the coding block is determined. For example, in some embodiments, a PB can be divided into one, two, four, or other numbers of TBs.
[0122] In some embodiments, to partition a base block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and the chrominance channel may be treated differently. For example, in some embodiments, coding blocks of the luminance channel may be allowed to be partitioned into prediction blocks and / or transform blocks, while coding blocks of the chrominance channel may not be allowed to be partitioned into prediction blocks and / or transform blocks. In such embodiments, the transformation and / or prediction of luminance blocks may thus be performed only at the coding block level. For another example, the minimum transform block size of the luminance channel and the chrominance channel may be different. For example, coding blocks of the luminance channel may be allowed to be partitioned into smaller transform blocks and / or prediction blocks than those of the chrominance channel. For yet another example, the maximum depth of partitioning a coding block into transform blocks and / or prediction blocks may be different between the luminance channel and the chrominance channel. For example, coding blocks of the luminance channel may be allowed to be partitioned into deeper transform blocks and / or prediction blocks than those of the chrominance channel. For a specific example, a luminance coding block may be partitioned into transform blocks of multiple sizes, which may be represented by a recursive partitioning down to up to 2 levels, and transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4×4 to 64×64 may be allowed. However, for chrominance blocks, only the maximum possible transform blocks specified for luminance blocks may be allowed.
[0123] In some example embodiments for partitioning coding blocks into PBs, the depth, shape, and / or other characteristics of partitioning the PBs may depend on whether the PB is intra-coded or inter-coded.
[0124] In various example scenarios, partitioning a coding block (or prediction block) into transform blocks may be implemented, which includes but is not limited to quadtree splitting and predefined pattern splitting, either recursively or non-recursively, and additional consideration of transform blocks at the boundaries of the coding block or prediction block. Generally, the resulting transform blocks may be at different splitting levels, may not have the same size, and may not need to be square-shaped (e.g., they may be rectangles with some allowed sizes and aspect ratios).
[0125] In some embodiments, a coded partition tree scheme or structure may be used. The coded partition tree schemes for the luminance and chrominance channels need not be the same. In other words, the luminance and chrominance channels may have separate coded tree structures. Further, whether the luminance and chrominance channels use the same or different coded partition tree structures and the actual coded partition tree structure to be used may depend on whether the slice being coded is a P, B, or I slice. For example, for an I slice, the chrominance channel and the luminance channel may have separate coded partition tree structures or coded partition tree structure patterns, while for a P or B slice, the luminance and chrominance channels may share the same coded partition tree scheme. When separate coded partition tree structures or patterns are applied, the luminance channel may be partitioned into CUs by one coded partition tree structure, and the chrominance channel may be partitioned into chrominance CUs by another coded partition tree structure.
[0126] Specific example embodiments of coded block and transform block partitioning are described below. In such example embodiments, a base coded block may be partitioned into coded blocks using the recursive quadtree splitting described above. At each level, whether further quadtree splitting of a particular partition should continue may be determined by local video data characteristics. The resulting CUs may be at various quadtree splitting levels of various sizes. A decision may be made at the CU level (or CU, for all three color channels) as to whether to use inter-picture (temporal) or intra-picture (spatial) prediction to code a picture region. Each CU may be further partitioned into one, two, four, or some other number of PBs depending on the PB partition type. Within one PB, the same prediction process may be applied, and relevant information may be sent to the decoder based on the PB. After obtaining the residual block by applying the prediction process based on the PB partition type, the CU may be partitioned into TUs according to another quadtree structure similar to the coding tree of the CU. In this particular embodiment, a CU or TU may be, but is not limited to, square. Further, in this particular example, a PB may be square or rectangular in shape for inter-frame prediction and only square for intra-frame prediction. A coded block may be further partitioned into, for example, four square-shaped TUs. Each TU may be further recursively partitioned (using quadtree splitting) into smaller TUs, referred to as a Residual Quad-Tree (RQT).
[0127] Another specific example for partitioning a base coded block into CUs and other PBs and / or TUs is described below. For example, a quadtree with a nested multi-type tree having a binary and ternary split segmentation structure, rather than using, such as Figure 10The multiple partition unit types shown in. The separation of the CB, PB, and TB concepts (i.e., partitioning the CB into PB and / or TB, and partitioning the PB into TB) can be abandoned, unless when a CB with a size too large for the maximum transform length is needed, in which case such a CB may need to be further segmented. This example partition scheme can be designed to support more flexibility in the CB partition shape, such that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, a CU can have a square or rectangular shape. Specifically, a coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a multi-type tree structure. Figure 11 An example of a multi-type tree structure is shown in. Specifically, Figure 11 The example multi-type tree structure includes four splitting types, which are referred to as vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). Then, the CB corresponds to the leaf of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, this segmentation is used for prediction and transformation processing without any further partitioning. This means that, in most cases, the CB, PB, and TB have the same block size in a quadtree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is less than the width or height of the color component of the CB.
[0128] Figure 12 An example of a quadtree with a nested multi-type tree coding block structure for block partitioning of one CTB is shown in. More specifically, Figure 12 It shows that the CTB 1200 is quadtree split into four square partitions 1202, 1204, 1206, and 1208. It is decided to further use Figure 11 the multi-type tree structure to split each of the quadtree split partitions. In Figure 12 the example, the partition 1204 is not further split. The partitions 1202 and 1208 each adopt another quadtree split. For the partition 1202, the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split respectively adopt the third-level split of the quadtree, Figure 11 1104 of, no split, and Figure 11 1108 of. The partition 1208 adopts another quadtree split, and the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split respectively adopt Figure 11 the third-level split of 1106 of, no split, no split, and Figure 11of 1104. Two in the sub - partitions of the upper - left third - level partition of 1208 are further divided according to 1104 and 1108. Partition 1206 is divided into two partitions using the second - level segmentation pattern of 1102 according to Figure 11 and these two partitions are further third - level segmented according to 1108 and 1102 of Figure 11 . According to 1104 of Figure 11 , a fourth - level segmentation is further applied to one of them.
[0129] For the above - specific example, the maximum luminance transform size can be 64×64, and the maximum supported chrominance transform size can be different from the luminance at, for example, 32×32. When the width or height of a luminance - coded block or a chrominance - coded block is greater than the maximum transform width or height, the luminance - coded block or the chrominance - coded block can be automatically segmented along the horizontal and / or vertical directions to meet the transform - size limit in that direction.
[0130] In a specific example for partitioning a base - coded block into the above - mentioned CBs, the coding - tree scheme can support the ability for luminance and chrominance to have separate block - tree structures. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can share the same coding - tree structure. For example, for I slices, luminance and chrominance can have separate coded - block - tree structures. When applying the separate - block - tree mode, the luminance CTB is partitioned into CBs through one coding - tree structure, and the chrominance CTB is partitioned into chrominance CBs through another coding - tree structure. This means that a CU in an I slice can be composed of coded blocks of the luminance component or coded blocks of two chrominance components, and a CU in a P or B slice is always composed of coded blocks of all three color components, unless the video is monochromatic.
[0131] Example embodiments for partitioning coded blocks or prediction blocks into transform blocks and the coding order of transform blocks are described in further detail below. In some example embodiments, transform partitioning can support multiple shapes, e.g., transform blocks of 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the transform - block size ranges from, for example, 4×4 to 64×64. In some embodiments, if the coded block is less than or equal to 64×64, the transform - block partitioning can be applied only to the luminance component, such that for chrominance blocks, the transform - block size is the same as the coded - block size. Otherwise, if the coded - block width or height is greater than 64, then the luminance and chrominance coded blocks can be implicitly segmented into multiple min(W,64)×min(H,64) and min(W,32)×min(H,32) transform blocks, respectively.
[0132] In some example embodiments, for both intra-coded blocks and inter-coded blocks, the coded blocks can be further partitioned into a plurality of transform blocks having a partition depth of up to a predetermined number of levels (e.g., 2 levels). The transform block partition depth and size can be related. An example mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below. Table 1: Transform Partition Size Settings
[0133] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform split can create four 1:1 square sub-transform blocks. The transform partition can stop, for example, at 4×4. Thus, the transform size of the current depth 4×4 corresponds to the same size 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform split will create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next-level transform split will create two 1:2 / 2:1 sub-transform blocks.
[0134] In some example embodiments, additional restrictions can be applied to the luminance component of intra-coded blocks. For example, for each level at which transform partitioning is performed, all sub-transform blocks can be restricted to have the same size. For example, for a 32×16 coded block, the level 1 transform split creates two 16×16 sub-transform blocks, and the level 2 transform split creates eight 8×8 sub-transform blocks. In other words, the second-level split must be applied to all first-level sub-blocks to keep the transform unit sizes equal. An Figure 13 example of the transform block partition for an intra-coded square block shown in Table 1, and the coding order indicated by the arrow diagram, is shown. Specifically, 1302 shows the square coded block. The first-level split into 4 equal-sized transform blocks according to Table 1 is shown in 1304, where the coding order is indicated by the arrows. In 1306, the second-level split of all first-level equal-sized blocks into 16 equal-sized transform blocks according to Table 1 is shown, where the coding order is indicated by the arrows.
[0135] In some example embodiments, the above restrictions on intra-coding may not be applied to the luminance component of inter-coded blocks. For example, after the first-level transform split, any one of the sub-transform blocks can be further independently split one more level. Thus, the resulting transform blocks may or may not have the same size. An Figure 14 example of the transform blocks into which an inter-coded block is split, with the coding order, is shown. In Figure 14In the example of Figure 14 , according to Table 1, the inter-frame coded block 1402 is divided into transform blocks at the second level. At the first level, the inter-frame coded block is divided into four transform blocks of equal size. Then, as shown in 1404, only one (not all) of the four transform blocks is further divided into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes. The example coding order of these 7 transform blocks is shown by the arrows in
[0136] In some example embodiments, for one or more chrominance components, some additional restrictions on the transform blocks may be applied. For example, for one or more chrominance components, the transform block size may be as large as the coded block size, but not less than a predefined size, such as 8×8.
[0137] In some other example embodiments, for coded blocks with a width (W) or height (H) greater than 64, both the luminance and chrominance coded blocks may be implicitly divided into multiple min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively.
[0138] Figure 15 Another alternative example scheme for partitioning a coded block or a prediction block into transform blocks is further shown. As Figure 15 shown, instead of using recursive transform partitioning, a predefined set of partitioning types may be applied to the coded block according to the transform type of the coded block. In the specific example shown in Figure 15 , one of 6 example partitioning types may be applied to divide the coded block into various numbers of transform blocks. Such a scheme may be applied to coded blocks or prediction blocks.
[0139] More specifically, Figure 15 the partitioning scheme in Figure 15 provides up to 6 partitioning types for any given transform type as shown in Figure 15 . In this scheme, for example, a transform type may be assigned to each coded block or prediction block based on the rate-distortion cost. In the example, the partitioning type assigned to a coded block or a prediction block may be determined based on the transform partitioning type of the coded block or the prediction block. A specific partitioning type may correspond to the transform block division size and pattern (or partitioning type), as shown by the 6 partitioning types illustrated in
[0140] ·PARTITION_NONE: Assign a transform size equal to the block size.
[0141] ·PARTITION_SPLIT: The allocated transform size is 1 / 2 of the width of the block size and 1 / 2 of the height of the block size.
[0142] ·PARTITION_HORZ: The allocated transform size has the same width as the block size and 1 / 2 of the height of the block size.
[0143] ·PARTITION_VERT: The allocated transform size has 1 / 2 of the width of the block size and the same height as the block size.
[0144] ·PARTITION_HORZ4: The allocated transform size has the same width as the block size and 1 / 4 of the height of the block size.
[0145] ·PARTITION_VERT4: The allocated transform size has 1 / 4 of the width of the block size and the same height as the block size.
[0146] In the above example, all partition types as shown in Figure 15 include a unified transform size for the partition transform block. This is merely an example and not a limitation. In some other embodiments, mixed transform block sizes may be used for the partition transform blocks of a particular partition type (or mode).
[0147] Some example embodiments directed to specific types of signaling and entropy coding / decoding for various blocks / units, for each intra-frame and inter-frame coded block / unit, a flag, i.e., the skip_txfm flag, can be signaled in the coded bitstream as shown in the example syntax of Table 2 below and is represented by the read_skip() function which is used to retrieve these flags from the bitstream. This flag can indicate whether the transform coefficients are all zero in the current coding unit. The skip_txfm flag can also be referred to as the zero-skip flag. In some example embodiments, if this flag is signaled with a value such as 1, then for any coded block in the coding unit, no other transform coefficient related syntax (e.g., EOB (end of block)) needs to be signaled, and the zero-skip flag can be derived as: a value or data structure predefined for and associated with zero transform coefficient blocks. For inter-frame coded blocks, as shown in the example of Table 2, this flag can be signaled after the skip_mode flag which indicates that the coding unit can be skipped for various reasons. When the skip_mode flag value is true, the coding unit should be skipped and no skip_txfm flag needs to be signaled, and the skip_txfm flag is inferred to be 1. Otherwise, if the skip_mode flag is false, then more information about the coded block / unit will be included in the bitstream and the skip_txfm flag will additionally be signaled to indicate whether the coded block / unit is all zero. Table 2: Skip mode and skip syntax Intra-frame mode information syntax Inter-frame mode information syntax Skip syntax
[0148] In this way, the skip_txfm flag can be used to indicate whether the block associated with the flag is all-zero and can be skipped. For example, if the flag is signaled as 1, then the block may contain all-zero coefficients. Otherwise, if the flag is signaled as zero, then at least some of the coefficients in the block are non-zero. In some embodiments, the block can be partitioned or divided into transform blocks in the various ways described above. Each of the t transform blocks can additionally be associated with an EOB skip flag that can be signaled in the encoded bitstream. If the EOB skip flag for a particular transform block is signaled as 1, then the EOB for the transform block is not included in the bitstream, and the transform block can be skipped because it has no non-zero coefficients. However, when the EOB skip flag for the transform block is signaled as 0, then this indicates that there is at least one non-zero coefficient for the transform block and that the transform block cannot be skipped. In the exemplary embodiments described below, the focus is on signaling and encoding the skip_txfm flag. For simplicity, the above skip_txfm flag may be referred to as the skip flag.
[0149] A context set for entropy coding of the skip_txfm flag (or zero skip flag) in the bitstream can be designed. A context for encoding the skip_txfm flag of a particular current coding block can be selected from the context set. The selection of the context can be indicated by an index. The selection from the context set can depend on various factors. For example, the selection of the coding context for the skip_txfm of the current block can depend on the skip_txfm flag value of at least one of the neighboring blocks of the current block. For example, at least one neighboring block can include the upper and / or left neighboring blocks of the current block. For example, the selection can be made from a total of three different contexts. In a specific exemplary embodiment, if neither the upper nor the left neighboring block is encoded with a non-zero skip_txfm flag, then the context index value 0 is used to encode the skip_txfm flag of the current block. If one of the upper and left neighboring blocks is encoded with a non-zero skip_txfm flag, then the context index value 1 can be used to encode the skip_txfm flag of the current block. If both the upper and left neighboring blocks are encoded with non-zero skip_txfm flags, then the context index value 2 is used to encode the skip_txfm flag of the current block. This embodiment is based on the statistical observation that the more neighboring blocks with all-zero coefficients, the more likely the current block is also to have all-zero coefficients. An example process for selection from the context set for encoding the current skip_txfm flag is shown in Table 3 below. Table 3. Neighboring-block-related context derivation for skip_txfm flag; cdf is given by TileSkipCdf[ctx]
[0150] The following additional example embodiments focus on schemes for further improving the entropy coding efficiency of the skip_txfm flag and some other flags in a bitstream by exploiting some statistical correlations between these flags and other flags within or across coding blocks (e.g., between adjacent blocks). Specifically, the context for encoding these flags of a coding block is designed to depend on the prediction mode (intra-frame or inter-frame prediction mode) of the coding block and / or its adjacent coding blocks.
[0151] For example, some of these example embodiments can be based on the following observation: in addition to the correlation of skip_txfm flag values between adjacent blocks, the probability that the skip_txfm value of the current block is equal to 1 can also depend to a large extent on whether the current block is an intra-coded block or an inter-coded block. By considering the prediction mode when determining the context for encoding the skip_txfm flag, higher coding efficiency can be obtained due to the possible correlation between the skip_txfm flag value and the prediction mode.
[0152] The following various example embodiments can be used individually or in any order of combination. Further, each of these embodiments can be embodied as part of an encoder and / or a decoder, and can be implemented in hardware or software. For example, these embodiments can be hard-coded in a dedicated processing circuit (e.g., one or more integrated circuits). In another example, these embodiments can be implemented by one or more processors executing a program stored in a non-volatile computer-readable medium. Additionally, the term "block" can refer to any unit of video information, and the term "block size" can refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width: height, or height: width).
[0153] In some general example embodiments, multiple contexts can be designed to encode the skip_txfm flag of the current block of video in a bitstream. The context for signaling and encoding the skip_txfm flag of a particular current block can depend on the prediction mode of the current block, e.g., whether the current block is intra-coded or inter-coded.
[0154] In some example embodiments, when the current block is an inter-coded block, the context for signaling and coding the skip_txfm flag may depend on two conditions: a) whether at least one neighboring block of the current block (e.g., the upper and / or left block) is intra-coded; and b) the skip_txfm flag of at least one neighboring block of the current block (e.g., the upper or left neighboring block).
[0155] In a specific example embodiment using both the upper and left neighboring blocks, if the current block is an inter-coded block, then the context selection (context index value) embodiment for coding the current skip_txfm flag based on the neighboring skip_txfm flags above can be further modified according to the following equation:
[0156] above_skip_txfm = is_inter(above_mi)? above_mi->skip_txfm : 0 (1)
[0157] left_skip_txfm = is_inter(left_mi)? left_mi->skip_txfm : 0 (2)
[0158] context = above_skip_txfm + left_skip_txfm (3)
[0159] where is_inter(left_mi) or is_inter(above_mi) returns a boolean value indicating whether the left block or the upper block is an inter-coded block, above_mi->skip_txfm represents the transform skip flag (skip_txfm) of the upper block, and left_mi->skip_txfm represents the transform skip flag of the left block. These neighboring block skip_txfm flags can be referred to as reference skip_txfm flags, as references for generating "above_skip_txfm" and "above_skip_txfm", which are called prediction mode-related neighboring skip_txfm flags. The context selection index can be the sum of the prediction mode-related neighboring skip_txfm flags.
[0160] After the above example, there can be three example contexts represented by contexts 0, 1, and 2. When the current block is an intra-prediction block, determining the context for encoding the current skip_txfm flag can follow the example process in Table 3 or other processes. However, when the current block is an inter-prediction block, the context selection can be modified according to the above equation such that the skip_txfm flags of the neighboring blocks considered in Table 3 are replaced by the prediction-mode-dependent neighboring skip_txfm flags in the above equation. The prediction-mode-dependent neighboring skip_txfm flags essentially represent the actual neighboring skip_txfm flags modified depending on the prediction mode of the neighboring blocks. In particular, if the neighboring block is an inter-coded block, then the actual skip_txfm flag of the neighboring block will be kept as the prediction-mode-dependent skip_txfm flag, and otherwise, if the neighboring block is an intra-coded block, then the prediction-mode-dependent skip_txfm flag of the neighboring block will be set to 0 (as if the neighboring block is not all-zero coefficients, regardless of whether this is actually true). Essentially, the above embodiments utilize the following observation: when one or more neighboring blocks are intra-predicted, then it is more likely that the current block is not all non-zero.
[0161] The above equation is only for the example of generating three different context indices pointing to three different contexts. Other mathematical formulas can be used for this discrimination, and there can be more than three different contexts and corresponding context indices.
[0162] In some other or further example embodiments, the context for signaling and encoding the skip_txfm flag can depend on the current block prediction mode, for example, depending on the following conditions: a) whether the current block is an intra- or inter-coded block; b) the skip_txfm value of at least one neighboring block (e.g., the upper and / or left neighboring block).
[0163] In some example embodiments of the above example embodiments, if the current block (or neighboring block) is in the intra-block copy mode, then when signaling the skip_txfm flag, the current block (or neighboring block) is regarded as an inter-coded block rather than an intra-prediction block.
[0164] In some of the above example embodiments, a flag called the is_inter flag can be used to indicate whether the current block is an inter-coded block. Such a flag can be signaled prior to the skip_txfm flag in the bitstream so that the decoder can first parse its value and use it to determine the decoding context for the skip_txfm flag following in the bitstream.
[0165] In some example embodiments, if the current block is an intra-coded block, then a set of contexts can be used to signal the skip_txfm flag. Otherwise, another set of contexts can be used to signal the skip_txfm flag. In other words, two sets of contexts can be designed, and the set of contexts for signaling can be first selected by the prediction mode of the current block, and then each set of contexts can be further selected based on the other information mentioned above (e.g., adjacent skip_txfm flags and adjacent prediction modes).
[0166] In the above embodiments, the context for signaling the skip_txfm flag of the current block and for encoding it can depend on its prediction mode and / or the prediction mode of its adjacent blocks, and / or the skip_txfm flags of its adjacent blocks. In some other example embodiments, the signaled is_inter flag and its encoding can depend on the value of the skip_txfm flag of the current block, where the is_inter flag is used to indicate whether the current block is an inter-coded block.
[0167] For example, if the skip_txfm flag of the current block is equal to 0 (indicating that at least one coefficient in the current block is non-zero), then a set of contexts can be used to signal the is_inter flag. Otherwise, another set of contexts can be used to signal the is_inter flag. This can be based on the following observation: Statistically, blocks with non-zero coefficients can generally have similar prediction modes, and similarly, blocks with all-zero coefficients can also generally have similar prediction modes.
[0168] For another further example, the context for signaling the is_inter flag can alternatively or additionally depend on the adjacent skip_txfm. For example, such a context can depend on the following example conditions: a) whether the value of the skip_txfm flag of the current block is zero; b) whether the value of the skip_txfm flag of at least one adjacent block (e.g., the upper / left block) is zero. This scheme substantially utilizes the statistical correlation between the prediction mode and the coefficient characteristics (all-zero or otherwise).
[0169] For these embodiments, due to the above mutual dependencies, the skip_txfm flag can be signaled earlier than the is_inter flag in the bitstream.
[0170] In some further embodiments, the is_inter flag may not be included in the bitstream, and when the skip_txfm flag is equal to 1, the is_inter flag may be implied as true (or 1) on the decoder side. This can be useful because even if coefficients that are not transmitted in the bitstream due to the skip_txfm flag being 1 (all-zero coefficients), when those adjacent blocks are associated with flags encoded based on the prediction mode of the current block, inter-frame prediction will be implied or attributed to that block and used to decode its adjacent blocks.
[0171] Figure 16 A flowchart 1600 of an example method is shown, and the example method follows the basic principles of the above-described embodiments for zero-skip flag encoding and decoding. The example method flow starts at 1601. In S1610, the zero-skip flag of at least one adjacent block of the current block is determined as a reference zero-skip flag. In S1620, the prediction mode for the current block is determined as an intra-frame prediction mode or an inter-frame prediction mode. In S1630, based on the reference zero-skip flag and the prediction mode of the current block, at least one context for decoding the zero-skip flag of the current block is derived. In S1640, according to at least one context, the zero-skip flag of the current block is decoded. The example method flow ends at S1699. The above method flow also applies to encoding.
[0172] The embodiments in the present disclosure can be used alone or in any order combination. In addition, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments of the present disclosure can be applied to luminance blocks or chrominance blocks.
[0173] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 17 A computer system (1700) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0174] The computer software can be encoded using any suitable machine code or computer language, and the machine code or computer language can be created through mechanisms such as assembly, compilation, linking, or the like to include code that includes instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, etc.
[0175] The instructions can be executed on various types of computers or their components, including for example personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0176] Figure 17 The components shown for the computer system (1700) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency on or requirement for any one component or combination of components illustrated in the exemplary embodiments of the computer system (1700).
[0177] The computer system (1700) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (such as keystrokes, swipes, data glove movements), audio input (such as voice, taps), visual input (such as gestures), and olfactory input (not shown). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as voice, music, ambient sounds), images (such as scanned images, photographic images obtained from a still-image camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0178] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1701), mouse (1702), touchpad (1703), touch screen (1710), data glove (not shown), joystick (1705), microphone (1706), scanner (1707), camera (1708).
[0179] The computer system (1700) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as the tactile feedback of a touch screen (1710), data glove (not shown), or joystick (1705), but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1709), headphones (not depicted)), visual output devices (such as a screen (1710), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities and each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and fog machines (not depicted)), and printers (not depicted).
[0180] The computer system (1700) may also include human-accessible storage devices and their associated media, such as optical media including media (1721) such as CD / DVD ROM / RW (1720) with CD / DVD, thumb drives (1722), removable hard disk drives or solid state drives (1723), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0181] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves or other transitory signals.
[0182] The computer system (1700) may also include an interface (1754) to one or more communication networks (1755). The network may be wireless, wired, optical, for example. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV and terrestrial broadcast TV, vehicular and industrial networks including CAN bus, etc. Some networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1749) (such as, for example, the USB port of the computer system (1700)); other networks are typically integrated into the core of the computer system (1700) by attaching to the system bus described below (such as an Ethernet interface into a PC computer system or a cellular network interface into a smart phone computer system). Using any of these networks, the computer system (1700) can communicate with other entities. Such communication can be unidirectional only receiving (such as broadcast TV), unidirectional only sending (such as CANbus to certain CANbus devices), or bidirectional, such as to other computer systems using local digital networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.
[0183] The above-mentioned human-machine interface devices, human-accessible storage devices and network interfaces can be attached to the core (1740) of the computer system (1700).
[0184] The core (1740) may include one or more central processing units (CPUs) (1741), a graphics processing unit (GPU) (1742), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (1743), a hardware accelerator (1744) for certain tasks, a graphics adapter (1750), etc. These devices, together with a read-only memory (ROM) (1745), a random access memory (1746), and an internal mass storage (1747) such as an internal non-user-accessible hard disk drive, SSD, etc., may be connected via a system bus (1748). In some computer systems, the system bus (1748) may be accessible in the form of one or more physical plugs to enable expansion by attaching additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1749) to the system bus (1748) of the core. In one example, a screen (1710) may be connected to the graphics adapter (1750). The architecture of the peripheral bus includes PCI, USB, etc.
[0185] The CPU (1741), GPU (1742), FPGA (1743), and accelerator (1744) may execute certain instructions, and the combination of these instructions may constitute the aforementioned computer code. The computer code may be stored in the ROM (1745) or the RAM (1746). Transitional data may also be stored in the RAM (1746), while permanent data may be stored in, for example, the internal mass storage (1747). Fast storage and retrieval of any memory device may be enabled by using a cache memory, which may be closely associated with one or more CPUs (1741), GPUs (1742), mass storage (1747), ROM (1745), RAM (1746), etc.
[0186] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of this disclosure, or they may be of the type well-known and available to those skilled in the computer software art.
[0187] As a non - limiting example, a computer system (1700) having an architecture, and in particular a core (1740), can provide functionality as a result of one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer - readable media. Such computer - readable media can be media associated with the user - accessible mass storage introduced above, as well as certain memories of the core (1740) having a non - volatile nature, such as on - core mass storage (1747) or ROM (1745). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1740). Depending on specific needs, the computer - readable media can include one or more memory devices or chips. The software can cause the core (1740) and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of being logically hard - wired or otherwise embodied in circuitry (e.g., accelerator (1744)), which can operate in place of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, references to software can include logic and vice versa. In appropriate cases, references to computer - readable media can include circuitry (such as an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0188] Although the present disclosure describes several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of the present disclosure. In the above - described embodiments and implementations, any operations of the processes can be combined or arranged in any number or order as needed. Additionally, two or more of the above - described operations of the process can be executed in parallel. Thus, it is understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within its spirit and scope. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit HDR: high dynamic range SDR: standard dynamic range JVET: Joint Video Exploration Team MPM: most probable mode WAIP: Wide-Angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Transform Unit CTU: Coding Tree Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub-Partitions SPS: Sequence Parameter Set PPS: Picture Parameter Set APS: Adaptation Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-Component Sample Offset LSO: Local Sample Offset LR: Loop Restoration Filter AV1: AOMedia Video 1 AV2: AOMedia Video 2
Claims
1. A method for video decoding, characterized in that, the method comprises: determining the value of a first reference zero skip flag of a first neighboring block of a current block; deriving at least one first context for decoding a zero skip flag of the current block based on the value of the first reference zero skip flag; and decoding the zero skip flag of the current block according to the at least one first context; deriving at least one second context for decoding a prediction mode is_inter flag of the current block based on the value of the zero skip flag; decoding the prediction mode is_inter flag of the current block according to the at least one second context; and decoding the current block according to the zero skip flag and the prediction mode is_inter flag.
2. The method according to claim 1, characterized in that, further comprising: determining the value of a second reference zero skip flag of a second neighboring block of the current block; wherein, deriving at least one first context for decoding the zero skip flag of the current block is further based on the value of the second reference zero skip flag.
3. The method according to claim 2, characterized in that, deriving at least one first context for decoding the zero skip flag of the current block based on the first reference zero skip flag and the second reference zero skip flag includes: deriving the at least one first context based on the sum of the first reference zero skip flag and the second reference zero skip flag.
4. The method according to claim 2, characterized in that, the first neighboring block and the second neighboring block respectively include a top neighboring block and a left neighboring block of the current block.
5. The method according to claim 1, characterized in that, at least one second context for decoding the prediction mode is_inter flag is further based on the first reference zero skip flag of the first neighboring block.
6. The method according to claim 5, characterized in that, the first neighboring block includes a top neighboring block or a left neighboring block.
7. The method according to claim 1, characterized in that, deriving at least one second context for decoding the prediction mode is_inter flag of the current block based on the value of the zero skip flag includes: when the value of the zero skip flag is 0, selecting a first set of one or more contexts as the at least one second context; when the value of the zero skip flag is not 0, selecting a second set of one or more contexts as the at least one second context.
8. The method according to claim 1, characterized in that, in the video bitstream, the zero skip flag is signaled before the prediction mode is_inter flag.
9. A method for video encoding, characterized in that, the method comprises: determining the value of a first reference zero skip flag of a first neighboring block of a current block; deriving at least one first context for determining a zero skip flag of the current block based on the value of the first reference zero skip flag; and Determine a zero skip flag of the current block according to the at least one first context; Derive at least one second context for determining a prediction mode is_inter flag of the current block based on a value of the zero skip flag; Determine the prediction mode is_inter flag of the current block according to the at least one second context; and Encode the current block according to the zero skip flag and the prediction mode is_inter flag.
10. An electronic device, characterized in that, comprising: a memory for storing computer instructions; a processor for calling the computer instructions stored in the memory to implement the method according to any one of claims 1-9.
11. A non-volatile computer-readable storage medium for storing instructions, characterized in that, when the instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
12. A method for transmitting or storing a video bitstream, characterized in that, the video bitstream is generated by encoding according to the method of claim 9, or the video bitstream is decoded according to the method of any one of claims 1-8.