Method, apparatus, and computer program for signaling skip mode flag

The method improves video encoding efficiency by decoding prediction information to determine IBC mode and skip mode flags, addressing challenges in existing techniques and enhancing compression ratios.

JP2025085742AActive Publication Date: 2025-06-05TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025041468
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-11-11
Filing Date
2025-03-14
Publication Date
2025-06-05
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

Existing video encoding techniques face challenges in efficiently signaling and applying skip mode flags, particularly in intra-block copy (IBC) mode, which affects coding efficiency and compression ratios.

Method used

The proposed method and apparatus for video encoding/decoding involve a processing circuit that decodes prediction information for blocks in an I-slice to determine the applicability of the IBC mode and a skip mode flag, allowing for efficient reconstruction of blocks based on these flags.

Benefits of technology

This approach enhances coding efficiency by accurately determining the applicability of IBC mode and skip mode flags, leading to improved compression ratios and reduced redundancy in video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025085742000001_ABST
    Figure 2025085742000001_ABST
Patent Text Reader

Abstract

To provide methods and apparatuses for video encoding / decoding.SOLUTION: In some examples, an apparatus for video decoding includes receiving circuitry and processing circuitry. In some embodiments, the processing circuitry decodes prediction information for a block in an I slice from a coded video bitstream, and determines whether an intra block copy (IBC) mode is possible for the block in the I slice. In response to a slice type parameter indicating I slice and at least a width or height of the block being greater than 64, the processing circuitry sets a current mode type parameter to MODE_TYPE_INTRA. Further, in an embodiment, the processing circuitry decodes a flag that indicates whether a skip mode is applied on the block from the coded video bitstream. Then, the processing circuitry reconstructs the block at least partially based on the flag.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Patent Application No. 17 / 095,583, entitled "METHOD AND APPARATUS FOR SIGNALING SKIP MODE FLAG," filed November 11, 2020. That U.S. patent application claims the benefit of priority to U.S. Provisional Application No. 62 / 959,621, entitled "METHODS FOR SIGNALING OF SKIP MODE FLAG," filed January 10, 2020. The entire disclosures of these earlier applications are incorporated herein by reference.

[0002] [Technical field] This disclosure describes embodiments generally related to video encoding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the background of the present disclosure. The work of the inventors of the present application is not admitted explicitly or implicitly as prior art to the present disclosure to the extent that such work is described in this background section, nor are any aspects of the description that do not specifically qualify as prior art at the time of filing.

[0004] Video encoding and decoding can be performed using motion compensation and inter-picture prediction. Uncompressed digital video can include a sequence of pictures, each with spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures can have a fixed or variable picture rate (also informally known as frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution with a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.

[0005] One goal of video encoding and decoding is to be able to reduce redundancy in the input video signal through compression. Compression can help reduce the bandwidth or storage space requirements, in some cases by more than one order of magnitude. Both lossless and lossy compression, as well as combinations of these, can be used. Lossless compression refers to techniques where an exact copy of the original signal can be restored from the compressed original signal. When lossy compression is used, the restored signal may not be identical to the original signal, but the distortion between the original and restored signals is small enough to make the restored signal useful for the intended application. For video, lossy compression is widely used. The amount of acceptable distortion depends on the application. For example, users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio can reflect higher acceptable distortion / acceptable distortion can result in higher compression ratios.

[0006] Video encoders and decoders may utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007] Video codec techniques can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized before entropy coding. Intra prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are needed at a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example as known from MPEG-2 generation encoding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to predict intra-prediction from surrounding sample data and / or metadata, for example obtained during encoding / decoding of a previous block of data that is spatially adjacent and in decoding order. Such techniques are referred to below as "intra prediction" techniques. It should be noted that at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and does not use reference data from a reference picture.

[0009] There may be many forms of intra prediction. If one or more of such techniques are available in a given video coding technique, the technique used may be coded in intra prediction mode. In some cases, the mode may have sub-modes and / or parameters, which may be coded separately or may be included in the mode codeword. The codeword used for a given mode / sub-mode / parameter combination may affect the coding efficiency gain through intra prediction, which may in turn affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain intra prediction modes were introduced in H.264, improved in H.265, and further improved in newer coding techniques such as joint exploration model (JEM), versatile video coding (VVC) and benchmark set (BMS). A prediction block can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are copied to the prediction block according to a direction. The reference to the direction in use can be coded in the bitstream or it may be predicted itself.

[0011] Referring to FIG. 1, shown at the bottom right is a subset of 9 known predictor directions from the 33 possible predictor directions (corresponding to the 33 angle modes of the 35 intra modes) of H.265. The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] Still referring to FIG. 1, at the top left is shown a square block (104) of 4×4 samples (indicated by a dashed bold line). The square block (104) contains 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the size of the block is 4×4 samples, S44 is at the bottom right. Additionally, reference samples are shown following a similar numbering scheme. The reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are adjacent to the block being reconstructed, so there is no need for negative values ​​to be used.

[0013] Intra-picture prediction can work by copying reference sample values ​​from neighboring samples according to the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction consistent with the arrow (102). That is, assume that the samples are predicted from one or more prediction samples to the upper right at an angle of 45 degrees from the horizontal. In this case, samples S41, S32, S23 and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0014] In some cases, particularly when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.

[0015] As video coding techniques develop, the number of possible directions increases. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and at the time of disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been carried out to identify the most likely directions, and certain techniques in entropy coding have been used to represent these more likely directions with a small number of bits, accepting a certain penalty for the less likely directions. Furthermore, in some cases, the direction itself can be predicted from the neighboring directions used in adjacent already decoded blocks.

[0016] FIG. 2 shows a schematic diagram (201) illustrating 65 intra prediction directions according to JEM, showing the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits to represent directions in a coded video bitstream can vary between video coding techniques, ranging, for example, from simple direct mapping of prediction directions to complex adaptation schemes involving intra-prediction modes, codewords, most probable modes, and similar techniques. In all cases, however, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in well-performing video coding techniques, these less likely directions are represented by a greater number of bits than the more likely directions.

[0018] Motion compensation is a lossy compression technique and can be associated with a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are used to predict a newly reconstructed picture or part thereof after being spatially shifted in a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or it may have three dimensions, with the third dimension indicating the reference picture in use (the latter can indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, e.g., from MVs associated with other regions of sample data that are spatially adjacent to the region being restored and that precede that MV in decoding order. This can significantly reduce the amount of data required to encode the MV, thereby removing redundancy and increasing compression. For example, when encoding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical possibility that regions larger than the region to which a single MV is applicable move in a similar direction and therefore can be predicted using similar motion vectors, possibly derived from MVs of neighboring regions. As a result, the detected MV for a given region will be similar or identical to the MV predicted from the surrounding MVs, which, after entropy coding, can be represented with fewer bits than those used when encoding the MV directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, e.g., due to rounding errors when calculating a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms provided by H.265, there is a technique called “spatial merge” in this specification.

[0021] Referring to Figure 3, the current block (301) contains samples that the encoder found during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of directly encoding the MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the most recent reference picture (in decoding order) using the MV associated with any one of the five surrounding samples denoted A0, A1 and B0, B1, B2 (302-306, respectively). In H.265, the MV prediction can use predictors from the same reference picture as the neighboring blocks use. Summary of the Invention

[0022] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a receiving circuit and a processing circuit. In some embodiments, the processing circuit decodes prediction information for a block in an I-slice from the encoded video bitstream, and determines whether an IBC mode is applicable for the block in the I-slice based on the prediction information and a size limit of an intra block copy (IBC) mode. Furthermore, in one embodiment, the processing circuit decodes a flag indicating whether a skip mode is applied to the block from the encoded video bitstream in response to the IBC mode being applicable for the block in the I-slice. The processing circuit then reconstructs the block based at least in part on the flag.

[0023] Further, in some examples, the processing circuit infers a flag in response to which the IBC mode is not applicable for blocks in an I slice.

[0024] In some embodiments, the processing circuit determines that the IBC mode is not applicable for a block in an I slice in response to a size of the block in the I slice being greater than a threshold. In one embodiment, the processing circuit determines that the IBC mode is not applicable for a block in an I slice in response to at least one of a block width and a block height being greater than a threshold. In some examples, the processing circuit determines the presence of the flag in the encoded video bitstream based on an additional condition that compares the size of the block to a threshold. The additional condition precludes the presence of the flag in response to a size of the block in an I slice being greater than a threshold.

[0025] In some examples, the processing circuit determines the presence of the flag in the coded video bitstream based on a pre-existing condition that is modified based on a comparison of the size and the threshold. In one example, the pre-existing condition determines whether the block is an inter-coded block. In another example, the pre-existing condition determines whether an IBC mode is enabled.

[0026] Aspects of the present disclosure also provide a non-transitory computer-readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform a method for video decoding. [Brief description of the drawings]

[0027] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Diagram 2] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 3] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 5 is a schematic diagram of a simplified block diagram of a communication system (500) according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 7] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 8] 4 shows a block diagram of an encoder according to another embodiment; [Figure 9] 4 shows a block diagram of a decoder according to another embodiment; [Figure 10] 1 illustrates an example of intra block copy according to one embodiment of the present disclosure. [Figure 11] 1 shows a table of example syntax for signaling coding unit level prediction modes. [Figure 12A] 1 shows a table of an example syntax for the coding tree unit level. [Figure 12B] 1 shows a table of an example syntax for the coding tree unit level. [Figure 12C] 1 shows a table of an example syntax for the coding tree unit level. [Figure 12D] 1 shows a table of an example syntax for the coding tree unit level. [Figure 12E] 1 shows a table of an example syntax for the coding tree unit level. [Figure 13] 1 illustrates a table of example coding unit level syntax according to some embodiments of the present disclosure. [Figure 14] 1 illustrates a table of example syntax for coding tree levels according to some embodiments of the present disclosure. [Figure 15] 1 shows a flowchart outlining an example process according to some embodiments of the present disclosure. [Figure 16]FIG. 1 is a schematic diagram of a computer system according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] FIG. 4 illustrates a simplified block diagram of a communication system (400) according to an embodiment of the present disclosure. The communication system (400) includes, for example, a plurality of terminal devices that can communicate with each other via a network (450). For example, the communication system (400) includes a first pair of terminal devices (410) and (420) interconnected via the network (450). In the example of FIG. 4, the first pair of terminal devices (410) and (420) perform a one-way transmission of data. For example, the terminal device (410) may encode video data (e.g., a stream of video pictures captured by the terminal device (410)) for transmission to another terminal device (420) via the network (450). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (420) may receive the encoded video data from the network (450), decode the encoded video data, reconstruct the video pictures, and display the video pictures according to the reconstructed video data. One-way data transmission may be common, such as in media presentation applications.

[0029] In another example, the communication system (400) includes a second pair of terminal devices (430) and (440) performing bidirectional transmission of encoded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (430) and (440) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (430) and (440) over the network (450). Each of the terminal devices (430) and (440) may also receive the encoded video data transmitted by the other of the terminal devices (430) and (440), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0030] In the example of FIG. 4, the terminal devices (410), (420), (430), and (440) may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure are not limited thereto. The embodiments of the present disclosure may be applied to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. The network (450) represents any number of networks that convey encoded video data between the terminal devices (410), (420), (430), and (440), including, for example, wired (hardwired) and / or wireless communication networks. The communication network (450) may exchange data in circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (450) is not important to the operation of the present disclosure, unless otherwise described herein below.

[0031] 5 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is similarly applicable to other video-enabled applications including, for example, videoconferencing, digital TV, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.

[0032] The streaming system may include a capture subsystem (513), which may include, for example, a video source (501) (e.g., a digital camera) that generates a stream of uncompressed video pictures (502). In one example, the stream of video pictures (502) includes samples taken by a digital camera. The stream of video pictures (502), depicted as a thick line to highlight its higher amount of data when compared to the encoded video data (504) (or encoded video bitstream), may be processed by an electronic device (520) that includes a video encoder (503) coupled to the video source (501). The video encoder (503) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (504) (or encoded video bitstream (504)), depicted as a thin line to highlight its lower amount of data when compared to the stream of video pictures (502), may be stored in a streaming server (505) for future use. One or more streaming client subsystems, such as client subsystems (506) and (508) in FIG. 5, may access the streaming server (505) to obtain copies (507) and (509) of the encoded video data (504). The client subsystem (506) may include a video decoder (510), for example, within an electronic device (530). The video decoder (510) decodes an input copy (507) of the encoded video data and generates an output stream (511) of video pictures that can be rendered on a display (512) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (504), (507), and (509) (e.g., a video bitstream) may be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, a developing video coding standard is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0033] It should be noted that electronic devices 520 and 530 may include other components (not shown). For example, electronic device 520 may include a video decoder (not shown) and electronic device 530 may include a video encoder (not shown).

[0034] 6 shows a block diagram of a video decoder (610) according to one embodiment of the present disclosure. The video decoder (610) may be included in an electronic device (630). The electronic device (630) may include a receiver (631) (e.g., a receiving circuit). The video decoder (610) may be used in place of the video decoder (510) in the example of FIG. 5.

[0035] The receiver (631) may receive one or more coded video sequences to be decoded by the video decoder (610), and in the same or other embodiments, may receive one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (601), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (631) may receive the coded video data together with other data (e.g., coded audio data and / or auxiliary data streams), which may be forwarded to respective usage entities (not shown). The receiver (631) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (615) may be coupled between the receiver (631) and the entropy decoder / parser (620), hereinafter referred to as the "parser (620)". In certain applications, the buffer memory (615) is part of the video decoder (610). In other cases, it may be external to the video decoder (610) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (610), e.g., to prevent network jitter, plus another buffer memory (615) internal to the video decoder (610), e.g., to handle playback timing. If the receiver (631) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (615) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (615) may be needed and may be relatively large, advantageously adaptively sized, and may be at least partially implemented in an operating system or similar element (not shown) external to the video decoder (610).

[0036] The video decoder (610) may include a parser (620) for recovering symbols (621) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (610), potentially including information for controlling a rendering device such as a rendering device (612) (e.g., a display screen), which may not be an integral part of the electronic device (630) as shown in FIG. 6, but may be coupled to the electronic device (630). The rendering device control information may be in the form of Supplemental Enhancement Information (SEI) (SEI message) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (620) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (620) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transformation unit (TU), a prediction unit (PU), etc. The parser (620) may also extract information from the coded video sequence such as transform coefficients, quantization parameter values, motion vectors, etc.

[0037] The parser (620) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (615) to generate symbols (621).

[0038] The reconstruction of the symbols (621) may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (620). The flow of such subgroup control information between the parser (620) and the following units is not shown for clarity.

[0039] In addition to the functional blocks described above, the video decoder (610) may be conceptually subdivided into a number of functional units, as described below. In a practical implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0040] The first unit is a scalar / inverse transform unit (651), which receives the quantized transform coefficients as symbols (621) from the parser (620), along with control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.). The scalar / inverse transform unit (651) may output a block containing sample values ​​that can be input to an aggregator (655).

[0041] In some cases, the output samples of the scaler / inverse transform (651) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (652). In some cases, the intra-picture prediction unit (652) generates blocks of the same size and shape of the block being reconstructed using surrounding already reconstructed information retrieved from the current picture buffer (658). The current picture buffer (658) buffers, for example, a partially reconstructed and / or a fully reconstructed current picture. In some cases, the aggregator (655) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (652) to the output sample information provided by the scaler / inverse transform unit (651).

[0042] In other cases, the output samples of the scaler / inverse transform unit (651) may relate to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (653) may access the reference picture memory (657) to retrieve samples used for prediction. After motion compensating the retrieved samples according to the symbols (621) associated with the block, these samples may be added by the aggregator (655) to the output of the scaler / inverse transform unit (651) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (657) available to the motion compensation prediction unit (653) from which the motion compensation prediction unit (653) retrieves prediction samples may be controlled by a motion vector, for example in the form of a symbol (621) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory (657) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0043] The output samples of the aggregator (655) may be subjected to various loop filtering techniques in a loop filter unit (656). The video compression techniques may include in-loop filter techniques, controlled by parameters contained in the coded video sequence (also called coded video bitstream) and made available to the loop filter unit (656) as symbols (621) from the parser (620), that may be responsive to meta-information obtained during decoding of previous portions of the coded picture or coded video sequence (in decoding order), as well as to previously reconstructed loop filtered sample values.

[0044] The output of the loop filter unit (656) may be a sample stream that may be output to a rendering device (612) and also stored in a reference picture memory (657) for use in future inter-picture prediction.

[0045] Once a particular coded picture is fully reconstructed, it may be used as a reference picture for future prediction. For example, once a coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (620)), the current picture buffer (658) may become part of the reference picture memory (657), and a new current picture buffer may be reallocated before beginning reconstruction of the subsequent coded picture.

[0046] The video decoder (610) may perform decoding operations according to a given video compression technique in a standard, such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. In particular, the profile may select certain tools from all tools available in the video compression technique or standard as the only tools available for use in that profile. Also, what is required for compliance is that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited through a Hypothetical Reference Decoder (HRD) specification and metadata about HRD buffer management conveyed in the coded video sequence.

[0047] In one embodiment, the receiver (631) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (610) to properly decode the data and / or to more accurately recover the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0048] 7 shows a block diagram of a video encoder (703) according to one embodiment of the present disclosure. The video encoder (703) is included in an electronic device (720). The electronic device (720) includes a transmitter (740) (e.g., a transmission circuit). The video encoder (703) may be used in place of the video encoder (503) in the example of FIG. 5.

[0049] The video encoder (703) may receive video samples from a video source (701) (not part of the electronic device (720) in the example of FIG. 7), which may capture video images to be encoded by the video encoder (703). In other examples, the video source (701) is part of the electronic device (720).

[0050] The video source (701) may provide a source video sequence to be encoded by the video encoder (703) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (701) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (701) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0051] According to one embodiment, the video encoder (703) may encode and compress pictures of a source video sequence into an encoded video sequence (743) in real-time or under any other time constraint required by an application. Achieving an appropriate encoding rate is one function of the controller (750). In some embodiments, the controller (750) controls and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller (750) may include rate control related parameters (picture skip, quantization, lambda values ​​for rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (750) may be configured with other appropriate functions associated with the video encoder (703) that are optimized for a particular system design.

[0052] In some embodiments, the video encoder (703) is configured to operate in an encoding loop. As a very simplified description, in one example, the encoding loop may include a source coder (730) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (733) embedded in the video encoder (703). The decoder (733) reconstructs the symbols to generate sample data as the (remote) decoder would generate them (so that any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (734). Since the decoding of the symbol stream produces bit-wise accurate results independent of the location of the decoder (local or remote), the contents in the reference picture memory (734) are also bit-wise accurate between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values ​​as the reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of reference picture synchronization (including the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is used in some related technologies as well.

[0053] The operation of the "local" decoder (733) may be the same as a "remote" decoder, such as the video decoder (610), which has already been described in detail above in connection with Figure 6. However, with brief reference to Figure 6, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (745) and parser (620) may be lossless, the entropy decoding portion of the video decoder (610), including the buffer memory (615) and parser (620), may not be fully implemented in the local decoder (733).

[0054] An observation that can be made at this point is that any decoder technique, other than analysis / entropy decoding, present in the decoder must necessarily be present in the corresponding encoder in substantially the same functional form. For this reason, the subject matter of the disclosure focuses on the decoder operation. A description of the encoder techniques can be omitted, as they are the inverse of the decoder techniques, which are described generically. Only in certain areas are more detailed descriptions necessary, and are provided below.

[0055] In some examples, during operation, the source coder (730) may perform motion-compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from the video sequence designated as “reference pictures.” In this manner, the encoding engine (732) encodes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0056] The local video decoder (733) may decode the encoded video data of pictures that may be designated as reference pictures based on the symbols generated by the source coder (730). The operation of the encoding engine (732) may advantageously be a lossy process. If the encoded video data can be decoded in a video decoder (not shown in FIG. 7), the reconstructed video sequence may be a replica of the source video sequence, typically with some errors. The local video decoder (733) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (734). In this way, the video encoder (703) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures (without transmission errors) obtained by the far-end video decoder.

[0057] The predictor (735) may perform a predictive search for the coding engine (732). That is, for a new picture to be coded, the predictor (735) may search the reference picture memory (734) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (735) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, the input picture determined by the search results obtained by the predictor (735) may have prediction references drawn from multiple reference pictures stored in the reference picture memory (734).

[0058] The controller (750) may manage the encoding operations of the source coder (730), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0059] The output of all the above functional units may undergo entropy coding in an entropy coder (745), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0060] The transmitter (740) may buffer the encoded video sequence produced by the entropy coder (745) and prepare it for transmission over a communication channel (760), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (740) may merge the encoded video data from the video coder (703) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (not shown).

[0061] A controller (750) may manage the operation of the video encoder (703). During encoding, the controller (750) may assign each coded picture a particular coded picture type, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:

[0062] An Intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs allow different types of Intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize these variations of I-pictures and their respective uses and characteristics.

[0063] A predictive picture (P-picture) may be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0064] Bidirectionally predicted pictures (B-pictures) may be encoded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0065] In general, a source picture may be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8 or 16x16 samples, respectively) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the respective picture of the block. For example, blocks of an I-picture may be non-predictively coded or may be predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial or temporal prediction with reference to the reference picture coded one step before. Blocks of a B-picture may be predictively coded via spatial or temporal prediction with reference to the reference picture coded one or two steps before.

[0066] The video encoder (703) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (703) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.

[0067] In one embodiment, the transmitter (740) may transmit additional data along with the coded video. The source coder (730) may include such data as part of the coded video sequence. The additional data may include other types of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0068] A video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being coded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture, and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0069] In some embodiments, bidirectional prediction techniques may be used in inter-picture prediction. According to bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture, both preceding in decoding order to the current picture in the video (but may be past and future, respectively, in display order). A block in the current picture may be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0070] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0071] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PU, prediction unit) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB, prediction block) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0072] 8 shows a diagram of a video encoder (803) according to another embodiment of the present disclosure. The video encoder (803) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (803) is used in place of the video encoder (503) of the example of FIG. 5.

[0073] In an HEVC example, a video encoder (803) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8x8 samples. The video encoder (803) determines whether the processing block is best coded using an intra mode, an inter mode, or a bidirectional predictive mode, for example using rate-distortion optimization. If the processing block is coded in an intra mode, the video encoder (803) may use intra prediction techniques to code the processing block into a coded picture. If the processing block is coded in an inter mode or a bidirectional predictive mode, the video encoder (803) may use inter prediction techniques or bidirectional prediction techniques, respectively, to code the processing block into a coded picture. In certain video coding techniques, a merge mode may be an inter picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the benefit of any coded motion vector components other than the motion vector predictor. In certain other video coding techniques, there may be motion vector components applicable to the block of interest. In one example, the video encoder (803) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.

[0074] In the example of FIG. 8, the video encoder (803) includes an inter-encoder (830), an intra-encoder (822), a residual calculator (823), a switch (826), a residual encoder (824), an overall controller (821), and an entropy encoder (825) coupled together as shown in FIG.

[0075] The inter-encoder (830) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundant information due to an inter-coding technique, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a prediction block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0076] The intra encoder (822) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously encoded blocks in the same picture, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (822) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0077] The global controller (821) is configured to determine global control data and control other components of the video encoder (803) based on the global control data. In one example, the global controller (821) determines the mode of the block and provides a control signal to the switch (826) based on the mode. For example, if the mode is an intra mode, the global controller (821) controls the switch (826) to select an intra mode result to be used by the residual calculator (823) and controls the entropy encoder (825) to select intra prediction information and include the intra prediction information in the bitstream. If the mode is an inter mode, the global controller (821) controls the switch (826) to select an inter prediction result to be used by the residual calculator (823) and controls the entropy encoder (825) to select inter prediction information and include the inter prediction information in the bitstream.

[0078] The residual calculator (823) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (822) or the inter-encoder (830). The residual encoder (824) operates on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (824) is configured to transform the residual data from a spatial domain to a frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (803) also includes a residual decoder (828). The residual decoder (828) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (822) and the inter-encoder (830) as appropriate. For example, the inter-encoder (830) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (822) may generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are suitably processed to generate a decoded picture, which may be buffered in a memory circuit (not shown) and used as a reference picture in some examples.

[0079] The entropy encoder (825) is configured to format the bitstream to include the coded block. The entropy encoder (825) is configured to include various information according to an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (825) is configured to include global control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, according to the disclosed subject matter, when encoding a block in a merged sub-mode of either the inter mode or the bi-prediction mode, the residual information is not present.

[0080] 9 shows a diagram of a video decoder (910) according to another embodiment of the present disclosure. The video decoder (910) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (910) is used in place of the video decoder (510) of the example of FIG. 5.

[0081] In the example of FIG. 9, the video decoder (910) includes an entropy decoder (971), an inter decoder (980), a residual decoder (973), a reconstruction module (974), and an intra decoder (972) coupled together as shown in FIG. 9.

[0082] The entropy decoder (971) may be configured to recover from the coded picture certain symbols that represent the syntax elements of which the coded picture is composed. Such symbols may include, for example, prediction information (e.g., intra-mode, inter-mode, bi-predictive mode, the latter two in merged submode or other submode, etc.) that can identify the mode in which the block is coded, the particular samples or metadata used for prediction by the intra decoder (972) or the inter decoder (980), respectively (e.g., intra-predictive information or inter-predictive information, etc.), residual information in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter mode or bi-predictive mode, the inter-predictive information is provided to the inter decoder (980), and if the prediction type is an intra-predictive type, the intra-predictive information is provided to the intra decoder (972). The residual information may undergo inverse quantization and is provided to the residual decoder (973).

[0083] The inter decoder (980) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0084] The intra decoder (972) is configured to receive intra prediction information and to generate a prediction result based on the intra prediction information.

[0085] The residual decoder (973) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (973) may also require certain control information, including a quantizer parameter (QP), which may be provided by the entropy decoder (971) (data path not shown as this may be only low volume control information).

[0086] The reconstruction module (974) is configured to combine, in the spatial domain, the residual output by the residual decoder (973) and the prediction result (possibly output by an inter-prediction module or an intra-prediction module) to form a reconstruction block, which may be part of a reconstructed picture or part of a reconstructed video. It should be noted that other suitable operations may be performed to improve the visual quality, such as a debooking operation.

[0087] It should be noted that the video encoders (503), (703), and (803) and the video decoders (510), (610), and (910) may be implemented using any suitable technology. In one embodiment, the video encoders (503), (703), and (803) and the video decoders (510), (610), and (910) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (503), (703), and (803) and the video decoders (510), (610), and (910) may be implemented using one or more processors executing software instructions.

[0088] Aspects of the present disclosure provide signaling techniques for a skip mode flag.

[0089] Block-based compensation can be used for inter-prediction and intra-prediction. For inter-prediction, block-based compensation from different pictures is known as motion compensation. For intra-prediction, block-based compensation may also be done from previously reconstructed regions in the same picture. Block-based compensation from reconstructed regions in the same picture is called intra-picture block compensation, current picture referencing (CPR) or intra block copy (IBC). A displacement vector indicating the offset between a current block and a reference block in the same picture is called a block vector (abbreviated BV). Unlike the motion vector in motion compensation, which can be any value (positive or negative, x-direction or y-direction), the block vector has some restrictions to ensure that the reference block is available and has already been reconstructed. Also, in some examples, some reference regions that are tile boundaries or wavefront ladder shape boundaries are excluded to allow for parallel processing.

[0090] The coding of block vectors can be either explicit or implicit. In explicit modes (or called advanced motion vector prediction (AMVP) modes in inter coding), the difference between the block vector and its predictor is signaled, while in implicit modes, the block vector is recovered from a predictor (called the block vector predictor), similar to the motion vector in merge mode. In some implementations, the resolution of the block vector is limited to integer positions, while in other systems, the block vector is allowed to indicate fractional positions.

[0091] In some examples, the use of block-level intra block copy may be signaled using a block-level flag called the IBC flag. In one embodiment, the IBC flag is signaled if the current block is not coded in merge mode. In another example, the use of block-level intra block copy is signaled by a reference index approach. In this case, the current picture being decoded is treated as a reference picture. In one example, such a reference picture is placed at the last position of a list of reference pictures. This special reference picture is managed together with other temporal reference pictures in a buffer such as a decoded picture buffer (DPB).

[0092] There are also several variations on intra block copying, such as flipped intra block copying (where the reference block is flipped horizontally or vertically before being used to predict the current block) or line-based intra block copying (where each compensation unit in an M×N coding block is an M×1 or 1×N line).

[0093] 10 illustrates an example of intra block copying according to one embodiment of the present disclosure. A current picture (1000) is being decoded. The current picture (1000) includes a reconstructed region (1010) (dotted region) and a region to be decoded (1020) (white region). A current block (1030) is being reconstructed by the decoder. The current block (1030) may be reconstructed from a reference block (1040) located in the reconstructed region (1010). The position offset between the reference block (1040) and the current block (1030) is called the block vector (1050) (or BV (1050)).

[0094] In some examples, intra block copy may be considered as another mode other than intra prediction mode or inter prediction mode, but the signaling of intra block copy may be specified at the block level using a combination of flags such as pred_mode_ibc_flag and pred_mode_flag.

[0095] In a related example, the coding block may be coded in an intra prediction mode or an inter prediction mode. In a related example, a 1-bit prediction mode flag "pred_mode_flag" may be signaled or inferred at the coding block level to distinguish the coding mode for the current block. For example, if pred_mode_flag is equal to 0, MODE_INTER (inter prediction mode) is used, otherwise (if pred_mode_flag is equal to 1), MODE_INTRA (intra prediction mode) is used.

[0096] If intra block copy is used, the mode decision may be made based on the values ​​of both pred_mode_flag and pred_mode_ibc_flag.

[0097] In some examples (e.g., some versions of VVC), the maximum block size for inter-coded CUs may be as large as 128×128, the maximum block size for intra-coded CUs may be as large as 128×128, and the maximum block size for intra block copy coded CUs may be as large as 64×64, all in luma coded samples. If separate luma / chroma coding trees are used, the maximum possible chroma CU size in 4:2:0 color format becomes 32×32.

[0098] In some examples, motion parameters may be predicted for an inter-coded block (a block coded using an inter prediction mode) or an IBC-coded block (a block coded using an IBC mode), e.g., in a merge mode or a skip mode. In a merge mode, the video coder uses candidate motion parameters from neighboring blocks, including spatial and temporal neighboring blocks, to build a candidate list of motion parameters (e.g., reference pictures and motion vectors). The selected motion parameters may be communicated from the video encoder to the video decoder by transmitting an index of the selected candidate from the candidate list. In the video decoder, when the index is decoded, the motion parameters of the corresponding block of the selected candidate may be inherited. The video encoder and the video decoder are configured to build the same list based on already coded blocks. Thus, based on the index, the video decoder may identify the motion parameters of the candidate selected by the video encoder.

[0099] In skip mode, motion parameters may be predicted similarly to merge mode. Furthermore, in skip mode, residual data is not added to the predicted block, whereas in merge mode, residual data is added to the predicted block. Also, the building of the list and the transmission of the index to identify the candidates in the list, as described above with reference to merge mode, are typically performed in skip mode. In some embodiments, a skip flag at the coding block level may be used to indicate whether the block is coded using skip mode or not.

[0100] Figure 11 shows an example syntax table (1100) for signaling coding unit (or coding block) level prediction modes. The syntax table (1100) conditionally extracts three flags, cu_skip_flag, pred_mode_flag, and pred_mode_ibc_flag, from the coded video bitstream, as shown by (1101), (1102), and (1103) in Figure 11.

[0101] Specifically, cu_skip_flag[x0][y0] is the skip flag. The skip flag cu_skip_flag[x0][y0] equal to 1 indicates that the current coding unit is coded in skip mode. Thus, if the current coding unit is in a P slice (coded by inter prediction) or a B slice (coded by bidirectional inter prediction), no syntax elements are parsed after cu_skip_flag[x0][y0], except for the IBC mode flag pred_mode_ibc_flag[x0][y0] and one or more syntax elements of the merge_data() syntax structure. If the current coding unit is in an I slice, no syntax elements are parsed after cu_skip_flag[x0][y0], except for merge_idx[x0][y0].

[0102] The skip flag cu_skip_flag[x0][y0] equal to 0 indicates that the coding unit is not in skip mode. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0103] If cu_skip_flag[x0][y0] is not present in the coded video bitstream, then cu_skip_flag[x0][y0] may be inferred to be equal to 0.

[0104] Additionally, pred_mode_flag equal to 0 specifies that the current coding unit is coded in inter-prediction mode, and pred_mode_flag equal to 1 specifies that the current coding unit is coded in intra-prediction mode.

[0105] In some embodiments, pred_mode_flag is not present in the encoded video bitstream and pred_mode_flag may be inferred. Inference of pred_mode_flag may be based on the variable modeType. In general, the variable modeType specifies whether intra prediction modes (MODE_INTRA), IBC modes (MODE_IBC), palette modes (MODE_PLT) and inter prediction modes can be used (MODE_TYPE_ALL), or only intra coding modes, IBC coding modes and palette coding modes can be used (MODE_TYPE_INTRA), or only inter coding modes can be used (MODE_TYPE_INTER) for the coding units in the coding tree node. For example, if the variable modeType is MODE_TYPE_ALL, all of intra prediction modes, IBC modes, palette modes and inter prediction modes can be used, if the variable modeType is MODE_TYPE_INTRA, intra prediction modes, IBC modes and palette modes can be used, and if the variable modeType is MODE_TYPE_INTER, only inter prediction modes can be used.

[0106] In some examples, the variable modeType is used to control the possible prediction mode types for a particular small CU size. At the CTU root (where the coding_tree() syntax is first called), the value of this variable is set to MODE_TYPE_ALL (meaning no restrictions). Thus, for large CUs such as 128×128, 128×64, 64×128, modeType should be set to MODE_TYPE_ALL. When a large CU is split into small CUs, the variable modeType for the small CUs may be refined to be based on the variable treeType. In one embodiment, the coding tree scheme supports the ability for luma components and corresponding chroma components to have separate block tree structures. For example, for P slices and B slices, the luma CTB and chroma CTB in a CTU share the same coding tree structure (e.g., single tree). For I slices, the luma CTB and chroma CTB in a CTU may have separate block tree structures (e.g., dual tree), and the case of splitting a CTU using separate block tree structures is called dual tree splitting. When a dual tree split is applied, the luma CTB may be split into luma CUs by a luma coding tree structure (e.g., DUAL_TREE_LUMA), and the chroma CTB may be split into chroma CUs by a chroma coding tree structure (e.g., DUAL_TREE_CHROMA). Thus, a CU in an I slice may include a coding block of a luma component and may include coding blocks of two chroma components, and a CU in a P or B slice includes coding blocks of all three chroma components unless the video is monochrome. In one example, the variable treeType specifies whether a single tree (SINGLE_TREE) or a dual tree is used for the coding units split from the coding tree unit. When a dual tree is used, the variable treeType specifies whether a luma component (DUAL_TREE_LUMA) or a chroma component (DUAL_TREE_CHROMA) is currently being processed.

[0107] In some embodiments, pre_mode_flag is inferred as follows: -If cbWidth (width of coding unit) is equal to 4 and cbHeight (height of coding unit) is equal to 4, pred_mode_flag is inferred to be equal to 1 because in some cases inter prediction modes do not apply to 4x4 blocks. Otherwise, if modeType is equal to MODE_TYPE_INTRA, pred_mode_flag is inferred to be equal to 1. Otherwise, if modeType is equal to MODE_TYPE_INTER, pred_mode_flag is inferred to be equal to 0. - Otherwise, pred_mode_flag is inferred to be equal to 1 when decoding an I slice and 0 when decoding a P or B slice, respectively.

[0108] The variable CuPredMode[chType][x][y] is a prediction mode derived as follows for x=x0..x0+cbWidth?1 and y=y0..y0+cbHeight?1. If -pred_mode_flag is equal to 0, CuPredMode[chType][x][y] is set equal to MODE_INTER (inter prediction mode). - Otherwise (if pred_mode_flag is equal to 1), CuPredMode[chType][x][y] is set equal to MODE_INTRA (intra prediction mode).

[0109] Additionally, pred_mode_ibc_flag equal to 1 specifies that the current coding unit is coded in IBC mode, and pred_mode_ibc_flag equal to 0 specifies that the current coding unit is not coded in IBC mode.

[0110] In some embodiments, pred_mode_ibc_flag is not present in the encoded bitstream and may be inferred. The inference of pred_mode_ibc_flag may be based on the variables modeType and treeType.

[0111] For example, pred_mode_ibc_flag may be inferred as follows: - If cu_skip_flag[x0][y0] is equal to 1, cbWidth is equal to 4, and cbHeight is equal to 4, then pred_mode_ibc_flag is inferred to be equal to 1. Otherwise, if cbWidth or cbHeight is equal to 128, pred_mode_ibc_flag is inferred to be equal to 0. Otherwise, if modeType is equal to MODE_TYPE_INTER, pred_mode_ibc_flag is inferred to be equal to 0. Otherwise, if treeType is equal to DUAL_TREE_CHROMA, pred_mode_ibc_flag is inferred to be equal to 0. - Otherwise, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag when decoding an I slice, and to 0 when decoding a P or B slice, respectively.

[0112] If pred_mode_ibc_flag is equal to 1, the variable CuPredMode[chType][x][y] is set equal to MODE_IBC for x=x0..x0+cbWidth-1 and y=y0..y0+cbHeight-1.

[0113] 12A-12E show a table of an example syntax for the coding tree unit level 1200. In this example, the input parameter modeTypeCurr is a modeType variable inherited from a parent node.

[0114] Note that the syntax and semantics in some examples, such as the syntax in FIG. 11, may not accurately reflect the block size restrictions applied among intra / inter / IBC modes. For example, the condition (1110) in FIG. 11 is equivalent to performing a logical OR (||) of the first condition (!(cbWidth==4&&cbHeight==4)&&(modeType!==MODE_TYPE_INTRA)) and the second condition (sps_ibc_enabled_flag). The first condition checks whether the logical value of (!(cbWidth==4&&cbHeight==4)&&(modeType!==MODE_TYPE_INTRA)) is 1 or 0. The second condition checks whether the logical value of (sps_ibc_enabled_flag) is 1 or 0. The first condition being 1 specifies that the current block being coded is an inter-coded block. The second condition being 1 specifies that the IBC mode is enabled. In the example of FIG. 11, if at least one of the first condition and the second condition is equal to 1 and the treeType is not equal to DUAL_TREE_CHROMA, the skip flag (e.g., cu_skip_flag[x0][y0]) may be presented in the coded video bitstream and decoded. However, in one example, a coding unit in an I-slice may have at least one size (width and / or height) greater than 64. The parameters of the coding unit may cause the first condition to be 1 and / or cause the second condition to be 1. However, the coding unit cannot be inter coded (because of the I-slice) and cannot be IBC coded (because at least one size is greater than 64). Thus, the condition (1110) cannot accurately reflect the block size restriction applied between intra / inter / IBC modes.

[0115] Some aspects of the present disclosure provide techniques for accurately reflecting the block size restrictions applied among intra / inter / IBC modes. The following description is based on the assumption that the intra block copy (IBC) mode is considered as a separate mode different from the intra or inter modes. The techniques may signal a skip flag (e.g., cu_skip_flag[x0][y0]) only if the IBC mode is applicable in an I-slice. If only the intra mode is applicable, the skip flag does not need to be signaled and the skip flag may be inferred to be 0.

[0116] In one embodiment, further conditions may be imposed on the signaling of the skip flag such that if the CU block size suggests that only intra mode is applicable (if the block size exceeds the maximum IBC block size), the skip flag is not signaled (not presented in the coded video bitstream). Instead, the skip flag may be inferred to be 0.

[0117] In one example, if a block in an I-slice is larger than 64 luma samples in either the width or height, the block cannot be inter-coded or IBC-coded, and therefore only intra modes (intra prediction modes) are applicable to the block.

[0118] FIG. 13 illustrates an example syntax table (1300) according to some embodiments of the present disclosure. As shown in FIG. 13 (1310), a third condition (!(slice_type=I&&(cbWidth>64||cbHeight>64))) is added to combine (using the logical AND operator &&) with the first and second conditions. If a block in an I slice is larger than 64 luma samples on either the width or height side, the third condition may be 0. In this case, the combination of the first condition, the second condition, and the third condition is equal to 0. Therefore, the skip flag (cu_skip_flag[x0][y0]) is not signaled (not presented in the coded video bitstream).

[0119] In other embodiments, the first condition and / or the second condition may be appropriately modified to apply size restrictions for IBC mode. In some examples, if the CU block size suggests that only intra mode is applicable to I slices (if the block size exceeds the maximum IBC block size), modeType is set to MODE_TYPE_INTRA and the modeType assignment is changed to make the first condition equal to 0. In some embodiments, if (slice_type==I&&(cbWidth>64||cbHeight>64)), modeType is set to MODE_TYPE_INTRA.

[0120] FIG. 14 illustrates a table (1400) of an example syntax for a coding tree level according to some embodiments of the present disclosure. As shown in (1410), if a block in an I-slice is larger than 64 luma samples in either the width or height side, the variable modeTypeCurrr is set to MODE_TYPE_INTRA. The variable modeTypeCurr is an input parameter to the coding_unit() syntax, as shown in the syntax of FIG. 14 and FIG. 13. Thus, in the table of the coding_unit() syntax, the result of the first condition is 0. In another example, in the coding_unit() syntax, if a block in an I-slice is larger than 64 luma samples in either the width or height side, the variable modeType is set to MODE_TYPE_INTRA before the first "if" statement. Thus, the result of the first condition is 0.

[0121] In some examples, the second condition may be modified to apply size restrictions for IBC mode. For example, the second condition is modified to (sps_ibc_enabled_flag&&cbWidth<=64&&cbHeight<=64). Thus, if a block in an I-slice is larger than 64 luma samples on either the width or height side, the result of the second condition is 0.

[0122] It is noted that the embodiments of the present disclosure may be used separately or in combination in any order.

[0123] FIG. 15 shows a flow chart outlining a process (1500) according to one embodiment of the present disclosure. The process (1500) may be used in the reconstruction of a block to generate a prediction block for the block being reconstructed. In various embodiments, the process (1500) is performed by a processing circuit such as the processing circuit of the terminal devices (410), (420), (430), and (440), a processing circuit performing the function of the video encoder (503), a processing circuit performing the function of the video decoder (510), a processing circuit performing the function of the video decoder (610), a processing circuit performing the function of the video encoder (703), etc. In some embodiments, the process (1500) is implemented with software instructions, and thus the processing circuit performs the process (1500) as the processing circuit executes the software instructions. The process starts at (S1501) and proceeds to (S1510).

[0124] At (S1510), prediction information for blocks in the I slice is decoded from the coded video bitstream.

[0125] In (S1520), whether the IBC mode is applicable for the block in the I slice is determined based on the prediction information and the size restriction of the IBC mode. In some embodiments, the IBC mode is determined to be inapplicable for the block in the I slice in response to the size of the block being greater than a threshold. In one embodiment, the IBC mode is determined to be inapplicable for the block if at least one of the width or height of the block is greater than a threshold.

[0126] In some embodiments, the applicability of the IBC mode for a block in an I-slice determines whether a flag (skip mode flag) for the block is in the coded video bitstream. In some examples, the syntax includes some existing conditions for determining the presence of the skip mode flag. In one embodiment, the presence of the flag in the coded video bitstream is also determined based on an additional condition that compares the size of the block to a threshold. The additional condition excludes the presence of the flag in response to the size of the block in an I-slice being greater than the threshold.

[0127] In other embodiments, the presence of the flag in the coded video bitstream is determined based on a pre-existing condition that is modified to compare a size of the block to a threshold. In one example, the pre-existing condition is a first condition that determines whether the block is an inter-coded block, and the first condition is modified based on a comparison of the size of the block to the threshold. In another example, the pre-existing condition is a second condition that determines whether an IBC mode is enabled, and the second condition is modified based on a comparison of the size of the block to the threshold.

[0128] At (S1530), in some examples, a flag indicating whether skip mode is applied to the block is decoded from the coded video bitstream in response to the IBC mode being applicable for the block in the I slice. Further, in some examples, the flag may be inferred in response to the IBC mode being inapplicable for the block of the I slice.

[0129] At (S1540), the block is restored based at least in part on the flag. In one example, if the flag indicates a skip mode, the block is restored based on the skip mode. The process then proceeds to (S1599) and ends.

[0130] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 16 illustrates a computer system (1600) suitable for implementing certain embodiments of the disclosed subject matter.

[0131] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking, or similar mechanisms to generate code including instructions, which may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through an interpreter, microcode execution, etc.

[0132] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0133] 16 for computer system (1600) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (1600).

[0134] The computer system (1600) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, through tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly associated with conscious human input, such as audio (e.g., speech, music, ambient sounds), images (scanned images, photographic images obtained from a still camera, etc.), and video (2D video, 3D video including stereoscopic pictures, etc.).

[0135] The input human interface devices may include one or more of a keyboard (1601), a mouse (1602), a trackpad (1603), a touch screen (1610), a data glove (not shown), a joystick (1605), a microphone (1606), a scanner (1607), and a camera (1608).

[0136] The computer system (1600) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (1610), data gloves (not shown), or joystick (1605), although there may be haptic feedback devices that do not function as input devices), audio output devices (speakers (1609), headphones (not shown), etc.), visual output devices (screens (1610), including CRT screens, LCD screens, plasma screens, OLED screens, each of which may or may not have touch screen input capability, each of which may or may not have haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three or more dimensional output through such means as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0137] The computer system (1600) may also include human accessible storage devices and associated media, such as optical media, including CD / DVD ROM / RW (1620) with CD / DVD or similar media (1621), thumb drives (1622), removable hard drives or solid state drives (1623), legacy magnetic media, such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices, such as security dongles (not shown), etc.

[0138] Additionally, those skilled in the art should understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other non-transitory signals.

[0139] The computer system (1600) may also include interfaces to one or more communication networks. The networks may be, for example, wireless, wired, optical. The networks may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial (including CANBus), etc. Certain networks typically require an external network interface adapter (e.g., a USB port on the computer system (1600)) that is attached to a specific general-purpose data port or peripheral bus (1649), while other network interface adapters are typically integrated into the core of the computer system (1600) by being attached to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network to a smartphone computer system) as described below. Using any of these networks, the computer system (1600) can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or bidirectional, for example, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces.

[0140] The above-mentioned human interface devices, human accessible storage devices and network interfaces may be attached to a core (1640) of the computer system (1600).

[0141] The core (1640) may include one or more central processing units (CPUs) (1641), graphics processing units (GPUs) (1642), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1643), hardware accelerators for specific tasks (1644), etc. These devices may be connected through a system bus (1648), along with read only memory (ROM) (1645), random access memory (1646), and internal mass storage (internal non-user accessible hard drive, SSD, etc.) (1647). In some computer systems, the system bus (1648) may be accessible in the form of one or more physical plugs to allow expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1648) or through a peripheral bus (1649). Peripheral bus architectures include PCI, USB, etc.

[0142] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) may execute certain instructions, which in combination may constitute the above computer code. The computer code may be stored in a ROM (1645) or a RAM (1646). Also, temporary data may be stored in the RAM (1646), while persistent data may be stored, for example, in an internal mass storage device (1647). Fast storage and retrieval in any of the memory devices may be enabled through the use of a cache memory, which may be closely associated with one or more of the CPU (1641), GPU (1642), mass storage device (1647), ROM (1645), RAM (1646), etc.

[0143] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0144] By way of example and not limitation, the architecture (1600), and in particular a computer system having a core (1640), may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with a user-accessible mass storage device as described above, as well as specific storage of the core (1640) of a non-transitory nature, such as mass storage device (1647) internal to the core or ROM (1645). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (1640). The computer-readable media may include one or more memory devices or chips according to particular needs. The software may cause the core (1640), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform certain operations or certain portions of certain operations described herein, including defining data structures stored in RAM (1646) and modifying such data structures according to operations defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator (1644)), which may operate in place of or in conjunction with software to perform particular operations or portions of particular operations described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure includes any appropriate combination of hardware and software.

[0145] [Appendix A: Abbreviations] JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be recognized that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

[Claim 1] 1. A method for video decoding performed by a decoder, comprising: decoding prediction information for blocks in an I slice from the encoded video bitstream; determining whether an intra block copy (IBC) mode is applicable for the block in the I-slice based on the prediction information and a size limitation of the IBC mode; setting a current mode type parameter to MODE_TYPE_INTRA in response to a slice type parameter indicating the I slice and at least a width or height of the block being greater than 64; decoding, from the encoded video bitstream, a flag indicating whether a skip mode is applied to the block; recovering the block based at least in part on the flag; The method includes:

Citation Information

Patent Citations

  • Dual prediction of weighting sample in video coding

    JP2023145592A

  • Image encoding / decoding method and device for performing prediction based on reset prediction mode type of leaf node, and method for transmitting bitstream

    JP2023509053A

  • Skip mode signaling

    WO2021047633A1