Method, device and storage medium for reconstructing video blocks in a video stream

By introducing intra-block copy (IBC) mode into video encoding, using block vector (BV) to predict intra-blocks, the problem of intra-block copy encoding mode in the prior art is solved, and more efficient video encoding is achieved.

CN116368800BActive Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202280006720.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-25
Filing Date
2022-04-13
Publication Date
2025-05-30
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The existing video encoding technology has problems of inefficiency in intra-block copy encoding mode, especially when processing complex video content, it is difficult to effectively utilize redundancy between intra-blocks.

Method used

A video encoding method is proposed, using block vector (BV) in the current frame through intra-block copy (IBC) mode to predict and encode using similarity between intra-blocks. This method implements the search and selection of block vectors in the encoder, and uses the reconstructed blocks as reference blocks for prediction.

Benefits of technology

Through the intra-block copy mode, the efficiency of video encoding is significantly improved, especially when processing video frames containing a large number of repetitive modes, the size of the bitstream can be effectively reduced and the encoding quality can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116368800B_ABST
    Figure CN116368800B_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to video coding, and more particularly, to an intra block copy coding mode. For example, a method for reconstructing a video block in a video stream is disclosed. The method may include: extracting at least one syntax element from the video stream, the at least one syntax element being associated with intra block copy (IBC) prediction of the video block; determining an IBC reference mode for IBC prediction of the video block, the IBC reference mode including one of the following: no IBC mode, local reference IBC mode, non-local reference IBC mode, and local and non-local reference IBC mode; and generating a reconstructed sample of the video block from the video stream based on the IBC reference mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims priority to U.S. Non - Provisional Patent Application No. 17 / 704,948, filed on March 25, 2022, which claims priority benefit to U.S. Provisional Application No. 63 / 245,665, titled "Method and Apparatus for Intra Block Copy (IntraBC) Mode Coding with Search Range Switching", filed on September 17, 2021. The entire contents of the two prior patent applications are incorporated by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to video coding, and more particularly to intra - block copy coding mode. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the content of the present disclosure. To the extent that the work of the currently named inventors is described in this background art section, the work of the named inventors and aspects that were not prior art at the time of filing of this application have never been expressly or implicitly recognized as prior art to the present disclosure.

[0005] Video coding and decoding can be performed using inter - picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial size of, for example, 1920x1080 luminance samples and associated full or subsampled chrominance samples. The series of pictures can have a fixed or variable picture rate (or frame rate), such as 60 pictures per second or 60 frames per second. Uncompressed video has specific bit - rate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920x1080, a frame rate of 60 frames per second, and a 4:2:0 chrominance subsampling of 8 bits per color channel per pixel requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.

[0006] One purpose of video encoding and decoding is to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the above bandwidth or storage space requirements, and in some cases can reduce them by two orders of magnitude or more than two orders of magnitude. Lossless compression and lossy compression, and combinations thereof, can be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to an encoding / decoding process where the original video information cannot be fully retained during the encoding process and cannot be fully recovered during the decoding process. When lossy compression is used, the reconstructed signal may be different from the original signal. Although there is some information loss, the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, lossy compression is widely used in many applications. The amount of tolerable distortion depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally enables the encoding algorithm to produce higher losses and higher compression ratios.

[0007] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0008] Video codec technology can include techniques known as intra-frame encoding. In intra-frame encoding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture can be called an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and a video session, or as a still image. Then the samples of the blocks after intra-frame prediction can be transformed to the frequency domain, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value and the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step.

[0009] Traditional intra coding (e.g., intra coding known from, for example, MPEG-2 generation coding techniques) does not use intra prediction. However, some newer video compression techniques include techniques that attempt block coding / decoding based on, for example, surrounding sample data and / or metadata that are obtained during spatially adjacent coding / decoding and whose decoding order is before the data blocks being intra-coded or decoded. Such techniques are hereafter referred to as "intra prediction" techniques. It should be noted that, in at least some cases, intra prediction uses only reference data from the currently being reconstructed picture and not reference data from other reference pictures.

[0010] There can be many different forms of intra prediction. When more than one such technique is available in a given video coding technique, the techniques used can be referred to as intra prediction modes. One or more intra prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes and / or can be associated with various parameters, and the mode / sub-mode information and the intra coding parameters for video blocks can be encoded either separately or jointly included in a mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination can have an impact on the coding efficiency gain through intra prediction, and entropy coding techniques can also be used to convert the codeword into a bitstream.

[0011] Certain intra prediction modes were introduced with H.264, improved in H.265, and further improved in newer coding techniques (e.g., Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS)). Generally, for intra prediction, the available adjacent sample values that are already available can be used to form a predictor block. For example, the available values of a particular set of adjacent samples along certain directions and / or lines can be copied into the predictor block. The reference to the direction in use can be encoded in the bitstream or can itself be predicted.

[0012] Referring Figure 1A , a subset of 9 predictor directions out of the 33 possible intra predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra modes specified in H.265) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the directions in which adjacent samples are used to predict the sample at 101. For example, arrow (102) represents predicting sample (101) from one or more adjacent samples to the upper right at a 45-degree angle to the horizontal direction. Similarly, arrow (103) represents predicting sample (101) from one or more adjacent samples to the lower left of sample (101) at a 22.5-degree angle to the horizontal direction.

[0013] Still referring Figure 1A, a square block (104) of 4x4 samples is depicted in the upper left (represented by the dashed thick line). The square block (104) includes 16 samples, and each sample is marked with "S" for its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension in the block (104). Since the size of the block is 4x4 samples, S44 is located in the lower right. Example reference samples following a similar numbering scheme are further shown. The reference samples are marked with "R" for their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, predictive samples adjacent to the block being reconstructed are used.

[0014] Intra-picture prediction of block 104 can start by copying reference sample values from adjacent samples according to the predicted direction signaled. For example, assume that the encoded video bitstream includes signaling that, for this block 104, indicates the prediction direction of arrow (102) - that is, to predict the sample direction from one or more predictive samples to the upper right at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0015] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a reference sample; especially when the direction cannot be evenly divisible by 45 degrees.

[0016] As video coding technologies continue to develop, the number of possible directions has increased. For example, in H.264 (in 2003), nine different directions are available for intra prediction. In H.265 (in 2013), it increased to 33 directions, and at the time of this application, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, thus accepting a certain bit penalty for the directions. In addition, sometimes the direction itself can be predicted from adjacent directions used in the intra prediction of already decoded adjacent blocks.

[0017] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to JEM is shown to illustrate the increase in the number of prediction directions in each coding technology over time.

[0018] The mapping of bits representing the intra prediction direction to the prediction direction in an encoded video bitstream can vary depending on the video coding technology; and the range can be, for example, from a simple direct mapping of the prediction direction to the intra prediction mode, to codewords, to complex adaptive schemes involving the most probable mode, and similar techniques. However, in all cases, there may be certain directions of intra prediction that are statistically less likely to occur in the video content compared to some other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technology, those less likely directions can be represented by more bits than the more likely directions.

[0019] Inter-picture prediction or inter prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used to predict a newly reconstructed picture or picture portion (e.g., block) after being spatially offset along a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture being used (similar to a temporal dimension).

[0020] In some video compression techniques, the current MV applicable to a certain region of sample data can be predicted based on other MVs, for example, based on MVs that are spatially adjacent to other regions of sample data and whose decoding order is prior to the current MV. Doing so can significantly reduce the overall amount of data required to encode the MVs by relying on eliminating redundancy in the related MVs, thereby increasing the compression ratio. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is the following statistical likelihood: a region larger than the region applicable to a single MV moves in a similar direction in the video sequence. Therefore, in some cases, a similar motion vector derived from the MVs of adjacent regions can be used to predict that larger region. This results in the actual MV of a given region being similar to or the same as the MV predicted from the surrounding MVs. After entropy coding, such an MV can in turn be represented by fewer bits than the number of bits used if the MV were directly encoded instead of being predicted from adjacent MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs.

[0021] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). In addition to the various MV prediction mechanisms specified in H.265, the technique hereinafter referred to as "spatial merge" is described below.

[0022] Specifically, still referring to Figure 2 , in spatial merge, the current block (201) includes samples that have been discovered by the encoder during the motion search process and are predictable based on the previous block of the same size that has been spatially shifted. An MV can be derived from the metadata associated with one or more reference pictures, for example, from the most recent (in decoding order) reference picture, using the MV associated with any one of five surrounding samples (denoted as A0, A1, and B0, B1, B2 (from 202 to 206) respectively), rather than directly encoding the MV. In H.265, MV prediction can use the predictors of the same reference pictures used by adjacent blocks. SUMMARY OF THE INVENTION

[0023] Aspects of the present disclosure generally relate to video coding, and more specifically, to an intra-block copy coding mode.

[0024] Aspects of the present disclosure also provide a video coding or decoding device or apparatus, including circuitry configured to perform any of the above-described method implementations.

[0025] Aspects of the present invention also provide a non-transitory computer-readable medium for storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform a method of video decoding and / or encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0027] Figure 1A A schematic diagram showing an exemplary subset of intra-prediction direction modes;

[0028] Figure 1B A schematic diagram showing an exemplary intra-prediction direction;

[0029] Figure 2 A schematic diagram showing a current block and its surrounding spatial merge candidates for motion vector prediction in an example;

[0030] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment;

[0031] Figure 4 A schematic diagram showing a simplified block diagram of a communication system according to another exemplary embodiment;

[0032] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an exemplary embodiment;

[0033] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an exemplary embodiment;

[0034] Figure 7 A block diagram showing a video encoder according to another exemplary embodiment;

[0035] Figure 8 A block diagram showing a video decoder according to another exemplary embodiment;

[0036] Figure 9 A scheme showing coding block partitioning according to an exemplary embodiment of the present disclosure;

[0037] Figure 10 Another scheme showing coding block partitioning according to an exemplary embodiment of the present disclosure;

[0038] Figure 11 Another scheme showing coding block partitioning according to an exemplary embodiment of the present disclosure;

[0039] Figure 12 An example showing the partitioning of a base block into coding blocks according to an exemplary partitioning scheme;

[0040] Figure 13 An exemplary ternary partitioning scheme is shown;

[0041] Figure 14 An exemplary quadtree - binary tree coding block partitioning scheme is shown;

[0042] Figure 15 A scheme showing the partitioning of a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present disclosure;

[0043] Figure 16 Another scheme showing the partitioning of a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the present disclosure;

[0044] Figure 17 Another scheme showing the partitioning of a coding block into multiple transform blocks according to an exemplary embodiment of the present disclosure;

[0045] Figure 18 The concept of intra - block copy (IBC) for predicting a current coding block using a reconstructed coding block in the same frame is shown;

[0046] Figure 19 Shows an example reconstructed sample that can be used as a reference sample for the IBC;

[0047] Figure 20 Shows an example reconstructed sample that can be used as a reference sample for the IBC with some example limitations;

[0048] Figure 21 Shows an example on - chip reference sample memory (RSM) update mechanism for the IBC;

[0049] Figure 22 Shows Figure 21 A spatial view of the example on - chip RSM update mechanism;

[0050] Figure 23 Shows another example on - chip RSM update mechanism for the IBC;

[0051] Figure 24 Shows a comparison of the spatial views of example RSM update mechanisms for the IBC for horizontally - split superblocks and vertically - split superblocks;

[0052] Figure 25 Shows example non - local and local search regions for an IBC reference block;

[0053] Figure 26 Shows exemplary limitations on the positions of reference blocks of the IBC using local and non - local reference block search regions;

[0054] Figure 27 Shows a flowchart of a method according to an example embodiment of the present disclosure; and

[0055] Figure 28 Shows a schematic diagram of a computer system according to an example embodiment of the present disclosure. Detailed Description of the Invention

[0056] The present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, it can be understood that the present invention can be implemented in various different forms. Therefore, any embodiment set forth below is intended to explain rather than limit the subject matter covered or claimed. It can also be understood that the present invention can be embodied as a method, device, component, or system. Thus, the embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.

[0057] Throughout the specification and claims, terms may have nuances of meaning that are not explicitly stated but are implied or implicit in the context. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, it is intended that the claimed subject matter include combinations of all or part of the exemplary embodiments / implementations.

[0058] Generally speaking, terms can be understood at least in part from their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can have various meanings, which can depend at least in part on the context in which these terms are used. Generally, "or" when used to relate a list such as A, B, or C is intended to mean A, B, and C (used here in the inclusive sense) as well as A, B, or C (used here in the exclusive sense). Additionally, the terms "one or more" or "at least one" as used herein, depending at least in part on the context, can be used to describe any feature, structure, or characteristic in the singular sense or can be used to describe a combination of features, structures, or characteristics in the plural sense. Similarly, terms such as "a", "an", or "the" can also be understood to convey singular usage or convey plural usage, which depends at least in part on the context. Furthermore, the terms "based on" or "determined by" can be understood to not necessarily imply an exclusive set of factors, but can allow for the existence of other factors that are not necessarily explicitly described, which also depends at least in part on the context.

[0059] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes pairs of terminal devices (310) and (320) interconnected via the network (350). In Figure 3In the example of, the first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (such as a video picture stream captured by the terminal device (310)) for transmission over the network (350) to another terminal device (320). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. The unidirectional data transmission can be implemented in a media service application or the like.

[0060] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conference. For the bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (such as a video picture stream captured by the terminal device) for transmission over the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device based on the recovered video data.

[0061] In Figure 3 the example of, the terminal devices (310), (320), (330), and (340) can be implemented as servers, personal computers, and smart phones, but the applicability of the basic principles of the present disclosure is not limited thereto. Embodiments of the present disclosure can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and / or similar devices. The network (350) represents any number or type of network that conveys encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless explicitly explained herein, the architecture and topology of the network (350) may be unimportant for the operation of the present disclosure.

[0062] As an example of an application of the disclosed subject matter, Figure 4Shows the placement of a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0063] A video streaming system may include a video capture subsystem (413), which may include a video source (401) such as a digital camera for creating an uncompressed video picture or image stream (402). In an example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402), depicted as a thick line to emphasize the high data volume, may be processed by an electronic device (420) compared to the encoded video data (404) (or encoded video bitstream), and the electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data volume compared to the uncompressed video picture stream (402), may be stored on a streaming server (405) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, e.g., Figure 4 the client subsystem (406) and the client subsystem (408) in, may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes an incoming copy (407) of the encoded video data and produces an output video picture stream (411) that is uncompressed and can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure. In some streaming systems, the encoded video data (404), the encoded video data (407), and the encoded video data (409) (e.g., video bitstreams) may be encoded according to certain video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC as well as other video coding standards.

[0064] It will be appreciated that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may include a video encoder (not shown).

[0065] Figure 5 A block diagram of a video decoder (510) according to any embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 the video decoder (410) in the example of

[0066] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source that transmits the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided outside and separated from the video decoder (510) (not depicted). Still in other applications, a buffer memory (not depicted) may be provided outside the video decoder (510) for purposes such as preventing network jitter, and another buffer memory (515) may be provided inside the video decoder (510) for purposes such as handling playback timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be required, or the buffer memory may be made smaller. For use on a best-effort packet network (e.g., the Internet), a buffer memory (515) of sufficient size may be required, and its size may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be implemented at least partially in an operating system or similar element (not shown) outside the video decoder (510).

[0067] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530) but may be coupled to the electronic device (530), as Figure 5 shown. The control information for one (or more) display devices may be in the form of Supplemental Enhancement Information (SEI messages) or a parameter set segment (not depicted) of Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include Group of Pictures (GOP), picture, tile, strip, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, and so on.

[0068] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0069] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter - picture and intra - picture, inter - block and intra - block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. The units involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and the multiple processing or functional units below are not depicted.

[0070] In addition to the functional blocks already mentioned, the video decoder (510) can conceptually be subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.

[0071] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive, as symbols (521), quantized transform coefficients and control information from the parser (520), including information indicating which inverse transform type, block size, quantization factor / parameter, quantization scaling matrix, etc. to use as symbols (521). The scaler / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).

[0072] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block having the same size and shape as the block being reconstructed using the reconstructed surrounding block information and the block information stored in the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some embodiments, the aggregator (555) may add the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.

[0073] In other cases, the output samples of the scaler / inverse transform unit (551) can belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) can access the reference picture memory (557) to extract samples for picture inter-prediction. After motion compensating the extracted samples according to the signs (521) belonging to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 can be referred to as residual samples or a residual signal), thereby generating output sample information. The motion compensation prediction unit (553) obtaining prediction samples from addresses within the reference picture memory (557) can be controlled by a motion vector, and the motion vector is in the form of signs (521) for use by the motion compensation prediction unit (553), and the signs (521) can have, for example, X, Y components (shifts) and reference picture components (times). Motion compensation can also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sampled accurate motion vectors, and can also be associated with a motion vector prediction mechanism, etc.

[0074] The output samples of the aggregator (555) can be subject to various loop filtering techniques in the loop filter unit (556). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video code bitstream) and available as signs (521) from the parser (520) for the loop filter unit (556). However, video compression techniques can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. Several types of loop filters can be included as part of the loop filter unit 556 in various orders, which will be described in further detail below.

[0075] The output of the loop filter unit (556) can be a sample stream that can be output to the rendering device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.

[0076] Once fully reconstructed, some encoded pictures can be used as reference pictures for future picture inter-prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0077] The video decoder (510) may perform decoding operations according to a predetermined video compression technique, such as that employed in the ITU-T H.265 recommendation standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile may select certain tools from all available tools in the video compression technique or standard as the only tools available under that profile. For compliance with the standard, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the encoded video sequence.

[0078] In some example embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to perform proper decoding and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0079] Figure 6 A block diagram of a video encoder (603) according to an example embodiment of the present disclosure is shown. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the example of

[0080] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example of Figure 6 ), and the video source may capture video images to be encoded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0081] A video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 YCrCb, RGB, XYZ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures or images, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0082] According to some example embodiments, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a controller (650). In some embodiments, as described below, the controller (650) can be functionally coupled to and control other functional units. For the sake of brevity, couplings are not depicted in the figures. Parameters set by the controller (650) can include rate control-related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions that relate to optimizing the video encoder (603) for a certain system design.

[0083] In some example embodiments, a video encoder (603) may be configured to operate in an encoding loop. As a simple description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and one (or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). Even though the embedded decoder 633 processes the video stream encoded by the source encoder 630 without entropy coding, the decoder (633) reconstructs the symbols in a manner similar to the way a (remote) decoder would create to create sample data (since in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols in the entropy coding and the encoded video bitstream can be lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exactly corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is used to improve the encoding quality.

[0084] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder of the video decoder (510) described above in conjunction with Figure 5 However, briefly referring additionally to Figure 5 , when symbols are available and the entropy encoder (645) and parser (520) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) of the encoder.

[0085] At this point, it can be observed that any decoder technology other than parsing / entropy decoding that may only exist in the decoder must also exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations, which are related to the decoding part of the encoder. Since the encoder technology is reciprocal to the decoder technology described comprehensively, the description of the encoder technology can be simplified. A more detailed description is provided only in certain areas or aspects below.

[0086] During operation, in some example embodiments, the source encoder (630) may perform motion-compensated predictive coding, referring to one or more previously encoded pictures designated as "reference pictures" in a video sequence, and this motion-compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (632) encodes the difference (or residue) in the color channels between the pixel blocks of the input picture and the pixel blocks of one (or more) reference pictures, and the reference picture can be selected as the prediction reference for the input picture. The term "residue" and its adjective form "residual" can be used interchangeably.

[0087] The local video decoder (633) can decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the coding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 6 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture, and can store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture, which has the same content (without transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0088] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be encoded, the predictor (635) can search in the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor (635) can operate block by block based on sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory (634).

[0089] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0090] The outputs of all the above functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc., thereby transforming the symbols into an encoded video sequence.

[0091] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over the communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0092] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may typically be assigned to any of the following picture types:

[0093] Intra pictures (I pictures), which may be pictures that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0094] Predictive pictures (P pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and reference index to predict the sample values of each block.

[0095] Bi - directional predictive pictures (B pictures), which may be pictures that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0096] Source pictures can typically be spatially subdivided into multiple sample-coded blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples), and coded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively coded, or the block can be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be predictively coded with reference to a previously coded reference picture either through spatial prediction or through temporal-domain prediction. Blocks of a B picture can be predictively coded with reference to one or two previously coded reference pictures either through spatial prediction or through temporal prediction. For other purposes, source pictures or pictures in intermediate processing can be subdivided into other types of blocks. As described in further detail below, the partitioning of coded blocks and other types of blocks can follow or not follow the same way.

[0097] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0098] In some example embodiments, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and stripes, SEI messages, VUI parameter set fragments, etc.

[0099] The captured video can be multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the temporal or other correlations between pictures. For example, a particular picture being encoded / decoded can be segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0100] In some example embodiments, bidirectional prediction techniques can be used for inter - picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be jointly predicted by a combination of the first reference block and the second reference block.

[0101] In addition, merge mode techniques can be used in inter - picture prediction to improve coding efficiency.

[0102] According to some example embodiments of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a per - block basis. For example, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU can include three coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Each CTU can be recursively split into one or more coding units (CUs) in a quadtree manner. For example, a 64x64 - pixel CTU can be split into a 64x64 - pixel CU or 4 32x32 - pixel CUs. Each of one or more of the 32x32 blocks can be further split into 4 16x16 - pixel CUs. In some example embodiments, each CU can be analyzed during encoding to determine the prediction type for the CU in its respective prediction type, such as an inter - prediction type or an intra - prediction type. Depending on temporal and / or spatial predictability, a CU can be split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, prediction operations in encoding (encoding / decoding) are performed on a per - prediction - block basis. The splitting of a CU into PUs (or PBs of different color channels) can be performed in various spatial patterns. For example, a luminance or chrominance PB can include a matrix of values (e.g., luminance values) of samples such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 samples, etc.

[0103] Figure 7 A diagram of a video encoder (703) according to another example embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into an encoded picture that is part of an encoded video sequence. An example video encoder (703) can be used in place ofFigure 4 the video encoder (403) in the example of

[0104] For example, the video encoder (703) receives a matrix of sample values for processing a block, which is a prediction block of, for example, 8x8 samples. Then, the video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to optimally encode the processing block. When determining the processing block to be encoded in the intra mode, the video encoder (703) can use intra prediction techniques to encode the processing block into an encoded picture; and when determining the processing block to be encoded in the inter mode or the bi-prediction mode, the video encoder (703) can use inter prediction or bi-prediction techniques to encode the processing block into an encoded picture, respectively. In some example embodiments, the merge mode can be used as a sub-mode of inter-picture prediction, where a motion vector is derived from one or more motion vector predictors without resorting to encoded motion vector components external to the predictor. In some other example embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) may include Figure 7 components (such as a mode decision module) not explicitly shown in

[0105] In Figure 7 the example of Figure 7 the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the example setup of

[0106] The inter encoder (730) is configured to receive samples of a current block (such as a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundancy information, a motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (such as a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information using a decoding unit 633 (shown as Figure 6 the residual decoder 728 of Figure 7 in the example encoder 620 of

[0107] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) may calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.

[0108] The general controller (721) may be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is an intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the prediction mode of the block is an inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.

[0109] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and a prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo quantization processing to obtain quantized transform coefficients. In various example embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture may be buffered in a memory circuit (not shown) and used as a reference picture.

[0110] The entropy encoder (725) may be configured to format a bitstream to produce an encoded block. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in an inter mode or a merge submode of a bi-prediction mode, there may be no residual information.

[0111] Figure 8 FIG. shows an example video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (810) may be used in place of Figure 4 the video decoder (410) in the example of

[0112] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in the example arrangement of

[0113] The entropy decoder (871) may be configured to reconstruct certain symbols based on the encoded picture, where these symbols represent syntax elements that make up the encoded picture. Such symbols may include, for example, the mode used to encode the block (e.g., intra mode, inter mode, bi-prediction mode, merge submode, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that identifies certain samples or metadata for use by the intra decoder (872) or the inter decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In the example, when the prediction mode is an inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information may be inverse quantized and provided to the residual decoder (873).

[0114] The inter decoder (880) may be configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0115] The intra decoder (872) may be configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.

[0116] The residual decoder (873) may be configured to perform inverse quantization to extract the dequantized transform coefficients, and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (used to include quantization parameter (QP)), which may be provided by the entropy decoder (871) (the data path is not depicted as this is merely low-data volume control information).

[0117] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module depending on the circumstances) to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture as part of the reconstructed video. It may be noted that other suitable operations such as deblocking operations may also be performed to improve the visual quality.

[0118] It may be noted that any suitable technology may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In some example embodiments, one or more integrated circuits may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810).

[0119] Returning to block partitioning for encoding and decoding, the general partitioning can start from a base block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After partitioning or splitting the base block according to any of the example partitioning processes described below or other processes or a combination thereof, the final partitions or groups of coded blocks can be obtained. Each of these partitions can be at one of the various partition levels in the partitioning hierarchy and can be of various shapes. Each partition can be referred to as a coded block (CB). For the various example partitioning embodiments described further below, each resulting CB can be of any allowed size and partitioning level. Since such partitioning can form units for which some basic encoding / decoding decisions can be made and the encoding / decoding parameters can be optimized, determined, and signaled in the coded video bitstream, such partitions are referred to as coded blocks. The highest or deepest level in the final partition represents the depth of the coded block partitioning structure of the tree. The coded blocks can be luminance coded blocks or chrominance coded blocks. The CB tree structure for each color can be referred to as a coded block tree (CBT).

[0120] The coded blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchical structures for all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures for the various color channels in a CTU can be the same or different.

[0121] In some embodiments, the partitioning tree scheme or structure for the luminance and chrominance channels may not need to be the same. In other words, the luminance and chrominance channels can have separate coding tree structures or patterns. Additionally, whether the luminance and chrominance channels use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being coded is a P, B, or I slice. For example, for an I slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P or B slice, the luminance and chrominance channels can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partitioning tree structure and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0122] In some example embodiments, a predefined partitioning pattern can be applied to the base block. As Figure 9As shown, an exemplary 4-way split tree may start from a first predefined level (e.g., 64x64 block level or other size, as the base block size), and the base block may be hierarchically split down to a predefined lowest level (e.g., 4x4 level). For example, the base block may be subject to four predefined split options or patterns indicated by 902, 904, 906, and 908, where the partition designated as R is allowed to be recursively split, i.e., the same split option as indicated in Figure 9 can be repeated at a lower scale until the lowest level (e.g., 4x4 level). In some embodiments, additional restrictions may be applied to Figure 9 the split scheme. In Figure 9 the embodiments, rectangular splits (e.g., 1:2 / 2:1 rectangular splits) may be allowed, but they may not be allowed to be recursive, while square splits are allowed to be recursive. If needed, the split according to Figure 9 and recursively generate the final coding block group. The coding tree depth may be further defined to indicate the split depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64x64 block) may be set to 0, and after the root block is further split once according to Figure 9 , the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from the 64x64 base block to the 4x4 minimum partition will be 4 (starting from level 0). This split scheme may be applied to one or more color channels. Each color channel may be split independently according to Figure 9 the scheme (e.g., for each color channel at each hierarchical level, the split pattern or option in the predefined pattern may be determined independently). Optionally, two or more color channels may share Figure 9 the same hierarchical pattern tree (e.g., the same split pattern or option in the predefined pattern may be selected for two or more color channels at each hierarchical level).

[0123] Figure 10 FIG. shows another example of a predefined split pattern that enables recursive splitting to form a split tree. As Figure 10 shown, an exemplary 10-way split structure or pattern may be predefined. The root block may start at a predefined level (e.g., from a base block at 128x128 level or 64x64 level). Figure 10 The exemplary split structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular splits. Figure 10 The split types with 3 sub-partitions indicated by 1002, 1004, 1006, and 1008 in the second row of Figure 10Any rectangular partition in the rectangular partition. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128x128 block) can be set to 0, and after the root block is further split once according to Figure 10 the coding tree depth increases by 1. In some embodiments, only the all-square partitions in 1010 are allowed to be recursively split into the next level of the split tree according to Figure 10 the pattern. In other words, for the square partitions within the T-shaped patterns 1002, 1004, 1006, and 1008, recursive splitting may not be allowed. If necessary, perform the splitting process according to Figure 10 and recursively generate the final coded block group. This splitting scheme can be applied to one or more color channels. In some embodiments, more flexibility can be added when using splits below the 8x8 level. For example, 2x2 chrominance inter prediction can be used in some cases.

[0124] In some other exemplary embodiments for coding block splitting, a quadtree structure can be used to split a base block or an intermediate block into quadtree partitions. This quadtree splitting can be applied hierarchically and recursively to any square partition. Whether to further perform quadtree splitting on the base block or intermediate block or partition can be adjusted according to various local characteristics of the base block or intermediate block / partition. The quadtree splitting at the picture boundary can be further adjusted. For example, an implicit quadtree splitting can be performed at the picture boundary such that the block will maintain the quadtree splitting until the size fits the picture boundary.

[0125] In some other exemplary embodiments, hierarchical binary splitting from the base block can be used. For such a scheme, the base block or intermediate-level block can be split into two partitions. The binary splitting can be horizontal or vertical. For example, a horizontal binary splitting can split the base block or intermediate block into an equal right partition and left partition. Similarly, a vertical binary splitting can split the base block or intermediate block into an equal upper partition and lower partition. This binary splitting can be hierarchical and recursive. A decision can be made at each base block or intermediate block in the base block or intermediate block as to whether the binary splitting scheme should continue, and if the scheme is further continued, a decision can be made as to whether to use horizontal or vertical binary splitting. In some embodiments, further splitting can stop at a predefined minimum partition size (one dimension or two dimensions). Optionally, once a predefined partition level or depth from the base block is reached, further splitting can be stopped. In some embodiments, the aspect ratio of the partition can be restricted. For example, the aspect ratio of the partition can be not less than 1:4 (or greater than 4:1). Thus, a vertical bar partition with a vertical-to-horizontal aspect ratio of 4:1 can only be further vertically binary split into an upper partition and a lower partition with a vertical-to-horizontal aspect ratio of 2:1.

[0126] In some other examples, such as Figure 13 As shown, the three-pronged segmentation scheme can be used to segment the base block or any intermediate block. The three-pronged pattern can be implemented vertically, such as Figure 13 1302, or horizontally implemented, such as Figure 13 1304. Although Figure 13 The example segmentation ratio in is shown as 1:2:1 vertically or horizontally, but other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. Since the ternary tree segmentation can capture objects located at the center of the block in one continuous partition, while the quadtree and binary tree always segment along the center of the block, thereby dividing the object into different partitions, this ternary segmentation scheme can be used to supplement the quadtree or binary segmentation structure. In some embodiments, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transformations.

[0127] The above partitioning schemes can be combined in any manner at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition the base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition can be quadtree partitioned or binary partitioned, if specified, subject to a predefined set of conditions. Specific examples such as Figure 14 As shown, in Figure 14 In the example of , as shown at 1402, 1404, 1406 and 1408, the base block is a first quadtree partitioned into four partitions. Thereafter, each of the resulting partitions is either quadtree partitioned into four further partitions (e.g., 1408), or binary partitioned at the next level into two further partitions (horizontally or vertically, such as 1402 or 1406, both of which are symmetrical), or is not partitioned (e.g., 1404). For square partitions, binary or quadtree partitioning can be allowed to be performed recursively, as shown in the overall example partitioning pattern of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree partitioning and dashed lines represent binary tree partitioning. A flag can be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, consistent with the partition structure of 1410, the flag "0" may indicate horizontal binary partitioning, and the flag "1" may indicate vertical binary partitioning. Since quadtree partitioning always partitions a block or partition horizontally and vertically to generate 4 sub-blocks / partitions of equal size, it is not necessary to indicate the partition type for quadtree partitioning. In some embodiments, the flag "1" may indicate horizontal binary partitioning, and the flag "0" may indicate vertical binary partitioning.

[0128] In some example embodiments of QTBT, the quadtree and binary split rule sets can be represented by the following predefined parameters and their associated respective functions:

[0129] - CTU size: The size of the root node of the quadtree (the size of the base block)

[0130] - MinQTSize: The minimum allowed size of the quadtree leaf nodes

[0131] - MaxBTSize: The maximum allowed size of the binary tree root nodes

[0132] - MaxBTDepth: The maximum allowed depth of the binary tree

[0133] - MinBTSize: The minimum allowed size of the binary tree leaf nodes

[0134] In some example embodiments of the QTBT splitting structure, the CTU size can be set to (when considering and using example chroma subsampling) 128x128 luma samples with two corresponding 64x64 chroma sample blocks, MinQTSize can be set to 16x16, MaxBTSize can be set to 64x64, and MinBTSize (for both width and height) can be set to 4x4, and MaxBTDepth can be set to 4. Quadtree splitting can be first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its minimum allowed size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the node is 128x128, since the size exceeds MaxBTSize (i.e., 64x64), this node will not be first split by the binary tree. Otherwise, nodes not exceeding MaxBTSize can be split by the binary tree. In Figure 14 the example, the base block is 128x128. According to the predefined rule set, the base block can only be split by the quadtree. The splitting depth of the base block is 0. Each of the four resulting partitions is 64x64, which does not exceed MaxBTSize and can be further split by the quadtree or binary tree at level 1. Continue the process. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of the binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of the binary tree node equals MinBTSize, further vertical splitting is not considered.

[0135] In some example embodiments, the above QTBT scheme can be configured to support the flexibility that luminance and chrominance have the same QTBT structure or independent QTBT structures. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be split into CUs by a QTBT structure, and the chrominance CTB can be split into chrominance CUs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can include coding blocks of a luminance component or coding blocks of two chrominance components, and a CU in a P or B slice can include coding blocks of all three color components.

[0136] In some other embodiments, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. This embodiment can be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary splitting of nodes, one of the ternary splitting patterns can be selected. Figure 13 In some embodiments, only square nodes can be ternary split. An additional flag can be used to indicate whether the ternary split is horizontal or vertical.

[0137] The design of two-level or multi-level trees, such as the QTBT embodiment and the QTBT embodiment supplemented by ternary splitting, can be mainly driven by complexity reduction. Theoretically, the complexity of traversing a tree is T D , where T represents the number of splitting types, and D is the depth of the tree. A full trade-off can be made by using multiple types (T) with a reduced depth (D).

[0138] In some embodiments, a coding block (CB) can be further divided. For example, for intra-frame or inter-frame prediction during encoding and decoding processes, the CB can be further divided into multiple prediction blocks. In other words, the CB can be further divided into different sub-partitions where separate prediction decisions / configurations can be made. In parallel, to describe the level at which the transformation or inverse transformation of video data is performed, the CB can be further divided into multiple transform blocks (TBs). The scheme for dividing the CB into prediction blocks (PBs) and TBs can be the same or different. For example, each division scheme can use its own process based on various characteristics of the video data, such as. In some example embodiments, the PB and TB division schemes can be independent. In some other example embodiments, the PB and TB division schemes and boundaries can be related. In some embodiments, for example, the TBs can be divided after the PB division, and specifically, each PB is determined after dividing the coding block and then can be further divided into one or more TBs. For example, in some embodiments, a PB can be divided into one, two, four, or other numbers of TBs.

[0139] In some embodiments, to divide a base block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and chrominance channels can be processed differently. For example, in some embodiments, for the luminance channel, dividing a coding block into prediction blocks and / or transform blocks can be allowed, while for one (or more) chrominance channels, dividing a coding block into prediction blocks and / or transform blocks can be not allowed. In such embodiments, thus, the transformation and / or prediction of luminance blocks can be performed only at the coding block level. As another example, the minimum transform block size of the luminance channel and one (or more) chrominance channels can be different. For example, a coding block of the luminance channel can be divided into smaller transform and / or prediction blocks than those of the chrominance channel. As another example, the maximum depth of dividing a coding block into transform blocks and / or prediction blocks can be different between the luminance channel and chrominance channels. For example, a coding block for the luminance channel can be divided into deeper transform blocks and / or prediction blocks than one (or more) chrominance channels. For a specific example, a luminance coding block can be divided into transform blocks of multiple sizes, which can be represented by recursive division, and can have transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4x4 to 64x64. However, for chrominance blocks, only the maximum possible transform blocks specified for luminance blocks are allowed.

[0140] In some example embodiments, for dividing a coding block into PBs, the depth, shape, and / or other characteristics of the PB division can depend on whether the PB is intra-frame encoded or inter-frame encoded.

[0141] The splitting of a coding block (or prediction block) into transform blocks can be implemented in various exemplary scenarios, including but not limited to recursive or non-recursive quadtree splitting and predetermined pattern splitting, with additional consideration given to transform blocks at the boundaries of the coding block or prediction block. Generally, the resulting transform blocks can be at different splitting levels, can have different sizes, and do not need to be square (e.g., can be rectangles with some possible sizes and aspect ratios). Further examples are described in more detail below in conjunction with Figure 15 , 16 , 17.

[0142] However, in some other embodiments, the CB obtained through any of the above splitting schemes can be used as the basic or minimum coding block for prediction and / or transformation. In other words, no further splitting is performed for the purposes of performing inter-frame prediction / intra-frame prediction and / or transformation. For example, the CB obtained from the above QTBT scheme can be directly used as the unit for performing prediction. Specifically, such a QTBT structure eliminates the concept of multiple partition types, i.e., the separation of CU, PU, and TU, and supports greater flexibility in the CU / CB partition shape as described above. In this QTBT block structure, the CU / CB can be square or rectangular in shape. The leaf nodes of this QTBT are used as the units for prediction and transformation processing without any further splitting. This means that in this exemplary QTBT coding block structure, the CU, PU, and TU have the same block size.

[0143] The above various CB splitting schemes and the further splitting of the CB into PB and / or TB (including no PB / TB splitting) can be combined in any way. The following specific embodiments are provided as non-limiting examples.

[0144] Specific example embodiments of coding block and transform block splitting are described below. In such example embodiments, the above recursive quadtree splitting can be used or (as Figure 19 and Figure 10Those) predefined partitioning patterns in divide the base block into coding blocks. At each level, whether a particular partition should be further quadtree partitioned can be determined by local video data characteristics. The resulting CBs can be at various quadtree partition levels and have various sizes. The decision on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture region can be made at the CB level (or CU level, for all three color channels). Each CB can be further partitioned into one, two, four, or some other number of PBs according to a predefined PB partitioning type. Within one PB, the same prediction process can be applied and relevant information can be sent to the decoder based on the PB. After obtaining the residual blocks by applying the prediction process based on the PB partitioning type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this particular embodiment, a CB or a TB can be, but is not limited to, square. Additionally, in this particular example, for inter-prediction, a PB can be square or rectangular, and for intra-prediction, a PB can be only square. A coding block can be partitioned into, for example, four square TBs. Each TB can be further recursively partitioned (using quadtree partitioning) into smaller TBs, called Residual Quadtree (RQT).

[0145] Another exemplary embodiment of dividing the base block into CBs, PBs, and / or TBs is further described below. For example, instead of using multiple partition unit types as shown in Figure 9 or Figure 10 , a quadtree with a nested multi-type tree (a partitioning structure using binary and ternary partitioning (e.g., QTBT as described above or QTBT with ternary partitioning)) can be used. The separation of CBs, PBs, and TBs can be waived (i.e., dividing a CB into PBs and / or TBs, and dividing a PB into TBs), unless, when needed, the size of a CB is too large for the maximum transform length, and such a CB may need to be further partitioned. This example partitioning scheme can be designed to support greater flexibility in the CB partitioning shape so that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, a CB can be square or rectangular in shape. Specifically, a Coding Tree Block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary partitioning is shown in Figure 11 Specifically, Figure 11The example multi-type tree structure includes four splitting types, called vertical binary splitting (SPLIT_BT_VER) (1102), horizontal binary splitting (SPLIT_BT_HOR) (1104), vertical ternary splitting (SPLIT_TT_VER) (1106), and horizontal ternary splitting (SPLIT_TT_HOR) (1108). Then, the CB corresponds to the leaf of the multi-type tree. In this example embodiment, unless the CB is too large for the maximum transform length, this partition is used for prediction and transform processing without any further splitting. This means that, in most cases, in a quadtree with a nested multi-type tree coding block structure, the CB, PB, and TB have the same block size. An exception occurs when the maximum supported transform length is less than the width or height of the color components of the CB. In some embodiments, in addition to binary or ternary splitting, Figure 11 the nested pattern may also include quadtree splitting.

[0146] Figure 12 FIG. shows a specific example of a quadtree with a nested multi-type tree coding block structure having block splitting (including quadtree, binary, and ternary splitting options) for one base block. More specifically, Figure 12 FIG. shows that the base block 1200 is split into four square partitions 1202, 1204, 1206, and 1208 by a quadtree. For each quadtree split partition, a decision is made to further use Figure 11 the multi-type tree structure and the quadtree for further splitting. In Figure 12 the example, partition 1204 is not further split. Partitions 1202 and 1208 each adopt another quadtree split. For partition 1202, the third-level split of the quadtree for the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split respectively adopts Figure 11 the third-level split of the quadtree, horizontal binary splitting 1104, no splitting, and Figure 11 the horizontal ternary splitting 1108. Partition 1208 adopts another quadtree split, and the upper left, upper right, lower left, and lower right splits of the second-level quadtree split respectively adopt Figure 11 the third-level split of the vertical ternary splitting 1106, no splitting, no splitting, and Figure 11 the horizontal binary splitting 1104. The two sub-partitions of the third-level upper left partition of 1208 are further split respectively according to Figure 11 the horizontal binary splitting 1104 and the horizontal ternary splitting 1108. Partition 1206 adopts the second-level split mode according to Figure 11 the vertical binary splitting 1102, and is divided into two partitions, and these two partitions are further split at the third level according to Figure 11 the horizontal ternary splitting 1108 and the vertical binary splitting 1102. According to Figure 11The horizontal binary split 1104, and the fourth-level split is further applied to one of these two partitions.

[0147] For the specific example above, the maximum luminance transform size can be 64x64, and the maximum supported chrominance transform size can be different from that of luminance. For example, it can be 32x32. Even though Figure 12 in the above example in , the CBs are generally not further split into smaller PBs and / or TBs, when the width or height of a luminance coding block or a chrominance coding block is greater than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically split in the horizontal and / or vertical directions to conform to the transform size limit in that direction.

[0148] As described above, for the specific example of splitting a base block into CBs, the coding tree scheme can support the ability for luminance and chrominance to have separate block tree structures. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can have the same coding tree structure. For example, for I slices, luminance and chrominance can have separate coded block tree structures. When applying separate block tree structures, the luminance CTB is split into luminance CBs through one coding tree structure, and the chrominance CTB is split into chrominance CBs through another coding tree structure. This means that a CU in an I slice can include coded blocks of the luminance component or coded blocks of both chrominance components, and a CU in a P or B slice always includes coded blocks of all three color components, unless the video is monochrome.

[0149] When a coding block is further split into multiple transform blocks, the transform blocks therein can be sorted in the bitstream in various orders or scan patterns. Example embodiments of splitting a coding block or a prediction block into transform blocks and the coding order of the transform blocks are described in further detail below. In some example embodiments, as described above, the transform split can support transform blocks of multiple shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the range of transform block sizes is from, for example, 4x4 to 64x64. In some embodiments, if the coding block is less than or equal to 64x64, then the transform block split can be applied only to the luminance component, so for chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, then both the luminance and chrominance coding blocks can be implicitly split into multiples of min(W, 64)xmin(H, 64) and min(W, 32)xmin(H, 32) transform blocks, respectively.

[0150] In some example embodiments of transform block partitioning, for intra - and inter - frame coded blocks, the coded block can be further divided into multiple transform blocks, with a partitioning depth up to a predetermined number of levels (e.g., 2 levels). The transform block partitioning depth and size can be related. For some example embodiments, the mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.

[0151] Table 1 Transform Partitioning Size Settings

[0152]

[0153] Based on the example mapping in Table 1, for a 1:1 square block, the next - level transform partitioning can create four 1:1 square sub - transform blocks. For example, the transform partitioning can stop at 4x4. Thus, a transform size of 4x4 at the current depth corresponds to the same size of 4x4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non - square block, the next - level transform partitioning can create two 1:1 square sub - transform blocks, and for a 1:4 / 4:1 non - square block, the next - level transform partitioning can create two 1:2 / 2:1 sub - transform blocks.

[0154] In some example embodiments, for the luminance component of an intra - frame coded block, additional restrictions can be applied with respect to the transform block partitioning. For example, for each level of transform partitioning, all sub - transform blocks can be restricted to have equal sizes. For example, for a 32x16 coded block, the first - level transform partitioning creates two 16x16 sub - transform blocks, and the second - level transform partitioning creates eight 8x8 sub - transform blocks. In other words, the second - level partitioning must be applied to all first - level sub - blocks to keep the transform unit sizes equal. Figure 15 An example of the transform block partitioning of an intra - frame coded square block according to Table 1, and the coding order shown by arrows, is presented. Specifically, 1502 shows the square coded block. The first - level partitioning into 4 transform blocks of equal size according to Table 1 is shown in 1504, where the coding order is indicated by the arrows. The second - level partitioning of all first - level equal - size blocks into 16 equal - size transform blocks according to Table 1 is shown in 1506, where the coding order is indicated by the arrows.

[0155] In some example embodiments, for the luminance component of an inter - frame coded block, the above - mentioned restrictions for intra - frame coding may not be applied. For example, after the first - level transform partitioning, any one of the sub - transform blocks can be further independently partitioned by more than one level. Thus, the resulting transform blocks can have the same size or not. Figure 16 An example of partitioning an inter - frame coded block into transform blocks and its coding order is shown. In Figure 16In the example of Figure 16 , the inter-frame encoded block 1602 is divided into transform blocks at two levels according to Table 1. At the first level, the inter-frame encoded block is divided into four transform blocks of equal size. Then, as shown in 1604, only one (not all) of the four transform blocks is further divided into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes. The example encoding order of these 7 transform blocks is shown by the arrows in

[0156] In some example embodiments, for one (or more) chrominance components, some additional restrictions for the transform blocks may be applied. For example, for one (or more) chrominance components, the transform block size may be the same as the coding block size, but not less than a predefined size, such as 8x8.

[0157] In some other example embodiments, for coding blocks with a width (W) or height (H) greater than 64, the luminance and chrominance coding blocks may be implicitly divided into transform units that are multiples of min(W, 64) x min(H, 64) and min(W, 32) x min(H, 32), respectively. Here, in the present disclosure, "min(a, b)" may return the smaller value between a and b.

[0158] Figure 17 Another alternative example scheme for dividing a coding block or a prediction block into transform blocks is further shown. As shown in Figure 17 , a predefined set of partitioning types may be applied to the coding block according to the transform type of the coding block, without using recursive transform partitioning. In the specific example shown in Figure 17 , one of 6 example partitioning types may be applied to divide the coding block into various numbers of transform blocks. This scheme for generating transform block partitions may be applied to coding blocks or prediction blocks.

[0159] More specifically, Figure 17 's partitioning scheme provides up to 6 example partitioning types for any given transform type (the transform type refers to, for example, the type of the main transform, such as ADST, etc.). In this scheme, a transform partition type may be assigned to each coding block or prediction block based on, for example, the rate-distortion cost. In the example, the transform partition type assigned to the coding block or prediction block may be determined based on the transform type of the coding block or prediction block. A specific transform partition type may correspond to a transform block partition size and pattern, as shown by the 6 transform partition types shown in Figure 17 . The correspondence between various transform types and various transform partition types may be predefined. The following examples are shown, where the capital labels indicate the transform partition types that may be assigned to the coding block or prediction block based on the rate-distortion cost:

[0160] ·PARTITION_NONE: Assign a transform size equal to the block size.

[0161] ·PARTITION_SPLIT: Allocate a transform size that is 1 / 2 the width of the block size and 1 / 2 the height of the block size.

[0162] ·PARTITION_HORZ: Allocate a transform size that is the same width as the block size and 1 / 2 the height of the block size.

[0163] ·PARTITION_VERT: Allocate a transform size that is 1 / 2 the width of the block size and the same height as the block size.

[0164] ·PARTITION_HORZ4: Allocate a transform size that is the same width as the block size and 1 / 4 the height of the block size.

[0165] ·PARTITION_VERT4: Allocate a transform size that is 1 / 4 the width of the block size and the same height as the block size.

[0166] In the above example, the transform split types as Figure 17 shown all include a unified transform size for the split transform blocks. This is only an example and not a limitation. In some other embodiments, mixed transform block sizes can be used for the split transform blocks in a particular split type (or mode).

[0167] Video blocks (PB or CB, also called PB when not further split into multiple prediction blocks) can be predicted in various ways instead of being directly encoded, thereby leveraging various correlations and redundancies in the video data to improve compression efficiency. Accordingly, such prediction can be performed in various modes. For example, a video block can be predicted by intra prediction or inter prediction. Particularly in the inter prediction mode, a video block can be predicted from one or more other frames by one or more other reference blocks or inter predictor blocks through single-reference or composite-reference inter prediction. To implement inter prediction, a reference block can be specified by its frame identifier (the temporal position of the reference block) and a motion vector (the spatial position of the reference block) indicating the spatial offset between the currently encoded or decoded block and the reference block. The reference frame identifier and the motion vector can be signaled in the bitstream. The motion vector as the spatial block offset can be directly signaled, or it can be predicted by another reference motion vector or predictor motion vector. For example, the current motion vector can be directly predicted by a reference motion vector (e.g., of a candidate neighboring block), or the current motion vector can be predicted by a combination of a reference motion vector and the motion vector difference (MVD) between the current motion vector and the reference motion vector. The latter can be referred to as the merge mode with motion vector difference (MMVD). The reference motion vector can be identified in the bitstream as a pointer pointing to, for example, a spatially neighboring block of the current block or a temporally neighboring but spatially co-located block.

[0168] In some other exemplary embodiments, intra block copy (IBC) prediction may be employed. In IBC, another block in the current frame (instead of a temporally different frame, thus called "intra") is used in combination with a block vector (BV) to predict a current block in the current frame, and the block vector is used to indicate the offset of the position of the intra predictor or the reference block relative to the position of the block to be predicted. The position of the coded block may be represented, for example, by the pixel coordinates of the upper left corner relative to the upper left corner of the current frame (or slice). Thus, the IBC mode uses an inter prediction concept within the current frame. For example, the BV may be predicted directly by other reference BVs or in combination with the BV difference between the current BV and the reference BV, which is consistent with using the reference MV and the MV difference to predict the MV in inter prediction. IBC is useful in providing improved coding efficiency, especially for encoding and decoding video frames with screen content, which has, for example, a large number of repeating patterns, such as text information, where the same text segments (letters, symbols, words, phrases, etc.) appear in different parts of the same frame and can be used to predict each other.

[0169] In some embodiments, IBC may be regarded as a separate prediction mode in addition to the normal intra prediction mode and the normal inter prediction mode. Thus, the prediction mode of a particular block can be selected and signaled among three different prediction modes: intra prediction, inter prediction, and IBC mode. In these embodiments, flexibility can be established in each of these modes to optimize the coding efficiency of each of these modes. In some other embodiments, similar motion vector determination, reference, and coding mechanisms may be used, and IBC may be regarded as a sub-mode or a branch within the inter prediction mode. In such embodiments (integrating the inter prediction mode and the IBC mode), the flexibility of IBC may be limited to coordinate the general inter prediction mode and the IBC mode. However, this embodiment is less complex while still being able to utilize IBC to improve the coding efficiency of video frames characterized by, for example, screen content. In some exemplary embodiments, by using the existing pre-specified mechanisms for the separate inter prediction mode and intra prediction mode, the inter prediction mode can be extended to support IBC.

[0170] The selection of these prediction modes can be made at various levels, including but not limited to the sequence level, frame level, picture level, slice level, CTU level, CT level, CU level, CB level, or PB level. For example, for the purpose of IBC, the decision on whether to adopt the IBC mode can be made and signaled at the CTU level. If a CTU is signaled to adopt the IBC mode, all the coding blocks in that entire CTU can be predicted by IBC. In some other embodiments, the IBC prediction can be determined at the super block (SB) level. Each SB can be divided into multiple CTUs or partitions in various ways (e.g., quadtree splitting). Examples are provided further below.

[0171] Figure 18 An example snapshot showing a portion of the current frame containing multiple CTUs from the perspective of the decoder is shown. Each square block (such as 1802) represents a CTU. As detailed above, the size of a CTU can be one of various predefined sizes. Each CTU can include one or more coding blocks (or prediction blocks for a particular color channel). The CTUs with horizontal shading represent those that have been reconstructed. CTU 1804 represents the current CTU being reconstructed. In the current CTU 1804, the coding blocks with horizontal shading represent those that have been reconstructed in the current CTU, the coding block 1806 with diagonal shading is currently being reconstructed, and the coding blocks without shading in the current CTU 1804 are waiting to be reconstructed. The other CTUs without shading have not been processed yet.

[0172] The position or offset of the reference block (relative to the current block) for predicting the current coding block in IBC can be indicated by a BV, as Figure 18 shown by the example arrows in. For example, a BV can indicate the position difference between the reference block (labeled "Ref" in Figure 18 ) and the upper left corner of the current block in vector form. While Figure 18 is shown using the CTU as the basic IBC unit. The basic principle applies to embodiments using the SB as the basic IBC unit. In such embodiments, as described in more detail below, each super block can be divided into multiple CTUs, and each CTU can be further divided into multiple coding blocks.

[0173] As will be disclosed further in more detail below, depending on the position of the reference CTU / SB relative to the current CTU / SB of the IBC, the reference CTU / SB can be referred to as a local CTU / SB or a non-local CTU / SB. A local CTU / SB can refer to a CTU / SB that coincides with the current CTU / SB, or a CTU / SB that is close to the current CTU / SB and has been reconstructed (e.g., the left adjacent CTU / SB of the current CTU / SB). A non-local CTU / SB can refer to a CTU / SB that is farther away from the current CTU / SB. When performing IBC prediction for the current coding block, either one or both of the local CTU / SB and the non-local CTU / SB can be searched for the reference block. Since the on-chip and off-chip storage management (e.g., off-chip picture buffer (DPB) and / or on-chip memory) of the reconstructed samples for the local or non-local CTU / SB reference can be different, the specific manner of implementing IBC can depend on whether the reference CTU / SB is local or non-local. For example, the reconstructed local CTU / SB samples can be suitable for storage in the on-chip memory of the IBC encoder or decoder. For example, the reconstructed non-local CTU / SB samples can be stored in the off-chip DPB memory.

[0174] In some embodiments, the position of the reconstructed blocks that can be used as reference blocks for the current coding block 1804 can be restricted. Such a restriction can be the result of various factors and can depend on whether IBC is implemented as an integrated part of the general inter prediction mode, a special extension of the inter prediction mode, or a separate and independent IBC mode. In some examples, only the currently reconstructed CTU / SB samples can be searched to identify the IBC reference block. In some other examples, as shown by the thick dashed box 1808 in Figure 18 , the currently reconstructed CTU / SB samples and another adjacent reconstructed CTU / SB sample (e.g., the left adjacent CTU / SB) can be used for reference block search and selection. For such an embodiment, only the local reconstructed CTU / SB samples can be used for IBC reference block search and selection. In some other examples, for various other reasons, certain CTU / SB may not be available for IBC reference block search and selection. For example, as will be described further below, Figure 18 the CTU / SB 1810 marked with crosshairs in

[0175] may be used for special purposes (e.g., wavefront parallel processing), so they may not be available for the search and selection of the reference block for the current block 1804. Figure 19Shows an example where each square represents a CTU / SB. As Figure 19 shown by the CTU / SB with diagonal shading, parallel decoding can be implemented, where multiple consecutive rows and multiple CTU / SBs in every other column (every two columns) can be reconstructed in a parallel processing manner. Other CTU / SBs with horizontal shading have been reconstructed, and the CTU / SBs without shading are the CTU / SBs to be reconstructed. In such parallel processing, for the currently parallel processed CTU / SB with its upper left coordinates being (x 0 , y 0 ), only when the vertical coordinate y is less than y 0 and the horizontal coordinate x is less than x 0 + 2(y 0 - y), can the reconstructed samples at (x, y) be accessed to predict the current CTU / SB in IBC. Therefore, the reconstructed CTU / SBs with horizontal shading can be used as references for the current block in parallel processing.

[0176] In some embodiments, the write-back latency of writing immediately reconstructed samples to off-chip DPB may further limit the CTU / SBs that can be used to provide IBC reference samples for the current block, especially when off-chip DPB is used to store IBC reference samples. Figure 19 Shows an example where, based on the limitations shown in Figure 19 , other limitations can also be applied. Specifically, to allow for hardware write-back latency, IBC prediction may not access the immediately reconstructed region to search for and select reference blocks. The number of immediately reconstructed regions that are restricted or prohibited can be 1 to n CTU / SBs (n is a positive integer). Therefore, based on the specific parallel processing limitations in Figure 19 , the coordinates of the upper left position of a current CTU / SB are (x 0 , y 0 ), if the vertical coordinate y is less than y 0 and the horizontal coordinate is less than x 0 + 2(y 0 - y) - D, then the prediction at position (x, y) can be accessed through IBC, where D represents the number of immediately reconstructed regions that restrict / prohibit us from using as IBC references (e.g., to the left of the current CTU / SB). Figure 20 Shows that for D = 2, such additional CTU / SBs are restricted as IBC reference samples. These additional CTU / SBs that cannot be used as IBC references are represented by anti-diagonal shading.

[0177] In some embodiments, as further described in detail below, both local and non-local CTU / SB search regions can be used for IBC reference block search and selection. Additionally, when on-chip memory is used, some restrictions on write-back latency can be relaxed or removed regarding the availability of already constructed CTUs / SBs as IBC references. In some further embodiments, due to differences in buffer management of reference blocks using, for example, on-chip or off-chip memory, the usage of local CTUs / SBs and non-local CTUs / SBs can be different when they coexist. These embodiments are described in more detail in the disclosure below.

[0178] In some embodiments, IBC can be implemented as an extension of the inter-frame prediction mode, treating the current frame as a reference frame in the inter-frame prediction mode such that blocks within the current frame can be used as prediction references. Thus, even if the IBC process only involves the current frame, such IBC embodiments can follow the encoding path of inter-frame prediction. In such embodiments, the reference structure of the inter-frame prediction mode can be adapted for IBC, where the representation of the addressing mechanism for reference samples using BV can be similar to the motion vector (MV) in inter-frame prediction. Thus, IBC can be implemented as a special inter-frame prediction mode, relying on similar or identical syntax structures and decoding processes as the inter-frame prediction mode with the current frame as the reference frame.

[0179] In such embodiments, since IBC can be regarded as an inter-frame prediction mode, only the intra-predicted slices must become slices that allow the use of IBC for prediction. In other words, only intra-predicted slices are not inter-frame predicted (since the intra-prediction mode does not invoke any inter-frame prediction processing path), and thus IBC cannot be used for prediction in such only intra slices. When IBC is applicable, the encoder will extend the reference picture list by an entry that points to the current picture. Thus, the current picture can occupy up to one picture-sized buffer in the shared decoded picture buffer (DPB). The signaling of using IBC can be implicit in the selection of the reference frame in the inter-frame prediction mode. For example, when the selected reference picture points to the current picture, if needed and available, the coding unit will employ IBC with a coding path similar to inter-frame prediction with a special IBC extension. In some specific embodiments, contrary to conventional inter-frame prediction, the reference samples in the IBC process may not be loop-filtered before being used for prediction. Additionally, the corresponding reference current picture can be a long-term reference frame as it will be close to the next frame to be encoded or decoded. In some embodiments, to minimize memory requirements, the encoder can release the buffer immediately after reconstructing the current picture. When the reconstructed picture becomes a reference image for subsequent frames in true inter-frame prediction, the encoder can fill the filtered version of the reconstructed picture back into the DPB as a short-term reference, even if it may be unfiltered when used for IBC.

[0180] In the above exemplary embodiments, even if IBC can be merely an extension of the inter-frame prediction mode, several special processes that may deviate from normal inter-frame prediction can be used to process IBC. For example, IBC reference samples can also be unfiltered. In other words, the reconstructed samples before in-loop filtering processing can be used for IBC prediction, and the in-loop filtering processing includes deblocking filtering, Sample Adaptive Offset (SAO) filtering, Cross-Component Sample Offset (CCSO) filtering, etc., while the normal inter-frame prediction mode uses the filtered samples for prediction. For another example, no luma sample interpolation for IBC may be performed, and chroma sample interpolation may be necessary only when the chroma BV is non-integer when derived from the luma BV. For another example, when the chroma BV is non-integer and the reference block for IBC is near the boundary of the available region of IBC reference, the surrounding reconstructed samples can be outside the boundary to perform chroma interpolation. A BV pointing to a single adjacent boundary cannot avoid this situation.

[0181] In such an embodiment, the prediction of the current block by IBC can reuse the prediction and encoding mechanisms of the inter-frame prediction process, including using the reference BV to predict the current BV and, for example, additional BV differences. However, in some specific embodiments, the luma BV can be implemented at integer resolution instead of at fractional precision as in the MV of conventional inter-frame prediction.

[0182] In some embodiments, in addition to the two CTUs ( Figure 18 indicated by the cross lines in Figure 18 as shown in 1810 in Figure 18 ), which are to the right and above the current CTU for Wavefront Parallel Processing (WPP),

[0183] all CTUs and SBs indicated by the horizontal hatching in Figure 18The thick dashed box 1808 indicates an example. In this example, the CTU / SB to the left of the current CTU can be used as a reference sample region for IBC at the beginning of the reconstruction process of the current CTU. When using this local reference region, on-chip memory space can be allocated to hold the local CTU / SB for IBC reference, rather than allocating additional external memory space in the DPB. In some embodiments, fixed on-chip memory can be used for IBC, thus reducing the complexity of implementing IBC in the hardware architecture. Therefore, a dedicated IBC mode independent of normal inter prediction can be implemented using on-chip memory, rather than just being implemented as an extension of the inter prediction mode.

[0184] For example, for each color component, the fixed on-chip memory size for storing local IBC reference samples (e.g., the left CTU or SB) can be 128x128. In some embodiments, the maximum CTU size can also be 128x128. In this case, the reference sample memory (RSM) can hold samples with a single CTU size. In some other alternative embodiments, the CTU size can be smaller. For example, the CTU size can be 64x64. Thus, the RSM can hold multiple (4 in this example case) CTUs simultaneously. In some other embodiments, the RSM can hold multiple SBs, each SB can include one or more CTUs, and each CTU can include multiple coding blocks.

[0185] In some embodiments of local on-chip IBC reference, the on-chip RSM holds one CTU and can implement a continuous update mechanism to replace the reconstructed samples of the left adjacent CTU with the reconstructed samples of the current CTU. Figure 21 A simplified example of this continuous RSM update mechanism at four intermediate times during the reconstruction process is shown. At Figure 21 In the example, the RSM has a fixed size that can accommodate one CTU. The CTU can include implicit partitions. For example, the CTU can be implicitly divided into four separate regions (e.g., quadtree partitioning). Each region can include multiple coding blocks. The size of the CTU can be 128x128, while for the example quadtree partitioning, the size of each example region or partition can be 64x64. The region / partition of the RSM with horizontal line shading at each intermediate time holds the corresponding reconstructed reference samples of the left adjacent CTU, and the region / partition with gray vertical line shading holds the corresponding reconstructed reference samples of the current CTU. The coding blocks of the RSM with diagonal shading represent the current coding blocks within the current region being encoded / decoded / reconstructed.

[0186] At a first intermediate time indicating the start of the current CTU reconstruction, as shown in 2102, the RSM for each of the four example regions may include only the reconstructed reference samples of the left adjacent CTU. At the other three intermediate times, the reconstruction process gradually replaces the reconstructed reference samples of the left adjacent CTU with the reconstructed samples of the current CTU. When the encoder processes the first coded block of the region / partition, the 64x64 region / partition in the RSM is reset. When resetting the region of the RSM, the region is considered blank and is considered to have not saved any reconstructed reference samples for IBC (in other words, this region of the RSM is not ready to be used as an IBC reference sample). When processing the corresponding current coded block in the region, the corresponding block in the RSM is classified into the reconstructed samples of the corresponding block of the current CTU to be used as a reference sample for the IBC of the next current block, as shown in Figure 21 the intermediate times 2104, 2106, 2108 shown in. Once the region / partition corresponding to the RSM has been processed for all coded blocks, the entire region is filled with the reconstructed samples of these current coded blocks as IBC reference samples, as shown in Figure 21 the regions fully shaded with vertical lines at each intermediate time shown in. Thus, at intermediate times 2104 and 2106, some regions / partitions in the RSM save IBC reference samples from adjacent CTUs, some other regions / partitions fully save reference samples from the current CTU, and some regions / partitions partially save reference samples from the current CTU and are partially blank (not used for IBC reference due to the above reset process). When the last region (e.g., the bottom right region) is processed, all the other three regions will save the reconstructed samples of the current CTU as reference samples for IBC, and the last region / partition partially saves the reconstructed samples of the corresponding coded block in the current CTU and is partially blank until the last coded block of the CTU is reconstructed, at which time the entire RSM saves the reconstructed samples of the current CTU and if still coded in IBC mode, the RSM is ready for the next CTU.

[0187] Figure 22 Illustrates the above continuous update implementation of the RSM in space at a specific intermediate time, i.e., both the left adjacent CTU and the current CTU with the current coded block (the slant-shaded block) are shown. The corresponding reconstructed samples of these two CTUs that are in the RSM and effectively serve as IBC reference samples for the current coded block are shown by horizontal and vertical hatching. At the specific reconstruction time in this example, in the RSM, the process has replaced the samples covered by the unshaded region in the left adjacent CTU with the region of the current CTU covered by vertical hatching. The remaining effect samples from the adjacent CTU are shown as horizontal hatching.

[0188] In the above example embodiment, when the fixed RSM size is the same as the CTU size, the RSM is implemented to contain one CTU. In some other embodiments where the CTU size is smaller, the RSM may contain more than one CTU. For example, the size of the CTU may be 32x32, while the size of the fixed RSM may be 128x128. Thus, the RSM can store samples of 16 CTUs. Following the same basic RSM update principle described above, the RSM can store 16 adjacent CTUs of the current 128x128 patch before being reconstructed. Once the processing of the first coded block of the current 128x128 patch begins, the first 32x32 region of the RSM, which stores the reconstructed samples of a single adjacent CTU as described above, can be initially filled. The remaining 15 32x32 regions contain 15 adjacent CTUs as reference samples for IBC. Once the CTU corresponding to the first 32x32 region of the currently decoded 128x128 patch is reconstructed, the first 32x32 region of the RSM is updated with the reconstructed samples of that CTU. Then, the CTU corresponding to the second 32x32 region of the current 128x128 patch can be processed and ultimately updated with the reconstructed samples. This process continues until the 16 32x32 regions of the RSM contain the reconstructed samples (all 15 CTUs) of the current 128x128 patch. The decoding process then moves to the next 128x128 patch.

[0189] In some other embodiments, as an Figure 21 and 22 extension, the RSM can store a set of adjacent CTUs. Processing one current CTU at a time, the part of the RSM that stores the farthest adjacent CTU is updated with the reconstructed current CTU in the above manner. For the next current CTU, similarly, the farthest adjacent CTU in the RSM is updated and replaced. Thus, the multiple CTUs stored in the fixed-size RSM are updated as a moving window of adjacent CTUs for IBS.

[0190] Another specific example embodiment of using on-chip RSM for local IBC is as Figure 23 shown. In this example, the maximum block size of the IBC mode may be limited. For example, the maximum IBC block size may be 64x64. The on-chip RSM can be configured to have a fixed size corresponding to a superblock (SB), such as 128x128. Figure 23 The RSM embodiment of Figure 21 and Figure 22 uses a basic principle similar to the embodiments of Figure 23 . In Figure 23In the example, the SB can be a quadtree segmentation. Accordingly, the RSM can be segmented into 4 regions or units by the quadtree, and each region or unit is 64x64. Each of these regions can store one or more coded blocks. Optionally, each of these regions can store one or more CTUs, and each CTU can store one or more coded blocks. The coding order of the quadtree regions can be predefined. For example, the coding order can be top left, top right, bottom left, bottom right. Figure 23 The quadtree segmentation of the SB is only an example. In some other alternative embodiments, the SB can be segmented according to any other scheme. The RSM update embodiments described herein for local IBC are applied to those alternative segmentation schemes.

[0191] In this local SBC embodiment, the local reference blocks available for SBC prediction can be restricted. For example, it can be required that the reference block and the current block should be in the same SB row. Specifically, the local reference block can be located only in the current SB or in an SB to the left of the current SB. An example current block predicted by another allowed coded block in the SBC is shown by Figure 23 the dashed arrow in. When the current SB or the left SB is used as the SBC reference, the reference sample update process in the RSM can follow the above reset process. For example, when any one of the 64x64 unit reference sample memories starts to be updated with the reconstructed samples from the current SB, the previously stored reference samples (from the left SB) in the entire 64x64 unit are marked as unavailable for generating IBC prediction samples and are gradually updated with the reconstructed samples of the current block.

[0192] Figure 23 shows 5 example states of the RSM during the local IBC decoding of the current SB in panel 2302. Similarly, in each example state, the region / partition of the RSM shaded horizontally stores the corresponding reference samples of the quadtree of the corresponding left adjacent SB, and the region / partition shaded vertically in gray stores the corresponding reference samples of the current SB. The coded blocks of the RSM shaded diagonally represent the current coded blocks within the current region being encoded / decoded. At the start of the encoding of each current SB, the RSM stores the samples of the previously encoded SB ( Figure 23 RSM state (0) of). When the current block is in one of the four 64x64 quadtree regions in the current SB, the corresponding region in the RSM is reset and used to store the samples of the current 64x64 encoding region. In this way, the samples in each 64x64 quadtree region of the RSM are gradually updated with the samples in the current SB (states (1) - state (3)). When the current SB has been fully encoded, the entire RSM is filled with all the samples of the current SB (state (4)).

[0193] Figure 23Each of the 64x64 regions in the panel 2302 is marked with a spatially encoded sequence number. Sequence numbers 0 - 3 represent the four 64x64 quadtree regions of the left neighbor SB, while sequence numbers 4 - 7 represent the four 64x64 quadtree regions of the current SB panel. In Figure 23 For the RSM states (1), (2), and (3) of the panel 2302 in Figure 23 , the panel 2304 further shows the corresponding spatial distribution of the reference samples in the left adjacent and current SBs in the 128x28 RSM. The shaded regions without cross - lines represent the regions in the RSM with reconstructed samples. The shaded regions with cross - lines represent the regions in the RSM where the reconstructed samples of the left SB are reset (and thus not available as reference samples for the local SBC).

[0194] The encoding order of the 64x64 regions and the corresponding RSM update order can follow a horizontal scan (as shown above in Figure 23 ) or a vertical scan. The horizontal scan starts from the upper - left, goes to the upper - right, lower - left, and lower - right. The vertical scan starts from the upper - left, goes to the lower - left, upper - right, and lower - left. Figure 24 The panels 2402 and 2404 in Figure 24 respectively show the left adjacent SB and current SB reference sample update processes for horizontal and vertical scans for comparison when reconstructing each of the four 64x64 regions of the current SB. In Figure 24 , the 64x64 regions shaded with horizontal lines without cross - lines represent the regions with samples available for the SBC. The regions shaded with horizontal lines with cross - lines represent the regions of the left adjacent SB that have been updated to the corresponding reconstructed samples of the current SB. The unshaded regions represent the unprocessed regions of the current SB. The diagonally shaded blocks represent the currently encoded blocks being processed.

[0195] As Figure 24 shown, depending on the position of the current encoded block relative to the current SB, the following restrictions on the reference blocks for IBC can be applied.

[0196] If the current block falls into the upper - left 64x64 region of the current SB, then in addition to the reconstructed samples in the current SB, the reference samples in the lower - right, lower - left, and upper - right 64x64 blocks of the left SB can also be referred to, as shown in Figure 24 2412 (for horizontal scan) and 2422 (for vertical scan) in

[0197] If the current block falls into the upper - right 64x64 block of the current SB, then in addition to the reconstructed samples in the current SB, if the luminance sample located at (0, 64) relative to the current SB has not been reconstructed, the current block can also refer to the reference samples in the lower - left 64x64 block and lower - right 64x64 block of the left SB ( Figure 24of 2414). Otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB for SBC( Figure 24 of 2426).

[0198] If the current block falls into the lower left 64x64 block of the current SB, then in addition to the already reconstructed samples in the current SB, if the luminance position (64, 0) relative to the current SB has not been reconstructed, the current block can also refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left SB( Figure 24 of 2424). Otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB for SBC( Figure 24 of 2416).

[0199] If the current block falls into the lower right 64x64 block of the current SB, then the current block can only refer to the already reconstructed samples in the current SB for SBC( Figure 24 of 2418 and 2428).

[0200] As described above, in some example embodiments, one or both of local and non-local CTU / SBs can be used for IBC reference block search and selection. In addition, when on-chip RSM is used for local reference, some restrictions on the availability of already constructed CTU / SBs as IBC references regarding write-back latency can be relaxed or removed. Such embodiments can be applied regardless of whether parallel decoding is employed.

[0201] Example embodiments of local and non-local reference CTU / SBs available for IBC are as Figure 25 shown, again, where each square represents a CTU / SB. The CTU / SB with diagonal shading (labeled "0") represents the current CTU / SB, while the CTU / SB with horizontal shading (labeled "1"), the CTU / SB with vertical shading (labeled "2"), and the CTU / SB with backslash shading represent the already constructed regions. The CTU / SB without shading represents the region that has not been reconstructed. Assume the use of parallel decoding similar to Figure 19 and Figure 20 . Due to the write-back latency to the DPB when only off-chip memory is used for SBC reference (see Figure 20 ), thus the CTU / SBs with vertical ("2") and backslash ("3") shading represent example regions that are generally restricted as SBC references for the current CTU / SB. When on-chip RSM is used, then Figure 20 one or more restricted regions can be directly referenced from the RSM, and thus may not need to be restricted out. The number of restricted regions now accessible through the RSM for IBC reference can depend on the size of the RSM. In Figure 25In the example, it is assumed that the RSM can store one CTU / SB and the above RSM update mechanism is adopted. Thus, Figure 20 one of the two sets of adjacent CTU / SBs with backslash shading (marked as "3") can be used for local reference. Then, the RSM stores samples from the left CTU / SB and the current CTU / SB. Therefore, in Figure 25 the example, the search area available for non-local SBC reference blocks includes the CTU / SB marked as "1" (search area 1, or SA1), the scan area available for local SBC reference blocks includes the CTU / SBs marked as "2" and "0" (SA2), and the restricted output area of the SBC reference block includes the CTU / SB marked as "3" due to write-back delay. In some other embodiments, if the on-chip RSM size is large enough to accommodate the entire restricted CTU / SB, then all these potential restricted areas can be included in the RSM for local reference.

[0202] Figure 26 Further shown is a further restriction on the reference coding blocks that can be used for predicting the current coding block in the IBC when both local and non-local reference searches are allowed and enabled. In Figure 26 , again, each square represents a CTU / SB. The CTU / SBs with horizontal shading represent the CTU / SBs that have been constructed. The CTU / SBs without shading represent the areas that have not been reconstructed yet. The CTU / SBs with backslash shading are the CTU / SBs that cannot be used as IBC references (only the current CTU / SB is shown to be allowed for IBC reference here. However, the basic principle applies to the case where only the first of the two CTU / SBs with backslash shading is not allowed, as shown in Figure 25 . The coded block with diagonal shading is the current coding block. Coded blocks A, B, and C are potential reference blocks for the IBC of the current coding block. The other shaded coded blocks in the current CTU / SB have been constructed. In this embodiment, since coded block B is completely outside the restricted area and in SA2 (local search area) and has been reconstructed, the reference coded block B is allowed. Since coded block C is completely outside the restricted area and in SA1 (non-local search area) and has been reconstructed, coded block C is also allowed. Since coded block A covers SA1 and SA2, coded block A is not allowed to be used as a prediction block. In other words, since the processing of the IBC for SA1 and SA2 may be different and not easily coordinated, the reference coded block covering both SA1 and SA2 may not be allowed.

[0203] Returning to the encoding of the block vector (BV) in IBC, in some example embodiments, processes similar to those specified for inter prediction may be employed, but simpler rules for BV prediction candidate list construction may be used. For example, the candidate list construction for some inter prediction embodiments may consist of five spatial, one temporal, and six history-based candidates. In this inter prediction, multiple candidate comparisons may be performed on the history-based candidates to avoid duplicate entries in the final candidate list. Additionally, the list construction may include candidates that are pair-averaged. In some example embodiments of BV prediction, the IBC list construction process may consider multiple (e.g., two) spatially adjacent BVs and multiple (e.g., five) history-based BVs (HBVPs), where only the first HBVP may be compared with the spatial candidates when added to the candidate list. While conventional inter prediction may use two different candidate lists, one for the merge mode and another for the regular mode, the candidate list in IBC may be used for both cases regarding BVs. However, the merge mode may use at most six candidates in the list, while the regular mode only uses the first two candidates. In some example embodiments, the block vector difference (BVD) encoding may employ the motion vector difference (MVD) process, resulting in a final BV of any magnitude. The reconstructed BV may point to a region outside the reference sample area, and correction may be required by removing the absolute offset in each direction using modulo arithmetic on the width and height of the RSM.

[0204] In the above embodiments using one or both of local and non-local IBC references, a loop filter may be used in certain cases. For example, when using a non-local-based IBC search range (with or without a local-based IBC search range), for example for a picture, the loop filter may be disabled in IBC for the same picture. On the other hand, if only a local-based IBC search range is used (without using a non-local-based IBC search range), the loop filter may be used for the same picture. The loop filter may include, but is not limited to, a deblocking filter, a Constrained Directional Enhancement Filter (CDEF), a Sample Adaptive Offset (SAO) filter, a Cross-Component Sample Offset (CCSO) filter, and a Loop Restoration filter (LR). In this way, a second picture buffer dedicated to enabling IBC may be avoided.

[0205] Now, turning back to the IBC-related signaling, in some embodiments, for the current block, first, a flag indicating whether IBC is enabled for the current block is sent in the bitstream. This flag can be signaled at a higher level such as CTU, CU, sequence, slice, or picture level. Then, if the current block is in the IBC mode (either as a mode separate from the inter-prediction mode or as an integral part of the inter-prediction mode), reference blocks can be searched, and the corresponding BV can be determined by the encoder. For BV prediction, the BV difference can be derived in the decoder by subtracting the predicted BV from the current BV, and then the BV difference can be classified into multiple types (e.g., 4 types) according to the horizontal and vertical components of the BV difference. The BV different type information can be further signaled in the bitstream, and then the BV difference values of the two (horizontal and vertical) components can be signaled. In some example embodiments, a set of high-level syntax flags is further included in the bitstream and is used to indicate the allowable local and / or non-local reference ranges for IBC prediction. This flag set can be signaled at various levels (e.g., CTU, CU, sequence, slice, or picture level).

[0206] For example, a syntax flag called global_ibc_flag can be used to turn on / off the non-local-based region, while another syntax flag called local_ibc_flag can be used to turn on / off the local-based region for IBC prediction. These two syntax flags can be controlled independently of each other. In other words, these flags can have any combination of flag values. Each of these flags can be signaled at the same different levels. In one example, when both flags are turned off, in fact, IBC is disabled. In this case, where the non-local IBC flag and the local IBC flag are signaled independently for a specific level, the IBC enable flag at that level (e.g., picture level or sequence level) does not need to be signaled in the bitstream.

[0207] In some example embodiments, the non-local IBC syntax flag global_ibc_flag and the local IBC syntax flag local_ibc_flag can be configured to have certain dependencies. For example, the non-local global_ibc_flag can be signaled first. Depending on the value of the local_ibc_flag, the local_ibc_flag can be signaled or inferred. When the global_ibc_flag is equal to 0 (indicating not in use), in combination with the IBC being signaled as being in use (e.g., the above advanced IBC enable syntax), the local_ibc_flag can be inferred to be 1 (indicating in use) instead of being signaled. In this example, one or both of the local and non-local flags are only signaled when the IBC enable flag is on. Otherwise, neither flag needs to be signaled.

[0208] In some example embodiments, when using a non-local-based IBC search range, e.g., for a picture, the loop filter will be disabled for the same picture. On the other hand, if only a local-based IBC search range is used (but not a non-local-based IBC search range), the loop filter can be used for the same picture. Thus, under the condition of using IBC without using a non-local-based IBC search range, the loop filter enable flag in the IBC is signaled. In other words, if the above other flags indicate not using a non-local-based IBC, the loop filter enable flag can be signaled. The loop filter enable flag will indicate whether the local-based IBC should call the loop filter. Otherwise, when not using IBC or only using non-local IBC, it is inferred that loop filtering is disabled and the loop filter enable flag does not need to be signaled. Specifically, the condition for signaling the use of the loop filter can be the value of the global_ibc_flag. When the global_ibc_flag is enabled (on) (meaning using non-local IBC reference search), the loop filter enable flag for the picture can be inferred to be 0 (or off) and does not need to be signaled.

[0209] In the manner described above, the various flags or syntax elements above can indicate or signal the IBC reference mode of the current block, either individually or in various combinations. The IBC reference mode indicates how the IBC prediction block can access the local search area and the non-local search area. For example, a combination of these flags or syntax elements can indicate that only CTUs or SBs in the local search area are available for IBC reference, thus being the local reference IBC mode. As another example, a combination of these flags or syntax elements can indicate that only CTUs or SBs in the non-local search area are available for IBC reference, thus being the non-local IBC reference mode. As yet another example, a combination of these flags or syntax elements can indicate that CTUs or SBs in both the local and non-local search areas are available for IBC reference, thus being the local and non-local reference IBC mode. The decoder can extract these syntax elements independently or based on the dependencies described above to determine the IBC reference mode, thereby obtaining information about the search area for determining the IBC reference block.

[0210] Figure 27 FIG. 2700 is a flowchart of an example method that follows the basic principles of the IBC implementation described above. The example method flow starts at 2701. At 2710, at least one syntax element is extracted from the video stream, and the at least one syntax element is associated with intra block copy (IBC) prediction of a video block. At 2720, an IBC reference mode for IBC prediction of the video block is determined, and the IBC reference mode can include one of the following: no IBC mode, local reference IBC mode, non-local reference IBC mode, and local and non-local reference IBC mode. At 2730, reconstructed samples of the video block are generated from the video stream based on the IBC reference mode. The example method flow ends at 2799.

[0211] In the embodiments and implementations of the present disclosure, any steps and / or operations can be combined or arranged in any quantity or order as needed. Two or more of the steps and / or operations can be executed in parallel. The embodiments and implementations of the present disclosure can be used alone or in any order combination. Additionally, each method (or implementation), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term "block" can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term "block" herein can also be used to refer to a transform block. In the following items, when referring to the block size, it can refer to the width or height of the block, or the maximum of the width and height, or the minimum of the width and height, or the area size of the block (width * height), or the aspect ratio (width: height, or height: width).

[0212] The above technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 28 FIG. shows a computer system (2800) suitable for implementing certain embodiments of the disclosed subject matter.

[0213] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[0214] The instructions can be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0215] Figure 28 The components of the computer system (2800) shown in are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (2800).

[0216] The computer system (2800) can include certain human-machine interface input devices. Such human-machine interface input devices can respond to one or more human users through inputs such as the following: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices can also be used to capture certain media that are not necessarily directly related to human conscious inputs, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0217] The input human-machine interface devices can include one or more of the following (only one of each is shown): keyboard (2801), mouse (2802), touchpad (2803), touch screen (2810), data glove (not shown), joystick (2805), microphone (2806), scanner (2807), camera (2808).

[0218] A computer system (2800) may include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (such as haptic feedback of a touch screen (2810), a data glove (not shown), or a joystick (2805), but may also be haptic feedback devices that are not input devices), audio output devices (such as: speakers (2809), headphones (not depicted)), visual output devices (such as a screen (2810) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each screen having or not having touch screen input functionality, each screen having or not having haptic feedback functionality - some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0219] The computer system (2800) may also include human-accessible storage devices and their associated media, for example, including optical media such as a CD / DVD ROM / RW (2820) with media such as CD / DVD (2821), a thumb drive (2822), a removable hard disk drive or a solid-state drive (2823), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0220] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0221] The computer system (2800) may also include an interface (2854) to one or more communication networks (2855). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CAN bus, and so on. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (2800)) connected to certain general-purpose data ports or peripheral buses (2849); as described below, other network interfaces are typically integrated into the kernel of the computer system (2800) by connecting to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (2800) may communicate with other entities using any of these networks. Such communication may be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CAN bus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.

[0222] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface may be attached to the kernel (2840) of the computer system (2800).

[0223] The kernel (2840) may include one or more central processing units (CPUs) (2841), a graphics processing unit (GPU) (2842), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2843), a hardware accelerator (2844) for certain tasks, a graphics adapter (2850), etc. These devices, as well as a read-only memory (ROM) (2845), a random access memory (2846), and an internal mass storage such as an internal non-user-accessible hard disk drive, SSD, etc. (2847) may be connected via a system bus (2848). In some computer systems, the system bus (2848) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus of the kernel (2848) or connected to the system bus of the kernel (1848) via a peripheral bus (2849). In one example, a screen (2810) may be connected to the graphics adapter (2850). The architecture of the peripheral bus includes PCI, USB, etc.

[0224] The CPU (2841), GPU (2842), FPGA (2843), and accelerator (2844) can execute certain instructions, which can be combined to form the above computer code. The computer code can be stored in the ROM (2845) or RAM (2846). Transitional data can also be stored in the RAM (2846), while permanent data can be stored, for example, in the internal mass storage (2847). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more of the following: one or more CPUs (2841), GPUs (2842), mass storage (2847), ROM (2845), RAM (2846), etc.

[0225] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the type well-known and available to those skilled in the field of computer software.

[0226] As a non-limiting example, a computer system having an architecture (2800), particularly a core (2840), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, and certain non-transitory memories of the core (2840), such as the core internal mass storage (2847) or ROM (2845). The software implementing the embodiments of this disclosure can be stored in such devices and executed by the core (2840). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (2840), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM (2846) and modifying such data structures according to the processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hard-wired or otherwise embodied in a circuit (e.g., accelerator (2844)), which can replace the software or operate together with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or including both. This disclosure encompasses any suitable combination of hardware and software.

[0227] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

[0228] Appendix A: Acronyms

[0229] JEM: Joint Exploration Model

[0230] VVC: Versatile Video Coding

[0231] BMS: Benchmark Set

[0232] MV: Motion Vector

[0233] HEVC: High Efficiency Video Coding

[0234] SEI: Supplemental Enhancement Information

[0235] VUI: Video Usability Information

[0236] GOP: Group of Pictures

[0237] TU: Transform Unit

[0238] PU: Prediction Unit

[0239] CTU: Coding Tree Unit

[0240] CTB: Coding Tree Block

[0241] PB: Prediction Block

[0242] HRD: Hypothetical Reference Decoder

[0243] SNR: Signal-to-Noise Ratio

[0244] CPU: Central Processing Unit

[0245] GPU: Graphics Processing Unit

[0246] CRT: Cathode Ray Tube

[0247] LCD: Liquid Crystal Display

[0248] OLED: Organic Light-Emitting Diode

[0249] CD: Compact Disc

[0250] DVD: Digital Video Disc

[0251] ROM: Read-Only Memory

[0252] RAM: Random Access Memory

[0253] ASIC: Application Specific Integrated Circuit

[0254] PLD: Programmable Logic Device

[0255] LAN: Local Area Network

[0256] GSM: Global System for Mobile Communications

[0257] LTE: Long Term Evolution

[0258] CANBus: Controller Area Network Bus

[0259] USB: Universal Serial Bus

[0260] PCI: Peripheral Component Interconnect

[0261] FPGA: Field Programmable Gate Array

[0262] SSD: Solid State Drive

[0263] IC: Integrated Circuit

[0264] HDR: High Dynamic Range

[0265] SDR: Standard Dynamic Range

[0266] JVET: Joint Video Exploration Team

[0267] MPM: Most Probable Mode

[0268] WAIP: Wide Angle Intra Prediction

[0269] CU: Coding Unit

[0270] PU: Prediction Unit

[0271] TU: Transform Unit

[0272] CTU: Coding Tree Unit

[0273] PDPC: Position Dependent Prediction Combination

[0274] ISP: Intra Sub Partitions

[0275] SPS: Sequence Parameter Set

[0276] PPS: Picture Parameter Set

[0277] APS: Adaptive Parameter Set

[0278] VPS: Video Parameter Set

[0279] DPS: Decoding Parameter Set

[0280] ALF: Adaptive Loop Filter

[0281] SAO: Sample Adaptive Offset

[0282] CC-ALF: Cross-Component Adaptive Loop Filter

[0283] CDEF: Constrained Directional Enhancement Filter

[0284] CCSO: Cross-Component Sample Offset

[0285] LSO: Local Sample Offset

[0286] LR: Loop Restoration Filter

[0287] AV1: Alliance for Open Media Video 1

[0288] AV2: Alliance for Open Media Video 2

[0289] RPS: Reference Picture Set

[0290] DPB: Decoded Picture Buffer

[0291] MMVD: Merge Mode with Motion Vector Difference

[0292] IntraBC or IBC: Intra Block Copy

[0293] BV: Block Vector

[0294] BVD: Block Vector Difference

[0295] RSM: Reference Sample Memory

Claims

1. A method for reconstructing a video block in a video stream, characterized in that, comprising: receiving the video stream; extracting at least one syntax element from the video stream, the at least one syntax element being associated with intra-block copy (IBC) prediction of the video block; based on the at least one syntax element, determining an IBC reference mode for IBC prediction of the video block, the at least one syntax element indicating that the IBC reference mode is one of a plurality of predefined IBC reference modes, the video block belonging to a current IBC prediction unit including a plurality of video blocks, the plurality of predefined IBC reference modes including: no IBC mode, local reference IBC mode, non-local reference IBC mode, and local and non-local reference IBC mode, wherein, for the local reference IBC mode, the reference block for IBC prediction of the video block includes reference samples in a predefined adjacent unit group of the current IBC prediction unit or a reconstructed video block in the current IBC prediction unit, for the non-local reference IBC mode, the reference block for IBC prediction of the video block includes reference samples not adjacent to the current IBC prediction unit in the coding direction of the current IBC prediction unit, and for the local and non-local reference IBC mode, the reference block for IBC prediction of the video block includes the reference samples in the adjacent unit group and the non-adjacent reference samples; and generating a reconstructed sample of the video block from the video stream based on the IBC reference mode.

2. The method according to claim 1, characterized in that, the predefined adjacent unit group includes a single left adjacent unit of the current IBC prediction unit.

3. The method according to claim 1, characterized in that, for the local reference IBC mode, the reference samples for IBC prediction are maintained in a on-chip reference sample memory (RSM) of a fixed size.

4. The method according to claim 3, characterized in that, the fixed size of the RSM corresponds to the size of one IBC prediction unit.

5. The method according to claim 4, characterized in that, a first part of the RSM includes corresponding samples of the reconstructed video blocks in the current IBC prediction unit; and a second part of the RSM includes corresponding reconstructed samples from the predefined adjacent unit group.

6. The method according to claim 5, characterized in that, further comprising: replacing the reconstructed samples of the adjacent units in the RSM corresponding to the video block in the current IBC prediction unit with the reconstructed samples of the video block.

7. The method according to claim 3, characterized in that, the current IBC prediction unit is divided into a predefined partition group; the video block is the first coded block to be reconstructed in the current partition of the predefined partition group; and and the method further comprises: before reconstructing the video block, resetting the partition of the RSM corresponding to the current partition to be unavailable for IBC reference.

8. The method according to claim 1, characterized in that, The at least one syntax element includes a first flag and a second flag, where the first flag is used to indicate that local IBC reference is enabled when set, and the second flag is used to indicate that non-local reference IBC is enabled when set.

9. The method according to claim 8, wherein, further comprising: In response to the first flag being set and the second flag not being set, determining that the IBC reference mode is the local reference IBC mode; In response to the second flag being set and the first flag not being set, determining that the IBC reference mode is the non-local reference IBC mode; In response to both the first flag and the second flag being set, determining that the IBC reference mode is the local and non-local reference IBC mode; and In response to both the first flag and the second flag not being set, determining that the IBC reference mode is the no-IBC mode.

10. The method according to claim 8, wherein, The first flag and the second flag are signaled in the video stream at the coded block level, coded unit level, coding tree unit level, slice level, picture level, or sequence level.

11. The method according to claim 1, wherein, The at least one syntax element includes a first flag for indicating whether IBC is used for the video block.

12. The method according to claim 11, wherein, further comprising: In response to the first flag indicating that IBC is not used for the video block, determining that the IBC reference mode is the no-IBC mode.

13. The method according to claim 12, wherein, further comprising: In response to the first flag indicating that IBC is used for the video block, further extracting a second flag for indicating whether non-local IBC reference is used as part of the at least one syntax element; In response to the second flag indicating that non-local IBC reference is not used, inferring that the IBC reference mode of the video block is the local reference IBC mode.

14. The method according to claim 13, wherein, further comprising: In response to the second flag indicating that non-local IBC reference is used, further extracting a third flag for indicating whether local IBC reference is used as part of the at least one syntax element; In response to the third flag indicating that local IBC reference is used, determining that the IBC reference mode is the local and non-local reference IBC mode; and In response to the third flag indicating that local IBC reference is not used, determining that the IBC reference mode is the non-local reference IBC mode.

15. The method according to claim 1, wherein, When the IBC reference mode is the local reference IBC mode, enabling the loop filtering process; and When the IBC reference mode is the non-local reference IBC mode or the local and non-local reference IBC mode, disabling the loop filtering process.

16. The method according to claim 15, wherein, Derive whether the loop filtering process is enabled from the at least one syntax element for signaling the IBC reference mode.

17. A video processing device for reconstructing a video block in a video stream, characterized in that it includes a memory for storing computer instructions and a processor for executing the computer instructions to perform the method according to any one of claims 1 to 16.

18. A non-transitory computer-readable medium, characterized in that it is for storing instructions which, when executed by a computer for video decoding, cause the computer to perform the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method and apparatus for video coding

    CN113273201A