Method for encoding video block, method for storing video stream, device and storage medium
By employing frame-intra block copy prediction modes, the method addresses inefficiencies in existing video coding technologies, enhancing compression efficiency and reducing storage requirements while maintaining video quality.
Patent Information
- Application Number
- CN202510739420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-25
- Filing Date
- 2022-04-13
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video encoding technology has limited directions in intra prediction, resulting in low encoding efficiency, and the redundancy of motion vector prediction in inter prediction is not fully utilized, affecting compression efficiency.
Intra-block copy prediction (IBC) technology using multi-reference mode, the mode is identified when encoding the video block by searching for reference blocks in the current frame and determining the appropriate IBC reference mode, including local and non-local reference modes.
The encoding efficiency of intra prediction is improved, the bit requirements in the encoding direction are reduced, the motion vector prediction in inter prediction is optimized, and the video compression rate is improved.
Smart Images

Figure CN120321404A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with an application date of April 13, 2022, a Chinese patent application number of 202280006720.1, and an invention title of "Method, Device, and Storage Medium for Reconstructing Video Blocks in a Video Stream". Technical Field
[0002] This application generally relates to video encoding, video decoding, and specifically relates to an intra block copy coding mode. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the content of the embodiments of this application. To the extent that the work of the currently named inventors is described in this background art section, the work of the named inventors and aspects that were not prior art at the time of filing this application have never been explicitly or implicitly recognized as prior art for this application.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of pictures, each picture having a spatial size of, for example, 1920x1080 luminance samples and associated full or subsampled chrominance samples. The sequence of pictures can have a fixed or variable picture rate (or frame rate), such as 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920x1080, a frame rate of 60 frames per second, and a 4:2:0 chrominance subsampling of 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.
[0005] One purpose of video encoding and decoding is to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the above bandwidth or storage space requirements, and in some cases can reduce them by two orders of magnitude or more than two orders of magnitude. Lossless compression and lossy compression, and combinations thereof, can be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal through a decoding process. Lossy compression refers to an encoding / decoding process where the original video information cannot be fully retained during encoding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may be different from the original signal. Although there is some information loss, the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, lossy compression is widely used in many applications. The amount of distortion that can be tolerated depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally enables the encoding algorithm to produce higher losses and higher compression ratios.
[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0007] Video codec technology can include techniques referred to as intra-frame encoding. In intra-frame encoding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, the picture can be referred to as an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and can thus be used as the first picture in an encoded video bitstream and video session, or as a still image. Then the samples of the blocks after intra-frame prediction can be transformed to a domain, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction represents techniques for minimizing sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after transformation, the fewer bits are required to represent the block after entropy coding for a given quantization step.
[0008] Traditional intra - frame coding (e.g., intra - frame coding known from, for example, MPEG - 2 generation coding techniques) does not use intra - frame prediction. However, some newer video compression techniques include techniques that attempt block encoding / decoding based on, for example, surrounding sample data and / or metadata that are obtained during spatially adjacent encoding / decoding and whose decoding order is prior to the data blocks being intra - frame encoded or decoded. Such techniques are hereafter referred to as "intra - frame prediction" techniques. It should be noted that, in at least some cases, intra - frame prediction uses only reference data from the currently being reconstructed picture and not reference data from other reference pictures.
[0009] There can be many different forms of intra - frame prediction. When more than one such technique is available in a given video coding technique, the techniques used can be referred to as intra - frame prediction modes. One or more intra - frame prediction modes can be provided in a particular codec. In some cases, a mode can have sub - modes and / or can be associated with various parameters, and the mode / sub - mode information and the intra - frame coding parameters for a video block can be encoded either separately or jointly included in a mode codeword. Which codeword is used for a given mode / sub - mode / parameter combination can have an impact on the coding efficiency gain through intra - frame prediction, and entropy coding techniques can also be used to convert the codewords into a bitstream.
[0010] Certain intra - frame prediction modes were introduced with H.264, improved in H.265, and further improved in newer coding techniques (e.g., Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS)). Generally, for intra - frame prediction, the available adjacent sample values that are already available can be used to form a predictor block. For example, the available values of a particular set of adjacent samples along certain directions and / or lines can be copied into the predictor block. The reference to the direction in use can be encoded in the bitstream or can itself be predicted.
[0011] Reference Figure 1A , a subset of 9 predictor directions out of the 33 possible intra - frame predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra - frame modes specified in H.265) is depicted in the lower right. The point (101) where the arrows converge represents the sample being predicted. The arrows represent the directions in which adjacent samples are used to predict the sample at 101. For example, arrow (102) represents predicting sample (101) from one or more adjacent samples to the upper right at a 45 - degree angle to the horizontal direction. Similarly, arrow (103) represents predicting sample (101) from one or more adjacent samples to the lower left of sample (101) at a 22.5 - degree angle to the horizontal direction.
[0012] Still referring to Figure 1A, a square block (104) of 4x4 samples is depicted in the upper left (represented by a dashed thick line). The square block (104) includes 16 samples, and each sample is marked with "S" for its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension in the block (104). Since the size of the block is 4x4 samples, S44 is located in the lower right. Also shown are example reference samples following a similar numbering scheme. The reference samples are marked with "R" for their Y position (e.g., row index) and X position (column index) relative to the block (104). In H.264 and H.265, prediction samples adjacent to the block being reconstructed are used.
[0013] Intra picture prediction of block 104 can start by copying reference sample values from adjacent samples according to a signalized prediction direction. For example, assume that the encoded video bitstream includes signaling that, for this block 104, indicates the prediction direction of arrow (102) - that is, to predict the sample direction from one or more prediction samples to the upper right at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.
[0014] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a reference sample; especially when the direction cannot be evenly divisible by 45 degrees.
[0015] As video coding technology continues to develop, the number of possible directions has increased. For example, in H.264 (in 2003), nine different directions were available for intra prediction. In H.265 (in 2013), it increased to 33 directions, and at the time of this application, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, thus accepting a certain bit penalty for the directions. Additionally, sometimes the direction itself can be predicted from adjacent directions used in the intra prediction of already decoded adjacent blocks.
[0016] Figure 1B A schematic diagram (180) depicting 65 intra prediction directions according to JEM is shown to illustrate the increase in the number of prediction directions in each coding technology over time.
[0017] The mapping of bits representing an intra prediction direction to a prediction direction in an encoded video bitstream can vary depending on the video coding technology; and the range can be, for example, from a simple direct mapping of the prediction direction to an intra prediction mode, to codewords, to complex adaptive schemes involving most probable modes, and similar techniques. However, in all cases, there may be certain intra prediction directions that are statistically less likely to occur in the video content compared to some other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technology, those less likely directions can be represented by more bits than the more likely directions.
[0018] Inter-picture prediction or inter prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used to predict a newly reconstructed picture or picture portion (e.g., block) after being spatially offset along a direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture being used (similar to a temporal dimension).
[0019] In some video compression techniques, the current MV applicable to a certain region of sample data can be predicted based on other MVs, for example, based on MVs that are spatially adjacent to other regions of sample data and whose decoding order is prior to the current MV. Doing so can significantly reduce the overall amount of data required to encode the MVs by relying on eliminating redundancy in the relevant MVs, thereby improving the compression ratio. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is the following statistical likelihood: a region larger than the region applicable to a single MV moves in a similar direction in the video sequence, and thus, in some cases, similar motion vectors derived from MVs of adjacent regions can be used to predict that larger region. This results in the actual MV of a given region being similar to or the same as the MV predicted from surrounding MVs. After entropy coding, such an MV can in turn be represented by fewer bits than the number of bits used if the MV were directly encoded instead of being predicted from adjacent MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from an original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). In addition to the various MV prediction mechanisms specified in H.265, the technique hereinafter referred to as "spatial merge" is described below.
[0021] Specifically, still referring to Figure 2 , in spatial merge, the current block (201) includes samples that have been discovered by the encoder during the motion search process and that are predictable based on a previous block of the same size that has been spatially shifted. An MV associated with any of five surrounding samples (denoted as A0, A1, and B0, B1, B2 (from 202 to 206, respectively)) can be used to derive the MV from metadata associated with one or more reference pictures, for example, from the nearest (in decoding order) reference picture, rather than encoding the MV directly. In H.265, MV prediction can use the predictors of the same reference pictures used by adjacent blocks. Summary of the Invention
[0022] Embodiments of the present application provide a method for encoding a video block, a method for storing a video stream, a device, and a storage medium.
[0023] A method for encoding a video block provided by an embodiment of the present application, the method comprising:
[0024] Receiving the video block;
[0025] Searching for a reference block for IBC prediction of the video block from a current video frame;
[0026] Determining a position of the reference block relative to the video block, wherein types of the reference block include: a local reference IBC mode, a non-local reference IBC mode;
[0027] Based on the position of the reference block, determining an IBC reference mode for IBC prediction of the video block, wherein the IBC reference mode is one selected from a plurality of predefined IBC reference modes, the video block belongs to a current IBC prediction unit including a plurality of video blocks, and the plurality of predefined IBC reference modes include:
[0028] The IBC modes include a non-IBC mode, a local reference IBC mode, a non-local reference IBC mode, and a local and non-local reference IBC mode. In the local reference IBC mode, the reference blocks for IBC prediction of the video block include reference samples in a predefined adjacent unit group of the current IBC prediction unit or video blocks that have been reconstructed in the current IBC prediction unit. In the non-local reference IBC mode, the reference blocks for IBC prediction of the video block include reference samples that are not adjacent to the current IBC prediction unit in the encoding direction of the current IBC prediction unit. The local and non-local reference IBC mode is used to indicate that the reference blocks for IBC prediction of the video block include the reference samples in the adjacent unit group and the non-adjacent reference samples; and
[0029] Encode the video block in the video stream based on the reference block, and use the identification of the IBC reference mode as an identification in at least one syntax element of the video stream.
[0030] In some embodiments, the embodiments of the present application further provide a video processing device for encoding a video block, including a memory for storing computer instructions and a processor, and the processor is configured to execute the computer instructions to perform the method described in the embodiments of the present application.
[0031] In some embodiments, the embodiments of the present application further provide a non-transitory computer-readable medium for storing instructions, which, when executed by a computer for video encoding, cause the computer to perform the method described in the embodiments of the present application.
[0032] In some embodiments, the embodiments of the present application further provide a method for storing a video stream, where the video stream is obtained by encoding through the method described in the embodiments of the present application. Description of the Drawings
[0033] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0034] Figure 1A A schematic diagram showing an example subset of intra prediction direction modes;
[0035] Figure 1B A schematic diagram showing an exemplary intra prediction direction;
[0036] Figure 2 A schematic diagram showing a current block in an example and its surrounding spatial merge candidates for motion vector prediction;
[0037] Figure 3 A schematic diagram showing a simplified block diagram of a communication system according to an example embodiment;
[0038] Figure 4 A schematic diagram showing a simplified block diagram of a communication system according to another exemplary embodiment;
[0039] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an exemplary embodiment;
[0040] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an exemplary embodiment;
[0041] Figure 7 A block diagram showing a video encoder according to another exemplary embodiment;
[0042] Figure 8 A block diagram showing a video decoder according to another exemplary embodiment;
[0043] Figure 9 A scheme for coding block segmentation according to an exemplary embodiment of the embodiment of the present application;
[0044] Figure 10 Another scheme for coding block segmentation according to an exemplary embodiment of the embodiment of the present application;
[0045] Figure 11 Another scheme for coding block segmentation according to an exemplary embodiment of the embodiment of the present application;
[0046] Figure 12 An example of dividing a base block into coding blocks according to an example segmentation scheme;
[0047] Figure 13 An example of a ternary segmentation scheme;
[0048] Figure 14 An example of a quadtree - binary tree coding block segmentation scheme;
[0049] Figure 15 A scheme for dividing a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the embodiment of the present application;
[0050] Figure 16 Another scheme for dividing a coding block into multiple transform blocks and the coding order of the transform blocks according to an exemplary embodiment of the embodiment of the present application;
[0051] Figure 17 Another scheme for dividing a coding block into multiple transform blocks according to an exemplary embodiment of the embodiment of the present application;
[0052] Figure 18 The concept of intra - block copy (IBC) for predicting a current coding block using a reconstructed coding block in the same frame;
[0053] Figure 19 Shows an example reconstructed sample that can be used as a reference sample for IBC;
[0054] Figure 20 Shows an example reconstructed sample that can be used as a reference sample for IBC with some example limitations;
[0055] Figure 21 Shows an example on - chip reference sample memory (RSM) update mechanism for IBC;
[0056] Figure 22 Shows Figure 21 A spatial view of an example on - chip RSM update mechanism;
[0057] Figure 23 Shows another example on - chip RSM update mechanism for IBC;
[0058] Figure 24 Shows a comparison of the spatial views of example RSM update mechanisms for IBC for horizontally - split superblocks and vertically - split superblocks;
[0059] Figure 25 Shows example non - local and local search regions for IBC reference blocks;
[0060] Figure 26 Shows exemplary limitations on the positions of reference blocks of IBC using local and non - local reference block search regions;
[0061] Figure 27 Shows a flowchart of a method according to an example embodiment of the present application; and
[0062] Figure 28 Shows a schematic diagram of a computer system according to an example embodiment of the present application. Detailed Description of the Invention
[0063] The present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, it can be understood that the present invention can be implemented in various different forms. Therefore, any embodiment set forth below is intended to explain rather than limit the subject matter covered or claimed. It can also be understood that the present invention can be embodied as a method, device, component, or system. Thus, embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.
[0064] Throughout the specification and claims, terms may have nuanced meanings that go beyond the explicitly stated meanings, implied or implicit in the context. The phrases "in one embodiment" or "in some embodiments" as used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" as used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" as used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" as used herein do not necessarily refer to different implementations. For example, it is intended that the claimed subject matter include combinations of all or part of the exemplary embodiments / implementations.
[0065] Generally speaking, terms can be understood at least in part from their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can include a variety of meanings, which can depend at least in part on the context in which these terms are used. Typically, "or" if used to relate a list such as A, B, or C is intended to mean A, B, and C (used herein in an inclusive sense) as well as A, B, or C (used herein in an exclusive sense). In addition, the terms "one or more" or "at least one" as used herein, depending at least in part on the context, can be used to describe any feature, structure, or characteristic in a singular sense or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as "a", "an", or "the" can also be understood to convey a singular usage or to convey a plural usage, which depends at least in part on the context. In addition, the terms "based on" or "determined by" can be understood to not necessarily intend to convey an exclusive set of factors, but can allow for the existence of other factors that are not necessarily explicitly described, which also depends at least in part on the context.
[0066] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present application is shown. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes pairs of terminal devices (310) and (320) interconnected via a network (350). In Figure 3In the example of, the first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (such as a video picture stream captured by the terminal device (310)) for transmission over the network (350) to another terminal device (320). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. The unidirectional data transmission can be implemented in a media service application or the like.
[0067] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (such as a video picture stream captured by the terminal device) for transmission over the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device based on the recovered video data.
[0068] In Figure 3 the example of, the terminal devices (310), (320), (330), and (340) can be implemented as servers, personal computers, and smart phones, but the applicability of the basic principles of the embodiments of the present application is not limited thereto. The embodiments of the present application can be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and / or similar devices. The network (350) represents any number or type of network that conveys encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless explicitly explained herein, the architecture and topology of the network (350) may be unimportant for the operation of the embodiments of the present application.
[0069] As an example of the application of the disclosed subject matter, Figure 4Shows the placement of a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0070] A video streaming system may include a video capture subsystem (413), which may include a video source (401) such as a digital camera, for creating an uncompressed video picture or image stream (402). In an example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402), depicted as a thick line to emphasize the high data volume, may be processed by an electronic device (420) compared to the encoded video data (404) (or encoded video bitstream), and the electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower data volume, may be stored on a streaming server (405) for future use or directly stored to a downstream video device (not shown) compared to the uncompressed video picture stream (402). One or more streaming client subsystems, for example, Figure 4 the client subsystem (406) and the client subsystem (408) in, may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an output video picture stream (411) that is uncompressed and can be presented on a display (412) (such as a display screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in embodiments of the present application. In some streaming systems, the encoded video data (404), the encoded video data (407), and the encoded video data (409) (such as video bitstreams) may be encoded according to certain video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC and other video coding standards.
[0071] It is understood that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may include a video encoder (not shown).
[0072] Figure 5 A block diagram of a video decoder (510) according to any embodiment of the following embodiments of the present application is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used to replace Figure 4 the video decoder (410) in the example of.
[0073] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source that transmits the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, that may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided outside and separated from the video decoder (510) (not depicted). Still in other applications, a buffer memory (not depicted) may be provided outside the video decoder (510) for purposes such as preventing network jitter, and another buffer memory (515) may be provided inside the video decoder (510) for purposes such as handling playback timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory may be made smaller. For use on a best-effort packet network (e.g., the Internet), a buffer memory (515) of sufficient size may be required, and its size may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be at least partially implemented in an operating system or similar element (not shown) outside the video decoder (510).
[0074] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530), but may be coupled to the electronic device (530), as Figure 5 shown. The control information for one (or more) display devices may be in the form of Supplemental Enhancement Information (SEI messages) or a parameter set segment (not depicted) of Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract a subgroup parameter set for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a strip, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, and so on.
[0075] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0076] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. The units involved and the way they are involved may be controlled by the subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and the multiple processing or functional units below are not depicted.
[0077] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.
[0078] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients as symbols (521) and control information, including information indicating which inverse transform type, block size, quantization factor / parameter, quantization scaling matrix, etc. to use, as symbols (521) from a parser (520). The scaler / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).
[0079] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block having the same size and shape as the block being reconstructed using the reconstructed surrounding block information and the block information stored in the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some embodiments, the aggregator (555) may add the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0080] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for picture inter-prediction. After motion compensating the extracted samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or a residual signal), thereby generating output sample information. The motion compensation prediction unit (553) obtaining the prediction samples from an address within the reference picture memory (557) may be controlled by a motion vector, and the motion vector is in the form of symbols (521) for use by the motion compensation prediction unit (553), and the symbols (521) may have, for example, X, Y components (shifts) and reference picture components (temporal). Motion compensation may also include interpolation of the sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with a motion vector prediction mechanism, etc.
[0081] The output samples of the aggregator (555) may be subject to various loop filtering techniques in the loop filter unit (556). The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video code bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). However, the video compression technique may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, which will be described in further detail below.
[0082] The output of the loop filter unit (556) may be a sample stream that may be output to the rendering device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.
[0083] Once fully reconstructed, some encoded pictures may be used as reference pictures for future picture inter-prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.
[0084] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted, for example, in the ITU-T H.265 recommendation standard. In the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence may conform to the syntax specified by the video compression technique or standard used. Specifically, the profile may select certain tools from all available tools in the video compression technique or standard as the only tools available under that profile. For compliance with the standard, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0085] In some example embodiments, the receiver (531) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to perform proper decoding of the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant strips, redundant pictures, forward error correction codes, etc.
[0086] Figure 6 A block diagram of a video encoder (603) according to an example embodiment of the present application is shown. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the example of
[0087] The video encoder (603) may receive video samples from a video source (601) (not Figure 6 part of the electronic device (620) in the example of
[0088] A video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, XYZ, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures or images, which are given motion when viewed in sequence. The pictures themselves can be constructed as a spatial pixel array, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0089] According to some example embodiments, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a controller (650). In some embodiments, as described below, the controller (650) can be functionally coupled to and control other functional units. For simplicity, the couplings are not depicted in the figure. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions that relate to optimizing the video encoder (603) for a certain system design.
[0090] In some example embodiments, the video encoder (603) may be configured to operate in an encoding loop. As a simple description, in the example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and one (or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). Even though the embedded decoder 633 processes the video stream encoded by the source encoder 630 without entropy coding, the decoder (633) reconstructs the symbols in a manner similar to the way a (remote) decoder would create to create sample data (because in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols in the entropy coding and the encoded video bitstream can be lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is used to improve the encoding quality.
[0091] The operation of the “local” decoder (633) may be the same as that of the “remote” decoder of the video decoder (510) described above in conjunction with Figure 5 However, briefly referring to Figure 5 further, when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) of the encoder.
[0092] At this point, it can be observed that any decoder technique other than parsing / entropy decoding that may only exist in the decoder must also exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter sometimes focuses on the decoder operation, which is related to the decoding part of the encoder. Since the encoder technique is reciprocal to the decoder technique described comprehensively, the description of the encoder technique can be simplified. A more detailed description is provided only in certain areas or aspects below.
[0093] During operation, in some example embodiments, the source encoder (630) may perform motion-compensated predictive coding, with reference to one or more previously encoded pictures designated as "reference pictures" in the video sequence, where the motion-compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (632) encodes the difference (or residue) in the color channels between a pixel block of the input picture and a pixel block of one (or more) reference pictures, which may be selected as the prediction reference for the input picture. The term "residue" and its adjective form "residual" may be used interchangeably.
[0094] The local video decoder (633) may decode the encoded video data of a picture that may be designated as a reference picture, based on the symbols created by the source encoder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the encoded video data is decoded at a video decoder (not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture, and may cause the reconstructed reference picture to be stored in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture, which has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by the remote video decoder. Figure 6 The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new picture. The predictor (635) may operate on a per-pixel-block basis of the sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0095] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0096] The outputs of all the above functional units may be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby transforming the symbols into an encoded video sequence.
[0097]
[0098] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over a communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0099] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may typically be assigned to any of the following picture types:
[0100] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0101] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and reference index to predict the sample values of each block.
[0102] A bi-predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0103] Source pictures can generally be spatially subdivided into multiple sample - coded blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples), and coded block - by - block. These blocks can be predictively coded with reference to other (already - coded) blocks, which are determined by the coding assignment of the corresponding picture applied to the block. For example, blocks of an I - picture can be non - predictively coded, or the block can be predictively coded with reference to already - coded blocks of the same picture (spatial prediction or intra - prediction). Pixel blocks of a P - picture can be predictively coded with reference to a previously - coded reference picture either through spatial prediction or through temporal - domain prediction. Blocks of a B - picture can be predictively coded with reference to one or two previously - coded reference pictures either through spatial prediction or through temporal prediction. For other purposes, source pictures or pictures during intermediate processing can be subdivided into other types of blocks. As described in further detail below, the subdivision of coded blocks and other types of blocks may or may not follow the same way.
[0104] The video encoder (603) can perform encoding operations according to a predetermined video - coding technique or standard such as, for example, the ITU - T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive - coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video - coding technique or standard being used.
[0105] In some example embodiments, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, etc., SEI messages, VUI parameter - set fragments, and the like.
[0106] The captured video can be multiple source pictures (video pictures) in a time series. Intra - picture prediction (often simplified to intra - prediction) exploits the spatial correlation within a given picture, while inter - picture prediction exploits the temporal or other correlations between pictures. For example, a particular picture being encoded / decoded can be segmented into blocks, and the particular picture being encoded / decoded is called the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded with a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0107] In some example embodiments, bidirectional prediction techniques can be used for inter-picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be jointly predicted by a combination of the first reference block and the second reference block.
[0108] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0109] According to some example embodiments of the embodiments of the present application, predictions such as inter-picture prediction and intra-picture prediction are performed on a per-block basis. For example, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU can include three coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Each CTU can be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64x64 pixel CTU can be split into a 64x64 pixel CU, or 4 32x32 pixel CUs. Each of one or more of the 32x32 blocks can be further split into 4 16x16 pixel CUs. In some example embodiments, each CU can be analyzed during encoding to determine a prediction type for the CU in its respective prediction type, such as an inter prediction type or an intra prediction type. Depending on temporal and / or spatial predictability, a CU can be split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, prediction operations in encoding (encoding / decoding) are performed on a per-prediction block basis. Splitting a CU into PUs (or PBs of different color channels) can be performed in various spatial patterns. For example, a luminance or chrominance PB can include a matrix of values (e.g., luminance values) of samples such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 samples, etc.
[0110] Figure 7 A diagram of a video encoder (703) according to another example embodiment of the embodiments of the present application is shown. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence, and encode the processing block into an encoded picture that is part of an encoded video sequence. An example video encoder (703) can be used to replaceFigure 4 the video encoder (403) in the example of
[0111] For example, the video encoder (703) receives a matrix of sample values for processing a block, such as a prediction block of 8x8 samples. Then, the video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to best encode the processing block. When determining the processing block to be encoded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into an encoded picture; and when determining the processing block to be encoded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction or bi-prediction techniques to encode the processing block into an encoded picture, respectively. In some example embodiments, the merge mode can be used as a sub-mode of inter-picture prediction, where a motion vector is derived from one or more motion vector predictors without relying on encoded motion vector components external to the predictor. In some other example embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) may include Figure 7 components (such as a mode decision module) not explicitly shown in
[0112] In Figure 7 the example of Figure 7 the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the example settings of
[0113] The inter encoder (730) is configured to receive samples of a current block (such as a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundant information, a motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (such as a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information using a decoding unit 633 (described in detail below and shown as Figure 6 the residual decoder 728 of Figure 7 the example encoder 620 of
[0114] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) may calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0115] The general controller (721) may be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an example, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is an intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the prediction mode of the block is an inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and add the inter prediction information to the bitstream.
[0116] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and a prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo quantization processing to obtain quantized transform coefficients. In various example embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture may be buffered in a memory circuit (not shown) and used as a reference picture.
[0117] The entropy encoder (725) can be configured to format the bitstream to produce an encoded block. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) can be configured to include general control data, selected prediction information (such as intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in the merge submode of the inter mode or the bi - directional prediction mode, there may be no residual information.
[0118] Figure 8 FIG. shows an example video decoder (810) according to another embodiment of an embodiment of the present application. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (810) can be used instead of Figure 4 the video decoder (410) in the example of
[0119] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter - frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra - frame decoder (872) coupled together as shown in the example of
[0120] The entropy decoder (871) can be configured to reconstruct certain symbols from the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode used to encode the block (such as the intra mode, the inter mode, the bi - directional prediction mode, the merge submode, or another submode), prediction information (such as intra prediction information or inter prediction information) that can identify certain samples or metadata for the intra - frame decoder (872) or the inter - frame decoder (880) to use for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In the example, when the prediction mode is the inter - frame or bi - directional prediction mode, the inter - frame prediction information is provided to the inter - frame decoder (880); and when the prediction type is the intra - frame prediction type, the intra - frame prediction information is provided to the intra - frame decoder (872). The residual information can be inverse - quantized and provided to the residual decoder (873).
[0121] The inter - frame decoder (880) can be configured to receive the inter - frame prediction information and generate an inter - frame prediction result based on the inter - frame prediction information.
[0122] The intra - frame decoder (872) can be configured to receive the intra - frame prediction information and generate a prediction result based on the intra - frame prediction information.
[0123] The residual decoder (873) may be configured to perform inverse quantization to extract the dequantized transform coefficients, and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (for including quantization parameter (QP)), which may be provided by the entropy decoder (871) (the data path is not depicted as this is merely low data volume control information).
[0124] The reconstruction module (874) may be configured to combine, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module depending on the situation) to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture as part of the reconstructed video. It may be noted that other suitable operations such as deblocking operations may also be performed to improve the visual quality.
[0125] It may be noted that any suitable technology may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In some example embodiments, one or more integrated circuits may be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810).
[0126] Returning to block partitioning for encoding and decoding, the general partitioning can start from a base block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After partitioning or splitting the base block according to any of the example partitioning processes described below or other processes or combinations thereof, the final partitions or groups of coded blocks can be obtained. Each of these partitions can be at one of various partition levels in the partitioning hierarchy and can be of various shapes. Each partition can be referred to as a coded block (CB). For the various example partitioning embodiments described further below, each resulting CB can be of any allowed size and partitioning level. Since such partitioning can form units for which some basic encoding / decoding decisions can be made and the encoding / decoding parameters can be optimized, determined, and signaled in the coded video bitstream, such partitions are referred to as coded blocks. The highest or deepest level in the final partitions represents the depth of the coded block partitioning structure of the tree. The coded blocks can be luminance coded blocks or chrominance coded blocks. The CB tree structure for each color can be referred to as a coded block tree (CBT).
[0127] The coded blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchical structures for all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures for the various color channels in a CTU can be the same or different.
[0128] In some embodiments, the partitioning tree schemes or structures for the luminance and chrominance channels may not need to be the same. In other words, the luminance and chrominance channels can have separate coding tree structures or patterns. Additionally, whether the luminance and chrominance channels use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being coded is a P, B, or I slice. For example, for an I slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P or B slice, the luminance and chrominance channels can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be split into CBs by one coding partitioning tree structure and the chrominance channel can be split into chrominance CBs by another coding partitioning tree structure.
[0129] In some example embodiments, a predefined partitioning pattern can be applied to the base block. As Figure 9As shown, an exemplary 4-way split tree can start from a first predefined level (e.g., 64x64 block level or other size, as the base block size), and the base block can be hierarchically split down to a predefined lowest level (e.g., 4x4 level). For example, the base block can be subject to four predefined split options or patterns indicated by 902, 904, 906, and 908, where the partition designated as R is allowed to be recursively split, i.e., the same split option as indicated in Figure 9 can be repeated at a lower scale until the lowest level (e.g., 4x4 level). In some embodiments, additional restrictions can be applied to the Figure 9 split scheme. In the Figure 9 embodiment, rectangular splits (e.g., 1:2 / 2:1 rectangular splits) can be allowed, but they can be made non-recursive, while square splits are allowed to be recursive. If needed, the split according to Figure 9 and recursive generation of the final coded block group can be performed. The coding tree depth can be further defined to indicate the split depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64x64 block) can be set to 0, and after the root block is further split once according to Figure 9 , the coding tree depth is increased by 1. For the above scheme, the maximum or deepest level from the 64x64 base block to the 4x4 minimum partition will be 4 (starting from level 0). This split scheme can be applied to one or more color channels. Each color channel can be split independently according to the Figure 9 scheme (e.g., for each color channel at each hierarchical level, the split pattern or option in the predefined pattern can be determined independently). Optionally, two or more color channels can share the Figure 9 same hierarchical pattern tree (e.g., the same split pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).
[0130] Figure 10 FIG. shows another example of a predefined split pattern that enables recursive splitting to form a split tree. As shown in Figure 10 , an exemplary 10-way split structure or pattern can be predefined. The root block can start at a predefined level (e.g., from a base block of 128x128 level or 64x64 level). Figure 10 The exemplary split structure includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular splits. Figure 10 The split type with 3 sub-partitions indicated by 1002, 1004, 1006, and 1008 in the second row of Figure 10Any rectangular partition in the rectangular partition. The coding tree depth can be further defined to indicate the splitting depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128x128 block) can be set to 0, and after the root block is further split once according to Figure 10 the coding tree depth increases by 1. In some embodiments, only the full square partitions in 1010 are allowed to be recursively split into the next level of the split tree according to Figure 10 the pattern. In other words, for the square partitions within the T-shaped patterns 1002, 1004, 1006, and 1008, recursive splitting may not be allowed. If needed, perform the splitting process and recursively generate the final coded block group according to Figure 10 This splitting scheme can be applied to one or more color channels. In some embodiments, more flexibility can be added when using splits below the 8x8 level. For example, 2x2 chrominance inter prediction can be used in some cases.
[0131] In some other example embodiments for coding block splitting, a quadtree structure can be used to split a base block or an intermediate block into quadtree partitions. This quadtree splitting can be applied hierarchically and recursively to any square partition. Whether to further perform quadtree splitting on the base block or intermediate block or partition can be adjusted according to various local characteristics of the base block or intermediate block / partition. The quadtree splitting at the picture boundary can be further adjusted. For example, implicit quadtree splitting can be performed at the picture boundary so that the blocks will remain quadtree split until the size fits the picture boundary.
[0132] In some other example embodiments, hierarchical binary splitting from the base block can be used. For such a scheme, the base block or intermediate-level block can be split into two partitions. The binary splitting can be horizontal or vertical. For example, horizontal binary splitting can split the base block or intermediate block into equal right and left partitions. Similarly, vertical binary splitting can split the base block or intermediate block into equal upper and lower partitions. This binary splitting can be hierarchical and recursive. A decision can be made at each base block or intermediate block in the base block or intermediate block as to whether the binary splitting scheme should continue, and if the scheme is further continued, a decision can be made as to whether to use horizontal or vertical binary splitting. In some embodiments, further splitting can stop at a predefined minimum partition size (one dimension or two dimensions). Optionally, once a predefined partition level or depth from the base block is reached, further splitting can be stopped. In some embodiments, the aspect ratio of the partition can be restricted. For example, the aspect ratio of the partition can be not less than 1:4 (or greater than 4:1). Thus, a vertical bar partition with a vertical-to-horizontal aspect ratio of 4:1 can only be further vertically binary split into upper and lower partitions with a vertical-to-horizontal aspect ratio of 2:1.
[0133] In some other examples, such as Figure 13 shown, a ternary splitting scheme can be used to split a base block or any intermediate block. The ternary pattern can be implemented vertically, as Figure 13 shown in 1302 of Figure 13 , or horizontally, as Figure 13 shown in 1304 of
[0134] . Although the example splitting ratio in Figure 13 is shown as 1:2:1 vertically or horizontally, other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. Since ternary tree splitting can capture an object located at the center of a block in a continuous partition, while quadtree and binary tree always split along the center of the block, thus dividing the object into different partitions, this ternary splitting scheme can be used to complement a quadtree or binary splitting structure. In some embodiments, the width and height of the partitions of an example ternary tree are always powers of 2 to avoid additional transformations.
[0134] The above splitting schemes can be combined in any way at different splitting levels. As an example, the above quadtree and binary splitting schemes can be combined to split a base block into a quadtree - binary - tree (QTBT) structure. In such a scheme, a base block or an intermediate block / partition can be either quadtree - split or binary - split, if specified, to conform to a predefined set of conditions. A specific example is as Figure 14 shown, in the example of Figure 14 , as shown in 1402, 1404, 1406, and 1408, the base block is a first quadtree that is split into four partitions. Thereafter, each of the resulting partitions is either quadtree - split into four further partitions (e.g., 1408), or binary - split into two further partitions at the next level (horizontally or vertically, e.g., 1402 or 1406, both are symmetric), or is not split (e.g., 1404). For square partitions, recursive binary tree or quadtree splitting can be allowed, as shown in the overall example partition pattern of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree splitting and dashed lines represent binary tree splitting. A flag can be used for each binary - split node (non - leaf binary partition) to indicate whether the binary split is horizontal or vertical. For example, as shown in 1420, consistent with the splitting structure of 1410, the flag "0" can represent a horizontal binary split, and the flag "1" can represent a vertical binary split. Since quadtree splitting always splits a block or partition horizontally and vertically to produce 4 sub - blocks / partitions of equal size, for quadtree - split partitions, the splitting type does not need to be indicated. In some embodiments, the flag "1" can represent a horizontal binary split, and the flag "0" can represent a vertical binary split.
[0135] In some example embodiments of QTBT, the quadtree and binary splitting rule sets can be represented by the following predefined parameters and their associated respective functions:
[0136] CTU size: The size of the root node of the quadtree (the size of the base block)
[0137] MinQTSize: The minimum allowable size of the quadtree leaf nodes
[0138] MaxBTSize: The maximum allowable size of the root node of the binary tree
[0139] MaxBTDepth: The maximum allowable depth of the binary tree
[0140] MinBTSize: The minimum allowable size of the binary tree leaf nodes
[0141] In some example embodiments of the QTBT splitting structure, the CTU size can be set to (when considering and using the example chroma subsampling) 128x128 luma samples with two corresponding 64x64 chroma sample blocks, MinQTSize can be set to 16x16, MaxBTSize can be set to 64x64, and MinBTSize (for both width and height) can be set to 4x4, and MaxBTDepth can be set to 4. Quadtree splitting can be first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its minimum allowable size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the node is 128x128, since the size exceeds MaxBTSize (i.e., 64x64), this node will not be first split by the binary tree. Otherwise, nodes not exceeding MaxBTSize can be split by the binary tree. In Figure 14 the example, the base block is 128x128. According to the predefined rule set, the base block can only be split by the quadtree. The splitting depth of the base block is 0. Each of the four resulting partitions is 64x64, which does not exceed MaxBTSize and can be further split by the quadtree or binary tree at level 1. Continue the process. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of the binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of the binary tree node equals MinBTSize, further vertical splitting is not considered.
[0142] In some example embodiments, the above QTBT scheme can be configured to support the flexibility that luminance and chrominance have the same QTBT structure or independent QTBT structures. For example, for P and B slices, the luminance and chrominance CTBs in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be partitioned into CUs by a QTBT structure, and the chrominance CTB can be partitioned into chrominance CUs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can include coding blocks of the luminance component or coding blocks of two chrominance components, and a CU in a P or B slice can include coding blocks of all three color components.
[0143] In some other embodiments, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. Such an embodiment can be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary partitioning of nodes, one of the ternary partitioning patterns can be selected. Figure 13 In some embodiments, only square nodes can be ternary partitioned. An additional flag can be used to indicate whether the ternary partition is horizontal or vertical.
[0144] The design of two-level or multi-level trees, such as the QTBT embodiment and the QTBT embodiment supplemented by ternary partitioning, can be mainly driven by reducing complexity. Theoretically, the complexity of traversing the tree is T D , where T represents the number of partitioning types, and D is the depth of the tree. A full trade-off can be made by using multiple types (T) while reducing the depth (D).
[0145] In some embodiments, a coding block (CB) can be further divided. For example, for intra or inter prediction during encoding and decoding processes, the CB can be further divided into multiple prediction blocks. In other words, the CB can be further divided into different sub - partitions where separate prediction decisions / configurations can be made. In parallel, to describe the level at which the transformation or inverse transformation of video data is performed, the CB can be further divided into multiple transform blocks (TBs). The scheme for dividing the CB into prediction blocks (PBs) and TBs can be the same or not the same. For example, each division scheme can use its own process based on various characteristics of the video data, such as. In some example embodiments, the PB and TB division schemes can be independent. In some other example embodiments, the PB and TB division schemes and boundaries can be related. In some embodiments, for example, the TB can be divided after PB division, and specifically, each PB is determined after dividing the coding block and then can be further divided into one or more TBs. For example, in some embodiments, a PB can be divided into one, two, four, or other numbers of TBs.
[0146] In some embodiments, to divide a base block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and chrominance channels can be processed differently. For example, in some embodiments, for the luminance channel, dividing the coding block into prediction blocks and / or transform blocks can be allowed, while for one (or more) chrominance channels, dividing the coding block into prediction blocks and / or transform blocks can be not allowed. In such embodiments, thus, the transformation and / or prediction of the luminance block can be performed only at the coding block level. As another example, the minimum transform block size of the luminance channel and one (or more) chrominance channels can be different. For example, the coding block of the luminance channel can be divided into smaller transform and / or prediction blocks than the chrominance channel. As another example, the maximum depth of dividing the coding block into transform blocks and / or prediction blocks can be different between the luminance channel and chrominance channels. For example, the coding block for the luminance channel can be divided into deeper transform blocks and / or prediction blocks than one (or more) chrominance channels. For a specific example, the luminance coding block can be divided into transform blocks of multiple sizes, which can be represented by recursive division, and can have transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4x4 to 64x64. However, for the chrominance block, only the maximum possible transform block specified for the luminance block is allowed.
[0147] In some example embodiments, for dividing a coding block into PBs, the depth, shape, and / or other characteristics of PB division can depend on whether the PB is intra - coded or inter - coded.
[0148] The splitting of a coding block (or prediction block) into transform blocks can be implemented in various exemplary scenarios, including but not limited to recursive or non - recursive quadtree splitting and predetermined pattern splitting, with additional consideration given to transform blocks at the boundaries of the coding block or prediction block. Generally, the resulting transform blocks can be at different splitting levels, can have different sizes, and do not need to be square (e.g., can be rectangles of some possible sizes and aspect ratios). Further examples are described in more detail below in conjunction with Figure 15 , 16 , 17.
[0149] However, in some other embodiments, the CB obtained through any of the above - mentioned splitting schemes can be used as the basic or minimum coding block for prediction and / or transformation. In other words, no further splitting is performed for the purpose of performing inter - prediction / intra - prediction and / or transformation. For example, the CB obtained from the above - mentioned QTBT scheme can be directly used as the unit for performing prediction. Specifically, such a QTBT structure eliminates the concept of multiple partition types, i.e., the separation of CU, PU, and TU, and supports greater flexibility in the CU / CB partition shape as described above. In this QTBT block structure, the CU / CB can be square or rectangular in shape. The leaf nodes of this QTBT are used as units for prediction and transformation processing without any further splitting. This means that in this exemplary QTBT coding block structure, the CU, PU, and TU have the same block size.
[0150] The above - mentioned various CB splitting schemes and the further splitting of the CB into PB and / or TB (including no PB / TB splitting) can be combined in any way. The following specific embodiments are provided as non - limiting examples.
[0151] Specific example embodiments of coding block and transform block splitting are described below. In such example embodiments, the above - mentioned recursive quadtree splitting can be used or (as Figure 19 and Figure 10Those) predefined partitioning patterns in divide the base block into coding blocks. At each level, whether a particular partition should be further quadtree partitioned can be determined by local video data characteristics. The resulting CBs can be at various quadtree partition levels and have various sizes. The decision on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture region can be made at the CB level (or CU level, for all three color channels). Each CB can be further divided into one, two, four, or some other number of PBs according to a predefined PB partitioning type. Within a PB, the same prediction process can be applied and related information can be sent to the decoder based on the PB. After obtaining the residual blocks by applying the prediction process based on the PB partitioning type, the CB can be divided into TBs according to another quadtree structure similar to the coding tree of the CB. In this particular embodiment, the CB or TB can be, but is not limited to, square. Additionally, in this particular example, for inter-picture prediction, the PB can be square or rectangular, and for intra-picture prediction, the PB can be only square. The coding block can be divided into, for example, four square TBs. Each TB can be further recursively divided (using quadtree partitioning) into smaller TBs, called Residual Quadtree (RQT).
[0152] Another exemplary embodiment of dividing the base block into CBs, PBs, and / or TBs is further described below. For example, instead of using multiple partition unit types as shown in Figure 9 or Figure 10 , a quadtree with a nested multi-type tree (a partitioning structure using binary and ternary partitioning (e.g., QTBT as described above or QTBT with ternary partitioning)) can be used. The separation of CBs, PBs, and TBs can be dispensed with (i.e., dividing the CB into PBs and / or TBs, and dividing the PB into TBs), unless, when needed, the size of the CB is too large for the maximum transform length, in which case such a CB may need to be further divided. This example partitioning scheme can be designed to support greater flexibility in the CB partitioning shape so that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, the CB can be square or rectangular in shape. Specifically, the Coding Tree Block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using binary or ternary partitioning is shown in Figure 11 . Specifically, Figure 11The example multi-type tree structure includes four splitting types, called vertical binary splitting (SPLIT_BT_VER) (1102), horizontal binary splitting (SPLIT_BT_HOR) (1104), vertical ternary splitting (SPLIT_TT_VER) (1106), and horizontal ternary splitting (SPLIT_TT_HOR) (1108). Then, the CB corresponds to the leaf of the multi-type tree. In this example embodiment, unless the CB is too large for the maximum transform length, this partition is used for prediction and transform processing without any further splitting. This means that, in most cases, in a quadtree with a nested multi-type tree coding block structure, the CB, PB, and TB have the same block size. An exception occurs when the maximum supported transform length is less than the width or height of the color components of the CB. In some embodiments, in addition to binary or ternary splitting, Figure 11 the nested pattern may also include quadtree splitting.
[0153] Figure 12 FIG. shows a specific example of a quadtree with a nested multi-type tree coding block structure having block splitting (including quadtree, binary, and ternary splitting options) for one base block. More specifically, Figure 12 FIG. shows that the base block 1200 is quadtree split into four square partitions 1202, 1204, 1206, and 1208. For each quadtree split partition, a decision is made to further use Figure 11 the multi-type tree structure and the quadtree for further splitting. In Figure 12 the example, partition 1204 is not further split. Partitions 1202 and 1208 each employ another quadtree split. For partition 1202, the upper left, upper right, lower left, and lower right partitions of the second-level quadtree split respectively employ Figure 11 the third-level split of the quadtree, horizontal binary splitting 1104, no splitting, and Figure 11 the horizontal ternary splitting 1108. Partition 1208 employs another quadtree split, and the upper left, upper right, lower left, and lower right splits of the second-level quadtree split respectively employ Figure 11 the third-level split of the vertical ternary splitting 1106, no splitting, no splitting, and Figure 11 the horizontal binary splitting 1104. The two sub-partitions of the third-level upper left partition of 1208 are further split respectively according to Figure 11 the horizontal binary splitting 1104 and the horizontal ternary splitting 1108. Partition 1206 employs the second-level split pattern according to Figure 11 the vertical binary splitting 1102, and is divided into two partitions, and these two partitions are further split at the third level according to Figure 11 the horizontal ternary splitting 1108 and the vertical binary splitting 1102. According to Figure 11The horizontal binary split 1104, and the fourth-level split is further applied to one of these two partitions.
[0154] For the specific example above, the maximum luma transform size can be 64x64, and the maximum supported chroma transform size can be different from the luma, e.g., 32x32. Even though Figure 12 the above example CB in is usually not further split into smaller PBs and / or TBs, when the width or height of a luma coding block or a chroma coding block is greater than the maximum transform width or height, the luma coding block or the chroma coding block can be automatically split in the horizontal and / or vertical directions to conform to the transform size limit in that direction.
[0155] As described above, for the specific example of splitting a base block into CBs, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P and B slices, the luma and chroma CTBs in a CTU can have the same coding tree structure. For example, for I slices, luma and chroma can have separate coded block tree structures. When separate block tree structures are applied, the luma CTB is split into luma CBs by one coding tree structure, and the chroma CTB is split into chroma CBs by another coding tree structure. This means that a CU in an I slice can include coded blocks of the luma component or coded blocks of both chroma components, and a CU in a P or B slice always includes coded blocks of all three color components, unless the video is monochrome.
[0156] When a coding block is further split into multiple transform blocks, the transform blocks therein can be ordered in the bitstream in various orders or scan patterns. Example embodiments of splitting a coding block or a prediction block into transform blocks and the coding order of the transform blocks are described in further detail below. In some example embodiments, as described above, transform splitting can support transform blocks of multiple shapes, e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the range of transform block sizes is from, e.g., 4x4 to 64x64. In some embodiments, if the coding block is less than or equal to 64x64, then transform block splitting can be applied only to the luma component, so that for chroma blocks, the transform block size is the same as the coding block size. Otherwise, if the coding block width or height is greater than 64, then both the luma and chroma coding blocks can be implicitly split into multiples of min(W, 64) x min(H, 64) and min(W, 32) x min(H, 32) transform blocks, respectively.
[0157] In some example embodiments of transform block partitioning, for intra and inter coded blocks, the coded block may be further divided into multiple transform blocks, with a partitioning depth up to a predetermined number of levels (e.g., 2 levels). The transform block partitioning depth and size may be related. For some example embodiments, the mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.
[0158] Table 1 Transform Partitioning Size Settings
[0159] Transformation size at the current depth Transformation size at the next depth TX_4×4 TX_4×4 TX_8×8 TX_4×4 TX_16×16 TX_8×8 TX_32×32 TX_16×16 TX_64×64 TX_32×32 TX_4×8 TX_4×4 TX_8×4 TX_4×4 TX_8×16 TX_8×8 TX_16×8 TX_8×8 TX_16×32 TX_16×16 TX_32×16 TX_16×16 TX_32×64 TX_32×32 TX_64×32 TX_32×32 TX_4×16 TX_4×8 TX_16×4 TX_8×4 TX_8×32 TX_8×16 TX_32×8 TX_16×8 TX_16×64 TX_16×32 TX_64×16 TX_32×16
[0160] Based on the example mapping in Table 1, for a 1:1 square block, the next level of transform partitioning may create four 1:1 square sub-transform blocks. For example, the transform partitioning may stop at 4x4. Thus, a transform size of 4x4 at the current depth corresponds to the same size of 4x4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level of transform partitioning may create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next level of transform partitioning may create two 1:2 / 2:1 sub-transform blocks.
[0161] In some example embodiments, for the luminance component of an intra coded block, additional restrictions may be applied with respect to the transform block partitioning. For example, for each level of transform partitioning, all sub-transform blocks may be restricted to have equal sizes. For example, for a 32x16 coded block, the first level of transform partitioning creates two 16x16 sub-transform blocks, and the second level of transform partitioning creates eight 8x8 sub-transform blocks. In other words, the second level of partitioning must be applied to all first-level sub-blocks to keep the transform unit sizes equal. Figure 15 An example of the transform block partitioning of an intra coded square block according to Table 1 is shown, along with the coding order indicated by arrows. Specifically, 1502 shows the square coded block. The first level of partitioning into 4 equal-sized transform blocks according to Table 1 is shown in 1504, where the coding order is indicated by arrows. The second level of partitioning of all first-level equal-sized blocks into 16 equal-sized transform blocks according to Table 1 is shown in 1506, where the coding order is indicated by arrows.
[0162] In some example embodiments, for the luminance component of an inter coded block, the above restrictions for intra coding may not be applied. For example, after the first level of transform partitioning, any one of the sub-transform blocks may be further independently partitioned by more than one level. Thus, the resulting transform blocks may or may not have the same size. Figure 16 An example of partitioning an inter coded block into transform blocks and its coding order is shown. In Figure 16In the example of Figure 16 , the inter-frame encoded block 1602 is split into transform blocks at two levels according to Table 1. At the first level, the inter-frame encoded block is split into four transform blocks of equal size. Then, as shown at 1604, only one (not all) of the four transform blocks is further split into four sub-transform blocks, resulting in a total of 7 transform blocks of two different sizes. The example encoding order of these 7 transform blocks is shown by the arrows in
[0163] In some example embodiments, for one (or more) chrominance components, some additional restrictions on the transform blocks may be applied. For example, for one (or more) chrominance components, the transform block size may be the same as the coding block size, but not less than a predefined size, such as 8x8.
[0164] In some other example embodiments, for coding blocks with a width (W) or height (H) greater than 64, the luminance and chrominance coding blocks may be implicitly split into transform units that are multiples of min(W, 64) x min(H, 64) and min(W, 32) x min(H, 32), respectively. Here, in the embodiments of the present application, "min(a, b)" may return the smaller value between a and b.
[0165] Figure 17 Another alternative example scheme for splitting a coding block or a prediction block into transform blocks is further shown. As shown in Figure 17 , a predefined set of partitioning types may be applied to the coding block according to the transform type of the coding block, without using recursive transform partitioning. In the specific example shown in Figure 17 , one of 6 example partitioning types may be applied to split the coding block into various numbers of transform blocks. This scheme for generating transform block partitions may be applied to coding blocks or prediction blocks.
[0166] More specifically, Figure 17 's partitioning scheme provides up to 6 example partitioning types for any given transform type (the transform type refers to, for example, the type of the main transform, such as ADST, etc.). In this scheme, a transform partitioning type may be assigned to each coding block or prediction block based on, for example, rate-distortion cost. In an example, the transform partitioning type assigned to a coding block or prediction block may be determined based on the transform type of the coding block or prediction block. A specific transform partitioning type may correspond to a transform block partition size and pattern, as shown by the 6 transform partitioning types shown in Figure 17 . The correspondence between various transform types and various transform partitioning types may be predefined. The following examples are shown, where the capital labels indicate the transform partition types that may be assigned to coding blocks or prediction blocks based on rate-distortion cost:
[0167] ·PARTITION_NONE: Assign a transform size equal to the block size.
[0168] ·PARTITION_SPLIT: The partitioning transformation has a width that is 1 / 2 of the block size and a height that is 1 / 2 of the block size.
[0169] ·PARTITION_HORZ: The partitioning transformation has a width that is the same as the block size and a height that is 1 / 2 of the block size.
[0170] ·PARTITION_VERT: The partitioning transformation has a width that is 1 / 2 of the block size and a height that is the same as the block size.
[0171] ·PARTITION_HORZ4: The partitioning transformation has a width that is the same as the block size and a height that is 1 / 4 of the block size.
[0172] ·PARTITION_VERT4: The partitioning transformation has a width that is 1 / 4 of the block size and a height that is the same as the block size.
[0173] In the above examples, the transformation split types as Figure 17 shown all include a unified transformation size for the partitioned transformation blocks. This is only an example and not a limitation. In some other embodiments, mixed transformation block sizes may be used for the partitioned transformation blocks in a particular split type (or mode).
[0174] Video blocks (PB or CB, also referred to as PB when not further split into multiple prediction blocks) can be predicted in various ways instead of being directly encoded, thus leveraging various correlations and redundancies in the video data to improve compression efficiency. Accordingly, such prediction can be performed in various modes. For example, a video block can be predicted through intra prediction or inter prediction. Particularly in the inter prediction mode, a video block can be predicted from one or more other frames by one or more other reference blocks or inter predictor blocks through single-reference or composite-reference inter prediction. To implement inter prediction, a reference block can be specified by its frame identifier (the temporal position of the reference block) and a motion vector (the spatial position of the reference block) indicating the spatial offset between the currently encoded or decoded block and the reference block. The reference frame identification and the motion vector can be signaled in the bitstream. The motion vector as a spatial block offset can be directly signaled, or it can be predicted by another reference motion vector or predictor motion vector. For example, the current motion vector can be directly predicted by a reference motion vector (e.g., of a candidate neighboring block), or the current motion vector can be predicted by a combination of a reference motion vector and the motion vector difference (MVD) between the current motion vector and the reference motion vector. The latter can be referred to as the merge mode with motion vector difference (MMVD). The reference motion vector can be identified in the bitstream as a pointer pointing to, for example, a spatially neighboring block of the current block or a temporally neighboring but spatially co-located block.
[0175] In some other example embodiments, Intra Block Copy (IBC) prediction may be employed. In IBC, another block in the current frame (instead of a temporally different frame, thus called "intra") is used in combination with a Block Vector (BV) to predict a current block in the current frame, where the block vector is used to indicate the offset of the position of the intra predictor or the reference block relative to the position of the block to be predicted. The position of the coded block may be represented, for example, by pixel coordinates relative to the top-left pixel of the current frame (or slice). Thus, the IBC mode uses a concept similar to inter prediction within the current frame. For example, the BV may be predicted directly by other reference BVs or in combination with the BV difference between the current BV and the reference BV, which is consistent with using the reference MV and MV difference to predict the MV in inter prediction. IBC is useful in providing improved coding efficiency, especially for encoding and decoding video frames with screen content that has, for example, a large number of repeating patterns, such as text information, where the same text segments (letters, symbols, words, phrases, etc.) appear in different parts of the same frame and can be used to predict each other.
[0176] In some embodiments, IBC may be regarded as a separate prediction mode in addition to the normal intra prediction mode and the normal inter prediction mode. Thus, the prediction mode of a particular block can be selected and signaled among three different prediction modes: intra prediction, inter prediction, and IBC mode. In these embodiments, flexibility can be established in each of these modes to optimize the coding efficiency of each of these modes. In some other embodiments, similar motion vector determination, reference, and coding mechanisms may be used, and IBC may be regarded as a sub-mode or a branch within the inter prediction mode. In such embodiments (integrating the inter prediction mode and the IBC mode), the flexibility of IBC may be restricted to coordinate the general inter prediction mode and the IBC mode. However, this embodiment is less complex while still being able to utilize IBC to improve the coding efficiency of video frames characterized by, for example, screen content. In some example embodiments, by leveraging the existing pre-specified mechanisms for the separate inter prediction mode and intra prediction mode, the inter prediction mode can be extended to support IBC.
[0177] The selection of these prediction modes can be made at various levels, including but not limited to the sequence level, frame level, picture level, slice level, CTU level, CT level, CU level, CB level, or PB level. For example, for the purpose of IBC, the decision on whether to adopt the IBC mode can be made and signaled at the CTU level. If a CTU is signaled to adopt the IBC mode, all coding blocks in that entire CTU can be predicted by IBC. In some other embodiments, IBC prediction can be determined at the super block (SB) level. Each SB can be divided into multiple CTUs or partitions in various ways (e.g., quadtree partitioning). Examples are provided further below.
[0178] Figure 18 An example snapshot showing a portion of a current frame containing multiple CTUs from the perspective of a decoder is shown. Each square block (such as 1802) represents a CTU. As described in detail above, the size of a CTU can be one of various predefined sizes. Each CTU can include one or more coding blocks (or prediction blocks for a specific color channel). The CTUs with horizontal shading represent those that have been reconstructed. CTU 1804 represents the current CTU being reconstructed. In the current CTU 1804, the coding blocks with horizontal shading represent those that have been reconstructed in the current CTU, the coding block 1806 with diagonal shading is currently being reconstructed, and the coding blocks without shading in the current CTU 1804 are waiting to be reconstructed. The other CTUs without shading have not been processed yet.
[0179] The position or offset of the reference block (relative to the current block) for predicting the current coding block in IBC can be indicated by a BV, as Figure 18 shown by the example arrows in. For example, a BV can indicate the position difference between the reference block (labeled "Ref" in Figure 18 ) and the upper left corner of the current block in vector form. And Figure 18 is shown using the CTU as the basic IBC unit. The basic principle applies to embodiments using the SB as the basic IBC unit. In such embodiments, as described in more detail below, each super block can be divided into multiple CTUs, and each CTU can be further divided into multiple coding blocks.
[0180] As will be disclosed in further detail below, depending on the position of the reference CTU / SB relative to the current CTU / SB of the IBC, the reference CTU / SB may be referred to as a local CTU / SB or a non-local CTU / SB. A local CTU / SB may refer to a CTU / SB that coincides with the current CTU / SB, or a CTU / SB that is close to the current CTU / SB and has been reconstructed (e.g., the left adjacent CTU / SB of the current CTU / SB). A non-local CTU / SB may refer to a CTU / SB that is farther away from the current CTU / SB. When performing IBC prediction for the current coding block, either one or both of the local CTU / SB and the non-local CTU / SB may be searched for the reference block. Since the on-chip and off-chip storage management of the reconstructed samples for local or non-local CTU / SB references (e.g., off-chip picture buffer (DPB) and / or on-chip memory) may be different, the specific implementation of IBC may depend on whether the reference CTU / SB is local or non-local. For example, the reconstructed local CTU / SB samples may be suitable for storage in the on-chip memory of the IBC encoder or decoder. For example, the reconstructed non-local CTU / SB samples may be stored in the off-chip DPB memory.
[0181] In some embodiments, the position of the reconstructed blocks that can be used as reference blocks for the current coding block 1804 may be restricted. Such a restriction may be the result of various factors and may depend on whether the IBC is implemented as an integrated part of the general inter prediction mode, a special extension of the inter prediction mode, or a separate and independent IBC mode. In some examples, only the currently reconstructed CTU / SB samples may be searched to identify the IBC reference blocks. In some other examples, as Figure 18 shown by the thick dashed box 1808 of, the currently reconstructed CTU / SB samples and another adjacent reconstructed CTU / SB sample (e.g., the left adjacent CTU / SB) may be used for reference block search and selection. For such embodiments, only the local reconstructed CTU / SB samples may be used for IBC reference block search and selection. In some other examples, certain CTU / SB may not be available for IBC reference block search and selection due to various other reasons. For example, as will be described in further detail below, Figure 18 the CTU / SB 1810 marked with crosshairs in may be used for special purposes (e.g., wavefront parallel processing), so they may not be available for the search and selection of the reference block of the current block 1804.
[0182] In some embodiments, the restriction on the reconstructed CTU / SB that is allowed to provide the IBC reference block or reference sample may be the result of adopting parallel decoding, in which multiple coding blocks are decoded simultaneously. Figure 19Shows an example where each square represents a CTU / SB. As Figure 19 As shown by the CTU / SB with diagonal shading, parallel decoding can be implemented, where multiple consecutive rows and multiple CTU / SBs in every other column (every two columns) can be reconstructed in a parallel processing manner. Other CTU / SBs with horizontal shading have been reconstructed, and the CTU / SBs without shading are the CTU / SBs to be reconstructed. In such parallel processing, for the currently parallel processed CTU / SB with its upper left coordinates (x0, y0), only when the vertical coordinate y is less than y0 and the horizontal coordinate x is less than x0 + 2(y0 - y), can the reconstructed samples at (x, y) be accessed to predict the current CTU / SB in IBC. Therefore, the reconstructed CTU / SBs with horizontal shading can be used as references for the current block in parallel processing.
[0183] In some embodiments, the write-back latency of immediately reconstructed samples written to off-chip DPB may further limit the CTU / SBs available for providing IBC reference samples to the current block, especially when off-chip DPB is used to store IBC reference samples. Figure 19 Shows an example where, based on Figure 19 the limitations shown above, other limitations can also be applied. Specifically, to allow for hardware write-back latency, IBC prediction may not access the immediately reconstructed region to search for and select reference blocks. The number of immediately reconstructed regions restricted or prohibited can be 1 to n CTU / SBs (n is a positive integer). Therefore, based on Figure 19 the specific parallel processing limitations, for a current CTU / SB with the upper left position coordinates (x0, y0), if the vertical coordinate y is less than y0 and the horizontal coordinate is less than x0 + 2(y0 - y) - D, then the prediction at position (x, y) can be accessed through IBC, where D represents the number of immediately reconstructed regions restricted or prohibited from being used as IBC references (e.g., on the left side of the current CTU / SB). Figure 20 Shows that for D = 2, such additional CTU / SBs are restricted as IBC reference samples. These additional CTU / SBs that cannot be used as IBC references are represented by backslash shading.
[0184] In some embodiments, also described in further detail below, both local and non-local CTU / SB search regions can be used for IBC reference block search and selection. Additionally, when on-chip memory is used, some of the restrictions on write-back latency can be relaxed or removed regarding the availability of already constructed CTUs / SBs as IBC references. In some further embodiments, due to, for example, different buffer management of reference blocks using on-chip or off-chip memory, the usage of local CTU / SBs and non-local CTU / SBs can be different when co-existing. These embodiments are described in more detail in the disclosure below.
[0185] In some embodiments, IBC can be implemented as an extension of the inter-frame prediction mode, treating the current frame as a reference frame in the inter-frame prediction mode such that blocks within the current frame can be used as prediction references. Thus, even though the IBC process only involves the current frame, such an IBC embodiment can follow the encoding path of inter-frame prediction. In such an embodiment, the reference structure of the inter-frame prediction mode can be adapted for IBC, where the representation of the addressing mechanism of reference samples using BV can be similar to the motion vector (MV) in inter-frame prediction. Thus, IBC can be implemented as a special inter-frame prediction mode, relying on similar or identical syntax structures and decoding processes as the inter-frame prediction mode with the current frame as the reference frame.
[0186] In such an embodiment, since IBC can be regarded as an inter-frame prediction mode, only the intra-predicted strips must become strips that allow the use of IBC for prediction. In other words, only intra-predicted strips are not inter-frame predictions (because the intra-prediction mode does not call any inter-frame prediction processing paths), so IBC cannot be used for prediction in such intra-only strips. When IBC is applicable, the encoder will extend the reference picture list with an entry of a pointer to the current picture. Thus, the current picture can occupy up to one picture-sized buffer in the shared decoded picture buffer (DPB). The signaling of using IBC can be implicit in the selection of the reference frame in the inter-frame prediction mode. For example, when the selected reference picture points to the current picture, if needed and available, the coding unit will employ IBC with a coding path similar to inter-frame prediction with a special IBC extension. In some specific embodiments, contrary to conventional inter-frame prediction, the reference samples in the IBC process may not be loop-filtered before being used for prediction. Additionally, the corresponding reference current picture can be a long-term reference frame as it will be close to the next frame to be encoded or decoded. In some embodiments, to minimize memory requirements, the encoder can release the buffer immediately after reconstructing the current picture. When the reconstructed picture becomes a reference image for subsequent frames in true inter-frame prediction, the encoder can fill the filtered version of the reconstructed picture back into the DPB as a short-term reference, even though it may be unfiltered when used for IBC.
[0187] In the above example embodiment, even if the IBC can be merely an extension of the inter-frame prediction mode, the IBC can be processed with several special processes that may deviate from normal inter-frame prediction. For example, the IBC reference samples can also be unfiltered. In other words, the reconstructed samples before in-loop filtering processing can be used for IBC prediction, and the in-loop filtering processing includes deblocking filtering, Sample Adaptive Offset (SAO) filtering, Cross-Component Sample Offset (CCSO) filtering, etc., while the normal inter-frame prediction mode uses the filtered samples for prediction. For another example, the luminance sample interpolation for the IBC may not be performed, and the chrominance sample interpolation may be necessary only when the chrominance BV is non-integer when derived from the luminance BV. For another example, when the chrominance BV is non-integer and the reference block for the IBC is near the boundary of the available area of the IBC reference, the surrounding reconstructed samples can be outside the boundary to perform chrominance interpolation. The BV pointing to a single adjacent boundary cannot avoid this situation.
[0188] In such an embodiment, the prediction of the current block by the IBC can reuse the prediction and encoding mechanisms of the inter-frame prediction process, including using the reference BV to predict the current BV and, for example, the additional BV difference. However, in some specific embodiments, the luminance BV can be implemented at integer resolution instead of at fractional precision as in the MV of the conventional inter-frame prediction.
[0189] In some embodiments, in addition to the two CTUs ( Figure 18 indicated by the cross lines in Figure 18 as shown in 1810 in Figure 18 ), which are to the right and above the current CTU for Wavefront Parallel Processing (WPP),
[0190] all CTUs and SBs indicated by the horizontal hatching in Figure 18 can be used to search for and select the IBC reference block. Thus, except for some exceptions for parallel processing purposes, almost the entire reconstructed area of the current picture.
[0190] In some other embodiments, the area from which the IBC reference block can be searched for and selected can be restricted to the local CTU / SB. Figure 18The thick dashed box 1808 indicates an example. In this example, the CTU / SB to the left of the current CTU can be used as a reference sample area for IBC at the start of the reconstruction process of the current CTU. When using this local reference area, on-chip memory space can be allocated to hold the local CTU / SB for IBC reference instead of allocating additional external memory space in the DPB. In some embodiments, fixed on-chip memory can be used for IBC, thereby reducing the complexity of implementing IBC in the hardware architecture. Thus, a dedicated IBC mode independent of normal inter-frame prediction can be implemented using on-chip memory rather than just being implemented as an extension of the inter-frame prediction mode.
[0191] For example, for each color component, the fixed on-chip memory size for storing local IBC reference samples (e.g., the left CTU or SB) can be 128x128. In some embodiments, the maximum CTU size can also be 128x128. In this case, the reference sample memory (RSM) can hold samples of a single CTU size. In some other alternative embodiments, the CTU size can be smaller. For example, the CTU size can be 64x64. Thus, the RSM can hold multiple (in this example case, 4) CTUs at the same time. In some other embodiments, the RSM can hold multiple SBs, each SB can include one or more CTUs, and each CTU can include multiple coding blocks.
[0192] In some embodiments of local on-chip IBC reference, the on-chip RSM holds one CTU and can implement a continuous update mechanism to replace the reconstructed samples of the left adjacent CTU with the reconstructed samples of the current CTU. Figure 21 A simplified example of such a continuous RSM update mechanism at four intermediate times during the reconstruction process is shown. At Figure 21 the example, the RSM has a fixed size that can accommodate one CTU. The CTU can include implicit partitions. For example, the CTU can be implicitly divided into four separate regions (e.g., quadtree partitioning). Each region can include multiple coding blocks. The size of the CTU can be 128x128, while for the example quadtree partitioning, the size of each example region or partition can be 64x64. The regions / partitions of the RSM with horizontal line shading at each intermediate time hold the corresponding reconstructed reference samples of the left adjacent CTU, and the regions / partitions with gray vertical line shading hold the corresponding reconstructed reference samples of the current CTU. The coding blocks of the RSM with diagonal shading represent the current coding blocks within the current region being encoded / decoded / reconstructed.
[0193] At a first intermediate time indicating the start of the current CTU reconstruction, as shown in 2102, for each of the four example regions, the RSM may include only the reconstructed reference samples of the left adjacent CTU. At the other three intermediate times, the reconstruction process gradually replaces the reconstructed reference samples of the left adjacent CTU with the reconstructed samples of the current CTU. When the encoder processes the first coded block of the region / partition, the 64x64 region / partition in the RSM is reset. When resetting the region of the RSM, the region is considered blank and is considered to have not saved any reconstructed reference samples for IBC (in other words, this region of the RSM is not ready to be used as an IBC reference sample). When processing the corresponding current coded block in the region, the corresponding block in the RSM is incorporated into the reconstructed samples of the corresponding block of the current CTU for use as a reference sample for the IBC of the next current block, as shown in Figure 21 intermediate times 2104, 2106, 2108 of Figure 21 shown by the regions fully shaded with vertical lines at each intermediate time in
[0194] Figure 22 shows the above-described successive update implementation of the RSM spatially at specific intermediate times, i.e., both the left adjacent CTU and the current CTU with the current coded block (the slant-shaded block) are shown. The corresponding reconstructed samples of these two CTUs that are effectively the IBC reference samples for the current coded block in the RSM are shown by horizontal and vertical shading lines. At the specific reconstruction time in this example, in the RSM, the process has replaced the samples covered by the unshaded region in the left adjacent CTU with the region of the current CTU covered by the vertical shading lines. The remaining effect samples from the adjacent CTU are shown as horizontal line shading.
[0195] In the above example embodiment, when the fixed RSM size is the same as the CTU size, the RSM is implemented to contain one CTU. In some other embodiments where the CTU size is smaller, the RSM may contain more than one CTU. For example, the size of the CTU may be 32x32, while the size of the fixed RSM may be 128x128. Thus, the RSM can hold samples of 16 CTUs. Following the same basic RSM update principle described above, the RSM can hold 16 adjacent CTUs of the current 128x128 patch before being reconstructed. Once the processing of the first coded block of the current 128x128 patch begins, the first 32x32 region of the RSM, which is initially filled with the reconstructed samples of one adjacent CTU as in the case of saving a single CTU RSM update as described above, can be saved. The remaining 15 32x32 regions contain 15 adjacent CTUs as reference samples for IBC. Once the CTU corresponding to the first 32x32 region of the currently decoded 128x128 patch is reconstructed, the first 32x32 region of the RSM is updated with the reconstructed samples of that CTU. Then, the CTU corresponding to the second 32x32 region of the current 128x128 patch can be processed and ultimately updated with the reconstructed samples. This process continues until the 16 32x32 regions of the RSM contain the reconstructed samples (all 15 CTUs) of the current 128x128 patch. The decoding process then moves to the next 128x128 patch.
[0196] In some other embodiments, as Figure 21 and 22 an extension, the RSM can hold a set of adjacent CTUs. Processing one current CTU at a time, the part of the RSM that holds the farthest adjacent CTU is updated with the reconstructed current CTU in the manner described above. For the next current CTU, again, the farthest adjacent CTU in the RSM is updated and replaced. Thus, the multiple CTUs held in the fixed-size RSM are updated as a moving window of adjacent CTUs for IBS.
[0197] Another specific example embodiment of using the on-chip RSM for local IBC is as Figure 23 shown. In this example, the maximum block size of the IBC mode can be restricted. For example, the maximum IBC block size can be 64x64. The on-chip RSM can be configured to have a fixed size corresponding to a superblock (SB), such as 128x128. Figure 23 The RSM embodiment of Figure 21 and Figure 22 uses a basic principle similar to the embodiments of Figure 23 In Figure 23In the example of , SB can be a quadtree segmentation. Correspondingly, RSM can be segmented into 4 regions or units by a quadtree, and each region or unit is 64x64. Each of these regions can store one or more coded blocks. Optionally, each of these regions can store one or more CTUs, and each CTU can store one or more coded blocks. The coding order of the quadtree regions can be predefined. For example, the coding order can be top left, top right, bottom left, bottom right. Figure 23 The quadtree segmentation of SB in is only an example. In some other alternative embodiments, SB can be segmented according to any other scheme. The RSM update embodiments described herein for local IBC are applied to those alternative segmentation schemes.
[0198] In this local SBC embodiment, the local reference blocks available for SBC prediction can be restricted. For example, it can be required that the reference block and the current block should be in the same SB row. Specifically, the local reference block can be located only in the current SB or in an SB to the left of the current SB. An example of a current block predicted by another allowed coded block in SBC is shown by Figure 23 the dashed arrow in . When the current SB or the left SB is used as an SBC reference, the reference sample update process in RSM can follow the above reset process. For example, when any one of the 64x64 unit reference sample memories starts to be updated with the reconstructed samples from the current SB, the previously stored reference samples (from the left SB) in the entire 64x64 unit are marked as unavailable for generating IBC prediction samples and are gradually updated with the reconstructed samples of the current block.
[0199] Figure 23 FIG. shows 5 example states of RSM during local IBC decoding of the current SB in panel 2302. Similarly, in each example state, the horizontally shaded regions / partitions of RSM store the corresponding reference samples of the quadtree of the corresponding left adjacent SB, and the vertically shaded regions / partitions in gray store the corresponding reference samples of the current SB. The coded blocks of the RSM with diagonal shading represent the current coded blocks within the current region being encoded / decoded. At the start of the encoding of each current SB, RSM stores the samples of the previously encoded SB ( Figure 23 RSM state (0) of ). When the current block is in one of the four 64x64 quadtree regions in the current SB, the corresponding region in RSM is reset and used to store the samples of the current 64x64 encoding region. In this way, the samples in each 64x64 quadtree region of RSM are gradually updated with the samples in the current SB (states (1)-(3)). When the current SB has been fully encoded, the entire RSM is filled with all the samples of the current SB (state (4)).
[0200] Figure 23Each of the 64x64 regions in the panel 2302 is marked with a spatially encoded sequence number. Sequence numbers 0 - 3 represent the four 64x64 quadtree regions of the left neighbor SB, while sequence numbers 4 - 7 represent the four 64x64 quadtree regions of the current SB panel. In Figure 23 For the RSM states (1), (2), and (3) of the panel 2302 in Figure 23 , the panel 2304 further shows the corresponding spatial distributions of the reference samples in the left adjacent and the current SB in the 128x28 RSM. The shaded regions without cross - lines represent the regions in the RSM with reconstructed samples. The shaded regions with cross - lines represent the regions in the RSM where the reconstructed samples of the left SB are reset (and thus not available as reference samples for the local SBC).
[0201] The encoding order of the 64x64 regions and the corresponding RSM update order can follow a horizontal scan (as shown above in Figure 23 ) or a vertical scan. The horizontal scan starts from the upper left, goes to the upper right, lower left, and lower right. The vertical scan starts from the upper left, goes to the lower left, upper right, and lower left. Figure 24 The panels 2402 and 2404 in Figure 24 respectively show the left adjacent SB and current SB reference sample update processes for horizontal and vertical scans for comparison when reconstructing each of the four 64x64 regions of the current SB. In Figure 24 , the 64x64 regions shaded with horizontal lines without cross - lines represent the regions with samples available for the SBC. The regions shaded with horizontal lines with cross - lines represent the regions of the left adjacent SB that have been updated to the corresponding reconstructed samples of the current SB. The unshaded regions represent the unprocessed regions of the current SB. The diagonally shaded blocks represent the currently encoded blocks being processed.
[0202] As Figure 24 shown, depending on the position of the current encoded block relative to the current SB, the following restrictions regarding the reference blocks for IBC can be applied.
[0203] If the current block falls into the upper left 64x64 region of the current SB, then in addition to the reconstructed samples in the current SB, the reference samples in the lower right, lower left, and upper right 64x64 blocks of the left SB can also be referenced, as shown in Figure 24 2412 (for horizontal scan) and 2422 (for vertical scan) in
[0204] If the current block falls into the upper right 64x64 block of the current SB, then in addition to the reconstructed samples in the current SB, if the luminance sample at (0, 64) relative to the current SB has not been reconstructed, the current block can also reference the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left SB ( Figure 24of 2414). Otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB for SBC( Figure 24 of 2426).
[0205] If the current block falls into the lower left 64x64 block of the current SB, then in addition to the reconstructed samples in the current SB, if the luminance position (64, 0) relative to the current SB has not been reconstructed, the current block can also refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left SB( Figure 24 of 2424). Otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left SB for SBC( Figure 24 of 2416).
[0206] If the current block falls into the lower right 64x64 block of the current SB, then the current block can only refer to the reconstructed samples in the current SB for SBC( Figure 24 of 2418 and 2428).
[0207] As described above, in some example embodiments, one or both of the local and non-local CTU / SBs can be used for IBC reference block search and selection. In addition, when using on-chip RSM for local reference, for the availability of the already constructed CTU / SB as an IBC reference, some restrictions on the write-back delay can be relaxed or removed. Such embodiments can be applied regardless of whether parallel decoding is employed.
[0208] Example embodiments of local and non-local reference CTU / SBs available for IBC are as Figure 25 shown, again, where each square represents a CTU / SB. The CTU / SB with slanted shading (labeled "0") represents the current CTU / SB, while the CTU / SB with horizontal shading (labeled "1"), the CTU / SB with vertical shading (labeled "2"), and the CTU / SB with backslash shading represent the already constructed regions. The CTU / SB without shading represents the region that has not been reconstructed. Assume the use of parallel decoding similar to Figure 19 and Figure 20 . Due to the write-back delay to the DPB when only off-chip memory is used for SBC reference (see Figure 20 ), thus the CTU / SB with vertical ("2") and backslash ("3") shading represents an example region that is generally restricted as the SBC reference for the current CTU / SB. When using on-chip RSM, then Figure 20 one or more restricted regions can be directly referenced from the RSM, and thus may not need to be restricted out. The number of restricted regions now accessible via the RSM for IBC reference can depend on the size of the RSM. In Figure 25In the example, it is assumed that the RSM can store one CTU / SB and adopts the above RSM update mechanism. Thus, Figure 20 one of the two sets of adjacent CTU / SBs with backslash shading (marked as "3") can be used for local reference. Then, the RSM stores samples from the left CTU / SB and the current CTU / SB. Therefore, in Figure 25 the example, the search area available for non-local SBC reference blocks includes the CTU / SB marked as "1" (search area 1, or SA1), the scan area available for local SBC reference blocks includes the CTU / SBs marked as "2" and "0" (SA2), and the restricted output area of the SBC reference block includes the CTU / SB marked as "3" due to write-back latency. In some other embodiments, if the on-chip RSM size is large enough to accommodate the entire restricted CTU / SB, then all these potential restricted areas can be included in the RSM for local reference.
[0209] Figure 26 Further shown is a further restriction on the reference coding blocks that can be used for predicting the current coding block in IBC when both local and non-local reference searches are allowed and enabled. In Figure 26 , similarly, each square represents a CTU / SB. The CTU / SB with horizontal shading represents the CTU / SB that has been constructed. The non-shaded CTU / SB represents the area that has not been reconstructed yet. The CTU / SB with backslash shading is the CTU / SB that cannot be used as an IBC reference (here only the current CTU / SB is shown to be allowed for IBC reference, however, the basic principle applies to the case where only the first of the two CTU / SBs with backslash shading is not allowed, as shown in Figure 25 ). The coded block with slash shading is the current coding block. Coded blocks A, B, and C are potential reference blocks for the IBC of the current coding block. The other shaded coded blocks in the current CTU / SB have been constructed. In this embodiment, since coded block B is completely outside the restricted area and in SA2 (local search area) and has been reconstructed, the reference coded block B is allowed. Since coded block C is completely outside the restricted area and in SA1 (non-local search area) and has been reconstructed, coded block C is also allowed. Since coded block A covers SA1 and SA2, coded block A is not allowed to be used as a prediction block. In other words, since the processing of IBC for SA1 and SA2 may be different and not easily coordinated, a reference coded block that covers both SA1 and SA2 may not be allowed.
[0210] Returning to the encoding of the block vector (BV) in IBC, in some example embodiments, processes similar to those specified for inter prediction may be employed, but simpler rules for BV prediction candidate list construction may be used. For example, the candidate list construction in some inter prediction embodiments may consist of five spatial, one temporal, and six history-based candidates. In this inter prediction, multiple candidate comparisons may be performed on the history-based candidates to avoid duplicate entries in the final candidate list. Additionally, the list construction may include candidates that are pair-averaged. In some example embodiments of BV prediction, the IBC list construction process may consider multiple (e.g., two) spatially adjacent BVs and multiple (e.g., five) history-based BVs (HBVPs), where only the first HBVP may be compared with the spatial candidates when added to the candidate list. While conventional inter prediction may use two different candidate lists, one for the merge mode and another for the regular mode, the candidate list in IBC may be used for both cases regarding BVs. However, the merge mode may use at most six candidates in the list, while the regular mode only uses the first two candidates. In some example embodiments, the block vector difference (BVD) encoding may employ the motion vector difference (MVD) process, resulting in a final BV of any magnitude. The reconstructed BV may point to a region outside the reference sample area, and correction may be required by removing the absolute offset in each direction using modulo operations on the width and height of the RSM.
[0211] In the above embodiments using one or both of local and non-local IBC references, a loop filter may be used in certain cases. For example, when using a non-local-based IBC search range (with or without a local-based IBC search range), e.g., for a picture, the loop filter may be disabled in IBC for the same picture. On the other hand, if only a local-based IBC search range is used (without using a non-local-based IBC search range), the loop filter may be used for the same picture. The loop filter may include, but is not limited to, a deblocking filter, a Constrained Directional Enhancement Filter (CDEF), a Sample Adaptive Offset (SAO) filter, a Cross-Component Sample Offset (CCSO) filter, and a Loop Restoration filter (LR). In this way, a second picture buffer dedicated to enabling IBC may be avoided.
[0212] Now, turning back to the IBC-related signaling, in some embodiments, for the current block, a flag is first sent in the bitstream to indicate whether IBC is enabled for the current block. This flag can be signaled at a higher level such as CTU, CU, sequence, slice, or picture level. Then, if the current block is in the IBC mode (either as a mode separate from the inter prediction mode or as an integral part of the inter prediction mode), reference blocks can be searched and the corresponding BV can be determined by the encoder. For BV prediction, the BV difference can be derived in the decoder by subtracting the predicted BV from the current BV, and then the BV difference can be classified into multiple types (e.g., 4 types) according to the horizontal and vertical components of the BV difference. The BV different type information can be further signaled in the bitstream, and then the BV differences of the two (horizontal and vertical) components can be signaled. In some example embodiments, a set of high-level syntax flags is further included in the bitstream and is used to indicate the allowable local and / or non-local reference ranges for IBC prediction. This set of flags can be signaled at various levels (e.g., CTU, CU, sequence, slice, or picture level).
[0213] For example, a syntax flag called global_ibc_flag can be used to turn on / off the non-local based region, while another syntax flag called local_ibc_flag can be used to turn on / off the local based region for IBC prediction. These two syntax flags can be controlled independently of each other. In other words, these flags can have any combination of flag values. Each of these flags can be signaled at the same different levels. In one example, when both flags are turned off, in fact, IBC is disabled. In this case, where the non-local IBC flag and the local IBC flag are signaled independently for a specific level, the IBC enable flag at that level (e.g., picture level or sequence level) does not need to be signaled in the bitstream.
[0214] In some example embodiments, the non-local IBC syntax flag global_ibc_flag and the local IBC syntax flag local_ibc_flag can be configured to have certain dependencies. For example, the non-local global_ibc_flag can be signaled first. Depending on the value of the local_ibc_flag, the local_ibc_flag can be signaled or inferred. When the global_ibc_flag is equal to 0 (indicating not in use), in combination with the IBC being signaled as being in use (e.g., the above advanced IBC enable syntax), the local_ibc_flag can be inferred to be 1 (indicating in use) instead of being signaled. In this example, one or both of the local and non-local flags are only signaled when the IBC enable flag is on. Otherwise, neither flag needs to be signaled.
[0215] In some example embodiments, when using a non-local-based IBC search range, e.g., for a picture, the loop filter will be disabled for the same picture. On the other hand, if only using a local-based IBC search range (but not a non-local-based IBC search range), the loop filter can be used for the same picture. Thus, under the condition of using IBC without using a non-local-based IBC search range, the enable flag of the loop filter in the IBC is signaled. In other words, if the above other flags indicate not using a non-local-based IBC, the loop filter enable flag can be signaled. The loop filter enable flag will indicate whether the local-based IBC should call the loop filter. Otherwise, when not using IBC or only using non-local IBC, it is inferred that loop filtering is disabled and the loop filter enable flag does not need to be signaled. Specifically, the condition for signaling the use of the loop filter can be the value of the global_ibc_flag. When the global_ibc_flag is enabled (on) (meaning using non-local IBC reference search), the enable flag of the loop filter for the picture can be inferred to be 0 (or off) and does not need to be signaled.
[0216] In the manner described above, the various flags or syntax elements above can indicate or signal the IBC reference mode of the current block, either individually or in various combinations. The IBC reference mode indicates how the IBC prediction block can access the local search region and the non-local search region. For example, a combination of these flags or syntax elements can indicate that only the CTU or SB in the local search region can be used for IBC reference, thus being the local reference IBC mode. For another example, a combination of these flags or syntax elements can indicate that only the CTU or SB in the non-local search region can be used for IBC reference, thus being the non-local IBC reference mode. For yet another example, a combination of these flags or syntax elements can indicate that the CTU or SB in both the local and non-local search regions can be used for IBC reference, thus being the local and non-local reference IBC mode. The decoder can extract these syntax elements independently or based on the dependencies as described to determine the IBC reference mode, thereby obtaining information about the search region for determining the IBC reference block.
[0217] Figure 27 FIG. 2700 is a flowchart of an example method that follows the basic principles of the IBC implementation manner described above. The example method flow starts at 2701. At 2710, at least one syntax element is extracted from the video stream, and the at least one syntax element is associated with the intra block copy (IBC) prediction of the video block. At 2720, the IBC reference mode for the IBC prediction of the video block is determined, and the IBC reference mode can include one of the following: no IBC mode, local reference IBC mode, non-local reference IBC mode, and local and non-local reference IBC mode. At 2730, the reconstructed samples of the video block are generated from the video stream based on the IBC reference mode. The example method flow ends at 2799.
[0218] In the embodiments and implementations of the embodiments of the present application, any steps and / or operations can be combined or set in any quantity or order as needed. Two or more of the steps and / or operations can be executed in parallel. The embodiments and implementations of the embodiments of the present application can be used alone or in any order combination. In addition, each method (or implementation), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the embodiments of the present application can be applied to luminance blocks or chrominance blocks. The term "block" can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term "block" herein can also be used to refer to a transform block. In the following items, when referring to the block size, it can refer to the width or height of the block, or the maximum value of the width and height, or the minimum value of the width and height, or the area size of the block (width * height), or the aspect ratio (width: height, or height: width).
[0219] The above-described techniques may be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 28 FIG. shows a computer system (2800) suitable for implementing certain embodiments of the disclosed subject matter.
[0220] The computer software may be encoded using any suitable machine code or computer language, and any suitable machine code or computer language may be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.
[0221] The instructions may be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0222] Figure 28 The components of the computer system (2800) shown in are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should also not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (2800).
[0223] The computer system (2800) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as the following: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.
[0224] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (2801), mouse (2802), touchpad (2803), touch screen (2810), data glove (not shown), joystick (2805), microphone (2806), scanner (2807), camera (2808).
[0225] A computer system (2800) may include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include haptic output devices (such as haptic feedback of a touch screen (2810), a data glove (not shown), or a joystick (2805), but may also be haptic feedback devices that are not input devices), audio output devices (such as speakers (2809), headphones (not depicted)), visual output devices (such as a screen (2810) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input function, each with or without haptic feedback function - some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).
[0226] The computer system (2800) may also include human-accessible storage devices and their associated media, for example, optical media including CD / DVD ROM / RW (2820) with media such as CD / DVD (2821), thumb drives (2822), removable hard disk drives or solid state drives (2823), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.
[0227] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0228] The computer system (2800) may also include an interface (2854) to one or more communication networks (2855). The network can be, for example, a wireless network, a wired network, an optical network. The network can also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CAN bus, and so on. Some networks typically require an external network interface adapter connected to certain general data ports or peripheral buses (2849) (e.g., the USB port of the computer system (2800)); as described below, other network interfaces are typically integrated into the core of the computer system (2800) by connecting to the system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (2800) can communicate with other entities using any of these networks. Such communication can be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CAN bus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.
[0229] The above-mentioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the core (2840) of the computer system (2800).
[0230] The core (2840) can include one or more central processing units (CPUs) (2841), a graphics processing unit (GPU) (2842), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2843), a hardware accelerator (2844) for certain tasks, a graphics adapter (2850), etc. These devices, as well as a read-only memory (ROM) (2845), a random access memory (2846), and internal mass storage such as an internal non-user accessible hard disk drive, SSD, etc. (2847) can be connected through a system bus (2848). In some computer systems, the system bus (2848) can be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus of the core (2848) or connected to the system bus of the core (1848) through a peripheral bus (2849). In one example, a screen (2810) can be connected to the graphics adapter (2850). The architecture of the peripheral bus includes PCI, USB, etc.
[0231] A CPU (2841), a GPU (2842), an FPGA (2843), and an accelerator (2844) can execute certain instructions, which can be combined to form the above computer code. The computer code can be stored in a ROM (2845) or a RAM (2846). Transitional data can also be stored in the RAM (2846), while permanent data can be stored, for example, in an internal mass storage (2847). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more of the following: a CPU (2841), a GPU (2842), a mass storage (2847), a ROM (2845), a RAM (2846), etc.
[0232] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of the embodiments of this application, or the medium and the computer code can be of the type well-known and available to those skilled in the art of computer software.
[0233] As a non-limiting example, a computer system having an architecture (2800), particularly a core (2840), can provide functionality due to one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software included in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the core (2840), such as on-core mass memory (2847) or ROM (2845). The software implementing the embodiments of the present application can be stored in such devices and executed by the core (2840). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (2840), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (2846) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator (2844)), which can replace the software or operate in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or include both. Embodiments of the present application include any suitable combination of hardware and software.
[0234] Although embodiments of the present application have described multiple exemplary embodiments, there are modifications, permutations, and various replacement equivalents that fall within the scope of embodiments of the present application. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of embodiments of the present application and thus fall within the spirit and scope of embodiments of the present application.
[0235] Appendix A: Acronyms
[0236] JEM: Joint Exploration Model
[0237] VVC: Versatile Video Coding
[0238] BMS: Benchmark Set
[0239] MV: Motion Vector
[0240] HEVC: High Efficiency Video Coding
[0241] SEI: Supplementary Enhancement Information
[0242] VUI: Video Usability Information
[0243] GOP: Group of Pictures
[0244] TU: Transform Unit
[0245] PU: Prediction Unit
[0246] CTU: Coding Tree Unit
[0247] CTB: Coding Tree Block
[0248] PB: Prediction Block
[0249] HRD: Hypothetical Reference Decoder
[0250] SNR: Signal-to-Noise Ratio
[0251] CPU: Central Processing Unit
[0252] GPU: Graphics Processing Unit
[0253] CRT: Cathode Ray Tube
[0254] LCD: Liquid Crystal Display
[0255] OLED: Organic Light-Emitting Diode
[0256] CD: Compact Disc
[0257] DVD: Digital Versatile Disc
[0258] ROM: Read-Only Memory
[0259] RAM: Random Access Memory
[0260] ASIC: Application-Specific Integrated Circuit
[0261] PLD: Programmable Logic Device
[0262] LAN: Local Area Network
[0263] GSM: Global System for Mobile Communications
[0264] LTE: Long-Term Evolution
[0265] CANBus: Controller Area Network Bus
[0266] USB: Universal Serial Bus
[0267] PCI: Peripheral Component Interconnect
[0268] FPGA: Field-Programmable Gate Array
[0269] SSD: Solid State Drive
[0270] IC: Integrated Circuit
[0271] HDR: High Dynamic Range
[0272] SDR: Standard Dynamic Range
[0273] JVET: Joint Video Exploration Team
[0274] MPM: Most Probable Mode
[0275] WAIP: Wide Angle Intra Prediction
[0276] CU: Coding Unit
[0277] PU: Prediction Unit
[0278] TU: Transform Unit
[0279] CTU: Coding Tree Unit
[0280] PDPC: Position Dependent Prediction Combination
[0281] ISP: Intra Sub Partitions
[0282] SPS: Sequence Parameter Set
[0283] PPS: Picture Parameter Set
[0284] APS: Adaptive Parameter Set
[0285] VPS: Video Parameter Set
[0286] DPS: Decoding Parameter Set
[0287] ALF: Adaptive Loop Filter
[0288] SAO: Sample Adaptive Offset
[0289] CC-ALF: Cross Component Adaptive Loop Filter
[0290] CDEF: Constrained Directional Enhancement Filter
[0291] CCSO: Cross Component Sample Offset
[0292] LSO: Local Sample Offset
[0293] LR: Loop Restoration Filter
[0294] AV1: Alliance for Open Media Video 1
[0295] AV2: Alliance for Open Media Video 2
[0296] RPS: Reference Picture Set
[0297] DPB: Decoded Picture Buffer
[0298] MMVD: Merge Mode with Motion Vector Difference
[0299] IntraBC or IBC: Intra Block Copy
[0300] BV: Block Vector
[0301] BVD: Block Vector Difference
[0302] RSM: Reference Sample Memory.
Claims
1. A method for encoding a video block, characterized in that The method includes: Receiving the video block; Searching for a reference block for IBC prediction of the video block from the current video frame; Determining the position of the reference block relative to the video block, wherein the types of the reference block include: local reference IBC mode, non-local reference IBC mode; Determining an IBC reference mode for IBC prediction of the video block based on the position of the reference block, wherein the IBC reference mode is one selected from a plurality of predefined IBC reference modes, the video block belongs to a current IBC prediction unit including a plurality of video blocks, and the plurality of predefined IBC reference modes include: No IBC mode, local reference IBC mode, non-local reference IBC mode, and local and non-local reference IBC mode, wherein the local reference IBC mode is used to characterize that the reference block for IBC prediction of the video block includes reference samples in a predefined adjacent unit group of the current IBC prediction unit or a reconstructed video block in the current IBC prediction unit, the non-local reference IBC mode is used to characterize that the reference block for IBC prediction of the video block includes reference samples not adjacent to the current IBC prediction unit in the coding direction of the current IBC prediction unit, and the local and non-local reference IBC is used to characterize that the reference block for IBC prediction of the video block includes the reference samples in the adjacent unit group and the non-adjacent reference samples; and Encoding the video block based on the reference block in the video stream and using the identifier of the IBC reference mode as an identifier in at least one syntax element of the video stream.
2. The method according to claim 1, wherein The predefined adjacent unit group includes a single left adjacent unit of the current IBC prediction unit.
3. The method according to claim 1, wherein For the local reference IBC mode, the reference samples for IBC prediction are maintained in a on-chip reference sample memory RSM of a fixed size.
4. The method according to claim 3, characterized in that, The fixed size of the RSM corresponds to the size of one IBC prediction unit.
5. The method according to claim 4, wherein A first part of the RSM includes corresponding samples of the reconstructed video blocks in the current IBC prediction unit; and A second part of the RSM includes corresponding reconstructed samples from the predefined adjacent unit group.
6. The method according to claim 5, characterized in that, The method further includes: replacing the reconstructed samples of the adjacent units in the RSM corresponding to the video block in the current IBC prediction unit with the reconstructed samples of the video block.
7. The method according to claim 3, wherein The current IBC prediction unit is divided into a predefined partition group; The video block is the first encoded block to be reconstructed in the current partition of the predefined partition group; And The method further includes: before reconstructing the video block, resetting the partition of the RSM corresponding to the current partition to be unavailable for IBC reference.
8. The method according to claim 1, wherein The at least one syntax element includes: a first flag and a second flag, the first flag is used to indicate that local IBC reference is enabled when set, and the second flag is used to indicate that non-local reference IBC is enabled when set.
9. The method according to claim 8, wherein The method further includes: In response to the first flag being set and the second flag not being set, determine that the IBC reference mode is the local reference IBC mode; In response to the second flag being set and the first flag not being set, determine that the IBC reference mode is the non-local reference IBC mode; In response to both the first flag and the second flag being set, determine that the IBC reference mode is the local and non-local reference IBC mode; And In response to both the first flag and the second flag not being set, determine that the IBC reference mode is the no-IBC mode.
10. The method according to claim 8, wherein Signal the first flag and the second flag in the video stream at the coded block level, coded unit level, coding tree unit level, slice level, picture level, or sequence level.
11. The method according to claim 1, wherein The at least one syntax element includes a first flag for indicating whether IBC is used for the video block.
12. The method according to claim 11, wherein The method further includes: In response to the first flag indicating that IBC is not used for the video block, determine that the IBC reference mode is the no-IBC mode.
13. The method according to claim 12, wherein The method further includes: In response to the first flag indicating that IBC is used for the video block, further extract a second flag for indicating whether non-local IBC reference is used as part of the at least one syntax element; In response to the second flag indicating that non-local IBC reference is not used, infer that the IBC reference mode of the video block is the local reference IBC mode.
14. The method according to claim 13, wherein The method further includes: In response to the second flag indicating that non-local IBC reference is used, further extract a third flag for indicating whether local IBC reference is used as part of the at least one syntax element; In response to the third flag indicating that local IBC reference is used, determine that the IBC reference mode is the local and non-local reference IBC mode; And In response to the third flag indicating that local IBC reference is not used, determine that the IBC reference mode is the non-local reference IBC mode.
15. The method according to claim 1, characterized in that The method further includes: When the IBC reference mode is the local reference IBC mode, enable the loop filter process; and When the IBC reference mode is the non-local reference IBC mode or the local and non-local reference IBC mode, disable the loop filter process.
16. The method according to claim 15, wherein The method further includes: Derive whether to enable the loop filter process from the at least one syntax element for signaling the IBC reference mode.
17. A video processing device for encoding a video block, characterized in that, Include a memory for storing computer instructions and a processor for executing the computer instructions to perform the method according to any one of claims 1 to 16.
18. A non-transitory computer-readable medium, characterized in that, For storing instructions that, when executed by a computer for video coding, cause the computer to perform the method according to any one of claims 1 to 16.
19. A method for storing a video stream, characterized in that, The video stream is obtained by encoding according to the method according to any one of claims 1 to 16.