Video processing method, device, and program
Cross-component level reconstruction methods refine transform coefficients to enhance video encoding and decoding efficiency, addressing inefficiencies in intra-prediction and motion compensation, thereby reducing bandwidth and storage needs.
Patent Information
- Application Number
- JP2023551157
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2022-10-21
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing video encoding and decoding technologies face inefficiencies in reducing redundancy and distortion, particularly in intra-prediction and motion compensation methods, leading to suboptimal compression ratios and increased bandwidth requirements.
Implementing cross-component level reconstruction methods that refine transform coefficients using offset values derived from co-located blocks of different color components, enhancing the video decoding and encoding process.
Improves compression efficiency by reducing redundancy and distortion, leading to lower bandwidth and storage requirements while maintaining video quality.
Smart Images

Figure 0007724038000007 
Figure 0007724038000008 
Figure 0007724038000009
Abstract
Description
[Technical Field]
[0001] [Incorporated by reference] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 273,789, filed October 29, 2021, entitled "PRIMARY TRANSFORMS FOR CROSS-COMPONENT LEVEL RECONSTRUCTION," and U.S. Patent Application No. 17 / 966,390, filed October 14, 2022, entitled "PRIMARY TRANSFORMS FOR CROSS-COMPONENT LEVEL RECONSTRUCTION," both of which are incorporated herein by reference in their entireties.
[0002] [Technical field] This disclosure relates generally to a set of advanced video encoding / decoding techniques, and more specifically to primary transforms for offset-based cross-component methods that refine transform coefficients before or after inverse quantization. [Background technology]
[0003] This background description provided herein is intended to generally present the context for the present disclosure. The work of the presently named inventors, to the extent that that work is described in this background section, and any aspects of the description that may not otherwise qualify as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of pictures, each with spatial dimensions of, for example, 1920 x 1080 luminance samples and associated full or subsampled chrominance samples. The sequence of pictures can have a fixed or variable picture rate (alternatively called frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920 x 1080, a frame rate of 60 frames per second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.
[0005] One goal of video encoding and decoding can be the reduction of redundancy in an uncompressed input video signal through compression. Compression can help reduce the bandwidth and / or storage space requirements by more than two orders of magnitude in some cases. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal by the decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully preserved during encoding and is not fully recoverable during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to make the reconstructed signal useful for the intended application despite some information loss. In the case of video, lossy compression is widely used in many applications. The amount of acceptable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect different distortion tolerances, with higher tolerable distortion generally allowing for coding algorithms that result in higher losses and higher compression ratios.
[0006] Video encoders and decoders can utilize techniques from several broad categories and steps, including, for example, motion compensation, Fourier transforms, quantization, and entropy coding.
[0007] Video codec techniques may include a technique known as intra-coding, in which sample values are represented without reference to samples or other data from a previously reconstructed reference picture. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be called an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded bitstream and video session or as a still image. The samples of a block after intra-prediction may then undergo a transform to the frequency domain, and the transform coefficients so produced may be quantized before entropy coding. Intra-prediction can refer to a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are required for a given quantization step size to represent the block after entropy coding.
[0008] Conventional intra-coding, such as that known from the MPEG-2 generation of coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to encode / decode blocks based on surrounding sample data and / or metadata obtained during the encoding and / or decoding of blocks of data that are spatially adjacent to the data and that precede it in decoding order. Such techniques are hereinafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses reference data only from the current picture being reconstructed, and not from other reference pictures.
[0009] Many different forms of intra-prediction may exist. When more than one such technique is available for a given video coding technique, the technique in use may be referred to as an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In certain cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and the intra-coding parameters of a block of video may be coded independently or collectively included in a mode codeword. Because which codeword to use for a given mode, sub-mode, and / or parameter combination may affect the coding efficiency gain through intra-prediction, entropy coding techniques may be used to convert the codeword into a bitstream.
[0010] Certain modes of intra prediction were introduced by H.264, refined in H.265, and further refined in newer coding techniques such as Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). In general, for intra prediction, a predictor block may be formed using neighboring sample values that are available. For example, available values of a particular set of neighboring samples along a particular direction and / or line may be copied into the predictor block. A reference to the direction in use may be coded into the bitstream or may itself be predicted.
[0011] Referring to FIG. 1A, the bottom right shows a subset of nine predictor directions specified by the 33 possible intra-predictor directions in H.265 (corresponding to the 33 angle modes among the 35 intra-modes specified in H.265). The point where the arrows converge (101) corresponds to the sample being predicted. The arrows represent the direction in which neighboring samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more neighboring samples to the upper right and at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more neighboring samples to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012] Still referring to FIG. 1A , a square block (104) of 4×4 samples (indicated by a thick dashed line) is depicted at the top left. The square block (104) contains 16 samples, each labeled with “S,” its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. Also shown are exemplary reference samples that follow a similar numbering scheme. The reference samples are labeled with “R” and their Y position (e.g., row index) and X position (column index) relative to the block (104). Both H.264 and H.265 use predicted samples that are adjacent to the block being reconstructed.
[0013] Intra-picture prediction of block 104 may begin by copying reference sample values from neighboring samples according to the signaled prediction direction. For example, suppose the coded video bitstream includes signaling for this block 104 indicating the prediction direction of arrow (102), i.e., that the sample is predicted from one or more prediction samples to the upper right at an angle of 45 degrees from the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, the values of multiple reference samples may be combined, for example through interpolation, to calculate the reference sample, especially when the direction is not equally divisible by 45 degrees.
[0015] The number of possible directions continues to grow as video coding technology continues to evolve. In H.264 (2003), for example, nine different directions are available for intra prediction. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions as of the time of this disclosure. Empirical studies have been conducted to help identify the most appropriate intra prediction directions, and specific techniques in entropy coding have been used to encode those most appropriate directions with a small number of bits while accepting a specific bit penalty for the direction. Furthermore, the direction itself can sometimes be predicted from neighboring directions used in intra prediction of decoded neighboring blocks.
[0016] FIG. 1B shows a schematic diagram (180) representing 65 intra prediction directions according to JEM to illustrate the increasing number of prediction directions in various coding techniques that have been developed over time.
[0017] The method of mapping bits representing intra-prediction directions to prediction directions in the coded video bitstream can vary from one video coding technique to another and can range, for example, from a simple direct mapping of prediction directions to intra-prediction modes to codewords to complex adaptive schemes including most-probable modes and similar techniques. In all cases, however, there may be certain directions for intra-prediction that are statistically less likely to occur in the video content than certain other directions. Because a goal of video compression is redundancy reduction, these less likely directions may be represented by more bits than more likely directions in a well-designed video coding technique.
[0018] Inter-picture prediction or inter-prediction may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used for predicting a newly reconstructed picture or picture part (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or may have three dimensions, with the third dimension being an indication of the reference picture in use (akin to the temporal dimension).
[0019] In some video compression techniques, the current MV applicable to a particular area of sample data can be predicted from other MVs, e.g., from other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and precede the current MV in decoding order. Doing so can significantly reduce the overall amount of data required to code the MV by relying on correlated MVs to remove redundancy, thereby increasing compression efficiency. For example, when coding an input video signal obtained from a camera (known as natural video), MV prediction can work effectively because there is a statistical possibility that areas larger than the area to which a single MV is applicable will move in similar directions in the video sequence, and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of neighboring areas. As a result, the actual MV of a given area is similar or identical to the MV predicted from surrounding MVs. Such an MV can then be represented, after entropy coding, with fewer bits than would be used if the MV were coded directly rather than predicted from neighboring MVs. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example due to rounding errors when computing the predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms specified by H.265, a technique hereinafter referred to as “spatial merge” will be described in this specification.
[0021] Specifically, referring to Figure 2, a current block (201) has samples that the encoder found during the motion search process to be predictable from a previous block of the same size that has been spatially shifted. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent reference picture (in decoding order) using the MV associated with any one of five surrounding samples denoted A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, MV prediction can use predictors from the same reference picture as neighboring blocks. Summary of the Invention
[0022] Aspects of the present disclosure provide cross-component methods and apparatus for selecting and employing transforms for transform coefficients refined at a cross-component level in video processing. In some example implementations, a video decoding method is disclosed. The method includes receiving a bitstream of video blocks including a first transform block of a first color component and a second transform block of a second color component, where the first transform block and the second transform block are collocated blocks; obtaining the first transform block of the first color component and the second transform block of the second color component from the bitstream of video blocks; determining a first flag indicating that all transform coefficients in the first transform block are zero; and performing cross-component level reconstruction. The method may include determining a second flag indicating that CCLR (Corrective Color Reconstruction) is applied to the first transform block; refining one or more of the transform coefficients in the first transform block by adding one or more offset values in response to determining that CCLR is applied to the first transform block to obtain a refined first transform block, the one or more offset values being derived based on transform coefficients in the second transform block that are co-located with one or more of the transform coefficients in the first transform block; determining a target transform kernel for the refined first transform block; performing an inverse transform on the refined first transform block based on the target transform kernel to obtain a target block; and reconstructing at least a first color component of the video block based on the target block.
[0023] Aspects of the present disclosure also provide a video encoding or decoding device or apparatus including circuitry configured to perform any of the implementations of the above methods.
[0024] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform a method of video decoding and / or encoding.
[0025] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0026] [Figure 1A] 1 illustrates a schematic diagram of an exemplary subset of intra-prediction direction modes. [Figure 1B] 1 shows an illustrative diagram of an exemplary intra-prediction direction. [Figure 2] 1 illustrates a schematic diagram of a current block and its surrounding spatial merge candidates for motion vector prediction in one example. [Figure 3] 1 shows a simplified block diagram of a communication system (300) according to an example embodiment. [Figure 4] 4 shows a simplified block diagram of a communication system (400) according to an example embodiment. [Figure 5] 1 shows a schematic block diagram of a decoder according to an example embodiment; [Figure 6] 1 shows a schematic block diagram of an encoder according to an example embodiment; [Figure 7] 1 shows a block diagram of a video encoder according to another exemplary embodiment; [Figure 8] 10 shows a block diagram of a video decoder according to another exemplary embodiment; [Figure 9] 1 illustrates a coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 10] 10 illustrates another coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 11] 10 illustrates another coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 12] 10 illustrates another coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 13]1 illustrates a scheme for partitioning a coding block into multiple transform blocks and a coding order of the transform blocks, according to an exemplary embodiment of the present disclosure. [Figure 14] 10 illustrates another scheme for dividing a coding block into multiple transform blocks and the coding order of the transform blocks, according to an exemplary embodiment of the present disclosure. [Figure 15] 10 illustrates another scheme for partitioning a coding block into multiple transform blocks according to an exemplary embodiment of the present disclosure. [Figure 16] 1 illustrates a planar rotation transform according to an exemplary embodiment of the present disclosure. [Figure 17] 10A-10C illustrate various DCT-2, DCT-4 partial butterfly lookup tables according to exemplary embodiments of the present disclosure. [Figure 18] 1 illustrates a DST-7 partial butterfly lookup table according to an exemplary embodiment of the present disclosure. [Figure 19] 1 illustrates an exemplary line graph transformation (LGT) according to an exemplary embodiment of the present disclosure. [Figure 20] 1 shows a flowchart of a method according to an example embodiment of the present disclosure. [Figure 21] 1 shows a schematic diagram of a computer system according to an example embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0027] Figure 3 illustrates a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes multiple terminal devices that can communicate with each other, e.g., via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) may perform unidirectional transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to recover the video pictures, and display the video pictures according to the recovered video data. One-way data transmission may be implemented in media serving applications, etc.
[0028] In another example, the communication system 300 includes a second pair of terminal devices 330 and 340 that perform bidirectional transmission of coded video data, such as may be performed during video conferencing applications. For the bidirectional transmission of data, in the example, each of the terminal devices 330 and 340 may code video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices 330 and 340 over the network 350. Each of the terminal devices 330 and 340 may also receive coded video data transmitted by the other of the terminal devices 330 and 340, decode the coded video data to recover video pictures, and display the video pictures on an accessible display device in accordance with the recovered video data.
[0029] In the example of FIG. 3 , terminal devices 310, 320, 330, and 340 may be implemented as a server, a personal computer, and a smartphone, although the applicability of the underlying principles of the present disclosure may not be so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated videoconferencing equipment, and / or the like. Network 350 represents any number or type of network that conveys coded video data between terminal devices 310, 320, 330, and 340, including, for example, wireline and / or wireless communication networks. Communication network 350 may exchange data over circuit-switched, packet-switched, and / or other types of channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 350 may be irrelevant to the operation of the present disclosure unless explicitly described herein.
[0030] 4 illustrates the placement of a video encoder and a video decoder in a video streaming environment as an example application of the disclosed subject matter, which may be similarly applicable to other video-enabled applications including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0031] A video streaming system may include a video capture subsystem (413) that may include a video source (401), such as a digital camera, that generates a stream of uncompressed video pictures or images (402). In an example, the stream of video pictures (402) includes samples recorded by the digital camera of the video source (401). The stream of video pictures (402) is represented by a bold line to emphasize its high data volume compared to the encoded video data (404) (or coded video bitstream), and may be processed by electronics (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data 404 (or encoded video bitstream 404), represented by thin lines to emphasize its lower data volume compared to the uncompressed video picture stream 402, may be stored on a streaming server 405 for future use or for direct downstream streaming to a video device (not shown). One or more streaming client subsystems, such as the client subsystems 406 and 408 of FIG. 4, can access the streaming server 405 to retrieve copies 407 and 409 of the encoded video data 404. The client subsystem 406 may include a video decoder 410, for example, in an electronic device 430. The video decoder 410 decodes the incoming copy of the encoded video data 407 and generates an outgoing stream of video pictures 411 that is uncompressed and can be rendered on a display 412 (e.g., a display screen) or other rendering device (not shown). The video decoder (410) may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data 404, 407, and 409 (e.g., video bitstreams) may be encoded according to a particular video coding / compression standard. An example of such a standard is ITU-T Recommendation H.265. In an example, the video coding standard under development is colloquially known as Versatile Video Coding (VVC). The disclosed subject matter may be used in conjunction with VVC and other video coding standards.
[0032] The electronics 420 and 430 may include other components (not shown). For example, the electronics 420 may include a video decoder (not shown), and the electronics 430 may also include a video encoder (not shown).
[0033] 5 shows a block diagram of a video decoder (510) according to any of the following embodiments of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of the video decoder (410) in the example of FIG. 4.
[0034] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or other embodiments, one coded video sequence may be decoded at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequences may be received from a channel (501), which may be a storage device storing the coded video data or a hardware / software link to a streaming source transmitting the coded video data. The receiver (531) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be located between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, it may be separate (not shown) outside the video decoder (510). In still other applications, there may be a buffer memory (not shown) outside the video decoder (510), for example, to combat network jitter, or there may be another additional buffer memory (515) within the video decoder (510), for example, to manipulate playback timing. When the receiver (531) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed, or may be small.For use with best-effort packet networks such as the Internet, a buffer memory (515) of sufficient size may be required, which may be relatively large. Such a buffer memory may be adaptively sized and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).
[0035] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and information for controlling a rendering device, such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronics (530) but may be coupled to the electronics (530) as shown in FIG. 5. Control information for the rendering device may take the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the coded video sequence received by the parser (520). The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser (520) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.
[0036] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to generate symbols (521).
[0037] The reconstruction of the symbols (521) can involve many different processing or functional units, depending on the type of coded video picture or portion thereof (e.g., inter and intra picture, inter and intra block) and other factors. The units involved, and how they participate, can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following processing or functional units is not shown for simplicity.
[0038] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units, which are described below. In an actual implementation operating under commercial constraints, many of these functional units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of clearly describing the various functions of the disclosed subject matter, a conceptual subdivision into functional units will be adopted hereinafter in this disclosure.
[0039] The first unit may include a scalar / inverse transform unit (551), which may receive quantized transform coefficients as symbols (521) from the parser (520) along with control information including information indicating what inverse transform to use, block size, quantization coefficients / parameters, quantization scaling matrices, etc. The scalar / inverse transform unit (551) may output blocks containing sample values that can be input to an aggregator (555).
[0040] In some cases, the output samples of the scaler / inverse transformer (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block of the same size and shape as the block being reconstructed using surrounding block information that has already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558), for example, buffers a partially reconstructed and / or fully reconstructed current picture. In some implementations, the aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transformer unit (551).
[0041] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, and potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (553) may access the reference picture memory (557) to fetch samples used for inter-picture prediction. After motion-compensating the fetched samples according to the symbols (521) related to the block, the samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (the output of unit 551 may also be referred to as a residual sample or residual signal) to generate output sample information. The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (553), for example, in the form of symbols (521), which may have X and Y components (shift) and a reference picture component (time). Motion compensation may also include interpolation of sample values fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, and may also be associated with motion vector prediction mechanisms, etc.
[0042] The output samples of the aggregator (555) may undergo various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques, which are controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, or may even respond to previously constructed loop-filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as described in more detail below.
[0043] The output of the loop filter unit (556) can be a sample stream that can be output to a rendering device (512) and further stored in a reference picture memory (557) for use in future inter-picture prediction.
[0044] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and any unused current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0045] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard, such as ITU-T Recommendation H.265. A coded video sequence may conform to the syntax prescribed by the video compression technique or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and a profile documented in the video compression technique or standard. Specifically, a profile may select specific tools from all tools available in the video compression technique or standard as the only tools available for use under that profile. For standard compliance, the complexity of the coded video sequence may be within the boundaries defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the levels may in some cases be further restricted through Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.
[0046] In some exemplary embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0047] 6 shows a block diagram of a video encoder (603) according to an exemplary embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) may be used in place of the video encoder (403) in the example of FIG. 4.
[0048] The video encoder (603) may receive video samples from a video source (601) (not part of the electronics (620) in the example of FIG. 6) that may capture video images to be coded by the video encoder (603). In other examples, the video source (601) may be implemented as part of the electronics (620).
[0049] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, XYZ, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures or images that, when viewed in sequence, impart motion. The picture itself may be organized as a spatial array of pixels, each of which may have one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. This specification will focus on samples hereafter.
[0050] According to some example embodiments, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Imposing an appropriate coding rate constitutes one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units, as described below. Coupling is not shown for simplicity. Parameters set by the controller (650) may include parameters related to rate control (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0051] In some example embodiments, the video encoder (603) may be configured to operate in a coding loop. As an overly simplified description, in an example, the coding loop may include a source coder (630) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in a manner similar to that which a (remote) decoder would generate (even if the embedded decoder (633) processes the video stream coded by the source coder (630) without entropy coding, in that any compression between the symbols and the coded video bitstream can be lossless in the video compression techniques contemplated in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding the symbol stream yields bit-exact results independent of the location of the decoder (local or remote), the contents in the reference picture memory (634) are also bit-perfect between the local and remote encoders. That is, the predictive portion of the encoder "sees" exactly the same sample values for reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, e.g., due to channel errors) is used to improve coding quality.
[0052] The operation of the "local" decoder (633) can be the same as a "remote" decoder, such as the video decoder (510), already described in detail above in conjunction with Figure 5. Referring also momentarily to Figure 5, however, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be implemented entirely in the local decoder (633) within the encoder, given the availability of symbols and the encoding / decoding of symbols into a coded video sequence by the entropy coder (645) and parser (520) can be lossless.
[0053] At this point, it can be observed that any decoder technique, with the exception of parsing / entropy decoding, which may only be present in the decoder, may necessarily need to be present in roughly the same functional form in the corresponding encoder. For this reason, the disclosed subject matter may at times focus on decoder operations related to the decoding portion of the encoder. Thus, descriptions of encoder techniques may be omitted, as they are the inverse of the decoder techniques, which are described generically. Only in certain areas or aspects will a more detailed description of the encoder be given below.
[0054] In operation, in some embodiments, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine (632) codes color channel differences (or residuals) between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture. The term "residue" or its adjective "residual" are sometimes used interchangeably.
[0055] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (630). The operation of the coding engine (632) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a copy of the source video sequence, with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that would be obtained by a far-end (distant) video decoder (without transmission errors).
[0056] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for specific metadata, such as reference picture motion vectors, block shapes, or sample data (as candidate reference pixel blocks) that can serve as suitable prediction references for the new picture. The predictor (635) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).
[0057] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0058] The output of all the above functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by lossless compression of the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0059] The transmitter (640) may buffer the coded video sequence produced by the entropy coder (645) to prepare it for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0060] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0061] An Intra Picture (I-picture) may be a picture that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow various types of Intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of such variations of I-pictures and their respective applications and characteristics.
[0062] A Predictive Picture (P-picture) may be a picture that can be coded and decoded by intra- or inter-prediction using at most one motion vector and reference index to predict the sample values of each block.
[0063] A Bi-directionally Predictive Picture (B-picture) may be a picture that can be coded and decoded by intra- or inter-prediction using at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0064] A source picture is generally spatially subdivided into multiple sample coding blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to each picture of the 'blocks'. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures. Source pictures or intermediate processed pictures may also be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same method, as described in more detail below.
[0065] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Accordingly, the coded video data may conform to a syntax defined by the video coding technique or standard being used.
[0066] In some example embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0067] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits temporal or other correlation between pictures. For example, a particular picture being encoded / decoded, called the current picture, may be partitioned into blocks. A block in the current picture may be coded by a vector, called a motion vector, if it is similar to a reference block in a reference picture that was previously coded in the video and is still buffered. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0068] In some example embodiments, bi-prediction techniques may be used for inter-picture prediction. According to such bi-prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture, both of which precede the current picture in the video in decoding order (but may be past or future, respectively, in display order). A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. The block is jointly predictable by a combination of the first and second reference blocks.
[0069] Furthermore, merge mode techniques may be used in inter-picture prediction to improve coding efficiency.
[0070] According to some exemplary embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, a picture in a sequence of video pictures is partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. In general, a CTU may include three parallel coding tree blocks (CTBs), i.e., one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels or four CUs of 32x32 pixels. One or more of the 32x32 blocks may be further divided into four CUs of 16x16 pixels. In some exemplary embodiments, each CU may be analyzed during encoding to determine a prediction type for the CU among various prediction types, such as an inter prediction type or an intra prediction type. A CU may be divided into one or more Prediction Units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. The division of a CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. A luma or chroma PB may include a matrix of sample values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0071] 7 shows a diagram of a video encoder (703) according to another exemplary embodiment of the disclosure. The video encoder (703) is configured to receive a processed block of sample values (e.g., a predictive block) in a current video picture included in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. The example video encoder (703) may be used in place of the example video encoder (403) of FIG. 4.
[0072] For example, the video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8x8 samples. The video encoder (703) then determines, for example, using rate-distortion optimization (RDO), whether the processing block is best coded in intra-mode, inter-mode, or bi-predictive mode. If it is determined that the processing block is coded in intra-mode, the video encoder (703) may use intra-prediction techniques to encode the processing block into a coded picture; if it is determined that the processing block is coded in inter-mode or bi-predictive mode, the video encoder (703) may use inter-prediction or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In some exemplary embodiments, merge mode may be used as a sub-mode of inter-picture prediction in which a motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other exemplary embodiments, there may be motion vector components applicable to the current block. Accordingly, the video encoder (703) may include components not explicitly shown in FIG. 7, such as a mode decision module, to determine the prediction mode of a processing block.
[0073] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled in the exemplary arrangement shown in Figure 7.
[0074] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures in a standard manner), generate inter-prediction information (e.g., a description of redundant information according to an inter-coding technique, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that has been decoded based on video information encoded using a decoding unit (633) embedded in the example encoder (620) of FIG. 6 (shown as the residual decoder 728 of FIG. 7, as described in more detail below).
[0075] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block to previously coded blocks in the same picture, generate transformed quantized coefficients, and, in some cases, also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) may calculate intra prediction results (e.g., prediction blocks) based on the intra prediction information and reference blocks in the same picture.
[0076] The general-purpose controller (721) may be configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In an example, the general-purpose controller (721) determines a prediction mode for a block and provides a control signal to the switch (726) based on the prediction mode. For example, if the prediction mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select intra-mode results for use by the residual calculation unit (723) and controls the entropy encoder (725) to select intra-prediction information and include the intra-prediction information in the bitstream. If the prediction mode for the block is inter-mode, the general-purpose controller (721) controls the switch (726) to select inter-prediction results for use by the residual calculation unit (723) and controls the entropy encoder (725) to select inter-prediction information and include the inter-prediction information in the bitstream.
[0077] The residual calculation unit (723) is configured to calculate the difference (residual data) between the received block and a prediction result of the block selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate the transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be used by the intra-encoder (722) and the inter-encoder (730) as appropriate. For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and intra-prediction information. The decoded blocks are processed appropriately to generate decoded pictures, which can be buffered in a memory circuit (not shown) and used as reference pictures.
[0078] The entropy encoder (725) may be configured to format a bitstream to include the encoded blocks and perform entropy coding. The entropy encoder (725) may be configured to include various information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Residual information may not be present when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode.
[0079] 8 shows a diagram of an example video decoder (810) according to another embodiment of the disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In an example, the video decoder (810) may be used in place of the example video decoder (410) of FIG. 4.
[0080] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), coupled as shown in the exemplary arrangement in Figure 8.
[0081] The entropy decoder (871) may be configured to reconstruct, from a coded picture, specific symbols representing syntax elements from which the coded picture is constructed. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information) that can identify the mode in which a block is coded (e.g., intra- or bi-prediction mode, merged submode, or other submode), specific samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), residual information in the form of, for example, quantized transform coefficients, etc. In an example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information may undergo inverse quantization and be provided to the residual decoder (873).
[0082] The inter decoder (880) may be configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0083] The intra decoder (872) may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0084] The residual decoder (873) may be configured to perform inverse quantization to retrieve dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (to include quantization parameters (QPs)), which may be provided by the entropy decoder (871) (datapath not shown as this is only low-volume control information).
[0085] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) and the prediction results (possibly output by the inter- or intra-prediction module) in the spatial domain to form reconstructed blocks that form part of the reconstructed picture as part of the reconstructed video. Note that other appropriate operations, such as deblocking operations, may be performed to improve visual quality.
[0086] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In some exemplary embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.
[0087] With reference to coding block partitioning, in some embodiments, a predetermined pattern may be applied. As shown in FIG. 9, an exemplary four-way partition tree may be used, starting from a first predefined level (e.g., the 64x64 block level) and descending to a second predefined level (e.g., the 4x4 level). For example, a base block may follow four partitioning options shown at 902, 904, 906, and 908, with the partitions labeled R allowing for recursive partitioning, in that the same partition tree shown in FIG. 9 may be repeated at a lower scale, down to the lowest level (e.g., the 4x4 level). In some implementations, additional restrictions may apply to the partitioning scheme of FIG. 9. In the implementation of FIG. 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they may not be recursive, whereas square partitioning can be recursive. Partitioning according to FIG. 9 by recursion generates a final set of coding blocks, as needed. Such a scheme may be applied to one or more of the color channels.
[0088] FIG. 10 illustrates another exemplary predefined partitioning pattern that enables recursive partitioning to form a partitioning tree. As shown in FIG. 10, ten exemplary partitioning structures or patterns may be predefined. The root block may start at a predefined level (e.g., the 128×128 level or the 64×64 level). The exemplary partitioning structure of FIG. 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. A partition type including three subpartitions, shown as 1002, 1004, 1006, and 1008 in the second row of FIG. 10, may be referred to as a “T-type” partition. The “T-type” partitions 1002, 1004, 1006, and 1008 may be referred to as Left T-Type, Top T-Type, Right T-Type, and Bottom T-Type. In some implementations, none of the rectangular partitions in FIG. 10 may be further subdivided. The coding tree depth may be further defined to indicate the division depth from the root node or root block. For example, the coding tree depth of the root node or root block, e.g., a 128x128 block, may be set to 0, and after the root block is further divided once according to FIG. 10, the coding tree depth is increased by 1. In some implementations, only full square partitions of 1010 may be allowed for recursive partitioning to the next level of the partitioning tree according to the pattern of FIG. 10. In other words, recursive partitioning may not be allowed for square partitions according to patterns 1002, 1004, 1006, and 1008. Partitioning according to FIG. 10 by recursion generates a final set of coding blocks, as needed. Such a scheme may be applied to one or more of the color channels.
[0089] After dividing or partitioning the base block according to any of the above partitioning procedures or other procedures, a final set of partitions or coding blocks may be obtained. Each of these partitions may be at one of various partitioning levels. Each partition may be referred to as a coding block (CB). For the various partitioning embodiments described above, each resulting CB may have any of the allowed sizes and partitioning levels. They are called coding blocks because they form units on which some basic encoding / decoding decisions can be made, encoding / decoding parameters can be optimized and determined, and signaled in the coded video bitstream. The highest level of the final partitions represents the depth of the coding block partitioning tree. A coding block may be a luma coding block or a chroma coding block.
[0090] In some other embodiments, a quadtree structure may be used to recursively partition the base luma and chroma blocks into coding blocks. Such partitioning structures may be referred to as coding tree units (CTUs), and the CTUs are partitioned into coding units (CUs) by using the quadtree structure to adapt the partitioning to various local characteristics of the base CTU. In such implementations, implicit quadtree partitioning may be performed at picture boundaries, such that blocks continue quadtree partitioning until their size fits the picture boundaries. The term CU is used to collectively refer to units of luma and chroma coding blocks (CBs).
[0091] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into multiple prediction blocks (PBs) for intra- or inter-frame prediction during the encoding and decoding processes. In other words, the CB may be further divided into different subpartitions, where individual prediction decisions / configurations may be made. In parallel, the CB may be further partitioned into multiple transform blocks (TBs) to delineate the levels at which transform or inverse transform of video data is performed. The partitioning schemes of the CB into PBs and TBs may be the same or different. For example, each partitioning scheme may be performed using its own procedure, for example, based on various characteristics of the video data. The PB and TB partitioning schemes may be independent in some embodiments. The PB and TB partitioning schemes and boundaries may be correlated in some implementations. In some implementations, for example, the TBs may be partitioned after PB partitioning, and in particular, each PB may be determined following the partitioning of the coding blocks and then further partitioned into one or more TBs. For example, in some implementations, the PB may be divided into one, two, four, or some other number of TBs.
[0092] In some implementations, the luma channel and the chroma channels may be treated differently for the partitioning of base blocks into coding blocks, and further into prediction blocks and / or transform blocks. For example, in some implementations, partitioning of coding blocks into prediction blocks and / or transform blocks may be allowed for the luma channel, while such partitioning of coding blocks into prediction blocks and / or transform blocks may not be allowed for the chroma channels. In such implementations, transform and / or prediction of luma blocks may therefore only be performed at the coding block level. As another example, the minimum transform block size for the luma channel and the chroma channels may be different, e.g., coding blocks of the luma channel may be allowed to be partitioned into smaller transform and / or prediction blocks than the chroma channels. As yet another example, the maximum depth of partitioning of coding blocks into transform and / or prediction blocks may differ between the luma channel and the chroma channels, e.g., coding blocks of the luma channel may be allowed to be partitioned into deeper transform and / or prediction blocks than the chroma channels. As a specific example, a luma coding block may be partitioned into transform blocks of multiple sizes, which can be repeated by recursive partitioning down to a maximum of two levels, allowing transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes from 4 x 4 to 64 x 64. In the case of a chroma block, however, the maximum number of transform blocks specified for the luma block may be allowed.
[0093] In some implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partitioning may depend on whether the PB is intra- or inter-coded.
[0094] Partitioning of coding blocks (or prediction blocks) into transform blocks may be performed recursively or non-recursively in various exemplary schemes, including but not limited to quadtree partitioning and predefined pattern partitioning, with further consideration of transform blocks at coding or prediction block boundaries. In general, the resulting transform blocks may be at different partitioning levels, may not be the same size, and need not be square in shape (e.g., they can be rectangular with any allowed size and aspect ratio).
[0095] In some implementations, a coding partition tree scheme or structure may be used. The coding partition tree schemes used for the luma channel and the chroma channels need not be the same. In other words, the luma channel and the chroma channel may have separate coding tree structures. Furthermore, whether the luma channel and the chroma channel use the same or different coding partition tree structures and the actual coding partition tree structure to be used may depend on whether the slice being coded is a P, B, or I slice. For example, in the case of an I slice, the chroma channel and the luma channel may have separate coding partition tree structures or coding partition tree structure modes, while in the case of a P or B slice, the luma channel and the chroma channel may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one coding partition tree structure, and the chroma channel may be partitioned into chroma CBs by another coding partition tree structure.
[0096] Specific examples of partitioning coding blocks and transform blocks are described below. In such examples, a base coding block may be partitioned into coding blocks using the recursive quadtree partitioning described above. At each level, whether further quadtree partitioning of a particular partition should continue may be determined according to local video data characteristics. The resulting CBs may be at various quadtree partition levels of various sizes. The decision as to whether to code a picture area with inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CB level (or at the CU level for all three color channels). Each CB may be further partitioned into one, two, four, or other number of PBs according to the PB partition type. Within one PB, the same prediction process may be applied, and related information is conveyed to the decoder on a PB-by-PB basis. After obtaining residual blocks by applying the prediction process based on the PB partition type, the CBs may be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific implementation, however, the CBs or TBs need not be restricted to a square shape. Furthermore, in this particular example, the PB may be square or rectangular in shape for inter prediction, and only square in shape for intra prediction. A coding block may be further divided, for example, into four square-shaped TBs. Each TB may be further divided recursively (by quad-tree division) into smaller TBs called Residual Quad-Tree (RQT).
[0097] Other specific examples for partitioning base coding blocks into CBs and other PBs and / or TBs are described below. For example, instead of using multiple partition unit types such as those shown in FIG. 10, a quadtree segmentation structure including nested multitype trees with bisection and ternary partitioning may be used. The separation of the CB, PB, and TB concepts (i.e., partitioning the CB into PBs and / or TBs, and partitioning the PB into TBs) may be abandoned except when required for CBs whose size is too large for the maximum transform length. However, such CBs may need to be further partitioned. This exemplary partitioning scheme may be designed to support more flexibility in CB partition shape, so that both prediction and transform can be performed at the CB level without further partitioning. In such a coding tree structure, the CB may have either a square or rectangular shape. Specifically, the coding tree block (CTB) may first be partitioned by a quadtree structure. Then, the quadtree leaf nodes may be further partitioned by a multitype tree. An example of a multitype tree structure is shown in FIG. 11. Specifically, the example multi-type tree structure of FIG. 11 includes four split types called vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical third split (SPLIT_TT_VER) (1106), and horizontal third split (SPLIT_TT_VER) (1108). CB then corresponds to the leaf of the multi-type tree. In this embodiment, as long as CT is not too large relative to the maximum transform length, this segmentation is used for both prediction and transform processing without further partitioning. This means that in most cases, CB, PB, and TB have the same block size in a quad-tree structure containing nested multi-type tree coding blocks. An exception occurs when the maximum supported transform length is smaller than the width or height of the color components of CB.
[0098] An example of a quadtree structure including nested multi-type tree coding blocks of block partitions for one CTB is shown in FIG. 12. More specifically, FIG. 12 shows that a CTB 1200 is quadtree-divided into four square partitions 1202, 1204, 1206, and 1208. A decision to further use the multi-tree structure of FIG. 11 for division is made for each quadtree-divided partition. In the example of FIG. 12, partition 1204 is not further divided. Partitions 1202 and 1208 each employ another quadtree division. For partition 1202, the upper-left, upper-right, lower-left, and lower-right partitions quadtree-divided at the second level employ quadtree division at the third level, respectively: horizontal binary division 1104 of FIG. 11, no division, and horizontal ternary division 1108 of FIG. 11. Partition 1208 employs another quadtree division, with the upper-left, upper-right, lower-left, and lower-right partitions quadtree-divided at the second level employing the vertical third division 1106 of FIG. 11 , no division, no division, and horizontal bisection 1104 of FIG. 11 at the third level, respectively. Two of the third-level subpartitions of the second-level upper-left partition of 1208 are further divided according to the horizontal bisection 1104 of FIG. 11 and the horizontal bisection 1108 of FIG. 11 , respectively. Partition 1206 is divided into two partitions using the second-level division pattern according to the vertical bisection 1102 of FIG. 11 , and these two partitions are further divided at the third level according to the horizontal bisection 1108 and vertical bisection 1102 of FIG. 11 . A fourth-level division is further applied to one of these two partitions according to the horizontal bisection 1104 of FIG. 11 .
[0099] For the above specific example, the maximum luma transform size may be 64x64, and the maximum supported chroma transform size may be different from the luma, for example, 32x32. If the width or height of a luma coding block or a chroma coding block is larger than the maximum transform width or height, the luma coding block or the chroma coding block may be automatically split in the horizontal and / or vertical directions to satisfy the transform size constraints in that direction.
[0100] In the specific example for partitioning the base coding blocks into CBs, as described above, the coding tree scheme may support the ability for luma and chroma to have separate block tree structures. For example, in the case of P slices and B slices, the luma CTB and chroma CTB in one CTU may share the same coding tree structure. In the case of an I slice, for example, the luma and chroma may have separate coding block tree structures. When the separate block tree mode is applied, the luma CTB may be partitioned into luma CBs by one coding tree structure, and the chroma CTB is partitioned into chroma CBs by another coding tree structure. This means that a CU in an I slice can consist of a coding block of the luma component or a coding block of two chroma components, and a CU in a P or B slice always consists of coding blocks of all three color components unless the video is monochrome.
[0101] Examples of partitioning coding or prediction blocks into transform blocks and the coding order of the transform blocks are described further below. In some embodiments, transform partitioning may support transform blocks of multiple shapes, e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with transform block sizes ranging from, e.g., 4x4 to 64x64. In some implementations, if the coding block is 64x64 or smaller, transform block partitioning may be applied only to the luma component, so that for chroma blocks, the transform block size is the same as the coding block size. Otherwise, if the width or height of the coding block is greater than 64, both the luma coding block and the chroma coding block may be implicitly divided into multiple transform blocks of min(W,64) x min(H,64) and min(W,32) x min(H,32), respectively.
[0102] In some embodiments, for both intra-coded and inter-coded blocks, the coding blocks may be further partitioned into multiple transform blocks with a partitioning depth up to a predefined number of levels (e.g., two levels). The depth and size of the transform block partitioning may be related. An example mapping from the transform size of the current depth to the transform size of the next depth is shown in Table 1 below. [Table 1]
[0103] Based on the example mapping of Table 1, in the case of a 1:1 square block, the next level transform partitioning may generate four 1:1 square sub-transform blocks. The transform partition may stop at, for example, 4x4. Thus, a transform size of 4x4 at the current depth corresponds to the same size of 4x4 at the next depth. In the example of Table 1, in the case of a 1:2 / 2:1 non-square block, the next level transform partitioning will generate two 1:1 square sub-transform blocks, while in the case of a 1:4 / 4:1 non-square block, the next level transform partitioning will generate two 1:2 / 2:1 sub-transform blocks.
[0104] In some embodiments, additional restrictions may be applied to the luma component of an intra-coded block. For example, for each level of transform partitioning, all sub-transform blocks may be constrained to have equal sizes. For example, for a 32x16 coding block, level 1 transform partitioning generates two 16x16 sub-transform blocks, and level 2 transform partitioning generates eight 8x8 sub-transform blocks. In other words, second-level partitioning should be applied to all first-level sub-blocks to keep the transform units equal in size. An example of transform block partitioning for an intra-coded square block according to Table 1 is shown in FIG. 13, with the coding order represented by the arrows. Specifically, 1302 denotes a square coding block. The first-level partitioning into four equally sized transform blocks according to Table 1 is shown at 1304, with the coding order indicated by the arrows. All second-level partitioning of the first-level equally sized blocks into 16 equally sized transform blocks according to Table 1 is shown at 1306, with the coding order indicated by the arrows.
[0105] In some embodiments, the above restrictions on intra-coding may not apply to the luma component of an inter-coded block. For example, after the first level of transform partitioning, any one of the sub-transform blocks may be further partitioned independently at one further level. Thus, the resulting transform blocks may or may not be of the same size. An example of the partitioning of an inter-coded block into transform blocks, along with their coding order, is shown in FIG. 14. In the example of FIG. 14, an inter-coded block 1402 is partitioned into transform blocks at two levels according to Table 1. At the first level, the inter-coded block is partitioned into four transform blocks of equal size. Then, only one of the four transform blocks (but not all of them) is further partitioned into four sub-transform blocks, resulting in a total of seven transform blocks with two different sizes, as indicated by 1404. An example of the coding order of these seven transform blocks is indicated by the arrows at 1404 in FIG. 14.
[0106] In some embodiments, for chroma components, some additional restrictions on transform blocks may apply, e.g., for chroma components, the transform block size can be as large as the coding block size, but cannot be smaller than a predefined size, e.g., 8x8.
[0107] In some other embodiments, for coding blocks with either width (W) or height (H) greater than 64, both the luma coding block and the chroma coding block may be implicitly divided into multiple transform units of min(W,64)×min(H,64) and min(W,32)×min(H,32), respectively.
[0108] Figure 15 further illustrates another alternative example of a scheme for partitioning coding blocks or predictive blocks into transform blocks. As shown in Figure 15, rather than using recursive transform partitioning, a predefined set of partitioning types may be applied to coding blocks depending on the transform type of the coding block. In the specific example shown in Figure 15, one of six example partitioning types may be applied to divide the coding block into various numbers of transform blocks. Such a scheme may be applied to either coding blocks or predictive blocks.
[0109] More specifically, the partitioning scheme of Figure 15 provides up to six partition types for any given transform type as shown in Figure 15. In this scheme, every coding block or predictive block may be assigned a transform type based on, for example, rate-distortion cost. In one example, the partition type assigned to a coding block or predictive block may be determined based on the transform partition type of the coding block or predictive block. A particular partition type may correspond to a transform block division size and pattern (or partition type), as indicated by the partition types depicted in Figure 15. The correspondence between various transform types and various partitions may be predefined. Using capital letter labels indicating transform partition types that may be assigned to a coding block or predictive block based on rate-distortion cost, an example correspondence is shown below: · PARTITION_NONE: Allocate a transformation size equal to the block size. · PARTITION_SPLIT: Allocate a transformation size that is 1 / 2 the width of the block size and 1 / 2 the height of the block size. PARTITION_HORZ: Allocates a transformation size with width equal to the block size and half the height of the block size. · PARTITION_VERT: Allocates a transformation size with a width half of the block size and a height equal to the block size. PARTITION_HORZ4: Allocates a transformation size with the same width as the block size and 1 / 4 of the height of the block size. · PATRITION_VERT4: Assigns a transformation size with a width of 1 / 4 of the block size and a height equal to the block size.
[0110] In the above example, all of the partition types shown in Figure 15 include uniform transform sizes for the partitioned transform blocks. This is by way of example only and not limitation. In some other implementations, mixed transform block sizes may be used for partitioned transform blocks in a particular partition type (or pattern).
[0111] With reference to the primary transform, an exemplary 2D (two-dimensional) transform process may involve the use of hybrid transform kernels (e.g., which may consist of a different 1D (one-dimensional) transform for each dimension of the coded residual block) in addition to using the same transform kernel for both dimensions. Exemplary primary 1D transform kernels may include, but are not limited to: a) 4-point (4p), 8-point (8p), 16-point (16p), 32-point (32p), and 64-point (64p); b) 4-point, 8-point, 16-point asymmetric DST and their inverse versions (DST stands for Discrete Sine Transform); c) 4-point, 8-point, 16-point, or 32-point identity transform; d) Incremental Distance Transforms (IDT). Therefore, the 2D transform process may involve the use of hybrid transforms or transform kernels (different transforms for each dimension of the coded residual block), where the selection of the transform or transform kernel used for each dimension may be based on a rate-distortion (RD) criterion. The term transform kernel may alternatively be referred to as a transform basis function. For example, the basis functions of 1D DCT-2, DST-4, and DST-7 that may be implemented as hybrids for 2D transforms are listed in Table 2 (DCT stands for Discrete Cosine Transform). [Table 2]
[0112] For example, because the DCT-2 (4p-64p), DST-4 (8p, 16p), and DST-7 (4p) transforms exhibit symmetric / asymmetric characteristics, a "partial butterfly" implementation may be supported in some embodiments to reduce the number of operation counts (multiplications, additions / subtractions, and shifts). The partial butterfly implementation may involve planar rotations using trigonometric cosine and sine functions at various angles, as shown in FIG. 16. Exemplary 12-bit lookup tables are shown in FIGS. 17 and 18 and may be used to generate the trigonometric function values. Specifically, FIG. 17 shows an exemplary DCT-2 (4p-64p) / DST-4 (8p, 16p) partial butterfly lookup table, and FIG. 18 shows an exemplary DST-7 (4p) partial butterfly lookup table.
[0113] In some embodiments, the transforms may include line graph transforms (LGTs) as shown in FIG. 19. A graph can be a general mathematical structure consisting of a set of vertices and edges used to model affinity relationships between objects of interest. In practice, weighted graphs (where a set of weights is assigned to edges and possibly to vertices) can yield sparse representations for robust modeling of signals / data. LGTs can improve coding efficiency by providing better adaptation to diverse block statistics. Separable LGTs can be designed and optimized by learning a line graph from data to model the underlying row- and column-wise statistics of block residual signals. The associated Generalized Graph Laplacian (GGL) matrix is then used to derive the LGT.
[0114] In one implementation, given a weighted graph G(W,V), GGL may be defined as LE=D−W+V, where W is the non-negative edge weight W c where D may be a diagonal order matrix, and V may be a weighted self-loop V c1 ,V c2The matrix L may be a diagonal matrix representing c can be expressed as:
number
[0115] Next, LGT is GGL L c can be derived by eigenvalue decomposition of:
number
[0116] The columns of the orthogonal matrix U are the basis vectors of the LGT, and Φ is the diagonal eigenvalue matrix. In practice, DCTs and DSTs, including DCT-2, DCT-8, and DST-7, are derived from a specific form of the GGL. The DCT-2 is c1 = 0, and DST-7 is derived by setting V c1 =W c The DCT-8 is derived by setting V c2 =W c DST-4 is derived by setting V c1 =2W c The DCT-4 is derived by setting V c2 =2W c is derived by setting
[0117] LGT may be implemented as a matrix multiplication. A 4p LGT core c In V c1 =2W c , i.e., it is DST-4. The 8pLGT core is c In V c1 =1.5W c and the 16p, 32p and 64p LGT cores can be derived by setting L c In c1 =W c , i.e., it is DST-7.
[0118] Referring to some examples of specific types of signaling for coding blocks / units, for each intra- and inter-coding unit, a flag, i.e., a skip_txfm flag, may be signaled in the coded bitstream, as shown in the example syntax of Table 3 below and represented by the read_skip() function for reading the flag from the bitstream. This flag may indicate whether the transform coefficients are all zero in the current coding unit. In some examples, when this flag is signaled, e.g., with a value of 1, other transform coefficient-related syntax, e.g., EOB (End of Block), need not be signaled for any of the color coding blocks within the coding unit and can be derived as a value or data structure predefined for and associated with the zero transform coefficient block. In the case of an inter-coding block, as shown by the example of Table 3, this flag may be signaled after a skip mode flag, which indicates that the coding unit may be skipped for various reasons. When skip_mode is true, the coding unit should be skipped and there is no need to signal the skip_txfm flag, which is inferred to be 1. Otherwise, if skip_mode is false, more information about the coding unit will be included in the bitstream and the skip_txfm flag will be further signaled to indicate whether the coding unit is all zeros or not. [Table 3]
[0119] Regarding the encoding and decoding (entropy coding) of the residual transform coefficients for each color component, for each transform block, transform coefficient coding begins with signaling the skip code, if the skip code is zero (indicating the presence of non-zero coefficients), followed by the transform kernel type and end-of-block (EOB) position. Each coefficient value is then mapped to multiple level maps (amplitude maps) and codes.
[0120] After the EOB position is coded, the low-level map and the mid-level map may be coded in reverse scan order, with the former indicating whether the magnitude of the coefficient is within the low level (e.g., between 0 and 2) and the latter indicating whether the range is within the mid-level (e.g., between 3 and 14). The next step is to code, in forward scan order, the signs of the coefficients and the residual values of the coefficients greater than a high level (for example, 14), for example by Exp-Golomb coding.
[0121] Regarding the use of context modeling, lower-level map coding can incorporate transform size and direction and up to five neighboring coefficient information, while mid-level map coding can follow a similar approach to lower-level amplitude coding, except that the number of neighboring coefficients is reduced to a smaller number (e.g., two). Exemplary Exp-Golomb codes for the residual level and AC coefficient signs are coded without a context model, while the DC coefficient sign is coded using the DC code of its neighboring transform block.
[0122] In some embodiments, the chroma residuals may be jointly coded. Such a coding scheme may be based on some statistical correlation between the chroma channels. For example, in many cases, Cr and Cb chroma coefficients may be similar in amplitude but opposite in sign, and thus may be jointly coded to improve coding efficiency, e.g., at the transform block level where transform coefficients are signaled, while only introducing small color distortion. The utilization (activation) of a joint chroma coding mode may be indicated, for example, by a joint chroma coding flag (e.g., a flag tu_joint_cbcr_residual_flag at the TU level), and the selected joint mode may be implicitly indicated by the chroma CBF.
[0123] Specifically, the flag tu_joint_cbcr_residual_flag may be present if one or both chroma CBFs of a TU (transform block) are equal to 1. In the PPS and slice header, chroma quantization parameter (QP) offset values may be signaled for the joint chroma residual coding mode to distinguish them from the chroma QP offset values signaled for the regular chroma residual coding mode. These chroma QP offset values may be used to derive the chroma QP value of a block coded using the joint chroma residual coding mode. When the corresponding joint chroma coding mode (mode 2 in Table 4 below) is active in a TU, this chroma QP offset may be added to the luma-derived chroma QP applied during quantization and decoding of that TU. For other modes (modes 1 and 3 in Table 4), the chroma QP may be derived in the same manner as for a conventional CB or Cr block. The reconstruction process of the chroma residual (resCb and resCr) from the transmitted transform block is shown in Table 4. When this mode is activated (mode 2), one single joint chroma residual block (resJointC[x][y]) may be signaled, and the residual block of Cb (resCb) and the residual block of Cr (resCr) may be derived taking into account information such as tu_cbf_cb, tu_cbf_cr, and CSign. CSign is a sign value specified, for example, in the slice header, rather than at the transform block level. In some implementations, CSign can be -1 most of the time.
[0124] The above three exemplary joint chroma coding modes may only be supported for intra-coded CUs. For inter-coded CUs, only mode 2 may be supported. Therefore, for inter-coded CUs, the syntax element tu_joint_cbcr_residual_flag is present only if both chroma CBFs are 1. [Table 4]
[0125] The above joint chroma coding schemes assume some correlation between the transform coefficients of the collocated Cr and Cb transform blocks. Such assumptions are usually statistical and may lead to distortions in some circumstances. In particular, if one of the color coefficients in a transform block is non-zero while the coefficient of the other color component is zero, some of the assumptions made in the joint chroma coding scheme will certainly be invalid, and such coding will not save any coding bits (because one of the chroma coefficients is zero anyway).
[0126] In various embodiments below, a coefficient-level (i.e., per transform coefficient) cross-component coding scheme is described, which exploits some correlation between collocated (collocated in the frequency domain) transform coefficients of color components. Such a scheme is particularly useful for transform blocks (or units) in which a coefficient of one color component is zero while the corresponding transform coefficient of the other color component is nonzero. For such pairs of zero and nonzero color coefficients, either before or after quantization, the nonzero color coefficient may be used to estimate or derive the original small value of the zero coded coefficient of the other color component (even a small value may not originally be zero before quantization in the coding process), thereby making it possible to recover some information lost in, for example, the quantization process during encoding. Some information lost during quantization to zero may be recovered due to statistically existing inter-color correlation. Such cross-component coding recovers the lost information to some extent without significant coding cost (for the zero coefficients). Such a coefficient information recovery process may also be referred to as a transform coefficient refinement process (or simply, a coefficient refinement process) because coefficients having a value of zero may be restored to, for example, a small non-zero value. In one implementation, during the coefficient refinement process, zero transform coefficients in the second transform block may be refined by adding an offset. The offset may be derived based on a corresponding (e.g., co-located) transform coefficient in the first transform block. The first transform block may be associated with a first color component, and the second transform block may be associated with a second color component different from the first color component. The color component may be any one of a luma component and a chroma component.
[0127] In an exemplary implementation, a cross-component coefficient sign coding method may be implemented that utilizes the coefficient sign value of a first color component to code the coefficient sign of a second color component. In one more specific example, the sign value of a Cb transform coefficient may be used as a context for coding the sign of a Cr transform coefficient. Such cross-component coding may be implemented coefficient-by-coefficient for a pair of transform coefficients of a color component. The principles underlying such an implementation, as well as other implementations described in more detail below, are not limited to the Cb and Cr components. They may be applied between any two of the three color components. In this context, the luma channel is considered one of the color components.
[0128] A method of deriving and refining collocated transform coefficients in a second transform block in a second component using transform coefficients in a first transform block in a first component may be called Cross Component Level Reconstruction (CCLR). In this method, CCLR is applied to a first transform block using a second transform block as a reference. For example, a level value of a Cb transform coefficient may be used to derive a level value of a corresponding (e.g., collocated) Cr transform coefficient, and vice versa. Using CCLR, at the decoder side, information of one color component may be refined or recovered by referring to another color component.
[0129] In the following examples, the term chroma channel may generally refer to both the Cb and Cr color components (or channels), or both the U and V color components (or channels). The term luma channel may include the luma component or the Y component. The luma component or channel may be referred to as the luma color component or channel. Y, U, and V are used below to represent the three color components. Furthermore, the terms "coded block" and "coding" block are used synonymously to mean either a block to be coded or a block that has already been coded. They may be blocks of any of the three color components. Three corresponding color-coded blocks / coding blocks may constitute a coded unit / coding unit.
[0130] In the following examples, a transform set refers to a group of transform kernel (or candidate) options. A transform set may include one or more of the following types of kernel (or candidate) options: DCT, ADST, FLIPADST, IDT, LGT, KLT, or RCT.
[0131] In the following examples, transform type refers to the type of primary and / or secondary transform. Examples of primary transform types include, but are not limited to, DCT, ADST, FLIPADST, IDT, LGT, KLT, and RCT. Examples of secondary transform types include, but are not limited to, KLT with different input sizes and different kernels.
[0132] In the following examples, the term transform can refer to a primary transform, or a secondary transform, or a combination of a primary transform and a secondary transform. The term inverse transform can refer to an inverse primary transform, or an inverse secondary transform, or a combination of an inverse primary transform and an inverse secondary transform.
[0133] The following examples may be used separately or combined in any order. The term block size may refer to either the width or height of a block, or the maximum value of the width and height, or the minimum value of the width and height, or the area size (width x height), or the aspect ratio of a block (width:height or height:width). The term "level value" or "level" may refer to the magnitude of a transform coefficient value.
[0134] In some embodiments, the level values and / or sign values of the transform coefficients of a first color component may be used to derive an offset value that is added to the level values of the transform coefficients of a second color component.
[0135] In some further implementations, the transform coefficients of the first color component used to generate the offset and transform coefficients of the second color component are collocated (same coordinates in the frequency domain, e.g., the estimates are non-crossing frequencies).
[0136] The above first and second color components may not be limited to specific color components, but in some embodiments, the first color component may be Cb (or Cr), while the second color component is Cr (or Cb).
[0137] In some specific embodiments, the first color component may be luma, and the second color component may be one of Cb and Cr.
[0138] In some specific embodiments, the first color component may be one of Cb and Cr, and the second color component may be luma.
[0139] In some embodiments, the quantized transform coefficients of a first color component may be non-zero, and the quantized transform coefficients of the second color component may be zero. As such, the original relatively small non-zero information of the original transform coefficients of the second color component may be lost due to quantization during the encoding process, and the embodiments described herein help recover some of the lost information using corresponding non-zero color components that may be statistically correlated with the zero-coefficient color components.
[0140] In some embodiments, the sign values of the transform coefficients of a first color component may be used to derive an offset value that is added to the dequantized transform coefficient level values of a second color component.
[0141] In some embodiments, the sign values of the transform coefficients of the first color component are used to derive an offset value that is added to the transform coefficient level values of the second color component before dequantization.
[0142] In some embodiments, when the sign value of the transform coefficient of the first color component is positive (or negative), a negative (or positive) offset value is added to the dequantized transform coefficient level value of the second color component to reconstruct the transform coefficient value of the second color component. In other words, the sign value of the transform coefficient of the first color component and the sign value of the offset value added to the transform coefficient level value of the second color component have different sign values. Such an implementation may be consistent with the statistical observation that two chroma components usually have opposite signs in transform.
[0143] In some embodiments, whether the sign values of the transform coefficients of a first color component and the sign values of the offset values added to the transform coefficient level values of a second color component have opposite sign values is signaled in a high-level syntax, including but not limited to, an SPS, VPS, PPS, APS, picture header, frame header, slice header, tile header, or CTU header, which is similar to the signaling scheme described above for the joint chroma coding scheme.
[0144] In some embodiments, the offset value added to the transform coefficient level value of the second color component may depend on both the sign and the level of the transform coefficient of the first color component.
[0145] In some embodiments, the magnitude of the offset value added to the transform coefficient level value of the second color component may depend on the transform coefficient level value of the first color component.
[0146] In some embodiments, the magnitude of the offset value added to the transform coefficient level value of the second color component may be predefined for each input value of the coefficient level of the transform coefficient of the first color component.
[0147] In some embodiments, the offset value added to the transform coefficient level value of the second color component may depend on the frequency that the transform coefficient is located in. For example, the offset value may be smaller for higher frequency coefficients.
[0148] In some embodiments, the offset value added to the transform coefficient level value of the second color component may depend on the block size of the block that the transform coefficient belongs to. For example, the offset value may generally be smaller for larger block sizes.
[0149] In some embodiments, the offset value added to the transform coefficient level value of the second color component may depend on whether the second color component is a luma (Y) or chroma (Cb or Cr) component. For example, the offset value may be smaller when the second color component is luma.
[0150] In some embodiments, for a transform block of a second color component, the selection of a transform kernel for this transform block may depend on whether a CCLR method is applied to refine the transform coefficients in this transform block. The selected transform kernel may be used for a primary transform or a secondary transform. Using the primary transform as an example, on the encoding (encoder) side, the selected transform kernel may be used to perform a transform on a prediction residual to obtain a transform block, and on the decoding (decoder) side, the selected transform kernel may be used to perform an inverse transform on the refined transform block to obtain a prediction residual when a CCLR method is applied to the transform block. It should be noted that before the inverse transform, a CCLR refinement process is performed on the transform block to obtain a refined transform block.
[0151] In some embodiments, whether CCLR is applied or enabled for a transform block may be signaled, for example, by a syntax value or flag.
[0152] In some embodiments, if an EOB signaling a relative end-of-block position relative to a transform block of a second color component is zero, thereby indicating that all transform coefficients in this transform block are zero, and CCLR is applied to this transform coefficient, a CCLR refinement process may be applied to each transform coefficient in this transform block, for example, by adding a corresponding offset value to each transform coefficient. The offset value may be derived based on the co-located transform coefficients in the transform block of the first color component. As a result of the refinement process, the refined transform coefficients in the transform block are no longer all zero. Following the refinement process, an inverse transform may be performed on the CCLR-refined transform block.
[0153] In some embodiments, rather than targeting the entire transform block, the CCLR refinement process may target only a portion of the transform coefficients within the transform block.
[0154] In some embodiments, the same transform kernel used for the inverse transform on the collocated transform block of the first color component may be selected when applying the inverse transform to the CCLR refined transform block of the second color component. Note that the collocated transform block of the first color component is used as a basis (or reference) to derive offset values used in the CCLR refinement process applied to the transform block of the second color component.
[0155] In some embodiments, when CCLR is applied to a transform block, the transform kernel for performing the inverse transform on the refined transform block (refined from this transform block) may be explicitly signaled. For example, an index indicating a selected transform kernel from a transform set (i.e., a set of candidate transform kernels) may be signaled. The transform set may be preset, predefined, derived, or signaled.
[0156] In some embodiments, when CCLR is applied to a transform block of a second color component and a coded block associated with this transform block is an intra-predicted block, the transform kernel for performing an inverse transform on the refined transform block (refined from this transform block) may be implicitly derived based on the intra-prediction mode. In one implementation, additional constraints may be imposed on the selection of the transform kernel, such that whether CCLR is applied to the transform block should also be taken into consideration. For example, the selected transform kernel for one transform block to which CCLR is applied should be different from the selected transform kernel for another transform block to which CCLR is not applied. The other transform blocks may be associated with the same coded block or different coded blocks that are intra-predicted.
[0157] In some embodiments, when CCLR is applied to a transform block of a certain color component (e.g., Cb or Cr), the transform kernel for performing the inverse transform on the refined transform block (refined from this transform block) may be the same as that applied to the co-located transform block of the luma component.
[0158] In some embodiments, a transform kernel for performing an inverse transform on the refined transform block may be selected from a transform set, and the selection may be based on the block size of the refined transform block. The transform set may be preset, predefined, derived, or signaled.
[0159] In some embodiments, there may be further constraints on whether the CCLR method can be applied to a transform block. The constraints may be based on the transform type. In one implementation, to apply the CCLR method, the transform types of the primary transform or secondary transform need to be restricted to a specific type or combination of types. For example, when the primary transform is a 2D transform, the two 1D transforms of the 2D transform must both be DCTs, both be IDTs, or one other combination of transform types.
[0160] In some embodiments, during the CCLR refinement process, the offset derivation may depend on the transform type selected for the transform block to be refined. The transform type may be applied to a primary transform or a secondary transform.
[0161] Although cross-component zero coefficient refinement has been described for a situation where the EOB is indicated as zero (i.e., the relative position of the end of the block is zero, i.e., the transform coefficients of the corresponding block are all zero), various implementations of cross-component refinement and the selection and signaling of transform kernel types and / or particular kernels are not so limited. For example, a particular transform block of one color component may have only a small number of non-zero coefficients. Nevertheless, other zero transform coefficients may be refined using transform coefficients in another color component. Furthermore, refinement may further be based on non-zero coefficients within the same transform block being refined. The selection of transform kernel types and / or kernels may be performed similarly to the implementations described above.
[0162] 20 shows a flowchart 2000 of an exemplary video decoding method in accordance with the principles underlying the implementations described above. Method 2000 may include some or all of the following steps: at step 2010, receive a bitstream of video blocks including a first transform block of a first color component and a second transform block of a second color component. The first transform block and the second transform block are co-located blocks. at step 2020, obtain the first transform block of the first color component and the second transform block of the second color component from the bitstream of video blocks. at step 2030, determine a first flag indicating that all transform coefficients in the first transform block are zero. at block 2040, determine a second flag indicating that cross-component level reconstruction (CCLR) is applied to the first transform block. at block 2050, in response to determining that CCLR is applied to the first transform block, refine one or more of the transform coefficients in the first transform block by adding one or more offset values to obtain a refined first transform block, wherein the one or more offset values are derived based on transform coefficients in a second transform block that are co-located with one or more of the transform coefficients in the first transform block; determining a target transformation kernel for the refined first transformation block; performing an inverse transform on the refined first transform block based on the target transform kernel to obtain a target block; A first color component of the video block is reconstructed based on at least the target block.
[0163] In the embodiments and implementations of the present disclosure, any steps and / or actions may be combined or arranged in any quantity or order as desired. Two or more of the steps and / or actions may be performed in parallel. The embodiments and implementations of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0164] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 21 illustrates a computer system (2800) suitable for implementing certain embodiments of the disclosed subject matter.
[0165] Computer software can be coded in any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly, or through interpretation, microcode execution, etc.
[0166] The instructions may be executable by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0167] 21 for computer system 2800 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 2800.
[0168] The computer system 2800 may include certain human interface input devices. Such human interface input devices may respond to input by one or more users through, for example, tactile input (e.g., keystrokes, swipes, dataglove movements), audio input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0169] The input human interface devices may include one or more of a keyboard (2801), a mouse (2802), a trackpad (2803), a touchscreen (2810), a data glove (not shown), a joystick (2805), a microphone (2806), a scanner (2807), and a camera (2808) (only one of each is shown).
[0170] The computer system 2800 may also include certain human interface output devices, such as those that stimulate one or more of the user's senses through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (2810), data gloves (not shown), or joystick (2805), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (2809), headphones (not shown)), visual output devices (e.g., screens (2810) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which are capable of outputting two-dimensional visual output or output in more than three dimensions by means of stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0171] The computer system (2800) may also include human-accessible storage devices and their associated media, such as CD / DVD ROM / RW (2820) with CD / DVD or similar media (2821), thumb drives (2822), removable hard disks or solid state drives (2823), legacy magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0172] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0173] The computer system 2800 may also include interfaces 2854 to one or more communications networks 2855. Networks may be, for example, wireless, wireline, or optical. Networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wireline or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and factory networks including CANBus. Particular networks generally require external network interface adapters attached to particular general-purpose data ports or peripheral buses 2849 (e.g., USB ports on the computer system 2800). Others are generally integrated into the core of the computer system 2800 by attachment to a system bus as described below (e.g., an Ethernet network interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 2800 can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV) or one-way transmit-only (e.g., CANBus to a specific CANBus device), or it can be two-way to other computer systems using, for example, a local or wide-area digital network. Specific protocols or protocol stacks can be used with each of the networks and network interfaces described above.
[0174] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core 2840 of the computer system 2800.
[0175] The cores 2840 may include one or more central processing units (CPUs) 2841, graphics processing units (GPUs) 2842, dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) 2843, hardware accelerators for specific tasks 2844, graphics adapters 2850, etc. These devices may be connected through a system bus 2848, along with read-only memory (ROM) 2845, random access memory (RAM) 2846, and internal mass storage devices such as internal non-user-accessible hard drives, SSDs, etc. 2847. In some computer systems, the system bus 2848 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus 2848 directly or through a peripheral bus 2849. In one example, a screen 2810 may be connected to a graphics adapter 2850. Architectures for peripheral buses include PCI, USB, and the like.
[0176] The CPU (2841), GPU (2842), FPGA (2843), and accelerator (2844) can execute specific instructions that, in combination, can constitute the above-mentioned computer code. The computer code can be stored in ROM (2845) or RAM (2846). Temporary data can also be stored in RAM (2846), while persistent data can be stored, for example, in an internal mass storage device (2847). Rapid storage and retrieval from any of the memory devices is enabled through the use of cache memory. Cache memory can be closely associated with one or more of the CPU (2841), GPU (2842), mass storage device (2847), ROM (2845), RAM (2846), etc.
[0177] The computer-readable medium can carry computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those of ordinary skill in the computer software arts.
[0178] As a non-limiting example, a computer system having the architecture (2800), and specifically the core (2840), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices introduced above, in addition to specific storage of the core (2840) that is non-transitory in nature, such as the core's internal mass storage (2847) or ROM (2845). Software implementing various embodiments of the present disclosure can be stored on such devices and executable by the core (2840). The computer-readable media can include one or more memory devices or chips, depending on particular needs. Software can cause the cores (2840), and specifically the processors (including CPUs, GPUs, FPGAs, etc.) therein, to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (2846) and modifying such data structures according to processes defined by the software. Additionally, or alternatively, the computer system can provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerators (2844)) that can operate in place of or in conjunction with software to perform particular processes or particular portions of particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure encompasses any appropriate combination of hardware and software.
[0179] While this disclosure has described several illustrative embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of the disclosure. It should thus be understood that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown and described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.
[0180] Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS:Benchmark Set MV: Motion Vector HEVC:High Efficiency Video Coding SEI:Supplementary Enhancement Information VUI:Video Usability Information GOP: Group of Picture(s) TU: Transform Unit(s) PU: Prediction Unit(s) CTU: Coding Tree Unit(s) CTB: Coding Tree Block(s) PB: Prediction Block(s) HRD:Hypothetical Reference Decoder SNR: Signal-to-Noise Ratio CPU:Central Processing Unit(s) GPU:Graphics Processing Unit(s) CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED:Organic Light-Emitting Diode CD:Compact Disc DVD:Digital Video Disc ROM:Read-Only Memory RAM:Random Access Memory ASIC:Application-Specific Integrated Circuit PLD:Programmable Logic Device LAN:Local Area Network GSM:Global System for Mobile communications LTE:Long-Term Evolution CANBus:Controller Area Network Bus USB:Universal Serial Bus PCI:Peripheral Component Interconnect FPGA:Field Programmable Gate Area(s) SSD:Solid-State Drive IC:Integrated Circuit HDR:High Dynamic Range SDR:Standard Dynamic Range JVET:Joint Video Exploration Team MPM:Most Probable Mode WAIP:Wide-Angle Intra Prediction CU:Coding Unit PU:Prediction Unit TU:Transform Unit CTU:Coding Tree Unit PDPC:Position Dependent Prediction Combination ISP:Intra Sub-Partitions SPS:Sequence Parameter Setting PPS:Picture Parameter Set APS:Adaptation Parameter Set VPS:Video Parameter Set DSP:Decoding Parameter Set ALF:Adaptive Loop Filter SAO:Sample Adaptive Offset CC-ALF:Cross-Component Adaptive Loop Filter CDEF:Constrained Directional Enhancement Filter CCSO:Cross-Component Sample Offset LSO:Local Sample Offset LR:Loop Restoration Filter AV1:AOMedia Video 1 AV2:AOMedia Video 2 DCT:Discrete Cosine Transform DST:Discrete Sine Transform ADST:Asymmetric DST FLIPADST:Flipped ADST IDT:Incremental Distance Transform LGT:Line Graph Transforms KLT:Karhunen Loeve Transform RCT:Row-Column Transform
Claims
1. A method of video processing performed by a decoder, comprising: receiving a bitstream of video blocks including a first transform block of a first color component and a second transform block of a second color component, the first transform block and the second transform block being co-located blocks; obtaining the first transform block of the first color component and the second transform block of the second color component from the bitstream of the video block; determining a first flag indicating that all transform coefficients in the first transform block are zero; determining a second flag indicating that cross-component level reconstruction (CCLR) is applied to the first transform block; In response to determining that a CCLR is applied to the first transform block, refining one or more of the transform coefficients in the first transform block by adding one or more offset values to obtain a refined first transform block, the one or more offset values being derived based on transform coefficients in the second transform block that are co-located with the one or more of the transform coefficients in the first transform block; determining a target transformation kernel for the refined first transformation block; performing an inverse transform on the refined first transform block based on the target transform kernel to obtain a target block; reconstructing the first color component of the video block based on at least the target block; A method having the following.
2. the first color component has one chroma component while the second color component has another chroma component, or the first color component comprises a luma component while the second color component comprises one chroma component; or the first color component has one chroma component, while the second color component has a luma component; The method of claim 1.
3. The step of determining the target transformation kernel includes: selecting the target transformation kernel of the refined first transformation block as the same transformation kernel of the second transformation block; The method of claim 1.
4. The step of determining the target transformation kernel includes: extracting an indicator signaled in the bitstream, the indicator specifying the target transform kernel, the indicator being signaled in response to determining that the CCLR is to be applied to the first transform block; selecting the target transformation kernel based on the indicator; having The method of claim 1.
5. The step of determining the target transformation kernel includes: deriving the target transform kernel based on a mode of the intra prediction in response to the video block being predicted under intra prediction. The method of claim 1.
6. the target transform kernel is different from the transform kernel of the second transform block if CCLR is not applied to the second transform block; The method of claim 5.
7. The step of determining the target transformation kernel includes: responsive to the video block being inter predicted, selecting the target transform kernel according to a luma transform block co-located with the first transform block. The method of claim 1.
8. The step of determining the target transformation kernel includes: selecting the target transformation kernel from a list of kernels based on a block size of the first transformation block; the list of kernels is predefined or signaled in the bitstream; The method of claim 1.
9. A CCLR is allowed to be applied to the first transform block only if the first transform block is associated with a predefined set of primary transform types. The method of claim 1.
10. the transform associated with each primary transform type in the predefined set of primary transform types is a two-dimensional transform, the two-dimensional transform being formed by two one-dimensional transforms, the two one-dimensional transforms both being discrete cosine transforms (DCTs) or both being incremental distance transforms (IDTs); 10. The method of claim 9.
11. deriving the one or more offset values based on transform coefficients in the second transform block that are co-located with the one or more of the transform coefficients in the first transform block. The method of claim 1.
12. 1. A device for video processing, comprising: a memory for storing computer instructions and a processor in communication with the memory; A device, wherein when the processor executes the computer instructions, the processor is configured to cause the device to perform the method of any one of claims 1 to 11.
13. A program comprising computer readable instructions, 12. A program comprising computer readable instructions which, when executed by a processor of a device for processing video data, cause the processor to perform the method of any one of claims 1 to 11.
14. 1. A method of video processing performed by an encoder, comprising: generating a bitstream of video blocks including a first transform block of a first color component and a second transform block of a second color component, the first transform block and the second transform block being co-located blocks; obtaining the first transform block of the first color component and the second transform block of the second color component from the bitstream of the video block; determining a first flag indicating that all transform coefficients in the first transform block are zero; determining a second flag indicating that cross-component level reconstruction (CCLR) is applied to the first transform block; In response to determining that a CCLR is applied to the first transform block, refining one or more of the transform coefficients in the first transform block by adding one or more offset values to obtain a refined first transform block, the one or more offset values being derived based on transform coefficients in the second transform block that are co-located with the one or more of the transform coefficients in the first transform block; determining a target transformation kernel for the refined first transformation block; performing an inverse transform on the refined first transform block based on the target transform kernel to obtain a target block; reconstructing the first color component of the video block based on at least the target block; A method having the following.
Citation Information
Patent Citations
Video signal encoding / decoding method and device therefor
JP2021509559A
Image processing apparatus and method
US20210021870A1
Image processing device and method
WO2019188466A1