Chroma prediction from luma based on merged chroma blocks
By averaging reconstructed neighboring luma samples and using this average for chroma from luma prediction, the method addresses inefficiencies in predicting small chroma blocks with differing transform unit depths, enhancing prediction accuracy and efficiency.
Patent Information
- Application Number
- JP2024514711
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2022-10-20
- Publication Date
- 2025-06-05
AI Technical Summary
Existing video coding techniques face challenges in efficiently predicting chroma blocks from luma blocks, particularly in scenarios where the size of chroma blocks is small or when there are differences in transform unit depths between luma and chroma blocks.
The proposed method involves determining whether a set of chroma blocks should be predicted in a chroma from luma (CfL) prediction mode based on their size and transform unit depth. These chroma blocks are then combined into a single block, and reconstructed neighboring luma samples are averaged to generate a neighboring luma average. This average is used for CfL prediction of the chroma blocks.
This approach improves the prediction accuracy and efficiency of chroma blocks from luma blocks, especially in cases where the chroma blocks are small or have different transform unit depths compared to the luma blocks.
Smart Images

Figure 2025517260000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims the benefit of priority to non-provisional application Ser. No. 17 / 951,952, entitled "CHROMA FROM LUMA PREDICTION BASED ON MERGED CHROMA BLOCKS," filed on September 23, 2022, and U.S. Provisional Application No. 63 / 342,427, entitled "CHROMA FROM LUMA INTRA PREDICTION MODE WITH BLOCK MERGING," filed on May 16, 2022, each of which is incorporated by reference in its entirety.
[0002] This disclosure describes a set of advanced video coding techniques. More specifically, the techniques disclosed involve chroma prediction from luma. [Background technology]
[0003] This background discussion provided herein is intended to generally present the context of the present disclosure. The work of the inventors cited herein is not admitted, expressly or impliedly, as prior art to the present disclosure, as are aspects of the description that may not have been admitted as prior art at the time of filing of this application, to the extent that their work is described in this background section.
[0004] Video coding and decoding may be performed using inter-picture prediction with motion compensation. Uncompressed digital video may include a sequence of pictures, each having a spatial dimension of, for example, 1920x1080 luma samples and associated full or subsampled chroma samples. The sequence of pictures may have a fixed or variable picture rate (alternatively called frame rate), for example, 60 pictures / second or 60 frames / second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with 1920x1080 pixel resolution, a frame rate of 60 frames / second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires more than 600 gigabytes of storage space.
[0005] One objective of video coding and decoding may be to reduce redundancy in an uncompressed input video signal due to compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations thereof, may be employed. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from the compressed original signal via a decoding process. Lossy compression refers to a coding / decoding process where the original video information is not fully preserved during coding and cannot be fully restored during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals will be small enough to make the reconstructed signal useful for its intended application, even with some information loss. For video, lossy compression has been widely adopted in many applications. The amount of tolerable distortion depends on the application. For example, a user of a particular consumer video streaming application may tolerate higher distortion than a user of a movie or television broadcast application. The compression ratio achievable by a particular coding algorithm can be selected or adjusted to reflect different distortion tolerances, and generally, higher distortion tolerances allow for coding algorithms that result in higher losses and higher compression ratios.
[0006] Video encoders and decoders can utilize techniques from a number of broad categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be called an intra-picture. Intra-pictures and their derived pictures, such as independent decoder refresh pictures, may be used to reset the decoder state and thus may be used as the first picture in the coded video bitstream and video session or as still images. Samples of the block after intra prediction may then be transformed to the frequency domain, and the transform coefficients so generated may be quantized before entropy coding. Intra prediction represents a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficients, the fewer bits are needed at a given quantization step size to represent the block after entropy coding.
[0008] Conventional intra-coding, for example as known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to code / decode a block based on surrounding sample data and / or metadata, for example obtained during encoding and / or decoding that are spatially neighbors and that precede in decoding order the block of data being intra-coded or intra-decoded. Such techniques are hereinafter referred to as "intra-prediction" techniques. It should be noted that in at least some cases, intra-prediction uses reference data only from the current picture being reconstructed, and not from other reference pictures.
[0009] Intra prediction may take many different forms. When more than one of such techniques is available in a given video coding technique, the technique used may be referred to as an intra prediction mode. One or more intra prediction modes may be provided in a particular codec. In certain cases, a mode may have sub-modes and / or may be associated with various parameters, and the mode / sub-mode information and intra coding parameters of a block of video may be coded individually or collectively included in the codeword of the mode. Which codeword is used for a given mode, sub-mode, and / or parameter combination may affect coding efficiency gains via intra prediction and may also affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further improved in newer coding techniques such as Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). In general, in intra prediction, the predictor block may be formed using neighboring sample values that become available. For example, available values of a particular set of neighboring samples along a particular direction and / or line may be copied to the predictor block. The reference to the direction in use may be coded in the bitstream or may itself be predicted.
[0011] Referring to FIG. 1A, at the bottom right, a subset of nine predictor directions specified in the 33 possible intra predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes specified in H.265) is shown. The point 101 where the arrows converge represents the sample to be predicted. The arrows represent the direction in which neighboring samples are used to predict the sample at 101. For example, arrow 102 indicates that sample 101 is predicted from a neighboring sample to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow 103 indicates that sample 101 is predicted from a neighboring sample to the lower left of sample 101 at an angle of 22.5 degrees from the horizontal.
[0012] Still referring to FIG. 1A, a square block 104 of 4×4 samples (indicated by a thick dashed line) is depicted at the top left. The square block 104 contains 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample of the block 104 in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. An exemplary reference sample is also shown that follows a similar numbering scheme. The reference sample is labeled R with its Y position (e.g., row index) and X position (column index) relative to the block 104. In both H.264 and H.265, predicted samples that are adjacent and in the neighborhood of the block being reconstructed are used.
[0013] Intra-picture prediction of block 104 can be started by copying reference sample values from neighboring samples according to the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block 104 indicating the prediction direction of arrow 102, i.e., the sample is predicted from the top right predicted sample at an angle of 45 degrees from the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.
[0014] In certain cases, to calculate a reference sample, especially when the orientation is not evenly divisible by 45 degrees, the values of multiple reference samples may be combined, for example by interpolation.
[0015] The number of possible directions has increased as video coding technology continues to develop. For example, in H.264 (2003), nine different directions are available for intra prediction. This increases to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and those most suitable directions can be encoded with a small number of bits using certain techniques of entropy coding, accepting certain bit penalties for the directions. Furthermore, the direction itself may in some cases be predictable from the neighboring directions used in the intra prediction of the decoded neighboring blocks.
[0016] FIG. 1B shows a diagram 180 depicting 65 intra prediction directions according to JEM to illustrate the increase in the number of prediction directions in various encoding techniques developed over time.
[0017] The manner in which bits representing intra-prediction directions in a coded video bitstream are mapped to prediction directions may vary across video coding techniques, and may range, for example, from simple direct mapping of prediction directions to intra-prediction modes to complex adaptive schemes involving codewords, most probable modes, and similar techniques. In all cases, however, there may be certain directions of intra-prediction that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, in a properly designed video coding technique, those less likely directions may be represented with more bits than the more likely directions.
[0018] Inter-picture prediction, or inter-prediction, may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or part thereof (reference picture) may be used to predict a newly reconstructed picture or picture part (e.g., block) after being spatially shifted in a direction indicated by a motion vector (hereinafter, MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (similar to the temporal dimension).
[0019] In some video compression techniques, the current MV applicable to a particular area of sample data can be predicted from other MVs, e.g., from those other MVs related to other areas of sample data that are spatially adjacent to the area being reconstructed and that precede the current MV in decoding order. Doing so can substantially reduce the overall amount of data required to code the MV by relying on the removal of redundancy of correlated MVs, thereby increasing compression efficiency. MV prediction can work effectively because, for example, when coding an input video signal derived from a camera (known as natural video), areas larger than the area where a single MV is applicable have a statistical likelihood to move in a similar direction in the video sequence, and therefore can potentially be predicted using similar motion vectors derived from MVs of nearby areas. As a result, the actual MV of a given area is similar or identical to the MV predicted from the surrounding MVs. Such MVs can then be represented with fewer bits after entropy coding than would be used if the MVs were directly coded instead of predicted from nearby MVs. In some cases, MV prediction may be an example of lossless compression of a signal (i.e., MV) derived from an original signal (i.e., a sample stream). In other cases, MV prediction itself may be lossy, for example due to rounding errors when computing a predictor from some surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms specified by H.265, the one described below is a technique hereafter referred to as "spatial merging".
[0021] Specifically, referring to FIG. 2, a current block (201) contains samples that are detected by the encoder during the motion search process as predictable from a spatially shifted previous block of the same size. Instead of coding its MV directly, the MV may be derived from metadata associated with one or more reference pictures, e.g., from the most recent reference picture (in decoding order), using MVs associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction may use predictors from the same reference picture that neighboring blocks use. Summary of the Invention [Means for solving the problem]
[0022] Aspects of the present disclosure provide methods and apparatus for chroma from luma (CfL) prediction.
[0023] In some implementations, a method for video processing includes determining that a plurality of chroma blocks from a video bitstream should be predicted in a chroma from luma (CfL) prediction mode, where a first width or a first height of each of the plurality of chroma blocks is less than or equal to a first predetermined threshold; combining the plurality of chroma blocks into a combined chroma block, where a second width or a second height of the combined chroma block is greater than or equal to a second predetermined threshold; determining a plurality of reconstructed neighboring luma samples of one or more luma blocks corresponding to the combined chroma block; averaging the plurality of reconstructed neighboring luma samples to generate a neighboring luma average; and performing CfL prediction of the plurality of chroma blocks based on at least the neighboring luma average.
[0024] In some other implementations, a method for video processing includes comparing at least one of a size of a chroma block with at least one size threshold, or a transform unit (TU) depth of the chroma block with a TU depth of a corresponding luma block, determining a type of chroma from luma (CfL) prediction process for the chroma block from among a plurality of types of CfL prediction processes based on the comparison, and performing a CfL prediction process for the chroma block according to the type of CfL prediction process.
[0025] In some other embodiments, a device for processing video information is disclosed. The device may include circuitry configured to perform any one of the method embodiments described above.
[0026] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding and / or video encoding.
[0027] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]
[0028] [Figure 1A] FIG. 13 is a schematic diagram of an example subset of intra-prediction directional modes. [Figure 1B] FIG. 2 illustrates an exemplary intra-prediction direction. [Diagram 2] 1 illustrates a schematic diagram of a current block and its surrounding spatial merging candidates for motion vector prediction in an example. [Diagram 3] 3 illustrates a simplified block diagram schematic of a communication system 300 in accordance with an exemplary embodiment. [Figure 4] 4 illustrates a simplified block diagram schematic of a communication system 400 in accordance with an exemplary embodiment. [Diagram 5] 1 shows a schematic diagram of a simplified block diagram of a video decoder according to an exemplary embodiment; [Figure 6] 1 shows a schematic diagram of a simplified block diagram of a video encoder according to an example embodiment; [Figure 7] 4 shows a block diagram of a video encoder according to another example embodiment. [Figure 8] 4 shows a block diagram of a video decoder according to another exemplary embodiment. [Figure 9] 1 illustrates a coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 10] 1 illustrates another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 11] 1 illustrates another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 12] 1 illustrates an example of partitioning a base block into coding blocks according to an exemplary partitioning scheme. [Figure 13] 1 illustrates an exemplary ternary partitioning scheme. [Figure 14] 1 illustrates an exemplary quadtree / binary tree coding block partitioning scheme. [Figure 15] 4 illustrates a scheme for partitioning a coding block into multiple transform blocks and a coding order for the transform blocks, according to an example embodiment of the present disclosure. [Figure 16] 4 illustrates another scheme for partitioning a coding block into multiple transform blocks and the coding order of the transform blocks according to an example embodiment of the present disclosure. [Figure 17] 1 illustrates another scheme for partitioning a coding block into multiple transform blocks, according to an example embodiment of this disclosure. [Figure 18] 1 illustrates an example fine angle in directional intra prediction. [Figure 19]13 shows the nominal angle for directional intra prediction. [Figure 20] Indicates the top, left, and top-left positions for the PAETH mode in the block. [Figure 21] 1 illustrates an exemplary recursive intra-filtering mode. [Figure 22A] 1 shows a block diagram of a luma-to-chroma (CfL) prediction unit configured to generate prediction samples of a chroma block based on luma samples of the luma block. [Figure 22B] 1 shows a block diagram of a CfL prediction unit configured to generate predicted samples of a chroma block based on neighboring luma samples of a co-located luma block. [Figure 22C] 1 shows a block diagram of a CfL prediction unit configured to generate predicted samples of a chroma block based on luma samples of a co-located luma block and neighboring luma samples. [Figure 23] 1 illustrates a flow diagram of an exemplary CfL prediction process. [Figure 24] 1 shows a block diagram of luma samples inside and outside a picture boundary. [Diagram 25] 1 shows a schematic diagram of neighboring luma samples of a luma block. [Figure 26] 13 shows a flowchart of another example of a CfL prediction process. [Figure 27] 13 shows a flowchart of another example of a CfL prediction process. [Figure 28] 13 illustrates a flow diagram of another example of a CfL prediction process. [Figure 29] 13 shows a flowchart of another example of a CfL prediction process. [Diagram 30] 1 illustrates a schematic diagram of luma samples mapped to neighboring chroma samples for CfL prediction of a chroma sample corresponding to the luma sample. [Diagram 31] 1 shows a schematic diagram of patch-based mapping of luma samples to neighboring luma samples. [Diagram 32]1 shows a schematic diagram of an exemplary four reference line intra-coding for a chroma block. [Diagram 33] 13 shows a flowchart of another example of a CfL prediction process. [Diagram 34] 13 shows a flowchart of another example of a CfL prediction process. [Diagram 35] 1 shows a schematic diagram of merged chroma blocks and neighboring luma samples. [Diagram 36] 1 shows a schematic diagram of a merged chroma block and corresponding luma block and neighboring luma and chroma samples. [Figure 37] 1 shows a schematic diagram of a combined chroma block and a luna block corresponding to neighboring luma samples. [Figure 38] 13 shows a flowchart of another example of a CfL prediction process. [Figure 39] 1 shows a schematic diagram of four 4×4 chroma blocks corresponding to an 8×8 luma region divided into four 4×4 luma blocks. [Diagram 40] 1 shows a schematic diagram of a luma coding unit with transform unit splitting and a chroma coding unit without transform unit splitting. [Diagram 41] 1 shows a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0029] FIG. 3 illustrates a simplified block diagram of a communication system 300 according to an embodiment of the present disclosure. The communication system 300 includes a plurality of terminals capable of communicating with each other, for example, via a network 350. For example, the communication system 300 includes a first pair of terminals 310 and 320 interconnected via the network 350. In the example of FIG. 3, the first pair of terminals 310 and 320 can perform unidirectional transmission of data. For example, the terminal 310 can code video data (e.g., of a stream of video pictures captured by the terminal 310) for transmission to another terminal 320 via the network 350. The encoded video data can be transmitted in the form of one or more coded video bitstreams. The terminal 320 can receive the coded video data from the network 350, decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. The unidirectional data transmission can be implemented in media serving applications, etc.
[0030] In another example, the communication system 300 includes a second pair of terminals 330 and 340 performing bidirectional transmission of coded video data, which may be implemented, for example, during video conferencing applications. In the case of bidirectional transmission of data, in one example, each of the terminals 330 and 340 can code video data (e.g., of a stream of video pictures captured by the terminal) for transmission to the other of the terminals 330 and 340 over the network 350. Also, each of the terminals 330 and 340 can receive coded video data transmitted by the other of the terminals 330 and 340, can decode the coded video data to recover the video pictures, and can display the video pictures on an accessible display device according to the recovered video data.
[0031] In the example of FIG. 3, terminals 310, 320, 330, and 340 may be implemented as a server, a personal computer, and a smartphone, although applicability of the underlying principles of the present disclosure may not be so limited. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated videoconferencing equipment, and the like. Network 350 represents any number or type of network that conveys coded video data between terminals 310, 320, 330, and 340, including, for example, wireline and / or wireless communication networks. Communications network 350 may exchange data over circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 350 may not be important to the operation of the present disclosure unless explicitly described herein.
[0032] 4 shows an arrangement of video encoders and video decoders in a video streaming environment as one example for an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video applications including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0033] The video streaming system may include a video source 401, e.g., a video capture subsystem 413 that may include a digital camera, for creating a stream of uncompressed video pictures or images 402. In one example, the stream of video pictures 402 includes samples recorded by the digital camera of the video source 401. The stream of video pictures 402, shown as a thick line to emphasize the large amount of data compared to the encoded video data 404 (or coded video bitstream), may be processed by an electronic device 420 that includes a video encoder 403 coupled to the video source 401. The video encoder 403 may include hardware, software, or a combination thereof for realizing or implementing aspects of the disclosed subject matter, as described in more detail below. The encoded video data 404 (or encoded video bitstream 404), shown as a thin line to emphasize the small amount of data compared to the uncompressed video picture stream 402, may be stored in a streaming server 405 for future use or may be directly stored in a downstream video device (not shown). One or more streaming client subsystems, such as client subsystems 406 and 408 of FIG. 4, can access the streaming server 405 to retrieve copies 407 and 409 of the encoded video data 404. The client subsystem 406 may include a video decoder 410, for example, within the electronic device 430. The video decoder 410 decodes the incoming copy of the encoded video data 407 and creates an outgoing stream of video pictures 411 that are uncompressed and can be rendered on a display 412 (e.g., a display screen) or other rendering device (not shown). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, the encoded video data 404, 407, and 409 (e.g., video bitstreams) may be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC and other video coding standards.
[0034] It should be noted that electronic devices 420 and 430 may include other components (not shown). For example, electronic device 420 may include a video decoder (not shown), and electronic device 430 may likewise include a video encoder (not shown).
[0035] 5 shows a block diagram of a video decoder 510 according to any embodiment of the present disclosure below. The video decoder 510 can be included in an electronic device 530. The electronic device 530 can include a receiver 531 (e.g., a receiving circuit). The video decoder 510 can be used in place of the video decoder 410 of the example of FIG. 4.
[0036] The receiver 531 may receive one or more coded video sequences to be decoded by the video decoder 510. In the same or another embodiment, one coded video sequence may be decoded at a time, and the decoding of each coded video sequence is independent of the other coded video sequences. Each video sequence may be associated with multiple video frames or images. The coded video sequences may be received from a channel 501, which may be a storage device that stores the encoded video data or a hardware / software link to a streaming source that transmits the encoded video data. The receiver 531 may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, and may be forwarded to respective processing circuits (not shown). The receiver 531 may separate the coded video sequences from the other data. To address network jitter, a buffer memory 515 may be disposed between the receiver 531 and the entropy decoder / parser 520 (hereinafter "parser 520"). In certain applications, the buffer memory 515 may be implemented as part of the video decoder 510. In other applications, the buffer memory 515 may be external to and separate from the video decoder 510 (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder 510, for example to deal with network jitter, and there may be another additional buffer memory 515 internal to the video decoder 510, for example to handle playback timing. If the receiver 531 is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory 515 may not be needed or may be small. For use with best-effort packet networks such as the Internet, a sufficient size of the buffer memory 515 may be required, and the size may be relatively large.Such a buffer memory may be implemented with an adaptive size and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder 510.
[0037] The video decoder 510 may include a parser 520 for reconstructing symbols 521 from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder 510 and potentially information for controlling a rendering device such as a display 512 (e.g., a display screen) that may or may not be an integral part of the electronic device 530, but may be coupled to the electronic device 530, as shown in FIG. 5. The control information for the rendering device may be in the form of a supplemental enhancement information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser 520 may parse / entropy decode the coded video sequence received by the parser 520. The entropy coding of the coded video sequence may be in accordance with a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 520 may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 520 may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.
[0038] Parser 520 may perform entropy decoding / parsing operations on the video sequence received from buffer memory 515 to produce symbols 521 .
[0039] The reconstruction of symbols 521 may involve a number of different processing or functional units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.) and other factors. The units involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by parser 520. The flow of such subgroup control information between parser 520 and the following processing or functional units is not shown for the sake of simplicity.
[0040] Besides the functional blocks already mentioned, the video decoder 510 may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these functional units may closely interact with each other and may be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, a conceptual subdivision into functional units is adopted in the following disclosure.
[0041] The first unit may include a scalar / inverse transform unit 551. The scalar / inverse transform unit 551 may receive quantized transform coefficients and control information from the parser 520, including information indicating which type of inverse transform to use, block size, quantization coefficients / parameters, quantization scaling matrices, and states (lies) as symbols 521. The scalar / inverse transform unit 551 may output a block including sample values that may be input to an aggregator 555.
[0042] In some cases, the output samples of the scalar / inverse transform 551 may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 552. In some cases, the intra-picture prediction unit 552 may generate a block of the same size and shape as the block being reconstructed using neighboring block information already reconstructed and stored in the current picture buffer 558. The current picture buffer 558 may, for example, buffer the partially reconstructed current picture and / or the fully reconstructed current picture. In some implementations, the aggregator 555 may add the prediction information generated by the intra-prediction unit 552 to the output sample information provided by the scalar / inverse transform unit 551 on a sample-by-sample basis.
[0043] In other cases, the output samples of the scalar / inverse transform unit 551 may relate to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 553 may access the reference picture memory 557 to fetch samples used for inter-picture prediction. After motion compensating the fetched samples according to the symbols 521 related to the block, these samples may be added to the output of the scalar / inverse transform unit 551 (the output of the unit 551 may be referred to as a residual sample or a residual signal) by the aggregator 555 to generate output sample information. The addresses in the reference picture memory 557 from which the motion compensation prediction unit 553 fetches prediction samples available to the motion compensation prediction unit 553 in the form of the symbols 521, which may have, for example, X, Y components (shift), and reference picture components (time), may be controlled by the motion vector. Motion compensation may also include interpolation of sample values fetched from the reference picture memory 557 when a sub-sample accurate motion vector is used, and may also be associated with a motion vector prediction mechanism, etc.
[0044] In loop filter unit 556, the output samples of aggregator 555 may be subjected to various loop filtering techniques. Video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also called coded video bitstream) and made available to loop filter unit 556 as symbols 521 from parser 520, but may also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, or to previously reconstructed, loop filtered sample values. As described in more detail below, several types of loop filters may be included as part of loop filter unit 556, in various orders.
[0045] The output of the loop filter unit 556 may be a sample stream that may be output to the rendering device 512 as well as stored in a reference picture memory 557 for use in future inter-picture prediction.
[0046] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future inter-picture prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 520), the current picture buffer 558 can become part of reference picture memory 557, and a new current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0047] The video decoder 510 may perform decoding operations according to a given video compression technique adopted in a standard, such as ITU-T Rec. H.265. The coded video sequence may conform to a syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and to a profile documented in the video compression technique or standard. Specifically, a profile may select a particular tool from all tools available in the video compression technique or standard as the only tool available for use under that profile. To conform to a standard, the complexity of the coded video sequence may be within a range defined by a level of the video compression technique or standard. In some cases, the level limits a maximum picture size, a maximum frame rate, a maximum reconstructed sample rate (e.g., measured in megasamples per second), a maximum reference picture size, etc. The limits set by the level may in some cases be further limited by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0048] In some example embodiments, the receiver 531 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder 510 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0049] 6 shows a block diagram of a video encoder 603 according to an exemplary embodiment of the present disclosure. The video encoder 603 may be included in an electronic device 620. The electronic device 620 may further include a transmitter 640 (e.g., a transmitting circuit). The video encoder 603 may be used in place of the video encoder 403 of the example of FIG. 4.
[0050] The video encoder 603 may receive video samples from a video source 601 (which is not part of the electronic device 620 in the example of FIG. 6) that may capture video images to be coded by the video encoder 603. In another example, the video source 601 may be implemented as part of the electronic device 620.
[0051] The video source 601 may provide a source video sequence to be coded by the video encoder 603 in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, XYZ ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media presentation system, the video source 601 may be a storage device capable of storing pre-prepared videos. In a video conferencing system, the video source 601 may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures or images that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc. being used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0052] According to some example embodiments, the video encoder 603 can code and compress pictures of a source video sequence into a coded video sequence 643 in real time or under other time constraints required by the application. Enforcing an appropriate coding rate constitutes one function of the controller 650. In some embodiments, the controller 650 can be functionally coupled to and control other functional units, as described below. For the sake of brevity, couplings are not shown. Parameters set by the controller 650 can include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller 650 can be configured with other appropriate functions related to the video encoder 603 optimized for a particular system design.
[0053] In some example embodiments, the video encoder 603 may be configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop may include a source coder 630 (responsible for creating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture, for example) and a (local) decoder 633 embedded in the video encoder 603. The decoder 633 reconstructs the symbols to create sample data in a similar manner as a (remote) decoder would create them, even if the embedded decoder 633 processes a video stream coded by the source coder 630 without entropy coding (since any compression between the symbols and the coded video bitstream in entropy coding may be lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory 634. Since the decoding of the symbol stream produces bit-exact results regardless of the location of the decoder (local or remote), the content in the reference picture memory 634 is also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values as the reference picture samples that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, e.g., due to channel error) is used to improve the quality of coding.
[0054] The operation of the "local" decoder 633 may be the same as that of a "remote" decoder, such as the video decoder 510 already described in detail above in connection with Figure 5. However, with brief reference also to Figure 5, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder 645 and parser 520 may be lossless, the entropy decoding portion of the video decoder 510, including the buffer memory 515, and the parser 520 may not be fully implemented in the local decoder 633 within the encoder.
[0055] At this point, it can be said that any decoder technology, except for parsing / entropy decoding, which may only exist in the decoder, may also necessarily need to exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may focus on the decoder operation, which is similar to the decoding part of the encoder. Thus, the description of the encoder technology may be omitted, since it is the reverse of the decoder technology described in general. Only in certain areas or aspects, a more detailed description of the encoder is provided below.
[0056] During operation in some example implementations, the source coder 630 may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine 632 codes color channel differences (or residuals) between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture. The terms "residual" and its adjective form "residual" may be used interchangeably.
[0057] The local video decoder 633 may decode the coded video data of the pictures that may be designated as reference pictures based on the symbols created by the source coder 630. The operation of the coding engine 632 may advantageously be a lossy process. When the coded video data is decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 633 may replicate the decoding process that may be performed by the video decoder on the reference pictures to store the reconstructed reference pictures in the reference picture cache 634. In this way, the video encoder 603 may locally store copies of reconstructed reference pictures that have a common content with the reconstructed reference pictures obtained by the far-end (remote) video decoder (without transmission errors).
[0058] The predictor 635 may perform a predictive search for the coding engine 632. That is, for a new picture to be coded, the predictor 635 may search the reference picture memory 634 for sample data (as candidate reference pixel blocks) that may serve as suitable predictive references for the new picture, or specific metadata, such as reference picture motion vectors, block shapes, etc. The predictor 635 may operate on one sample block per pixel block to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor 635, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 634.
[0059] The controller 650 may manage the coding operations of the source coder 630, including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0060] The output of all the aforementioned functional units may be subjected to entropy coding in the entropy coder 645. The entropy coder 645 converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0061] The transmitter 640 may buffer the coded video sequence created by the entropy coder 645 in preparation for transmission over a communication channel 660, which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter 640 may merge the coded video data from the video coder 603 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0062] The controller 650 may manage the operation of the video encoder 603. During coding, the controller 650 may assign a particular coded picture type to each coded picture, which may affect the coding techniques that may be applied to the respective picture. For example, pictures may often be assigned as one of the following picture types:
[0063] An intra picture (I-picture) may be a picture that can be coded and decoded without using any other picture in a sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.
[0064] A predictive picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which predicts the sample values of each block using at most one motion vector and reference index.
[0065] A bidirectionally predicted picture (B-picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which predicts sample values for each block using up to two motion vectors and reference indexes. Similarly, a multi-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0066] A source picture may generally be spatially subdivided into multiple sample coding blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture of the block. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). A pixel block of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A block of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures. A source picture or an intermediate processed picture may be subdivided into other types of blocks for other purposes. The division of coding blocks and other types of blocks may or may not follow the same scheme, as described in more detail below.
[0067] The video encoder 603 may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder 603 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.
[0068] In some example embodiments, the transmitter 640 can transmit additional data along with the encoded video. The source coder 630 can include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0069] A video may be captured in time sequence as multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits temporal or other correlation between pictures. For example, a particular picture being encoded / decoded, called the current picture, may be partitioned into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, it may be coded by a vector, called a motion vector. A motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0070] In some exemplary embodiments, bi-prediction techniques may be used for inter-picture prediction. According to such bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which advance the current picture in the video in decoding order (but may be in the past or future, respectively, in display order). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block may be jointly predicted by a combination of the first reference block and the second reference block.
[0071] Furthermore, merge mode techniques may be used to improve coding efficiency in inter-picture prediction.
[0072] According to some exemplary embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, a picture in a sequence of video pictures is partitioned into coding tree units (CTUs) for compression, and the CTUs in a picture may have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU may include three parallel coding tree blocks (CTBs), namely, one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels may be partitioned into one CU of 64×64 pixels, or four CUs of 32×32 pixels. Each of one or more of the 32×32 blocks may be further partitioned into four CUs of 16×16 pixels. In some exemplary embodiments, each CU may be analyzed during encoding to determine the prediction type of the CU from among various prediction types, such as an inter prediction type or an intra prediction type. The CU may be divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. In general, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. The division of the CU into PUs (or PBs of different color channels) may be performed in various spatial patterns. The luma PB or chroma PB may include a matrix of sample values (e.g., luma values), such as, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, etc.
[0073] 7 shows a diagram of a video encoder 703 according to another example embodiment of this disclosure. The video encoder 703 is configured to receive a processed block (e.g., a predictive block) of sample values in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. The example video encoder 703 may be used in place of the example video encoder (403) of FIG. 4.
[0074] For example, the video encoder 703 receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder 703 then determines whether the processing block is best coded using intra mode, inter mode, or bi-prediction mode, for example, using rate-distortion optimization (RDO). If it is determined that the processing block is coded in intra mode, the video encoder 703 may encode the processing block into a coded picture using intra prediction techniques, and if it is determined that the processing block is coded in inter mode or bi-prediction mode, the video encoder 703 may encode the processing block into a coded picture using inter prediction techniques or bi-prediction techniques, respectively. In some exemplary embodiments, as a sub-mode of inter-picture prediction, a merge mode may be used in which a motion vector is derived from one or more motion vector predictors without benefit of coded motion vector components outside the predictors. In some other exemplary embodiments, there may be motion vector components applicable to the current block. Thus, the video encoder 703 may include components not explicitly shown in FIG. 7, such as a mode decision module, to determine the perduction mode of the processing block.
[0075] In the example of FIG. 7, the video encoder 703 includes an inter-encoder 730, an intra-encoder 722, a residual calculator 723, a switch 726, a residual encoder 724, an overall controller 721, and an entropy encoder 725 coupled to each other as shown in the exemplary configuration of FIG. 7.
[0076] The inter-encoder 730 is configured to receive a sample of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures in display order), generate inter-prediction information (e.g., a description of redundant information due to inter-encoding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that has been decoded based on the encoded video information using a decoding unit 633 incorporated in the example encoder 620 of FIG. 6 (depicted as a residual decoder 728 of FIG. 7, as described in more detail below).
[0077] The intra encoder 722 is also configured to receive samples of a current block (e.g., a processing block), compare the block with blocks already encoded in the same picture, generate quantized coefficients after transformation, and possibly generate intra prediction information (e.g., intra prediction direction information by one or more encoding techniques). The intra encoder 722 can calculate an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.
[0078] The overall controller 721 may be configured to determine overall control data and control other components of the video encoder 703 based on the overall control data. In one example, the overall controller 721 determines a prediction mode of the block and provides a control signal to the switch 726 based on the prediction mode. For example, if the prediction mode is an intra mode, the overall controller 721 controls the switch 726 to select an intra mode result used by the residual calculator 723 and controls the entropy encoder 725 to select intra prediction information and include the intra prediction information in the bitstream. Also, if the prediction mode of the block is an inter mode, the overall controller 721 controls the switch 726 to select an inter prediction result used by the residual calculator 723 and controls the entropy encoder 725 to select inter prediction information and include it in the bitstream.
[0079] The residual calculator 723 may be configured to calculate a difference (residual data) between the received block and a prediction result of the block selected from the intra encoder 722 or the inter encoder 730. The residual encoder 724 may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder 724 may be configured to transform the residual data from a spatial domain to a frequency domain to generate transform coefficients. Then, the transform coefficients are subjected to a quantization process to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder 703 also includes a residual decoder 728. The residual decoder 728 is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be appropriately used in the intra encoder 722 and the inter encoder 730. For example, the inter encoder 730 may generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder 722 may generate a decoded block based on the decoded residual data and the intra prediction information. The decoded blocks are appropriately processed to generate a decoded picture, which may be buffered in a memory circuit (not shown) and used as a reference picture.
[0080] The entropy encoder 725 may be configured to format a bitstream to include the encoded block and perform entropy coding. The entropy encoder 725 is configured to include various information in the bitstream. For example, the entropy encoder 725 may be configured to include overall control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other suitable information in the bitstream. When coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, the residual information may not be present.
[0081] 8 shows a diagram of an example video decoder 810 according to another embodiment of the present disclosure. The video decoder 810 is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder 810 may be used in place of the video decoder 410 of the example of FIG. 4.
[0082] In the example of FIG. 8, the video decoder 810 includes an entropy decoder 871, an inter decoder 880, a residual decoder 873, a reconstruction module 874, and an intra decoder 872 coupled together as shown in the exemplary configuration of FIG.
[0083] The entropy decoder 871 may be configured to reconstruct from the coded picture certain symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information) that may identify the mode in which the block is coded (e.g., intra- or bi-prediction mode, merged submode or another submode), certain samples or metadata used for prediction by the intra decoder 872 or the inter decoder 880, residual information in the form of, for example, quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter decoder 880, and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder 872. The residual information is dequantized and provided to the residual decoder 873.
[0084] The inter decoder 880 may be configured to receive the inter prediction information and generate inter prediction results based on the inter prediction information.
[0085] The intra decoder 872 may be configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0086] The residual decoder 873 may be configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder 873 may also utilize certain control information (to include quantizer parameters (QP)) that may be provided by the entropy decoder 871 (data path not shown as this is likely to be low data amount control information).
[0087] The reconstruction module 874 may be configured to combine, in the spatial domain, the residual output by the residual decoder 873 and a prediction result (possibly output by an inter-prediction module or an intra-prediction module) to form a reconstructed block that forms part of the reconstructed picture as part of the reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may also be performed to improve visual quality.
[0088] It should be noted that the video encoders 403, 603, and 703 and the video decoders 410, 510, and 810 may be implemented using any suitable technique. In some exemplary embodiments, the video encoders 403, 603, and 703 and the video decoders 410, 510, and 810 may be implemented using one or more integrated circuits. In another embodiment, the video encoders 403, 603, and 603 and the video decoders 410, 510, and 810 may be implemented using one or more processors executing software instructions.
[0089] Turning to block partitioning for coding and decoding, a general partitioning can start from a base block and can follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. The partitioning can be hierarchical and recursive. After isolating or partitioning the base block according to any of the exemplary partitioning procedures described below or other procedures, or combinations thereof, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of various partitioning levels in the partitioning hierarchy and can be of various shapes of partitioning. Each of the partitions can be referred to as a coding block (CB). In various exemplary partitioning implementations described further below, each resulting CB can be a CB of any of the allowed sizes and partitioning levels. Such partitions are referred to as coding blocks because they form a unit for which some basic coding / decoding decisions can be made and coding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning tree structure. The coding block may be a luma coding block or a chroma coding block. The CB tree structure for each color may be referred to as a coding block tree (CBT).
[0090] The coding blocks of all color channels may be collectively referred to as a coding unit (CU). The hierarchical structure of all color channels may be collectively referred to as a coding tree unit (CTU). The partitioning pattern or structure of various color channels within a CTU may or may not be the same.
[0091] In some implementations, the partition tree scheme or structure used for the luma channel and the chroma channel may not need to be the same. In other words, the luma channel and the chroma channel may have separate coding tree structures or patterns. Furthermore, whether the luma channel and the chroma channel use the same coding partition tree structure or different coding partition tree structures, and the actual coding partition tree structure to be used may depend on whether the slice being coded is a P slice, a B slice, or an I slice. For example, for an I slice, the chroma channel and the luma channel may have separate coding partition tree structures or coding partition tree structure modes, while for a P slice or a B slice, the luma channel and the chroma channel may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, the luma channel may be partitioned into CBs by one coding partition tree structure, and the chroma channel may be partitioned into chroma CBs by another coding partition tree structure.
[0092] In some exemplary implementations, a predefined partitioning pattern may be applied to the base block. As shown in FIG. 9, an exemplary four-way partition tree may start at a first predefined level (e.g., 64×64 block level or other size as the base block size), and the base block may be partitioned hierarchically down to a predefined lowest level (e.g., 4×4 level). For example, the base block may follow four predefined partitioning options or patterns shown by 902, 904, 906, and 908, and the partition designated as R allows for recursive partitioning in that the same partitioning option shown in FIG. 9 may be repeated at a lower scale down to the lowest level (e.g., 4×4 level). In some implementations, additional restrictions may be added to the partitioning scheme of FIG. 9. In the implementation of FIG. 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed but not recursive, and square partitions are allowed to be recursive. Subsequent partitioning of FIG. 9 by recursion, if necessary, generates a final set of coding blocks. A coding tree depth may be further defined to indicate the partition depth from a root node or root block. For example, the coding tree depth for a root node or root block, e.g., a 64x64 block, may be set to 0, and the coding tree depth increases by 1 after the root block is further partitioned one time according to FIG. 9. The maximum or deepest level from the 64x64 base block to the 4x4 minimum partition is 4 (starting from level 0) in the above scheme. Such a partitioning scheme may be applied to one or more of the color channels. Each color channel may be partitioned independently according to the scheme of FIG. 9 (e.g., a partitioning pattern or option among predefined patterns may be determined independently for each of the color channels at each hierarchical level).Alternatively, two or more color channels may share the same hierarchical pattern tree of FIG. 9 (e.g., the same partitioning pattern or option among the predefined patterns may be selected for two or more color channels at each hierarchical level).
[0093] FIG. 10 illustrates another exemplary predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. As shown in FIG. 10, ten exemplary partitioning structures or patterns may be predefined. The root block may start from a predefined level (e.g., from a base block at a 128×128 level or a 64×64 level). The exemplary partitioning structure of FIG. 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. The partition type with three subpartitions shown at 1002, 1004, 1006, and 1008 in the second column of FIG. 10 may be referred to as a “T-type” partition. The “T-type” partitions 1002, 1004, 1006, and 1008 may be referred to as a left T-type, an upper T-type, a right T-type, and a lower T-type. In some exemplary implementations, none of the rectangular partitions of FIG. 10 may be further refined. A coding tree depth may be further defined to indicate the division depth from a root node or root block. For example, the coding tree depth for a root node or root block, e.g., a 128x128 block, may be set to 0, and the coding tree depth increases by 1 after the root block is further divided according to FIG. 10. In some implementations, only the full square partition of 1010 may allow recursive partitioning to the next level of the partitioning tree following the pattern of FIG. 10. In other words, recursive partitioning may not be possible for the square partitions in the T-type patterns 1002, 1004, 1006, and 1008. If necessary, the partitioning procedure following FIG. 10 by recursion generates a final set of coding blocks. Such a scheme may be applied to one or more of the color channels. In some implementations, more flexibility may be added to the use of partitions less than 8x8 levels. For example, 2x2 chroma inter prediction may be used in certain cases.
[0094] In some other exemplary implementations of coding block partitioning, a quadtree structure may be used to divide a base block or an intermediate block into quadtree partitions. Such quadtree division may be applied hierarchically and recursively to any square partition. Whether a base block or an intermediate block or partition is further quadtree divided may be adapted to various local characteristics of the base block or intermediate block / partition. The quadtree partitioning at the picture boundary may be further adapted. For example, an implicit quadtree division may be performed at the picture boundary such that a block continues to be quadtree divided until its size fits into the picture boundary.
[0095] In some other exemplary implementations, hierarchical bisection partitioning from the base block may be used. For such a scheme, the base block or mid-level block may be partitioned into two partitions. The bisection partitioning may be either horizontal or vertical. For example, horizontal bisection partitioning may divide the base block or mid-block into equal partitions on the left and right. Similarly, vertical bisection partitioning may divide the base block or mid-block into equal partitions on the top and bottom. Such bisection partitioning may be hierarchical and recursive. It may be determined at each of the base block or mid-block whether the bisection partitioning scheme should continue and whether horizontal or vertical bisection partitioning should be used if the scheme continues further. In some implementations, further partitioning may stop at a predefined minimum partition size (in one or both dimensions). Alternatively, further partitioning may stop when a predefined partitioning level or depth from the base block is reached. In some implementations, the aspect ratio of the partitions may be limited. For example, the aspect ratio of a partition may not be less than 1:4 (or greater than 4:1), so a vertical strip partition having a 4:1 vertical-to-horizontal aspect ratio may only be further bipartitioned vertically into upper and lower partitions, each having a 2:1 vertical-to-horizontal aspect ratio.
[0096] In yet some other examples, a ternary partitioning scheme may be used to partition the base block or any intermediate blocks, as shown in FIG. 13. The ternary pattern may be implemented vertically, as shown at 1302 in FIG. 13, or horizontally, as shown at 1304 in FIG. 13. The exemplary division ratio in FIG. 13 is shown as 1:2:1, either vertically or horizontally, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Such a ternary partitioning scheme may be used to complement quadtree or binary partitioning structures, in that such ternary partitioning is capable of capturing objects located at block centers in one contiguous partition, whereas quadtrees and binary trees always split along block centers, thus splitting objects into separate partitions. In some implementations, the width and height of the partitions of the exemplary ternary tree are always powers of two to avoid further transformations.
[0097] The above partitioning schemes may be combined in any manner at different partitioning levels. As an example, the above-mentioned quadtree and binary partitioning schemes may be combined to partition the base block into a quadtree-binary tree (QTBT) structure. In such a scheme, the base block or intermediate blocks / partitions may be either quadtree-partitioned or bipartitioned, if specified, according to a set of predefined conditions. A specific example is shown in FIG. 14. In the example of FIG. 14, the base block is first quadtree-partitioned into four partitions, as shown by 1402, 1404, 1406, and 1408. Each of the resulting partitions is then either quadtree-partitioned into four further partitions (such as 1408) or bipartitioned into two further partitions at the next level (e.g., either horizontally or vertically, such as 1402 or 1406, both of which are symmetric), or not partitioned (such as 1404). Bisection or quadtree partitioning may be recursively enabled for square partitions as shown by the overall exemplary partition pattern in 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree partitioning and dashed lines represent bisection. A flag may be used for each bisection node (non-leaf bisection partition) to indicate whether the bisection is horizontal or vertical. For example, flag "0" may represent horizontal bisection and flag "1" may represent vertical bisection, as shown in 1420, which is consistent with the partitioning structure in 1410. In the case of quadtree partitioning, there is no need to indicate the partition type, since quadtree partitioning always splits a block or partition both horizontally and vertically to generate four sub-blocks / partitions of equal size. In some implementations, flag "1" may represent horizontal bisection and flag "0" may represent vertical bisection.
[0098] In some example implementations of QTBT, the quadtree and bisection rule sets may be represented by the following predefined parameters and their associated corresponding functions: -CTU size: quadtree root node size (base block size) -MinQTSize: The minimum allowable quadtree leaf node size. -MaxBTSize: Maximum allowable binary tree root node size -MaxBTDepth: Maximum allowed binary tree depth -MinBTSize: The minimum allowable binary tree leaf node size. In some exemplary implementations of the QTBT partitioning structure, the CTU size may be set as 128×128 luma samples with two corresponding 64×64 blocks of chroma samples (when exemplary chroma subsampling is considered and used), the MinQTSize may be set as 16×16, the MaxBTSize may be set as 64×64, the MinBTSize may be set as 4×4 (for both width and height), and the MaxBTDepth may be set as 4. Quad-tree partitioning may be applied to the CTU first to generate quad-tree leaf nodes. The quad-tree leaf nodes may have a size from its minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If a node is 128×128, it will not be split initially by the binary tree because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes that do not exceed MaxBTSize may be partitioned by a binary tree. In the example of FIG. 14, the base block is 128×128. The base block can only be quadtree partitioned according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the resulting four partitions is 64×64, not exceeding MaxBTSize, and may be further quadtree or binary tree partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further partitions may be considered. When the width of a binary tree node is equal to MinBTSize (i.e., 4), no further horizontal partitions may be considered. Similarly, when the height of a binary tree node is equal to MinBTSize, no further vertical partitions may be considered.
[0099] In some example implementations, the above QTBT scheme may be configured to support flexibility for luma and chroma to have the same or separate QTBT structures. For example, for P slices and B slices, the luma CTB and chroma CTB in one CTU may share the same QTBT structure. However, for I slices, the luma CTB may be partitioned into CBs by a QTBT structure, and the chroma CTB may be partitioned into chroma CBs by another QTBT structure. This means that CUs may be used to refer to different color channels in an I slice, e.g., an I slice may consist of a coding block of a luma component or a coding block of two chroma components, and a CU in a P slice or B slice may consist of coding blocks of all three color components.
[0100] In some other implementations, the QTBT scheme may be complemented with the ternary scheme described above. Such implementations may be referred to as multi-type tree (MTT) structures. For example, in addition to bisection of nodes, one of the ternary partition patterns of FIG. 13 may be selected. In some implementations, only square nodes may undergo ternary division. An additional flag may be used to indicate whether the ternary partitioning is horizontal or vertical.
[0101] The design of two-level or multi-level trees, such as the QTBT implementation and the QTBT implementation complemented by trisection, may be motivated primarily by reducing complexity. In theory, the complexity of traversing a tree is T D where T denotes the number of split types and D is the depth of the tree. A tradeoff may be made by using multitypes (T) while reducing the depth (D).
[0102] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into multiple prediction blocks (PBs) for the purpose of intra-frame or inter-frame prediction during the coding and decoding process. In other words, the CB may be further divided into different sub-partitions, where individual prediction decisions / configurations may be made. In parallel, the CB may be further partitioned into multiple transform blocks (TBs) for the purpose of describing the level at which the transformation or inverse transformation of the video data is performed. The partitioning scheme of the CB into PBs and TBs may be the same or different. For example, each partitioning scheme may be performed using its own procedure, for example, based on various characteristics of the video data. The partitioning schemes of the PBs and TBs may be independent in some exemplary implementations. The partitioning schemes and boundaries of the PBs and TBs may be correlated in some other exemplary implementations. In some implementations, for example, the TBs may be partitioned after the PB partition, and in particular, each PB may be further partitioned into one or more TBs after being determined following the partitioning of the coding blocks. For example, in some implementations, the PB may be divided into one, two, four, or some other number of TBs.
[0103] In some implementations, the luma and chroma channels may be processed differently to partition base blocks into coding blocks and further into predictive and / or transform blocks. For example, in some implementations, partitioning of coding blocks into predictive and / or transform blocks may be allowed for the luma channel, but such partitioning of coding blocks into predictive and / or transform blocks may not be allowed for the chroma channels. In such implementations, therefore, transformation and / or prediction of luma blocks may only be performed at the coding block level. In another example, the minimum transform block size of the luma and chroma channels may be different, e.g., coding blocks of the luma channel may be allowed to be partitioned into smaller transform and / or predictive blocks than the chroma channels. In yet another example, the maximum depth of partitioning of coding blocks into transform and / or predictive blocks may differ between the luma and chroma channels, e.g., coding blocks of the luma channel may be allowed to be partitioned into deeper transform and / or predictive blocks than the chroma channels. As a specific example, a luma coding block may be partitioned into transform blocks of multiple sizes that can be represented by recursive partitions going down up to two levels, allowing transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, as well as transform block sizes from 4×4 to 64×64. However, for chroma blocks, only the largest possible transform block designated for the luma block may be allowed.
[0104] In some example implementations for partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partitioning may depend on whether the PB is intra-coded or inter-coded.
[0105] Partitioning of coding blocks (or prediction blocks) into transform blocks may be performed in various exemplary manners, including but not limited to quadtree partitioning and predefined pattern partitioning, recursively or non-recursively, further considering transform blocks at boundaries of coding blocks or prediction blocks. In general, the resulting transform blocks may be at different partitioning levels, may not be the same size, and may not need to be square in shape (e.g., they may be rectangular with some allowed size and aspect ratio). Further examples are described in more detail below with respect to Figures 15, 16, and 17.
[0106] However, in some other implementations, the CB obtained via any of the above partitioning schemes may be used as a basic or smallest coding block for prediction and / or transformation. In other words, no further division is performed for the purpose of performing inter / intra prediction and / or transformation. For example, the CB obtained from the above QTBT scheme may be used as it is as a unit for performing prediction. Specifically, such a QTBT structure removes the concept of multiple partition types, i.e., removes the separation of CU, PU, and TU, and supports more flexibility for CU / CB partition shapes as described above. In such a QTBT block structure, the CU / CB can have either a square or a rectangle. The leaf nodes of such a QTBT are used as units for prediction and transformation processing without any further partitioning. This means that the CU, PU, and TU have the same block size in such an exemplary QTBT coding block structure.
[0107] The various CB partitioning schemes described above, as well as further partitioning of the CB into PB and / or TB (including no PB / TB partitioning), may be combined in any manner. The following specific embodiments are provided as non-limiting examples.
[0108] Specific exemplary implementations of partitioning of coding blocks and transform blocks are described below. In such exemplary implementations, base blocks may be partitioned into coding blocks using recursive quadtree partitioning or predefined partitioning patterns described above (such as the partitioning patterns of Figures 9 and 10). At each level, whether further quadtree partitioning of a particular partition should be continued may be determined by local video data characteristics. The resulting CBs may be at different quadtree partitioning levels and different sizes. The decision on whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction may be made at the CB level (or at the CU level for all three color channels). Each CB may be further partitioned into one, two, four, or other number of PBs according to a predefined PB partitioning type. Within one PB, the same prediction process may be applied, and related information may be transmitted to the decoder on a PB basis. After obtaining the residual blocks by applying the prediction process based on the PB partitioning type, the CBs may be partitioned into TBs according to another quadtree structure similar to the coding tree for CBs. In this particular implementation, the CB or TB need not be limited to a square. Furthermore, in this particular example, the PB may be a square or rectangular shape for inter prediction, and may only be a square for intra prediction. A coding block may be divided, for example, into four square TBs. Each TB may be further divided recursively (using quadtree partitioning) into smaller TBs called residual quadtrees (RQTs).
[0109] Another exemplary embodiment for partitioning a base block into CB, PB, and / or TB is further described below. For example, rather than using multiple partition unit types such as those shown in FIG. 9 or FIG. 10, a quadtree with nested multi-type trees using bipartition and tripartite segmentation structures (e.g., QTBT or QTBT with tripartites as described above) may be used. Separation of CB, PB, and TB (i.e., partitioning of CB into PB and / or TB, and partitioning of PB into TB) may be abandoned except when necessary for CBs with sizes too large for the maximum transform length, when such CBs require further division. This exemplary partitioning scheme may be designed to support further flexibility on the shape of CB partitions, such that both prediction and transformation may be performed at the CB level without further partitioning. In such coding tree structures, CBs may have either square or rectangular shapes. Specifically, a coding tree block (CTB) may first be partitioned by a quadtree structure. The quadtree leaf nodes may then be further partitioned by a nested multi-type tree structure. An example of a nested multi-type tree structure using bisection or trisection is shown in Figure 11. Specifically, the exemplary multi-type tree structure in Figure 11 includes four split types called vertical bisection (SPLIT_BT_VER) 1102, horizontal bisection (SPLIT_BT_HOR) 1104, vertical trisection (SPLIT_TT_VER) 1106, and horizontal trisection (SPLIT_TT_HOR) 1108. Then, CB corresponds to the leaf of the multi-type tree. In this exemplary implementation, as long as CB is not too large for the maximum transform length, this segmentation is used for both prediction and transform processes without any further partitioning. This means that in most cases, CB, PB, and TB have the same block size in the quadtree with nested multi-type tree coding block structure.An exception occurs when the maximum supported transform length is smaller than the width or height of a color component of CB. In some implementations, in addition to bisection or trisection, the nested patterns of Figure 11 may further include a quadtree division.
[0110] One specific example of quadtree with nested multi-type tree coding block structure of block partitions (including quadtree, bisection, and trisection options) for one base block is shown in FIG. 12. More specifically, FIG. 12 shows that a base block 1200 is quadtree partitioned into four square partitions 1202, 1204, 1206, and 1208. A decision to further use the multi-type tree structure and quadtree of FIG. 11 for further partitioning is made for each of the quadtree partitioned partitions. In the example of FIG. 12, partition 1204 is not further partitioned. Partitions 1202 and 1208 each adopt another quadtree partitioning. For partition 1202, the second level quadtree partitioned top left, top right, bottom left, and bottom right partitions adopt third level partitioning of quadtree, horizontal bisection 1104 of FIG. 11, non-split, and horizontal trisection 1108 of FIG. 11, respectively. Partition 1208 adopts another quadtree division, and the second level quadtree divided top-left, top-right, bottom-left, and bottom-right partitions adopt the third level division of vertical trisection 1106, unsplit, unsplit, and horizontal bisection 1104 of FIG. 11, respectively. Two of the subpartitions of the top-left partition of the third level of 1208 are further divided according to horizontal bisection 1104 and horizontal trisection 1108 of FIG. 11, respectively. Partition 1206 adopts the second level division pattern into two partitions following vertical bisection 1102 of FIG. 11, and the two divisions are further divided at the third level according to horizontal bisection 1108 and vertical bisection 1102 of FIG. 11, respectively. A fourth level division is further applied to one of them according to horizontal bisection 1104 of FIG. 11.
[0111] In the above specific example, the maximum luma transform size may be 64×64, and the maximum supported chroma transform size may be different from the luma, e.g., 32×32. Even if the above example CB of FIG. 12 is not generally further divided into smaller PBs and / or TBs, when the width or height of a luma coding block or a chroma coding block is larger than the maximum transform width or maximum transform height, the luma coding block or the chroma coding block may be automatically divided in the horizontal and / or vertical directions to satisfy the transform size constraints in that direction.
[0112] In the specific example of partitioning the base blocks into CBs above, as mentioned above, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P slices and B slices, the luma CTB and chroma CTB in one CTU can share the same coding tree structure. For I slices, for example, luma and chroma can have separate coding block tree structures. When separate block tree structures are applied, the luma CTB may be partitioned into luma CBs by one coding tree structure, and the chroma CTB is partitioned into chroma CBs by another coding tree structure. This means that a CU in an I slice may be composed of a coding block of a luma component or a coding block of two chroma components, and a CU in a P slice or B slice is always composed of coding blocks of all three color components, unless the video is monochrome.
[0113] When a coding block is further partitioned into multiple transform blocks, the transform blocks therein may be ordered in the bitstream according to various orders or scanning schemes. Exemplary implementations for partitioning a coding block or a predictive block into transform blocks and the coding order of the transform blocks are described in further detail below. In some exemplary implementations, as described above, the transform partitioning may support transform blocks of multiple shapes, e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with transform block sizes ranging from 4×4 to 64×64, for example. In some implementations, if the coding block is less than or equal to 64×64, the transform block partitioning may be applied only to the luma component, such that for chroma blocks, the transform block size is identical to the coding block size. Otherwise, if the width or height of the coding block is greater than 64, both the luma coding block and the chroma coding block may be implicitly partitioned into a multiple of min(W,64)×min(H,64) transform blocks and min(W,32)×min(H,32) transform blocks, respectively.
[0114] In some exemplary implementations of transform block partitioning, for both intra-coded and inter-coded blocks, the coding blocks may be further partitioned into multiple transform blocks with a partitioning depth up to a predefined number of levels (e.g., two levels). The transform block partitioning depth and size may be related. In some exemplary implementations, the mapping from the transform size of the current depth to the transform size of the next depth is shown in Table 1 below. [Table 1]
[0115] Based on the example mapping in Table 1, for a 1:1 square block, the next level transform partition can create four 1:1 square sub-transform blocks. The transform partition may stop at, for example, 4×4. Thus, the transform size of the current depth of 4×4 corresponds to the same size of 4×4 of the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next level transform partition can create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next level transform partition can create two 1:2 / 2:1 sub-transform blocks.
[0116] In some example implementations, further restrictions may be applied regarding transform block partitioning for the luma components of intra-coded blocks. For example, for each level of transform partitioning, all sub-transform blocks may be restricted to have equal size. For example, for a 32×16 coding block, the level 1 transform partitioning creates two 16×16 sub-transform blocks, and the level 2 transform partitioning creates eight 8×8 sub-transform blocks. In other words, to keep the transform units equal in size, a second level partitioning must be applied to all first level sub-blocks. An example of transform block partitioning of an intra-coded square block according to Table 1 is shown in FIG. 15 with the coding order indicated by the arrows. Specifically, 1502 shows a square coding block. The first level partitioning according to Table 1 into four equal-sized transform blocks is shown in 1504 with the coding order indicated by the arrows. The second level partitioning of all first level equal-sized blocks according to Table 1 into 16 equal-sized transform blocks is shown in 1506 with the coding order indicated by the arrows.
[0117] In some exemplary implementations, the above restrictions on intra-coding may not apply to the luma components of an inter-coded block. For example, after the first level of transform partitioning, any one of the sub-transform blocks may be further partitioned independently at another level. Thus, the resulting transform blocks may or may not be of the same size. An exemplary partitioning of an inter-coded block into transform blocks according to their coding order is shown in FIG. 16. In the example of FIG. 16, an inter-coded block 1602 is partitioned into transform blocks at two levels according to Table 1. At the first level, the inter-coded block is partitioned into four transform blocks of equal size. Then, only one of the four transform blocks (but not all of them) is further partitioned into four sub-transform blocks, as indicated by 1604, resulting in a total of seven transform blocks with two different sizes. An exemplary coding order of these seven transform blocks is indicated by the arrows at 1604 in FIG. 16.
[0118] In some example implementations, for chroma components, some additional restrictions on the transform blocks may be applied, e.g., for chroma components, the transform block size may be as large as the coding block size, but cannot be smaller than a predefined size, e.g., 8×8.
[0119] In some other example implementations, for coding blocks with either width (W) or height (H) greater than 64, both luma coding blocks and chroma coding blocks may be implicitly divided into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively, where in this disclosure "min(a,b)" may return the smaller value between a and b.
[0120] FIG. 17 further illustrates another alternative exemplary scheme for partitioning coding blocks or predictive blocks into transform blocks. As illustrated in FIG. 17, instead of using recursive transform partitioning, a set of predefined partitioning types may be applied to the coding blocks according to the transform type of the coding blocks. In the particular example illustrated in FIG. 17, one of six exemplary partitioning types may be applied to split the coding blocks into various numbers of transform blocks. Such a scheme for generating transform block partitioning may be applied to either the coding blocks or the predictive blocks.
[0121] More specifically, the partitioning scheme of FIG. 17 provides up to six exemplary partition types for any given transform type (transform type refers to the type of first-order transform, such as ADST, for example). In this scheme, every coding block or predictive block may be assigned a transform partition type based on, for example, a rate-distortion cost. In one example, the transform partition type assigned to a coding block or predictive block may be determined based on the transform type of the coding block or predictive block. As shown by the six transform partition types illustrated in FIG. 17, a particular transform partition type may correspond to a transform block division size and pattern. The correspondence between various transform types and various transform partition types may be predefined. An example is shown below, where the capitalized labels indicate transform partition types that may be assigned to a coding block or predictive block based on a rate-distortion cost. ·PARTITION_NONE: Allocate the transformation size equal to the block size. ·PARTITION_SPLIT: Allocates a height of 1 / 2 the block size and a height transformation size of 1 / 2 the block size. ·PARTITION_HORZ: Allocates a transformation size with width equal to the block size and height equal to 1 / 2 the block size. ·PARTITION_VERT: Allocates a transformation size with a width half the block size and a height equal to the block size. ·PARTITION_HORZ4: Allocates a transformation size with width equal to the block size and height 1 / 4 of the block size. ·PARTITION_VERT4: Allocates a transformation size with a width of 1 / 4 of the block size and a height equal to the block size.
[0122] In the above example, all transform partition types as shown in Figure 17 include uniform transform sizes for the partitioned transform blocks. This is merely an example and not a limitation. In some other implementations, mixed transform block sizes may be used for the partitioned transform blocks in a particular partition type (or pattern).
[0123] A video block (PB or CB, also referred to as PB when not further partitioned into multiple predictive blocks) may be predicted in various manners rather than being directly encoded, thereby exploiting various correlations and redundancies in the video data to improve compression efficiency. Correspondingly, such prediction may be performed in various modes. For example, a video block may be predicted via intra prediction or inter prediction. In particular, in an inter prediction mode, a video block may be predicted by one or more other reference blocks or inter predictor blocks from one or more other frames via either single reference inter prediction or mixed reference inter prediction. To perform inter prediction, a reference block may be specified by its frame identifier (the temporal location of the reference block) and a motion vector indicating a spatial offset between a current block being encoded or decoded and the reference block (the spatial location of the reference block). The reference frame identification and the motion vector may be signaled in the bitstream. The motion vector as a spatial block offset may be directly signaled or may itself be predicted by another reference motion vector or predictor motion vector. For example, the current motion vector may be predicted directly by a reference motion vector (e.g., of a candidate neighboring block), or by a combination of the reference motion vector and a motion vector differential (MVD) between the current motion vector and the reference motion vector. The latter may be referred to as a merge mode with motion vector differential (MMVD). The reference motion vector may be identified in the bitstream, for example, as a pointer to a block that is spatially neighboring the current block or a block that is temporally neighboring but spatially co-located.
[0124] Returning to the intra prediction process, samples in a block (e.g., luma or chroma prediction block, or coding block if not further divided into prediction blocks) are predicted by samples of neighboring lines, next neighboring tines, or one or more other lines, or combinations thereof, to generate a prediction block. The residual between the actual block being coded and the prediction block may then be processed through a transform after quantization. Various intra prediction modes may be made available, and parameters related to intra mode selection and other parameters may be signaled in the bitstream. Various intra prediction modes may relate, for example, to one or more line positions for predicting samples, the direction in which the prediction samples are selected from predicting one or more lines, and other special intra prediction modes.
[0125] For example, the set of intra-prediction modes (interchangeably referred to as "intra modes") may include a predefined number of directional intra-prediction modes. As described above with respect to the exemplary embodiment of FIG. 1, these intra-prediction modes may correspond to a predefined number of directions to follow when selecting a sample outside the block as a destination for a sample being predicted within a particular block. In another particular exemplary implementation, eight main directional modes may be supported and predefined, corresponding to angles between 45 degrees and 207 degrees with respect to the horizontal axis.
[0126] In some other implementations of intra prediction, the directional intra modes may be further expanded to a set of angles with finer granularity to further exploit more diverse spatial redundancy in the directional texture. For example, the above 8-angle implementation may be configured to provide eight nominal angles, designated V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED, as shown in FIG. 19, and a predefined number (e.g., seven) of finer angles may be added to each nominal angle. Such expansion may result in a larger total number of directional angles (e.g., 56 in this example) that may be available for intra prediction, corresponding to an equal number of predefined directional intra modes. The prediction angle may be represented by the nominal intra angle plus an angle delta. In the above specific example with seven finer angle directions for each nominal angle, the angle delta may be -3 to 3 times the step size of 3 degrees. Any such angle scheme may be used, as shown in FIG. 18 with 65 different prediction angles.
[0127] In some implementations, instead of or in addition to the above directional intra modes, a predefined number of non-directional intra prediction modes may also be predefined and made available. For example, five non-directional intra modes called smooth intra prediction modes may be specified. These non-directional intra mode prediction modes may be specifically called DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H intra modes. Prediction of samples of a particular block under these exemplary non-directional modes is shown in FIG. 20. As an example, FIG. 20 shows that a 4×4 block 2002 is predicted by samples from an upper neighboring line and / or a left neighboring line. A particular sample 2010 in the block 2002 may correspond to a sample 2004 immediately above the sample 2010 of the upper neighboring line of the block 2002, a sample 2006 above and to the left of the sample 2010 as the intersection of the upper neighboring line and the left neighboring line, and a sample 2008 immediately to the left of the sample 2010 of the left neighboring line of the block 2002. In the exemplary DC intra prediction mode, the average of the left and top neighboring samples 2008 and 2004 may be used as a predictor for sample 2010. For the exemplary PAETH intra prediction mode, the top, left, and top left reference samples 2004, 2008, and 2006 may be fetched, and then the closest of these three reference samples to (top+left-top left) may be set as the predictor for sample 2010. For the exemplary SMOOTH_V intra prediction mode, sample 2010 may be predicted by quadratic interpolation in the vertical direction of the top left neighboring sample 2006 and the left neighboring sample 2008. For the exemplary SMOOTH_H intra prediction mode, sample 2010 may be predicted by quadratic interpolation in the horizontal direction of the top left neighboring sample 2006 and the top neighboring sample 2004. For the exemplary SMOOTH intra prediction mode, sample 2010 may be predicted by an average of quadratic interpolation in the vertical and horizontal directions. The above non-directional intra mode implementation is shown only as a non-limiting example: other neighboring lines and other non-directional selections of samples and methods of combining predicted samples to predict a particular sample within a predictive block are also possible.
[0128] The selection of a particular intra-prediction mode by the encoder from the above directional or non-directional modes at various coding levels (picture, slice, block, unit, etc.) may be signaled in the bitstream. In some exemplary implementations, the eight exemplary nominal directional modes may be signaled first, along with the five smooth modes without angles (a total of 13 options). Then, if the signaled mode is one of the eight nominal angle intra-modes, an index indicating the selected angle delta for the corresponding signaled nominal angle is further signaled. In some other exemplary implementations, all intra-prediction modes may be indexed together for signaling (e.g., 56 directional modes plus 5 non-directional modes to generate 61 intra-prediction modes).
[0129] In some example implementations, the example 56 or other number of directional intra-prediction modes may be implemented using a unified directional predictor that projects each sample of a block to a reference subsample position and interpolates the reference sample by a 2-tap bilinear filter.
[0130] In some implementations, additional filter modes, called FILTER INTRA modes, may be designed to capture the decaying spatial correlation with the references on the edges. In these modes, in addition to samples outside the block, samples predicted within the block may be used as intra prediction reference samples for some patches within the block. These modes may, for example, be predefined and made available for intra prediction of at least the luma block (or only the luma block). A predefined number (e.g., 5) of filter intra modes may be predesigned, each of which is represented by a set of n-tap filters (e.g., 7-tap filters) that reflect the correlation between a sample in a 4×2 patch and its adjacent n neighbors. In other words, the weighting coefficients of the n-tap filters may depend on the position. Taking an 8×8 block, a 4×2 patch, and 7-tap filtering as examples, an 8×8 block 2002 may be divided into eight 4×2 patches, as shown in FIG. 21. These patches are denoted by B0, B1, B1, B3, B4, B5, B6, and B7 in FIG. 21. For each patch, its seven neighbors, denoted by R0-R6 in FIG. 21, may be used to predict samples in the current patch. For patch B0, all neighbors may already be reconstructed. However, for other patches, some of the neighbors may not be reconstructed because they are in the current block, in which case the predicted values of the immediate neighbors are used as reference. For example, all neighbors of patch B7 as shown in FIG. 21 have not been reconstructed, so the predicted samples of the neighbors are used instead.
[0131] In some implementations of intra prediction, a color component may be predicted using one or more other color components, which may be any one of the components of the YCrCb color space, the RGB color space, the XYZ color space, etc.
[0132] One type of intra prediction, which predicts one color component using one or more other color components, is chroma from luma (CfL) prediction. In CfL prediction, a chroma component is predicted based on a luma component. The predicted chroma component may include a chroma block, which may include a sample or a chroma sample. The predicted sample is referred to as a prediction sample. Also, the predicted chroma block may correspond to a luma block. In this specification, unless otherwise specified, the correspondence between a luma block and a chroma block refers to the chroma block being co-located with the luma block.
[0133] Also, as used herein and described in further detail below, the luma components used to predict chroma components may include luma samples. The luma samples may include luma samples of the corresponding or co-located chroma block itself and / or may include neighboring luma samples that are luma samples of one or more neighboring luma blocks that are proximate or adjacent to the co-located luma block that corresponds to the predicted chroma block. Additionally, in at least some implementations, the luma samples used in the CfL prediction process are reconstructed luma samples, which may be copies of the original luma samples derived or reconstructed from a compressed version of the original luma samples using a decoding process.
[0134] In some implementations, an encoder (e.g., any of the encoders 403, 603, 703) and / or a decoder (e.g., any of the decoders 410, 510, 810) may be configured to perform CfL prediction via a CfL prediction process. Also, in at least some of these implementations, the encoder and / or the decoder may be configured in a CfL prediction mode to perform the CfL prediction process. As described in more detail below, the encoder and / or the decoder may be operable in at least one of a number of different CfL prediction modes. In the different CfL prediction modes, the encoder and / or the decoder may perform different respective CfL processes to generate chroma prediction samples.
[0135] 22A-22C, the encoder and / or decoder may include a CfL prediction unit 2202 configured to perform a CfL prediction process. In various implementations, the CfL prediction unit 2202 may be a standalone unit or a component or subunit of another unit of the encoder or decoder. For example, in any of the various implementations, the CfL prediction unit 2202 may be a component or subunit of the intra prediction unit 552, the intra encoder 722, or the intra decoder 872. The CfL prediction unit 2202 may also be configured to perform various or multiple operations or functions to perform or implement the CfL process. For simplicity, the CfL prediction unit 2202 is described as being a component of the encoder and / or decoder that performs each of these operations or functions. However, in any of the various embodiments, the CfL prediction unit 2202 may be further configured or organized into multiple subunits each configured to perform one or more of the operations or functions of the CfL prediction process, or one or more units other than the CfL prediction unit 2202 may be configured to perform one or more operations of the CfL prediction process. Also, in any of the various embodiments, the CfL prediction unit 2202 may be implemented in hardware or a combination of hardware or software to perform and / or execute the operations or functions of the CfL process. For example, the CfL prediction unit may be implemented as an integrated circuit, or a processor configured to execute software or firmware stored in a memory, or a combination thereof. Also, in any of the various embodiments, a non-transitory computer-readable storage medium may store computer instructions executable by a processor to perform the functions or operations of the CfL prediction unit 2202.
[0136] For CfL prediction, the encoder and / or decoder may determine to apply a CfL prediction mode to a luma block of a received coding bitstream, such as via CfL prediction unit 2202. The CfL prediction unit 2202 may then perform CfL prediction of the luma block according to the CfL prediction mode determined to be applied. Correspondingly, the encoder and / or decoder may reconstruct a chroma block that corresponds to or is co-located with the luma block by at least partially applying the CfL prediction mode, such as via CfL prediction unit 2202.
[0137] As described above, the CfL prediction unit 2202 may operate in at least one of a plurality of CfL prediction modes. With reference to FIG. 22A, in a first CfL prediction mode (or a first set of one or more CfL prediction modes), the CfL prediction unit 2202 may be configured to generate a plurality of prediction samples of a chroma block corresponding to a luma block based on a luma sample of the luma block. As described above, the chroma block corresponding to the luma block is a chroma block co-located with the luma block. Correspondingly, in FIG. 22A, in the first CfL prediction mode, the CfL prediction unit uses a luma sample of a luma block co-located with the chroma block for which the prediction sample is to be generated.
[0138] FIG. 23 illustrates a flow diagram of an example method 2300 of a CfL prediction process that the CfL prediction unit 2202 may perform in the first CfL prediction mode. In some implementations, a CfL prediction process such as that illustrated in FIG. 23 generates a number of chroma prediction samples based on an alternating current (AC) contribution of a luma sample and a direct current (DC) contribution of a chroma sample. Each of the AC and DC contributions may be a prediction of a chroma component, and are also referred to as an AC contribution prediction and a DC contribution prediction. In certain of these implementations, the chroma prediction samples are modeled as a linear function of the luma sample, such as according to the following equation: CfL(α)=α×LAC +DC (1) In the formula, L AC where α denotes the AC contribution of the luma component (luma sample), α denotes a scaling parameter of the linear model, and DC denotes the DC contribution of the chroma components. Also, in at least some implementations, the AC contributions are obtained for each of the samples of the block, while the DC contribution is obtained for the entire block.
[0139] In certain of these implementations shown in FIG. 23, in block 2302, multiple luma samples of a luma block may be subsampled (or downsampled) to a chroma resolution (e.g., 4:2:0, 4:2:2, or 4:4:4). In block 2304, the subsampled luma samples may be averaged to generate a luma average. In block 2306, the luma average may be subtracted from the luma sample to generate an AC contribution for the luma component. In block 2308, the AC contribution for the luma component may be multiplied by a scaling parameter α to generate a scaled AC contribution for the luma component. The scaled AC contribution for the luma component may also be an AC contribution prediction for the chroma component. In block 2310, a DC contribution prediction for the chroma component may be added to the AC contribution prediction according to a linear model to generate a chroma prediction sample. In at least some implementations, the scaling parameter α may be based on the original chroma samples or may be signaled in the bitstream. This may reduce decoder complexity and result in more accurate predictions. Additionally or alternatively, the DC contributions of the chroma components may be calculated using an intra-DC mode within the chroma components in some exemplary implementations.
[0140] Further, in some implementations of method 2300 or the first CfL mode, when some luma samples of a co-located luma block are outside a picture boundary, these luma samples may be padded and the padded luma samples may be used to calculate a luma average, such as block 2304. Figure 24 shows a schematic diagram of luma samples inside and outside a picture defined by a picture boundary. In at least some implementations, the outer picture luma samples may be padded by processing the value of the nearest available sample in the current block.
[0141] Additionally or alternatively, in some implementations, when performing CfL prediction, the subsampling performed in block 2302 may be combined with the averaging performed in block 2304 and / or the subtraction performed in block 2306. This, in turn, can simplify the linear modeling equations while eliminating the subsampling division and rounding errors. The following equation (2) corresponds to the combination of both steps and simplifies to equation (3). Both equations (2) and (3) use integer division. Also, MxN is a matrix of pixels in the luma plane.
number
[0142] Based on chroma subsampling, S x ×S y ∈{1,2,4}. Also, M and N may both be powers of 2, and similarly, M×N will also be a power of 2. For example, in the context of 4:2:0 chroma subsampling, instead of applying a box filter, one can use the sum of the four reconstructed luma pixels that match the chroma pixels. Correspondingly, the CfL prediction can be scaled by 2.
[0143] Further, as described above for equation (1), the CfL prediction process of FIG. 23 uses only one linear model between luma samples and chroma samples in an entire coding block. However, only one linear model may not be optimal for some relationships between luma samples and chroma samples in an entire coding block. Additionally or alternatively, the CfL prediction method 2300 uses luma samples of the corresponding luma block to calculate the average and therefore the AC contribution. However, in at least some embodiments, the DC contribution may be determined or calculated by averaging neighboring luma samples. The lack of alignment between the luma samples used for AC and DC contributions, respectively, and the neighboring luma samples may cause inaccurate or suboptimal prediction. Additionally or alternatively, signaling the scaling value α in the bitstream may undesirably consume some bits. Additionally or alternatively, the CfL prediction method 2300 of FIG. 23 predicts a chroma block from a co-located luma block, but does not use neighboring luma samples of the co-located luma block.
[0144] Furthermore, in some implementations of AV 1, the minimum block size supported for both luma and chroma planes is 4x4. Thus, for subsampled chroma planes (YUV 4:2:0 or YUV 4:2:2), multiple chroma blocks smaller than 4 x 4 may be combined into one chroma block. For example, referring to Figure 39, for YUV 4:2:0, an 8x8 luma region may be split into four 4x4 luma blocks. The four chroma blocks may then be combined into one single chroma block of size 4x4, covering the area of these four luma blocks.
[0145] Furthermore, in some implementations of AV 1, in the case of inter frames, luma blocks and chroma blocks may share the same coding tree structure. However, luma blocks and chroma blocks may have different TU depths. In general, TU depth is the number of times a CU is divided or partitioned into TUs. Also, for luma blocks, syntax in the bitstream may indicate the TU depth. However, for chroma blocks, the TU depth is always 0 except when the width or height of the chroma CU exceeds the maximum TU size. If the TU depth of a chroma block is smaller than the TU depth of the co-located luma block and the current chroma block is coded according to CfL intra prediction, all of the covering / matching co-located luma samples may be used to calculate the DC components for CfL prediction. This is shown in Figure 40.
[0146] However, for CfL prediction, when using neighboring reconstructed samples to calculate DC components of chroma prediction samples of small chroma blocks, such as 2x2, 2x4, or 4x2, the locations of neighboring samples of co-located luma blocks may be undefined. Additionally or alternatively, for CfL prediction, when using neighboring reconstructed samples to calculate DC components of CfL prediction and when the TU depth of the chroma block is smaller than the TU depth of the co-located luma block, the locations of neighboring samples of the co-located luma block may not be determined. The CfL prediction process described below can determine the locations of neighboring samples of the co-located luma block.
[0147] Referring to FIG. 22B, in addition to or as an alternative to operating in the first CfL prediction mode, in some implementations, the CfL prediction unit 2202 may operate in a second CfL prediction mode (or a second set of one or more CfL prediction modes) to generate a plurality of prediction samples of a chroma block corresponding to a luma block based on neighboring luma samples of the luma block. That is, the CfL prediction unit 2202 may generate a chroma prediction sample using neighboring luma samples of the luma block without using a luma sample of the luma block. Referring to FIG. 22C, in addition to or as an alternative to operating in the first CfL prediction mode and / or the second CfL prediction mode, in some implementations, the CfL prediction unit 2202 may operate in a third CfL prediction mode (or a third set of one or more CfL prediction modes) to generate a plurality of prediction samples of a chroma block corresponding to a luma block based on a luma sample of the luma block and neighboring luma samples of the luma block.
[0148] FIG. 25 illustrates an example of a luma block 2502 and a schematic diagram of neighboring luma samples of the luma block 2502. In general, neighboring luma samples of a given luma block are luma samples of neighboring or nearby luma blocks that are adjacent and / or in the vicinity of the given luma block. Each neighboring luma sample may be or have multiple types of neighboring luma samples of a particular type. Each type may correspond to a relative spatial relationship with the given luma block. Similarly, each neighboring or nearby luma block may have a particular type that matches the neighboring luma samples of a particular type contained therein. In at least some implementations, the multiple types of neighboring luma samples and / or blocks may include left, top left, top, top right, right, bottom right, bottom, and bottom left. Figure 25 shows where neighboring luma samples may be spatially located with respect to a given luma block 2502, including a left neighboring luma sample 2504 in a left neighboring luma block, an upper left neighboring luma sample 2506 in an upper left neighboring luma block, an upper neighboring luma sample 2508 in an upper neighboring luma block, an upper right neighboring luma sample 2510 in an upper right neighboring luma block, a right neighboring luma sample 2512 in a right neighboring luma block, a lower right neighboring luma sample 2514 in a lower right neighboring luma block, a lower neighboring luma sample 2516 in a lower neighboring luma block, and a lower left neighboring luma sample 2518 in a lower left neighboring luma block. The upper left, upper right, lower left, and lower right neighboring luma samples and blocks may also be generally and / or collectively referred to as corner neighboring luma samples and blocks, respectively. In any of the various implementations in the second CfL mode and / or the third CfL mode, the CfL prediction unit 2202 may use neighboring luma samples of all types, or at least one but less than all types, when performing CfL prediction.
[0149] FIG. 26 shows a flowchart of an example method 2600 of a CfL prediction process that the CfL prediction unit 2202 may perform when operating in the second CfL prediction mode (FIG. 22B) and / or the third CfL prediction mode (FIG. 22C). In block 2602, the CfL prediction unit 2202 may determine a plurality of neighboring luma samples of a luma block. In block 2604, the CfL prediction unit 2202 may generate a plurality of predicted samples of a chroma block corresponding to the luma block based on the plurality of neighboring luma samples.
[0150] FIG. 27 illustrates a flowchart of another example method 2700 of a CfL prediction process that the CfL prediction unit 2202 may perform when operating in the second CfL prediction mode (FIG. 22B) and / or the third CfL prediction mode (FIG. 22C). In block 2702, the CfL prediction unit 2202 may determine a plurality of neighboring luma samples of a luma block. In block 2704, the CfL prediction unit 2202 may generate AC and DC contributions of a plurality of predictive samples of a chroma block corresponding to the luma block. At least one of the AC or DC contributions is generated in block 2704 based on the set of luma samples including the plurality of neighboring luma samples determined in block 2702. In block 2706, the CfL prediction unit 2202 may generate a plurality of predictive samples of the chroma block based on the AC and DC contributions determined in block 2704. In at least some implementations, the CfL prediction unit 2202 may determine, in block 2706, a number of chroma prediction samples using AC and DC contributions according to the linear model of equation (1) above.
[0151] In some implementations, the example methods 2600 and 2700 may be combined. For example, block 2602 may include block 2702, and block 2604 may include block 2704 and / or block 2706.
[0152] Figure 28 shows a flow diagram of another example method 2800 of a CfL prediction process that the CfL prediction unit 2202 may perform when operating in the third CfL prediction mode (Figure 22C). The CfL prediction process 2800 may be similar to the CfL prediction process 2300 of Figure 23, except that instead of averaging luma samples of a luma block, the CfL prediction process 2800 may average neighboring luma samples to generate a neighboring luma average. Thus, the AC contribution is based on both the luma samples of the luma block and the neighboring luma samples.
[0153] In more detail, in block 2802, multiple luma samples of a luma block may be subsampled to a chroma resolution. In block 2804, multiple neighboring luma samples of a luma block may be subsampled to a chroma resolution. Furthermore, in at least some implementations, the same subsampling (or downsampling) method is used to subsample the luma samples and the neighboring luma samples (i.e., the same subsampling method is applied in both blocks 2802 and 2804). For example, if subsampling is performed according to a 4:2:0 format, two rows in the top neighboring region (e.g., region 2508 in FIG. 25), two columns in the left neighboring region (e.g., region 2504 in FIG. 25), and / or four pixels in the top-left region (e.g., region 2506 in FIG. 25) are subsampled (or downsampled). Correspondingly, when the CfL prediction unit 2202 determines the AC contribution (e.g., blocks 2806 and 2808 below), neighboring luma samples may be averaged and subtracted from the reconstructed luma sample value, as shown in FIG. 28 and further described below.
[0154] In more detail, in block 2806, the subsampled neighboring luma samples may be averaged to generate a neighboring luma average. In block 2808, the neighboring luma average may be subtracted from the subsampled luma sample to generate an AC contribution of the luma component. In block 2808, the AC contribution of the luma component may be multiplied by a scaling parameter α to generate a scaled AC contribution of the luma component. The scaled AC contribution of the luma component may also be an AC contribution prediction of the chroma component. In block 2812, a DC contribution prediction of the chroma component may be added to the AC contribution prediction to generate a chroma prediction sample, such as according to the linear model described above in equation (1). In at least some implementations, the scaling parameter α may be based on the original chroma sample or may be signaled in the bitstream. This may reduce decoder complexity and provide a more accurate prediction. Additionally or alternatively, the DC contribution of the chroma component may be calculated using an intra DC mode within the chroma component in some exemplary implementations.
[0155] In other implementations of the example method 2800, the neighboring luma samples may not be subsampled before being used to determine the neighboring luma average, i.e., in these other implementations, block 2804 may be skipped or otherwise not performed.
[0156] Also, in some implementations, all or some of the blocks of method 2800 may be combined with method 2600 and / or method 2700. For example, after the neighboring luma samples are determined in blocks 2602 and / or 2702, they may be subsampled in block 2804 and / or averaged in block 2806. Additionally or alternatively, generating a chroma prediction sample based on the neighboring luma samples in block 2604 may include one or more of subsampling the neighboring luma samples in block 2804, averaging the (subsampled) neighboring luma samples in block 2806, generating a luma AC contribution in block 2808, and / or scaling the luma AC contribution in block 2810. Additionally or alternatively, generating the AC contribution at block 2704 in method 2700 may include one or more of subsampling neighboring luma samples at block 2804, averaging the (subsampled) neighboring luma samples at block 2806, generating the luma AC contribution at block 2808, and / or scaling the luma AC contribution at block 2810. Additionally or alternatively, generating a chroma prediction sample based on the AC and DC contributions at block 2706 of method 2700 may include adding the DC contribution to the AC contribution at block 2812. Other ways of combining methods 2600, 2700, and / or 2800 may be possible.
[0157] FIG. 29 shows a flowchart of another example method 2900 of a CfL prediction process that the CfL prediction unit 2202 may perform when operating in the third CfL prediction mode (FIG. 22C). In block 2902, the CfL prediction unit 2202 may map multiple luma samples of a luma block to multiple neighboring chroma samples. For example, the CfL prediction unit 2202 may map each luma pixel PL from the i-th to the j-th at position (i,j) of the luma block. ij Let kth neighboring chroma pixel PCk In at least some implementations, in block 2902, the CfL prediction unit 2202 may also determine that a chroma block should be predicted in a CfL mode, where the chroma block corresponds to and / or is co-located with the luma block. The CfL prediction unit 2202 may then map the luma samples of the luma block to neighboring chroma samples in at least one neighboring chroma block in a neighborhood relative to the chroma block to be predicted. In block 2904, the CfL prediction unit 2202 may perform CfL prediction of the chroma block using the neighboring chroma samples as the prediction samples of the chroma block. For example, the CfL prediction unit 2202 may set the neighboring chroma samples as the prediction samples of the chroma block corresponding to the luma block. That is, the CfL prediction unit 2202 may determine that each k-th chroma pixel PC k , the corresponding i-th to j-th chroma sample PC ij The CfL prediction unit 2202 may use the same relationship or mapping between the (i,j) position and the neighborhood index k determined in block 2902 to predict each k-th chroma pixel PC k for each of the i-th to j-th chroma samples PC ij The CfL prediction method 2900 of FIG. 29 may avoid using the scaling parameter α and / or avoid consuming resources for signaling the scaling parameter α.
[0158] In at least some implementations, the CfL prediction unit 2202 may map multiple luma samples of a luma block to multiple neighboring chroma samples in two stages in block 2902. First, the CfL prediction unit 2202 may map multiple luma samples to multiple neighboring luma samples. For example, the CfL prediction unit 2202 may map multiple luma samples of each of the i-th to j-th pixels PL ijLet k be the neighboring luma pixel PL k For illustrative purposes, FIG. 30 illustrates the above k-th neighboring luma sample PL k A specific chroma pixel PL mapped to (112) ij (111). The CfL prediction unit 2202 can then map the nearby luma samples to corresponding nearby chroma samples or determine nearby chroma samples that correspond to and / or are co-located with the nearby luma samples. For example, FIG. 30 illustrates, via arrow 3004, the neighboring chroma pixels PC k The kth neighboring luma pixel PL k 30 shows a two-step process in which the luma samples can be mapped or correspond to the neighboring chroma samples. The CfL prediction unit 2202 can then use the neighboring chroma samples as the prediction samples of the corresponding chroma block for CfL prediction of the chroma block, as performed in block 2904. For example, FIG. 30 shows a kth neighboring chroma sample PC k is the i-th to j-th luma sample PL ij The i-th to j-th chroma sample PCs ij Indicates that the value of the
[0159] In addition, in at least some implementations of the two-stage process, when mapping luma samples to neighboring luma samples, the i-th to j-th luma samples PL ij Then, the CfL prediction unit 2202 predicts the i-th to j-th luma samples PL ij Also, for at least some of these embodiments, the CfL prediction unit 2202 may determine the neighboring luma sample that has the smallest difference from the i-th to j-th luma samples PL ijWhen the CfL prediction unit 2202 determines a plurality of neighboring luma samples having a minimum difference from the i-th to j-th luma samples PL ij Among the multiple minimum difference neighboring luma samples that have the smallest distance to the target k-th neighboring luma sample PL k The CfL prediction unit 2202 may select the kth neighboring luma sample of the target in the minimum difference luma sample used to determine the chroma prediction sample.
[0160] Additionally or alternatively, in some implementations, when determining the minimum difference between a luma sample and a neighboring luma sample, the CfL prediction unit 2202 may determine the minimum difference according to a sum of absolute differences (SAD) measurement and / or a sum of squared errors (SSE) measurement.
[0161] Additionally or alternatively, in some implementations, the CfL prediction unit 2202 may map multiple luma samples together to multiple neighboring luma samples. The multiple luma samples that are mapped together may be referred to as a patch. For example, the CfL prediction unit 2202 may identify a given patch of luma samples and find a patch of neighboring luma samples that has the smallest difference from the given patch of luma samples. In at least some of these implementations, the CfL prediction unit 2202 may use a center sample or pixel of the patch to determine the neighboring luma patch that has the smallest difference from the given luma path. Once the CfL prediction unit 2202 identifies the neighboring luma patch with the smallest difference, it may map each luma sample of that luma patch to each neighboring luma sample of the neighboring luma patch with the smallest difference. For at least some of these implementations, the CfL prediction unit 2202 may use SSE to determine the smallest difference. Also, in some implementations, the size of the sub-blocks may vary. Additionally or alternatively, in some implementations, only a subset of all neighboring luma samples is analyzed to determine the least different neighboring luma sample.
[0162] FIG. 31 illustrates an example of patch-based mapping of luma samples to neighboring luma samples, which may be used as part of mapping multiple luma samples to multiple neighboring chroma samples in block 2902. In this example, three luma samples 11, 100, and 21 may form a luma patch. The CfL prediction unit 2202 may analyze neighboring luma patches having three neighboring luma samples. For each candidate neighboring luma patch, the CfL prediction unit 2202 may compare the center luma sample with the center neighboring luma sample of the candidate neighboring luma patch to determine the difference of the neighboring luma patch. In the example of FIG. 31, the CfL prediction unit 2202 determines the neighboring luma sample with the smallest difference from the center luma sample 100 as the neighboring luma sample 101. Next, the CfL prediction unit 2202 may map luma sample 11 with neighboring luma sample 10, luma sample 100 with neighboring luma sample 101, and luma sample 21 with neighboring luma sample 20. As shown in Figure 31, after the luma samples of a luma block are mapped to neighboring luma samples on a patch basis, neighboring chroma samples corresponding to the neighboring luma samples may be determined and then set as or copied to the chroma samples of the corresponding chroma block, as described above with respect to Figures 29 and 30.
[0163] Also, in some implementations, all or some blocks of method 2900 may be combined with method 2600. For example, determining a neighboring luma sample in block 2602 may include mapping a luma sample to a neighboring luma sample performed in block 2902. Alternatively, the CfL prediction unit 2202 may determine a plurality of neighboring luma samples and then use the neighboring luma samples to map a luma sample to at least some of the neighboring luma samples in block 2602. Additionally or alternatively, in various implementations, generating a chroma prediction sample in block 2604 may include mapping a luma sample to a neighboring chroma sample in block 2902 and / or setting the neighboring chroma samples as a plurality of chroma prediction samples in block 2904.
[0164] Also, in any of the various implementations of the example methods 2600, 2700, 2800, 2900, the neighboring luma samples, sub-samples, averages, and / or uses that the CfL prediction unit 2202 determines to generate the plurality of chroma prediction samples may include all, or at least one and less than all, types of neighboring luma samples, including at least one of left, top-left, top, top-right, right, bottom-right, bottom, or bottom-left. For example, in some implementations, the at least one type may include at least one of left, top, or top-left. In some other implementations, the at least one type may include at least one of left, top, right, bottom, top-left, top-right, bottom-left, or bottom-right. In some other implementations, the at least one type may include at least one of top-right or bottom-left. Also, in any of the various implementations, including the implementations of methods 2600, 2700, and / or 2800, if the CfL prediction unit 2202 determines to use a certain type of neighboring luma sample for the CfL prediction process, the CfL prediction unit 2202 may use one or more neighboring luma samples having the determined type.
[0165] Additionally or alternatively, in some implementations of example methods 2600, 2700, 2800, and / or 2900, signaling may be used to explicitly indicate or instruct the CfL prediction unit 2202 as to which type of neighboring luma samples to use for the CfL prediction process. For example, the CfL prediction unit 2202 may receive a signal generated internally, such as by another unit or component of the electronic device (e.g., an encoder or decoder) in which the CfL prediction unit 2202 is implemented, or a signal generated externally or remotely, such as by another unit or component of a different or separate electronic device (e.g., an encoder or decoder) in which the CfL prediction unit 2202 is implemented. In some other implementations, code information and / or characteristics of the luma block and / or predicted chroma block may be used to implicitly indicate which of the types of neighboring luma samples to use for the CfL prediction process, non-limiting examples of which include the intra prediction mode, block shape, block size, or block aspect ratio of the luma block and / or chroma block.
[0166] Additionally or alternatively, in some implementations of example methods 2600, 2700, 2800, and / or 2900, only those types of available neighboring luma samples may be used for the CfL prediction process. In some implementations, neighboring luma samples are available when not at a picture boundary or superblock boundary. For example, the CfL prediction unit 2202 may use the top neighboring luma sample for CfL prediction (e.g., determine the top neighboring luma sample in block 2602 and / or generate the chroma prediction sample based on the above neighboring luma sample in block 2604) when only the top neighboring luma sample is available. As another example, the CfL prediction unit may use the left neighboring luma sample for CfL prediction (e.g., determine the left neighboring luma sample in block 2602 and / or generate the chroma prediction sample based on the left neighboring luma sample in block 2604) when only the left neighboring luma sample is available.
[0167] Additionally or alternatively, in some implementations of example methods 2600, 2700, 2800, and / or 2900, if it is determined that neighboring luma samples of a particular type are used in the CfL prediction process and all of the neighboring luma samples of the particular type are not available, the CfL prediction unit 2202 may pad the neighboring luma samples of the particular type and use the padded neighboring luma samples in the CfL prediction process. For example, if the left neighboring luma samples are not available, the padding may include copying the luma samples of the left column of the current luma block as the left neighboring luma samples. As another example, if the top neighboring luma samples are not available, the padding may include copying the luma samples of the top row of the current luma block as the top neighboring luma samples. For at least some of these implementations, the CfL prediction unit may pad the neighboring luma samples according to the same padding method used in the intra-angle prediction mode.
[0168] Additionally or alternatively, in some implementations of the exemplary methods 2600, 2700, 2800, and / or 2900, if it is determined that upper neighboring luma samples are used for the CfL prediction process, the CfL prediction unit 2202 may use only those upper neighboring luma samples in the nearest upper reference line in the CfL prediction process. In certain of these implementations, the CfL prediction unit 2202 may use the upper neighboring luma samples of the nearest upper reference line as the neighboring luma samples for the CfL prediction process when the co-located luma block is located on a superblock boundary. In general, the AoMediaVideo model (AVM) uses coding blocks of various sizes, the largest of which is called a superblock. In some implementations, the largest block or superblock is 128×128 pixels or 64×64 pixels. The size of the superblock is signaled in the sequence header and has a default size of 128×128 pixels. The size of the smallest coding block is 4×4.
[0169] Additionally or alternatively, in some implementations of the example methods 2600, 2700, 2800, and / or 2900, the neighboring luma samples used in the CfL prediction process may include, such as in some implementations including only neighboring luma samples of the nearest neighboring reference line. At least some of these implementations may be used in combination with one or more of the other above-mentioned conditions or aspects. For example, if the CfL prediction unit 2202 determines to use a neighboring luma sample of a particular type, the CfL prediction unit 2202 may use a neighboring luma sample of the particular type that is in a nearest neighboring reference line of the particular type. Figure 32 shows a diagram illustrating the top and left reference lines of a co-located chroma block. In this example, reference line 0 may be the nearest above-mentioned neighboring reference in that it is closer or nearer than the other above-mentioned reference lines 1, 2, and 3.
[0170] Additionally or alternatively, in some implementations of example methods 2600, 2700, 2800, and / or 2900, the CfL prediction unit 2202 may use one or more boundary luma samples of a luma block to perform the CfL prediction process. The boundary luma samples may be luma samples that define an outer boundary of a chroma block. For example, a left boundary luma sample of a given luma block at least partially defines a left boundary of the given chroma block. As another example, a top boundary at least partially defines an upper or top boundary of the given luma block. The CfL prediction unit 2202 may use one or more boundary luma samples in combination with neighboring luma samples. For example, in implementations of method 2800, the CfL prediction unit 2202 may determine a neighboring luma average by averaging the neighboring luma samples in combination with one or more boundary luma samples of the co-located chroma block. In various of these embodiments, the boundary luma samples used for averaging may be obtained from the subsampled luma samples before subsampling, such as before the subsampling in block 2802, or after subsampling, such as after the subsampling in block 2802. In certain of these embodiments, the one or more boundary luma samples may include one or more top boundary luma samples and / or one or more left boundary luma samples.
[0171] FIG. 33 shows a flowchart of another example method 3300 of a CfL prediction process that the CfL prediction unit 2202 may perform. In block 3302, the CfL prediction unit 2202 may determine a type of a CfL prediction process from among a plurality of types of CfL prediction processes. The plurality of types of CfL prediction processes may correspond to at least two of the CfL prediction processes 2300, 2600, 2700, 2800, or 2900. In block 3304, the CfL prediction unit 2202 may perform a CfL prediction process according to the type of CfL prediction process that it determined in block 3302.
[0172] In some implementations, the CfL prediction unit 2202 selects from only two of the CfL prediction processes 2300, 2600, 2700, 2800, 2900. The two processes selected may be any two of the processes 2300, 2600, 2700, 2800, 2900 in any of the various implementations. Additionally or alternatively, in at least some implementations, in block 3302, the CfL prediction unit 2202 may receive, identify, and / or use a flag to determine which type of CfL prediction process to use. In some of these implementations, the flag is bypass coded. In other implementations, the flag is context coded. In various embodiments, N contexts may be used, where N is one or more. In some other implementations, different CfL prediction processes are associated with different index values of the chroma intra prediction mode syntax that the CfL prediction unit 2202 may receive, identify, and / or use. The chroma intra prediction mode syntax may indicate to the CfL prediction unit 2202 whether to perform a CfL prediction process, and if so, what type of CfL prediction process to perform. In any of the various implementations, the index value may be one of two, three, or more possible different values to indicate one out of two, three, or more possible different CfL prediction processes.
[0173] FIG. 34 illustrates a flowchart of another example method 3400 of a CfL prediction process that the CfL prediction unit 2202 may perform when operating in the second CfL prediction mode (FIG. 22B) and / or the third CfL prediction mode (FIG. 22C). In block 3402, the CfL prediction unit 2202 may determine a combined or merged chroma block including multiple chroma blocks. In some example implementations, the CfL prediction unit 2202 may determine multiple chroma blocks and combine or merge them together to form a combined or merged chroma block. Also, in at least some implementations, in block 3402, the CfL prediction unit 2202 may determine that multiple chroma blocks from the video bitstream should be predicted in a CfL mode. The CfL prediction unit 2202 may then determine the combined or merged chroma block and / or combine or merge multiple chroma blocks as described above, such as for performing CfL prediction in a CfL prediction mode. In block 3404, the CfL prediction unit 2202 may determine multiple neighboring luma samples of at least one luma block corresponding to the combined chroma block. For example, the at least one luma block may be co-located with multiple chroma blocks that are merged or used to form the combined chroma block. In block 3406, the CfL prediction unit 2202 may generate a neighboring luma average of multiple prediction samples of the multiple chroma blocks based on the neighboring luma samples. For example, the CfL prediction unit 2202 may average the multiple neighboring luma samples to generate a neighboring luma average. In block 3408, the CfL prediction unit 2202 may perform CfL prediction of the multiple chroma blocks based at least on the neighboring luma average. For example, the CfL prediction unit 2202 may use the neighboring luma average as part of the CfL prediction process to generate multiple prediction samples of the multiple chroma blocks according to one or more of the CfL prediction processes described herein.In certain implementations, the CfL prediction unit 2202 may determine a chroma sample prediction using a neighboring luma average, such as by following a CfL prediction process 2800. For example, multiple neighboring luma samples determined in block 3404 may be subsampled in block 2804 and then averaged in block 2806, or averaged in block 2806 without subsampling, to generate a neighboring luma average, which is then used to generate a chroma sample prediction, as described above.
[0174] In some implementations, in block 3402, the CfL prediction unit 2202 may determine to merge or combine multiple chroma blocks if the multiple chroma blocks are determined to be small, such as by being smaller than (not greater than or equal to) at least one size threshold. In some of these implementations, the size threshold is 2xN and / or Nx2, where N is an integer equal to or greater than one. For example, each of the multiple chroma blocks may have a first width and / or a first height that is less than or equal to a first predetermined threshold, such as two as a non-limiting example. The CfL prediction unit 2202 may combine or merge together chroma blocks having a size less than or equal to the size threshold to form a combined chroma block. For at least some of these implementations, the combined chroma block has a second width and / or a second height that is greater than or equal to a second predetermined threshold. For example, the combined chroma blocks may have a second width and / or a second height that are each greater than or equal to four. The CfL prediction unit 2202 may then determine a neighboring luma sample of at least one luma block that corresponds to or is co-located with the merged chroma block in block 3404, and then determine or calculate a DC contribution for the chroma sample of the merged chroma block based on the neighboring luma sample in block 3406. Figure 35 shows an example including four co-located chroma blocks, each having a size of 2x4, combined together and a corresponding 4x8 luma block. Figure 35 also shows left, top-left, and left neighboring luma samples, at least one of which may be used to determine the DC contribution of a chroma prediction sample of the merged chroma block.
[0175] Additionally or alternatively, in some implementations, in block 3402, the CfL prediction unit 2202 may merge or combine two or more chroma blocks having respective transform unit (TU) depths smaller than their respective corresponding luma blocks. The combined chroma blocks may have neighboring chroma samples, and the corresponding luma blocks may have neighboring luma samples. In block 3404, the CfL prediction unit 2202 may determine multiple neighboring luma samples by determining those neighboring luma samples that spatially match, cover, or are co-located with the neighboring chroma samples. Then, in block 3406, the CfL prediction unit 2202 may determine or calculate DC contributions of chroma samples of the combined chroma block using those matching or co-located neighboring luma samples. Figure 36 illustrates an example in which four 2x2 chroma blocks are merged to form a combined co-located 4x4 chroma block. The shaded areas indicate top and left neighboring chroma samples. In the example of Figure 36, each 2x2 chroma block corresponds to a 4x4 luma block that is combined to correspond to the combined 2x2 chroma block. The combined luma block also has above and left neighboring luma samples. The CfL prediction unit 2202 can use those above and left neighboring luma samples that match, cover, and / or are co-located with the neighboring chroma samples of the combined 4x4 chroma block to determine the DC contribution of the chroma block.
[0176] Additionally or alternatively, in some implementations, in block 3402, the CfL prediction unit 2202 may determine to merge or combine multiple chroma blocks, each smaller (or not larger) than a size threshold, as described above. In some of these implementations, the size threshold is 2xN and / or Nx2, where N is an integer equal to or greater than 1. For at least some of these implementations, the combined chroma blocks each have a width and height equal to or greater than 4. In block 3404, the CfL prediction unit 2202 may determine neighboring luma samples of a luma block co-located with the chroma blocks combined in block 3402. In block 3406, the CfL prediction unit 2202 may use at least one of the neighboring luma samples to determine a DC contribution of a chroma prediction sample of the chroma blocks combined in block 3402.
[0177] Figure 37 shows an example of four luma blocks (Block 1, Block 2, Block 3, Block 4) identified as being co-located with four chroma blocks. As described above, each of the four chroma blocks may have a size smaller than a size threshold. Figure 37 also shows at least some neighboring luma samples of the four luma blocks. The neighboring CfL prediction unit 2202 may use one or more of these neighboring luma samples to determine a neighboring luma average.
[0178] In at least some of these implementations, one or more neighboring luma samples of one or more blocks may replace or be used in place of one or more neighboring luma samples of one or more other blocks for determining the DC contribution. To illustrate by way of example in FIG. 37, the CfL prediction unit 2202 may replace the left neighboring luma sample of block 2 with the left neighboring luma sample of block 1. Additionally or alternatively, the CfL prediction unit 2202 may replace the left neighboring luma sample of block 4 with the left neighboring luma sample of block 3. As another example, the upper neighboring luma sample of block 3 may replace the upper neighboring luma sample of block 1, and / or the upper neighboring luma sample of block 4 may replace the upper neighboring luma sample of block 2. After performing any replacement, the CfL prediction unit 2202 may then use one or more of the neighboring luma samples remaining after the replacement to determine the DC contribution.
[0179] FIG. 38 shows a flowchart of another example method 3800 of a CfL prediction process that the CfL prediction unit 2202 may perform. In block 3802, the CfL prediction unit may compare at least one of a size of a chroma block to at least one size threshold, or a transform unit (TU) depth of the chroma block to a TU depth of a corresponding luma block. In block 3804, the CfL prediction unit 2202 may determine a type of CfL prediction process for the chroma block from among a plurality of types of CfL prediction processes based on the comparison. In block 3806, the CfL prediction unit 2202 may perform a CfL prediction process for the chroma block according to the type of CfL prediction process determined in block 3604.
[0180] In at least some implementations of the CfL prediction process 3800, if the comparison at block 3802 indicates that the size of the chroma block is less than or equal to at least one size threshold and / or the TU depth of the chroma block is less than the TU depth of the corresponding luma block, the CfL prediction may determine the type of CfL prediction process to be one that does not use neighboring luma samples, such as CfL prediction process 2300 of Figure 23, as described above. Also, if the comparison at block 3802 indicates that the size of the chroma block is greater than at least one size threshold and / or the TU depth of the chroma block is greater than the TU depth of the corresponding luma block, the CfL prediction unit 2202 may determine the type of CfL prediction process 2600, 2700, 2800, 2900, 3300, or 3400 to be one that uses one or more neighboring luma samples, such as any of the CfL prediction processes, as described above. Additionally or alternatively, in at least some implementations, the at least one size threshold includes at least one of 2xN or Nx2. In certain of these implementations, N is one of 2, 4, 8, 16, 32, 64, or 128, although other numbers of N (including integers) may be possible in other implementations. Additionally or alternatively, in some implementations, if the size of the chroma block is equal to or smaller than the at least one size threshold, the CfL prediction unit 2202 may determine that the type of prediction process is not a type of prediction process. That is, the CfL prediction unit 2202 may determine not to perform any CfL prediction process or to disable all types of CfL prediction processes for chroma blocks equal to or smaller than the at least one size threshold.
[0181] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. The embodiments of the present disclosure may be applied to luma blocks or chroma blocks.
[0182] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 39 illustrates a computer system 3900 suitable for implementing certain embodiments of the disclosed subject matter.
[0183] Computer software can be coded using any suitable machine code or computer language that is amenable to mechanisms such as assembly, compilation, linking, etc., to create code that includes instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly, or via interpretation, microcode execution, etc.
[0184] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0185] 41 with respect to computer system 4100 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 4100.
[0186] The computer system 4100 may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, via tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), video (2D video, 3D video including stereoscopic video, etc.).
[0187] The input human interface devices may include one or more of a keyboard 4101, a mouse 4102, a trackpad 4103, a touch screen 4110, a data glove (not shown), a joystick 4105, a microphone 4106, a scanner 4107, and a camera 4108 (only one of each is shown).
[0188] The computer system 4100 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen 4110, data gloves (not shown), or joystick 4105, although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 4109, headphones (not shown)), visual output devices (such as screens 4110 including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or three- or more-dimensional output via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0189] The computer system 4100 may also include human accessible storage devices and media associated with the storage devices such as optical media including CD / DVD ROM / RW 4120 having media 4121 such as CD / DVDs, thumb drives 4122, removable hard drives or solid state drives 4123, legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), and the like.
[0190] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0191] The computer system 4100 may also include an interface 4154 to one or more communication networks 4155. The network may be, for example, wireless, wired, optical. The network may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial including CAN bus, etc. Certain networks generally require an external network interface adapter connected to a particular general-purpose data port or peripheral bus 4149 (e.g., a USB port of the computer system 4100, etc.), while others are generally integrated into the core of the computer system 4100 by connection to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (4100) can communicate with other entities. Such communications may be, for example, unidirectional, receive only (e.g., broadcast television), unidirectional transmit only (e.g., CANbus to a particular CANbus device), or bidirectional, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.
[0192] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core 4140 of the computer system 4100.
[0193] The core 4140 may include one or more central processing units (CPUs) 4141, graphics processing units (GPUs) 4142, dedicated programmable processing units in the form of field programmable gate areas (FPGAs) 4143, hardware accelerators for specific tasks 4144, graphics adapters 4150, etc. These devices may be connected through a system bus 4148 along with read only memory (ROM) 4145, random access memory 4146, internal mass storage 4147 such as an internal non-user accessible hard drive, SSD, etc. In some computer systems, the system bus 4148 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 4148 or through a peripheral bus 4149. In one example, a screen 4110 may be connected to the graphics adapter 4150. Architectures for peripheral buses include PCI, USB, etc.
[0194] The CPU 4141, GPU 4142, FPGA 4143, and accelerator 4144 can execute certain instructions that may combine to make up the aforementioned computer code. This computer code can be stored in ROM 4145 or RAM 4146. Transient data can also be stored in RAM 4146, while permanent data can be stored, for example, in internal mass storage 4147. Rapid storage and retrieval in any of the memory devices can be made possible by the use of cache memories that can be closely associated with one or more of the CPU 4141, GPU 4142, mass storage 4147, ROM 4145, RAM 4146, etc.
[0195] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the available kind well known to those skilled in the computer software arts.
[0196] As a non-limiting example, the architecture 4100, and specifically a computer system having the core 4140, can provide functionality as a result of the processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with a particular storage of the core 4140 that is non-transitory in nature, such as the core internal mass storage 4147 and ROM 4145. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 4140. The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core 4140, and specifically the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain portions of certain processes described herein, including defining data structures stored in RAM 4146 and modifying such data structures according to the processes defined by the software. Additionally, or alternatively, a computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator 4144) that may operate in place of or in conjunction with software to perform particular processes or particular portions of particular processes described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0197] The subject matter of the present disclosure may also relate to or include, among other aspects, the following aspects:
[0198] A first aspect includes a method for video processing, comprising the steps of determining that a plurality of chroma blocks from a video bitstream should be predicted in a chroma from luma (CfL) prediction mode, where a first width or a first height of each of the plurality of chroma blocks is less than or equal to a first predetermined threshold; combining the plurality of chroma blocks into a combined chroma block, where a second width or a second height of the combined chroma block is greater than or equal to a second predetermined threshold; determining a plurality of reconstructed neighboring luma samples of one or more luma blocks corresponding to the combined chroma block; averaging the plurality of reconstructed neighboring luma samples to generate a neighboring luma average; and performing CfL prediction of the plurality of chroma blocks based on at least the neighboring luma average.
[0199] A second aspect includes the first aspect, further including that the first predetermined threshold is two.
[0200] A third aspect includes any of the first or second aspects, further including that the second predetermined threshold is four.
[0201] A fourth aspect includes any of the first to third aspects, wherein each of the multiple chroma blocks has a transform unit depth that is shallower than a respective transform unit depth of a corresponding luma block of each of the one or more luma blocks.
[0202] A fifth aspect includes any of the first to fourth aspects, and further includes determining a plurality of neighboring chroma samples of the combined chroma block, wherein determining a plurality of reconstructed neighboring luma samples includes determining a plurality of reconstructed neighboring luma samples co-located with the plurality of neighboring chroma samples.
[0203] A sixth aspect includes any of the first to fifth aspects, wherein the one or more luma blocks include a plurality of luma blocks, and the step of determining the plurality of reconstructed neighboring luma samples includes replacing a first set of reconstructed neighboring luma samples of a first luma block of the plurality of luma blocks with a second set of reconstructed neighboring luma samples of a second luma block of the plurality of luma blocks.
[0204] A seventh aspect includes a method of video processing including: comparing at least one of a size of a chroma block to at least one size threshold, or a transform unit (TU) depth of the chroma block to a TU depth of a corresponding luma block; determining a type of chroma from luma (CfL) prediction process for the chroma block from among a plurality of types of CfL prediction processes based on the comparison; and performing a CfL prediction process for the chroma block according to the type of CfL prediction process.
[0205] An eighth aspect includes the seventh aspect, and further includes that the step of determining a type CfL prediction process based on the comparison includes determining a type of CfL prediction process that does not use neighboring luma samples in response to the comparison indicating at least one of: a size of the chroma block is less than or equal to at least one size threshold, or a TU depth of the chroma block is shallower than a TU depth of a corresponding luma block.
[0206] A ninth aspect includes any of the seventh or eighth aspects, and further includes that determining a type of the CfL prediction process based on the comparison includes determining a type of CfL prediction process using one or more neighboring luma samples in response to the comparison indicating at least one of: the size of the chroma block is greater than the at least one size threshold, or the TU depth of the chroma block is greater than the TU depth of the corresponding luma block.
[0207] A tenth aspect includes the ninth aspect, and further includes that the type of CFL prediction process uses one or more neighboring luma samples to determine alternating current (AC) contributions of multiple prediction samples of a chroma block.
[0208] An eleventh aspect includes the tenth aspect, wherein the type of CFL prediction process further includes using one or more nearby luma samples to determine a nearby luma average of the AC contribution.
[0209] A twelfth aspect includes the ninth aspect, wherein the type of CFL prediction process further includes mapping one or more neighboring luma samples to one or more neighboring chroma samples of a chroma block.
[0210] A thirteenth aspect includes any of the seventh or ninth to eleventh aspects, and further includes the type of CFL prediction process determining a neighboring luma average of multiple neighboring luma samples of one or more luma blocks corresponding to a combined chroma block that includes the chroma block.
[0211] A fourteenth aspect includes any of the seventh to thirteenth aspects, further including that the at least one size threshold includes 2xN or Nx2.
[0212] A fifteenth aspect includes the fourteenth aspect, further including that N comprises one of 2, 4, 8, 16, 32, 64, or 128.
[0213] A sixteenth aspect includes any of the seventh to fifteenth aspects, and further includes, in response to the size of the chroma block being less than or equal to at least one size threshold, the type of the CFL prediction process is not a type.
[0214] A seventeenth aspect includes a video processing device including a memory that stores a set of instructions and a processor configured to execute the set of instructions, the processor being configured, upon execution of the set of instructions, to determine that multiple chroma blocks from a video bitstream should be predicted in a chroma from luma (CfL) prediction mode, determine a combined chroma block including the multiple chroma blocks, determine multiple reconstructed neighboring luma samples of one or more luma blocks corresponding to the combined chroma block, generate direct current (DC) contributions of multiple prediction samples of the multiple chroma blocks based on the multiple reconstructed neighboring luma samples, and perform CfL prediction of the multiple chroma blocks based on at least a neighboring luma average.
[0215] An eighteenth aspect includes the seventeenth aspect, and further includes that the processor, upon execution of the instructions, is further configured to combine the plurality of chroma blocks to determine a combined chroma block.
[0216] A nineteenth aspect includes any of the seventeenth or eighteenth aspects, further including that each of the plurality of chroma blocks is less than or equal to at least one size threshold.
[0217] A twentieth aspect includes any of the seventeenth to nineteenth aspects, wherein each of the multiple chroma blocks has a transform unit depth that is shallower than a respective transform unit depth of a corresponding luma block of each of the one or more luma blocks.
[0218] In addition to the features mentioned in each of the independent aspects listed above, some examples may exhibit, alone or in combination, any features mentioned in the dependent aspects and / or disclosed in the above description and illustrated in the figures.
[0219] While this disclosure describes several exemplary embodiments, there are alterations, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Thus, it will be appreciated that those skilled in the art can devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.
[0220] Appendix A: Acronyms JEM joint exploration model VVC Versatile video coding BMS benchmark set MV Motion Vector HEVC High Efficiency Video Coding SEI Supplementary Enhancement Information VUI Video Usability Information GOP Groups of Pictures TU Transform Unit PU Prediction Units CTU Coding Tree Units CTB Coding Tree Blocks PB Prediction Block HRD Hypothetical Reference Decoder SNR Signal Noise Ratio CPU Central Processing Unit GPU Graphics Processing Unit CRT Cathode Ray Tube LCD Liquid Crystal Display OLED Organic Light-Emitting Diode CD Compact Disc DVD Digital Video Disc ROM Read-Only Memory RAM Random Access Memory ASIC Application-Specific Integrated Circuit PLD Programmable Logic Device LAN Local Area Network GSM Global System for Mobile communication LTE Long-Term Evolution CANBus Controller Area Network Bus USB Universal Serial Bus PCI Peripheral Component Interconnect FPGA Field Programmable Gate Area SSD Solid-state drive IC Integrated Circuit HDR high dynamic range SDR standard dynamic range JVET Joint Video Exploration Team MPM most probable mode WAIP Wide-Angle Intra Prediction CU Coding Unit PU Prediction Unit TU Transform Unit CTU Coding Tree Unit PDPC Position Dependent Prediction Combination ISP Intra Sub-Partitions SPS Sequence Parameter Setting PPS Picture Parameter Set APS Adaptation Parameter Set VPS Video Parameter Set DPS Decoding Parameter Set ALF Adaptive Loop Filter SAO Sample Adaptive Offset CC-ALF Cross-Component Adaptive Loop Filter CDEF Constrained Directional Enhancement Filter CCSO Cross-Component Sample Offset LSO Local Sample Offset LR Loop Restoration Filter AV1 AOMedia Video 1 AV2 AOMedia Video 2 LFNST Low-Frequency Non-Separable Transform IST Intra Secondary Transform [Explanation of symbols]
[0221] 100 luma samples 101 Point where arrows converge 101 nearby luma samples 101 Samples 102 Arrow 103 Arrow 104 Square Block 104 Block 111 ChromaPixel PL ij 112 kth neighboring luma sample PL k 180 Schematic diagram 201 Current Block 202 Ambient Samples 203 Ambient Samples 204 Ambient Samples 205 Ambient Samples 206 Ambient Samples 300 Communication Systems 310 Terminal 320 Terminal 330 Terminal 340 Terminal 350 Communication Network 350 Network 400 Communication Systems 401 Video Source 402 Stream 403 Video Encoder 403 Encoder 404 Video Bitstream 404 Video Data 405 Streaming Server 406 Client Subsystem 407 Copy 407 Incoming Copy 407 Video Data 408 Client Subsystem 409 Video Data 409 Copy 410 Decoder 410 Video Decoder 411 Video Picture 412 Display 413 Video Capture Subsystem 420 Electronic Devices 430 Electronic Devices 501 Channel 510 Video Decoder 512 Rendering Device 512 Display 515 Buffer Memory 520 Entropy Decoder / Parser 520 Parser 521 Symbols 530 Electronic Devices 531 Receiver 551 Scaler / Descaler Unit 551 Scaler / Inverse Transform 552 Intra Prediction Units 552 Intra-picture prediction unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Sources 603 Video Coder 603 Video Encoder 620 Electronic Devices 620 Encoder 630 Source Coder 632 Coding Engine 633 Local Decoder 633 Local Video Decoder 633 Decoding Unit 633 Decoder 634 Reference Picture Cache 634 Reference Picture Memory 635 Predictor 640 Transmitter 643 Video Sequences 645 Entropy Coder 650 Controller 660 Communication Channels 703 Video Encoder 721 General Controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 InterEncoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 902 Partitioning option or pattern 904 Partitioning option or pattern 906 Partitioning options or patterns 908 Partitioning options or patterns 1002 T-type partition 1004 T-type partition 1006 T-type partition 1008 T-type partition 1010 Full Square Partition 1102 Vertical bisection 1104 Horizontal bisection 1106 Vertical third division 1108 Horizontal third division 1200 Base Block 1202 Square Partition 1204 Square Partition 1206 Square Partition 1208 Square Partition 1302 Vertical 1304 horizontal 1402 Partition 1404 Partition 1406 Partition 1408 Partition 1410 Partition Pattern 1420 Corresponding tree structure / representation 1502 Square coding block 1504 First Level Split 1506 Second Level Split 1602 Block 1604 Block 2002 Block 2004 Sample above 2004 Upper Neighborhood Samples 2006 Upper left sample 2006 Upper left sample 2008 Neighborhood Samples 2008 Sample on the left 2008 Left Neighborhood Samples 2010 Sample 2202 CfL Prediction Unit 2300 CfL Prediction Method 2300 Process 2300 methods 2300 CfL forecasting process 2302 Block 2304 Block 2306 Block 2308 Block 2310 Block 2502 Luma Block 2504 area 2504 left neighbor luma samples 2506 top left luma samples 2506 area 2508 upper near-luma samples 2508 area 2510 upper right near luma samples 2512 right neighbor luma samples 2514 Luma samples near bottom right 2516 lower-neighborhood luma samples 2518 Luma samples near bottom left 2600 CfL Forecasting Process 2600 methods 2602 Block 2604 Block 2700 CfL Forecasting Process 2700 method 2702 Block 2704 Block 2706 Block 2800 methods 2800 CfL Forecasting Process 2802 Block 2804 Block 2806 Block 2808 Block 2810 Block 2812 Block 2900 CfL Prediction Method 2900 method 2902 Block 2904 Block 3002 Arrow 3004 Arrow 3006 Arrow 3300 method 3302 Block 3304 Block 3400 method 3402 Block 3404 Block 3406 Block 3408 Block 3604 Block 3800 CfL Forecasting Process 3800 method 3802 Block 3804 Block 3806 Block 3900 Computer Systems 4100 Architecture 4100 Computer Systems 4101 Keyboard 4102 Mouse 4103 Trackpad 4105 Joystick 4106 Microphone 4107 Scanner 4108 Camera 4109 Speaker 4110 Touch Screen 4121 Medium 4122 Thumb Drive 4123 Solid State Drive 4140 cores 4141 Central Processing Unit (CPU) 4142 Graphics Processing Unit (GPU) 4143 Field Programmable Gate Area (FPGA) 4144 Hardware accelerators for specific tasks 4144 Accelerator 4145 Read-Only Memory (ROM) 4146 Random Access Memory 4147 Internal Mass Storage 4147 Core Internal Mass Storage Unit 4147 Mass storage unit 4148 System Bus 4149 Surrounding bus 4150 Graphics Adapter 4154 Interface 4155 Communication Network
Claims
1. 1. A method for video processing, comprising: determining that a plurality of chroma blocks from a video bitstream should be predicted in a chroma from luma (CfL) prediction mode, wherein a first width or a first height of each of the plurality of chroma blocks is less than or equal to a first predetermined threshold; combining the plurality of chroma blocks into a combined chroma block, wherein a second width or a second height of the combined chroma block is greater than or equal to a second predetermined threshold; determining a plurality of reconstructed neighboring luma samples of one or more luma blocks corresponding to the combined chroma block; averaging the plurality of reconstructed neighboring luma samples to generate a neighboring luma average; performing CfL prediction for the plurality of chroma blocks based at least on the neighborhood luma average; A method comprising:
2. The method of claim 1 , wherein the first predetermined threshold is two.
3. The method of claim 1 , wherein the second predetermined threshold is four.
4. 2. The method of claim 1, wherein each of the plurality of chroma blocks has a transform unit depth that is shallower than a respective transform unit depth of a corresponding luma block of each of the one or more luma blocks.
5. determining a plurality of neighboring chroma samples of the combined chroma block; The method of claim 1 , wherein determining the plurality of reconstructed neighboring luma samples comprises determining a plurality of reconstructed neighboring luma samples co-located with the plurality of neighboring chroma samples.
6. The one or more luma blocks include a plurality of luma blocks, and the step of determining a plurality of reconstructed neighboring luma samples comprises:
2. The method of claim 1, comprising replacing a first set of reconstructed neighboring luma samples of a first luma block of the plurality of luma blocks with a second set of reconstructed neighboring luma samples of a second luma block of the plurality of luma blocks.
7. 1. A method of video processing, comprising: comparing at least one of a size of a chroma block to at least one size threshold or a transform unit (TU) depth of the chroma block to a TU depth of a corresponding luma block; determining a type of chroma (CfL) prediction process from luma for the chroma block from a plurality of types of CfL prediction processes based on the comparison; performing a CfL prediction process for the chroma block according to a type of the CfL prediction process; A method comprising:
8. determining a type of the CfL prediction process based on the comparison, 8. The method of claim 7, comprising determining a type of CfL prediction process that does not use neighboring luma samples in response to the comparison indicating at least one of: the size of the chroma block being less than or equal to the at least one size threshold; or the TU depth of the chroma block being less than the TU depth of the corresponding luma block.
9. determining a type of the CfL prediction process based on the comparison, 8. The method of claim 7, comprising: determining a type of CfL prediction process to use one or more neighboring luma samples in response to the comparison indicating at least one of: the size of the chroma block being greater than the at least one size threshold; or the TU depth of the chroma block being greater than the TU depth of the corresponding luma block.
10. 10. The method of claim 9, wherein the type of CFL prediction process uses the one or more neighboring luma samples to determine an AC contribution of multiple prediction samples of the chroma block.
11. The method of claim 10, wherein the type of CFL prediction process uses the one or more nearby luma samples to determine a nearby luma average of the AC contribution.
12. 10. The method of claim 9, wherein the type of CFL prediction process maps the one or more neighboring luma samples to one or more neighboring chroma samples of the chroma block.
13. 10. The method of claim 9, wherein the type of CFL prediction process determines a neighboring luma average of multiple neighboring luma samples of one or more luma blocks corresponding to a combined chroma block that includes the chroma block.
14. The method of claim 7 , wherein the at least one size threshold comprises 2×N or N×2.
15. 15. The method of claim 14, wherein N comprises one of 2, 4, 8, 16, 32, 64, or 128.
16. The method of claim 7 , further comprising: in response to the size of the chroma block being less than or equal to the at least one size threshold, the type of the CFL prediction process being a non-type.
17. 1. A video processing device, comprising: a memory for storing a set of instructions; and a processor configured to execute the set of instructions, said processor comprising: determining that a plurality of chroma blocks from the video bitstream are to be predicted in a chroma from luma (CfL) prediction mode; determining a combined chroma block comprising the plurality of chroma blocks; determining a plurality of reconstructed neighboring luma samples of one or more luma blocks corresponding to the combined chroma block; generating a direct current (DC) contribution for a plurality of prediction samples for the plurality of chroma blocks based on the plurality of reconstructed neighboring luma samples; 11. A video processing device comprising: a processor configured to perform CfL prediction of the plurality of chroma blocks based on at least a neighboring luma average generated by averaging the neighboring luma samples.
18. 20. The video processing device of claim 17, wherein the processor, upon execution of the set of instructions, is further configured to combine the multiple chroma blocks to determine the combined chroma block.
19. 20. The video processing device of claim 17, wherein each of the plurality of chroma blocks is less than or equal to at least one size threshold.
20. 20. The video processing device of claim 17, wherein each of the plurality of chroma blocks has a transform unit depth shallower than a respective transform unit depth of a corresponding luma block of each of the one or more luma blocks.
Citation Information
Patent Citations
Image decoding device, image decoding method and program
JP2021002813A
Method and system for processing luma and chroma signals
US20200404278A1
Intra-Prediction Using a Cross-Component Linear Model in Video Coding
US20210136409A1
Method and apparatus of local illumination compensation for predictive coding
US20210352277A1
Signaling of chroma residual scaling
US20220124340A1