Method, apparatus, and program for cross-component filtering

The method enhances video encoding/decoding by applying a cross-component filter to saturation blocks based on specific formats, addressing inefficiencies in existing techniques and improving compression efficiency and video quality.

JP7678754B2Active Publication Date: 2025-05-16TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2021549113
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-02
Filing Date
2020-09-10
Publication Date
2025-05-16
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

Existing video encoding techniques face challenges in efficiently reducing redundancy and improving compression efficiency, particularly in handling intra prediction and motion compensation across various video coding standards.

Method used

The proposed solution involves a method for video encoding/decoding that includes processing circuitry capable of applying a cross-component filter to saturation blocks, determining the filter shape based on saturation subsampling and sample formats, and generating intermediate blocks through loop filtering and cross-component filtering.

Benefits of technology

This approach enhances compression efficiency by effectively filtering saturation components based on determined filter shapes, leading to improved video quality and reduced bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678754000023
    Figure 0007678754000023
  • Figure 0007678754000024
    Figure 0007678754000024
  • Figure 0007678754000025
    Figure 0007678754000025
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a method and an apparatus including a processing circuit for video decoding. The processing circuit decodes coded information of a chroma coding block (CB) from a coded video bitstream. The coded information indicates that a cross-component filter is to be applied to the chroma CB and indicates a chroma subsampling format and a chroma sample format. The processing circuit determines a filter shape of the cross-component filter based on at least one of the chroma subsampling format and the chroma sample format. The processing circuit generates a first intermediate CB by applying a loop filter to the chroma CB, and generates a second intermediate CB by applying the cross-component filter having the determined filter shape to the corresponding luma CB. The processing circuit determines a filtered chroma CB based on the first intermediate CB and the second intermediate CB.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [Incorporated by reference] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 901,118, entitled "Of Cross-Component Adaptive Loop Filter," filed September 16, 2019, which claims the benefit of priority to U.S. Provisional Application No. 17 / 010,403, entitled "Method and Apparatus for Cross-Component Filtering," filed September 2, 2020. The entire disclosures of the prior applications are hereby incorporated by reference in their entireties.

[0002] This disclosure describes embodiments generally related to video encoding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. The inventors' work, to the extent that it is described in this background section, and aspects of the description that may not be admitted as prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of images, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of images can have a fixed or variable image rate (also informally known as frame rate) of, for example, 60 images per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GByte of storage space.

[0005] One of the goals of video encoding and decoding may be the reduction of redundancy in the input video signal through compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, possibly by two orders of magnitude or more. Both lossless and lossy compression, as well as combinations thereof, may be employed. Lossless compression refers to techniques where an exact copy of the original signal can be restored from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely adopted. The amount of acceptable distortion depends on the application, e.g., users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio may reflect that a higher acceptable / tolerable distortion may result in a higher compression ratio.

[0006] Video encoders and decoders may utilize techniques from a number of broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007] Video coding techniques can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to the samples or other data from a previously reconstructed reference picture. In some video coding, an image is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the image may be an intra-image. Intra-images and their derivatives, such as independent decoder refresh images, can be used to reset the decoder state and thus can be used as the first image in a coded video bitstream and video session or as still images. Samples of an intra-block may be subjected to a transform, and the transform coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are required for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, for example as known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to do so from surrounding sample data and / or metadata obtained during the encoding / decoding of blocks of spatially adjacent and preceding data in decoding order. Such techniques are hereafter referred to as "intra-prediction" techniques. It should be noted that in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from a reference picture.

[0009] Intra prediction can take many different forms. When two or more of such techniques can be used in a given video coding technique, the technique in use can be coded in an intra prediction mode. In some cases, a mode can have sub-modes and / or parameters, which can be coded separately or included in a mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination can affect the coding efficiency gains via intra prediction, and so can the entropy coding technique used to convert the codeword into a bitstream.

[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Generic Video Coding (VVC), and Benchmark Set (BMS). Predictor blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to the direction. The reference to the direction in use can be coded in the bitstream or it can be predicted itself.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of 9 known predictor directions from the 33 possible predictor directions of H.265 (corresponding to the 33 angle modes of the 35 intra modes). The point where the arrows converge (101) represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the top right sample at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted to the bottom left sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] Still referring to FIG. 1A, at the top left is shown a square block (104) of 4×4 samples (shown in dashed bold). The square block (104) includes 16 samples, each labeled with an “S”, its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). Since the size of the block is 4×4 samples, S44 is at the bottom right. Further shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed. Therefore, there is no need to use negative values.

[0013] Intra prediction can work by copying reference sample values ​​from adjacent samples as appropriate for the signaled prediction direction. For example, assume that the coded video bit stream includes a signal indicating for this block a prediction direction that coincides with the arrow (102), i.e., the top right sample is predicted from the prediction sample at an angle of 45 degrees from the horizontal. Then samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.

[0014] In some cases, especially when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.

[0015] The number of possible directions is increasing as video coding techniques develop. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions at the time of this disclosure. Experiments have been performed to identify the most likely directions, and certain techniques in entropy coding are used to represent those likely directions with a small number of bits, accepting certain penalties for less likely directions. Furthermore, the direction itself may be predictable from neighboring directions used in neighboring already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (180) illustrating 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits in the coded video bit stream to represent the directions may vary from video coding technique to video coding technique. For example, it can range from a simple direct mapping from prediction direction to intra-prediction mode, to complex adaptation schemes involving codewords, most likely modes, and similar techniques. In all cases, however, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is redundancy reduction, in a well-performing video coding technique, those less likely directions are represented with a greater number of bits than the more likely directions.

[0018] Motion compensation can be a lossy compression technique, and can refer to a technique in which blocks of sample data from a previously reconstructed image or part thereof (reference image) are used to predict a newly reconstructed image or image part after being spatially shifted in a direction indicated by a motion vector (MV or later). In some cases, the reference image can be the same as the image currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, with the third dimension being a representation of the reference image being used (the latter can indirectly be the temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular region of sample data can be predicted from other MVs, for example from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing compression. MV prediction can work effectively, for example, when encoding an input video signal derived from a camera (known as natural footage), because there is a statistical likelihood that regions larger than the region to which a single MV is applicable will move in similar directions and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of neighboring regions. This ensures that the MV found for a given region will be similar or the same as the MV predicted from the surrounding MVs, and after entropy encoding, can be represented with fewer bits than would be used to directly encode the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when computing a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Here, we will explain a technique called "spatial merging" among the many MV prediction mechanisms provided by H.265.

[0021] Referring to Figure 2, the current block (201) contains samples found by the encoder during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of directly encoding its MV, the MV can be derived from metadata associated with one or more reference pictures, for example the most recent (in decoding order) reference picture, using the MV associated with any one of the five surrounding samples denoted A0, A1, and B0, B1, B2 (202 to 206, respectively). In H.265, the MV prediction can use predictors from the same reference picture as the neighboring blocks are using. Summary of the Invention [Means for solving the problem]

[0022] An embodiment of the present disclosure provides a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit can decode coded information of a chroma coding block (CB) from a coded video bitstream. The coded information can indicate that a cross-component filter is applied to the chroma CB. The coded information can further indicate a chroma subsampling format and a chroma sample format indicating a relative position of the chroma sample with respect to at least one luma sample in a corresponding luma CB. The processing circuit can determine a filter shape of the cross-component filter based on at least one of the chroma subsampling format and the chroma sample format. The processing circuit can generate a first intermediate CB by applying a loop filter to the chroma CB. The processing circuit can generate a second intermediate CB by applying a cross-component filter having the determined filter shape to the corresponding luma CB. The processing circuit can determine a filtered chroma CB based on the first intermediate CB and the second intermediate CB.

[0023] In one example, the chroma sample format is signaled in the encoded video bitstream.

[0024] In one example, a number of filter coefficients of the cross-component filter are signaled in the encoded video bitstream, and the processing circuitry can determine a filter shape of the cross-component filter based on the number of filter coefficients and at least one of the chroma subsampling format and the chroma sample type.

[0025] In one example, the chroma subsampling format is 4:2:0. The at least one luma sample includes four luma samples, which are a top left sample, a top right sample, a bottom left sample, and a bottom right sample. The chroma sample format is one of six chroma sample formats 0-5 indicating six relative positions 0-5, respectively, where the six relative positions 0-5 of the chroma sample correspond to a left center position between the top left sample and the bottom left sample, a center position of the four luma samples, a top left position at the same position as the top left sample, a top center position between the top left sample and the top right sample, a bottom left position at the same position as the bottom left sample, and a bottom center position between the bottom left sample and the bottom right sample. The processing circuitry can determine a filter shape of the cross component filter based on the chroma sample format. In one example, the coded video bitstream includes a cross component linear model (CCLM) flag indicating that the chroma sample format is 0 or 2.

[0026] In one example, the cross-component filter is a cross-component adaptive loop filter (CC-ALF) and the loop filter is an adaptive loop filter (ALF).

[0027] In one embodiment, the range of the filter coefficients of the cross-component filter is equal to or less than K bits, where K is a positive integer. In one example, the filter coefficients of the cross-component filter are encoded using fixed-length coding. In one example, the processing circuit can shift the corresponding luma sample value of the luma CB to have a dynamic range of 8 bits based on the dynamic range of the luma sample value being greater than 8 bits, where K is 8 bits. The processing circuit can apply the cross-component filter having the determined filter shape to the shifted luma sample value.

[0028] In some examples, an apparatus for video decoding includes a processing circuit. The processing circuit can decode coded information of a chroma CB from the coded video bitstream. The coded information can indicate that a cross-component filter is applied to the chroma CB based on a corresponding luma CB. The processing circuit can generate a downsampled luma CB by applying a downsampling filter to the corresponding luma CB, where a chroma horizontal subsampling factor and a chroma vertical subsampling factor between the chroma CB and the downsampled luma CB are one. The processing circuit can generate a first intermediate CB by applying a loop filter to the chroma CB. The processing circuit can generate a second intermediate CB by applying a cross-component filter to the downsampled luma CB, where a filter shape of the cross-component filter is independent of a chroma subsampling format and a chroma sample format of the chroma CB. The chroma sample format can indicate a relative position of the chroma sample with respect to at least one luma sample in the corresponding luma CB. The processing circuit can determine a filtered chroma CB based on the first intermediate CB and the second intermediate CB. In one example, the downsampling filter corresponds to a filter applied to the co-located luma samples in CCLM mode. In one example, the downsampling filter is a {1,2,1;1,2,1} / 8 filter and the chroma subsampling format is 4:2:0. In one example, the filter shape of the cross-component filter is one of a 7x7 diamond shape, a 7x7 square shape, a 5x5 diamond shape, a 5x5 square shape, a 3x3 diamond shape, and a 3x3 square shape.

[0029] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video decoding.

[0030] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0031] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 4 is a block diagram of a decoder according to another embodiment. [Figure 9] 1A-1C are diagrams illustrating examples of filter shapes according to embodiments of the present disclosure. [Figure 10A] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10B] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10C] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10D] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 11A] FIG. 13 illustrates an example of virtual boundary filtering according to an embodiment of the present disclosure. [Figure 11B] FIG. 13 illustrates an example of virtual boundary filtering according to an embodiment of the present disclosure. [Figure 12A] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12B] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12C] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12D] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12E] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12F] 1 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 13] FIG. 2 is an exemplary functional diagram for generating luma and chroma components according to one embodiment of the present disclosure. [Figure 14] 14 illustrates an example of a filter 1400 according to an embodiment of the present disclosure. [Figure 15A] 4 illustrates an example location of chroma samples relative to luma samples, according to an embodiment of the present disclosure. [Figure 15B] 4 illustrates an example location of chroma samples relative to luma samples, according to an embodiment of the present disclosure. [Figure 16] 16 shows examples of filter shapes (1601) to (1603) of the cross-component adaptive loop filters (CC-ALF) according to an embodiment of the present disclosure. [Figure 17] 17 shows a flow chart outlining a process (1700) according to one embodiment of the present disclosure. [Figure 18] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0032] FIG. 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a stream of video images captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bit streams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to reconstruct the video images, and display the video images according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0033] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of encoded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a stream of video images captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), can decode the encoded video data to recover the video images, and can display the video images on an accessible display device according to the recovered video data.

[0034] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure need not be so limited. The embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that convey encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (cable) and / or wireless communication networks. The communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (350) may not be important to the operation of the present disclosure, unless otherwise described herein below.

[0035] 4 illustrates an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0036] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that generates an uncompressed video image stream (402). In one example, the video image stream (402) includes samples captured by a digital camera. The video image stream (402), shown as a thick line to emphasize the high amount of data compared to the encoded video data (404) (or encoded video bit stream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bit stream (404)), shown as a thin line to emphasize the lower amount of data compared to the video image stream (402), may be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to obtain copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example in an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video images that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bit stream) can be encoded according to a particular video encoding / compression standard, such as ITU-T Recommendation H.265. In one example, the video encoding standard under development is informally known as Versatile Video Coding (VVC). The disclosed The subject matter presented may be used in the context of VVC.

[0037] It should be noted that the electronics (420) and (430) may include other components (not shown). For example, the electronics (420) may include a video decoder (not shown), and the electronics (430) may also include a video encoder (not shown).

[0038] 5 shows a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used in place of the video decoder (410) of the example of FIG. 4.

[0039] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) may receive coded video data with other data, e.g., coded audio data and / or auxiliary data streams, which may be transferred to each using an entity (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / analyzer (520) (hereinafter "analyzer (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be external to the video decoder (510) (not shown). In still others, there may be a buffer memory (not shown) external to the video decoder (510), e.g., to combat network jitter, and yet another buffer memory (515) internal to the video decoder (510), e.g., to handle playback timing. When the receiver (531) is receiving data from a store-and-forward device of sufficient bandwidth and controllability, or from an asynchronous network, the buffer memory (515) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be needed and may be relatively large, advantageously of adaptive size, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0040] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device such as a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device may be in the form of additional extension information (SEI message) or a video usability information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence may follow a video encoding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The analyzer (520) can extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on the at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), images, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The analyzer (520) can also extract coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0041] The analyzer (520) can perform entropy decoding / analysis operations on the video sequence received from the buffer memory (515) to produce symbols (521).

[0042] The reconstruction of the symbols (521) may involve several different units, depending on the format of the coded video image or parts thereof (e.g., inter and intra images, inter and intra blocks), and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following units is not shown for clarity.

[0043] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0044] The first unit is a scalar / inverse transform unit (551). The scalar / inverse transform unit (551) receives quantized transform coefficients as well as control information from the analyzer (520) including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. as symbol(s) (521). The scalar / inverse transform unit (551) may output blocks comprising sample values ​​that may be input to an aggregator (555).

[0045] In some cases, the output samples of the scalar / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information may be provided by an intra-image prediction unit (552). In some cases, the intra-image prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information retrieved from a current image buffer (558). The current image buffer (558) buffers, for example, a partially reconstructed current image and / or a fully reconstructed current image. The aggregator (555) may append the prediction information generated by the intra-prediction unit (552) to the output sample information from the scalar / inverse transform unit (551) on a sample-by-sample basis.

[0046] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access a reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) fetches the prediction samples may be controlled by motion vectors available to the motion compensated prediction unit (553) in the form of symbols (521), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0047] The output samples of the aggregator (555) can be subjected to various loop filtering techniques in a loop filter unit (556). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called coded video bit stream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also be responsive to meta-information obtained during the decoding of a coded image or previous (in decoding order) part of the coded video sequence, or to previously reconstructed and loop filtered sample values.

[0048] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) and also stored in a reference image memory (557) for use in future inter-image prediction.

[0049] Once fully reconstructed, a particular coded image can be used as a reference image for future predictions. For example, once a coded image corresponding to a current image is fully reconstructed and the coded image is identified (e.g., by the analyzer (520)) as a reference image, the current image buffer (558) can become part of the reference image memory (557) and a new current image buffer can be relocated before commencing reconstruction of a subsequent coded image.

[0050] The video decoder (510) may perform decoding operations according to a given video compression technique in a standard, such as ITU-T Rec. H.265. The encoded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the encoded video sequence complies with both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, the profile may select a particular tool from all tools available in the video compression technique or standard as the only tool available under that profile. Also, what is required for compliance may be that the complexity of the encoded video sequence is within the bounds defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference image size, etc. The limits set by the level may be further limited in some cases by metadata for HRD buffer management and hypothetical reference decoder (HRD) specifications signaled in the encoded video sequence.

[0051] In one embodiment, the receiver (531) can receive additional (redundant) data with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.

[0052] 6 shows a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG. 4.

[0053] The video encoder (603) can receive video samples from a video source (601) (which in the example of FIG. 6 is not part of the electronic device (620)) that can capture video images to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0054] The video source (601) may provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared footage. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual images that give motion when viewed in succession. The image itself may be organized as a spatial array of pixels, each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art will readily appreciate the relationship between pixels and samples. The following description focuses on samples.

[0055] According to one embodiment, the video encoder (603) can encode and compress images of a source video sequence into an encoded video sequence (643) in real-time or under any other time constraint required by the application. Enforcing an appropriate encoding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. Couplings are not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions for the video encoder (603) optimized for a particular system design.

[0056] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an oversimplified explanation, in one example, the encoding loop can include a source encoder (630) (e.g., responsible for generating symbols, such as a symbol stream, based on an input image to be encoded and a reference image) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a similar manner as the (remote) decoder also creates (since in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bit stream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since the decoding of the symbol stream results in bit-exact results independent of the decoder location (local or remote), the content in the reference image memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" as reference image samples exactly the same sample values ​​that the decoder "sees" when using prediction during decoding. This basic principle of reference image synchrony (if synchrony cannot be maintained due to e.g. channel errors, resulting in drift) is also used in some related technologies.

[0057] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510) already described in detail in connection with Figure 5. However, with brief reference also to Figure 5, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy encoder (645) and analyzer (520) may be lossless, the entropy decoder of the video decoder (510), including the buffer memory (515), and the analyzer (520) may not be fully implemented in the local decoder (633).

[0058] An observation that can be made at this point is that any decoder techniques, except for parsing / entropy decoding, present in the decoder must also be present in substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the decoder operation. Descriptions of the encoder techniques can be omitted since they are the inverse of the decoder techniques described generically. Only in certain areas are more detailed descriptions required and are provided below.

[0059] During operation, in some examples, the source encoder (630) may perform motion compensated predictive encoding, which predictively encodes an input image with reference to one or more previously encoded images from a video sequence designated as “reference images.” In this manner, the encoding engine (632) encodes differences between pixel blocks of the input image and pixel blocks of reference images that may be selected as predictive references for the input image.

[0060] The local video decoder (633) may decode the encoded video data of the image that may be designated as the reference image based on the symbols generated by the source encoder (630). The operation of the encoding engine (632) may advantageously be a lossy process. When the encoded video data may be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may be a replica of the source video sequence, usually with some errors. The local video decoder (633) may replicate the decoding process that may be performed on the reference image by the video decoder and store the reconstructed reference image in a reference image cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference image that has a common content as the reconstructed reference image that will be obtained by the far-end video decoder (without transmission errors).

[0061] The predictor (635) can perform the prediction search of the coding engine (632). That is, for a new image to be encoded, the predictor (635) can search the reference image memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as motion vectors, block shapes, etc., of reference images that can serve as suitable prediction references for the new image. The predictor (635) can operate on a sample block-by-sample block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (635), the input image can have prediction references drawn from multiple reference images stored in the reference image memory (634).

[0062] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0063] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0064] The transmitter (640) may buffer the encoded video sequence produced by the entropy encoder (645) and prepare it for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0065] A controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a particular encoded image format to each encoded image, which can affect the encoding technique that can be applied to the respective image. For example, images are often assigned as one of the following image formats:

[0066] Note that an intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a predictor. Some video coding allows different forms of intra pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0067] A predicted image (P-image) may be an image that can be encoded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0068] A bidirectionally predicted image (B-image) may be one that can be encoded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predicted image may use more than two reference images and associated metadata for the reconstruction of a single block.

[0069] A source image may generally be spatially subdivided into a number of sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. The blocks may be predictively coded with reference to other (already coded) blocks as determined by a coding assignment applied to the respective image of the block. For example, the blocks of the I images may be non-predictively coded or predictively coded with reference to already coded blocks of the same image (spatial or intra prediction). The pixel blocks of the P images may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference image. The blocks of the B images may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference images.

[0070] The video encoder (603) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard being used.

[0071] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0072] A video may be captured as multiple source images (video images) in time sequence. Intra-image prediction (often abbreviated as intra-prediction) exploits spatial correlation in a given image, while inter-image prediction exploits correlation (temporal or other) between images. In one example, a particular image being encoded / decoded, called the current image, is divided into blocks. When a block in the current image is similar to a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference image and may have a third dimension that identifies the reference image if multiple reference images are used.

[0073] In some embodiments, bi-prediction techniques can be used for inter-image prediction. According to bi-prediction techniques, two reference images are used, such as a first reference image and a second reference image, both of which are prior to the decoding order of the current image in the video (but may be past and future in display order, respectively). A block in the current image can be coded by a first motion vector that points to a first reference block in the first reference image, and a second motion vector that points to a second reference block in the second reference image. A block can be predicted by a combination of the first reference block and the second reference block.

[0074] Furthermore, to improve coding efficiency, merge mode techniques can be used for inter-picture prediction.

[0075] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a sequence of video images are divided into coding tree units (CTUs) for compression, and the CTUs in an image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine the prediction format of the CU, such as an inter prediction format or an intra prediction format. The CU is partitioned into one or more prediction units (PUs) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using luma prediction block as an example of prediction block, prediction block includes a matrix of pixel values ​​(e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0076] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video image in a sequence of video images and to encode the processed block into an encoded image that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0077] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using intra mode, inter mode, or bi-predictive mode, for example using rate-distortion optimization. If the processing block is coded in intra mode, the video encoder (703) may use intra prediction techniques to code the processing block into a coded image. When the processing block is to be coded in inter mode or bi-predictive mode, the video encoder (703) may use inter prediction or bi-predictive techniques, respectively, to code the processing block into a coded image. In certain video coding techniques, the merge mode may be an inter image prediction submode in which motion vectors are derived from one or more motion vector predictors without benefit of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0078] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), all coupled together as shown in FIG.

[0079] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in reference images (e.g., blocks in previous and subsequent images), generate inter-prediction information (e.g., a description of redundant information through inter-coding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a prediction block) based on the inter-prediction information using any suitable technique. In some examples, the reference image is a decoded reference image that is decoded based on the coded video information.

[0080] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with blocks already encoded in the same image, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same image.

[0081] The general-purpose controller (721) is configured to determine general-purpose control data and control other components of the video encoder (703) based on the general-purpose control data. In one example, the general-purpose controller (721) determines a mode of the block and provides a control signal to the switch (726) based on the mode. For example, when the mode is an intra mode, the general-purpose controller (721) controls the switch (726) to select an intra mode result to be used by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information to be included in the bitstream. When the mode is an inter mode, the general-purpose controller (721) controls the switch (726) to select an inter prediction result to be used by the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information to be included in the bitstream.

[0082] The residual calculation unit (723) calculates the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate on the residual data to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be suitably used in the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and the intra-prediction information. In some examples, the decoded blocks may be appropriately processed to generate a decoded image, which may be buffered in a memory circuit (not shown) and used as a reference image.

[0083] The entropy encoder (725) is configured to format the bitstream to include the coding block. The entropy encoder (725) is configured to include various information in accordance with an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, according to the disclosed subject matter, when encoding a block in a merged sub-mode of either the inter-mode or the bi-prediction mode, the residual information is not present.

[0084] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive encoded images that are part of an encoded video sequence and decode the encoded images to generate reconstructed images. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0085] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872), which are coupled together as shown in FIG.

[0086] The entropy decoder (871) may be configured to reconstruct from the coded image certain symbols representing syntax elements of which the coded image is composed. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information, etc.) that may identify the mode in which the block is coded (e.g., intra-, inter-, or bi-prediction mode, the latter two being merged or separate submodes), certain samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively, residual information, e.g., in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880). If the prediction format is an intra-prediction format, the intra-prediction information is provided to the intra decoder (872). The residual information may undergo inverse quantization and is provided to the residual decoder (873).

[0087] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0088] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0089] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients and process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (the data path not shown as this may be only low volume control information).

[0090] The reconstruction module (874) is configured to combine, in the spatial domain, the residual as output by the residual decoder (873) and the prediction result (possibly as output by an inter- or intra-prediction module) to form a reconstructed block that may be part of the reconstructed image, which may be part of the reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0091] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0092] To reduce artifacts, an adaptive loop filter (ALF) with block-based filter adaptation can be applied by the encoder / decoder. For the luma component, for example, one of multiple filters (e.g., 25 filters) can be selected for a 4x4 luma block based on local gradient direction and activity.

[0093] The ALF may have any suitable shape and size. Referring to FIG. 9, the ALFs (910)-(911) have diamond shapes, such as a 5×5 diamond shape for the ALF (910) and a 7×7 diamond shape for the ALF (911). In the ALF (910), elements (920)-(932) may be used in a filter process to form the diamond shape. Seven values ​​(e.g., C0-C6) may be used for the elements (920)-(932). In the ALF (911), elements (940)-(964) may be used in a filter process to form the diamond shape. Thirteen values ​​(e.g., C0-C12) may be used for the elements (940)-(964).

[0094] Referring to FIG. 9, in some examples, two ALFs (910)-(911) with diamond filter shapes are used. A 5×5 diamond shaped filter (910) can be applied to the chroma components (e.g., chroma block, chroma CB) and a 7×7 diamond shaped filter (911) can be applied to the luma components (e.g., luma block, luma CB). Other suitable shapes and sizes can be used for the ALFs. For example, a 9×9 diamond shaped filter can be used.

[0095] The filter coefficients at the locations indicated by the values ​​(e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, if the ALF includes a clipping function, the clipping values ​​at those locations may be non-zero.

[0096] For block classification of the luminance component, a 4×4 block (or luminance block, luminance CB) can be classified or classified into one of a number of (e.g., 25) classes. The classification index C is calculated using Equation (1) by quantizing the directional parameter D and the activity value A.

number

number

number

number

[0097] To reduce the complexity of the block classification described above, a subsampled 1-D Laplacian calculation can be applied. Figures 10A-10D show the gradient g in the vertical direction (Figure 10A), horizontal direction (Figure 10B), and two diagonal directions d1 (Figure 10C) and d2 (Figure 10D). v , g h , g d1 , and g d2 10A shows an example of the subsampled positions used to compute the vertical gradient g vIn FIG. 10B, the label “H” indicates the subsampled positions for computing the horizontal gradient g h In FIG. 10C, the label “D1” indicates the subsampled positions for computing the d1 diagonal gradient g d1 In FIG. 10D, the label “D2” indicates the subsampled positions for computing the d2 diagonal gradient g d2 indicates the subsampled positions for computing

[0098] horizontal g v and the vertical direction g h The maximum value of the gradient of

number

number

number

[0099] Two diagonal directions g d1 and g d2 The maximum value of the gradient of

number

number

number

[0100] The directional parameter D can be derived based on the above values ​​and two thresholds t1 and t2 as follows: Step 1.(1)

number

number

number

number

number

[0101] The activity value A can be calculated as follows:

number

number

[0102] For chroma components in an image, no block classification is applied and therefore a single set of ALF coefficients can be applied for each chroma component.

[0103] A geometric transformation can be applied to the filter coefficients and the corresponding filter clip values ​​(also called clip values). Before filtering a block (e.g., a 4×4 luminance block), for example, the gradient values ​​(e.g., g v , g h , g d1and / or g d2 ), a geometric transformation such as a rotation or a diagonal and vertical flip can be applied to the filter coefficients f(k,l) and the corresponding filter clipping values ​​c(k,l). The geometric transformation applied to the filter coefficients f(k,l) and the corresponding filter clipping values ​​c(k,l) can be equivalent to applying a geometric transformation to the samples in the region supported by the filter. The geometric transformation can make the different blocks to which the ALF is applied more similar by aligning their respective directionality.

[0104] As described by equations (9)-(11), three geometric transformations can be performed, including a diagonal flip, a vertical flip, and a rotation. f D (k, l) = f(lk), c D (k,l)=c(l,k)Equation (9) f V (k,l)=f(k,Kl-1), c V (k,l)=c(k,Kl-1) (10) f R (k,l)=f(Kl-1,k), c R (k,l)=c(Kl-1,k) (11) where K is the size of the ALF or filter, and 0≦k, 1≦K-1 are the coordinates of the coefficients. For example, a filter f or clip value matrix (or clip matrix) c has position (0,0) in its upper left corner and position (K-1,K-1) in its lower right corner. Transformations can be applied to the filter coefficients f(k,l) and clip values ​​c(k,l) depending on the gradient values ​​calculated for the block. An example of the relationship between the transformations and the four gradients is summarized in Table 1.

[0105] [Table 1]

[0106] In some embodiments, the ALF filter parameters are signaled in an adaptive parameter set (APS) of the image. In the APS, one or more sets (e.g., up to 25 sets) of luma filter coefficients and clip value indexes can be signaled. In one example, the set of one or more sets can include luma filter coefficients and one or more clip value indexes. One or more sets (e.g., up to 8 sets) of chroma filter coefficients and clip value indexes can be signaled. To reduce signaling overhead, filter coefficients of different classifications (e.g., having different classification indices) of luma components can be merged. In the slice header, the index of the APS used for the current slice can be signaled.

[0107] In one embodiment, a clipping value index (also called a clipping index) can be decoded from the APS. The clipping value index can be used to determine the corresponding clipping value based on, for example, a relationship between the clipping value index and the corresponding clipping value. The relationship can be predefined and stored in the decoder. In one example, the relationship is described by a table, such as a luma table of clip value index and corresponding clip value (e.g., used for luma CB), a chroma table of clip value index and corresponding clip value (e.g., used for chroma CB). The clip value can depend on the bit depth B. The bit depth B can refer to the internal bit depth, the bit depth of the reconstructed sample in the CB to be filtered, etc. In some examples, the tables (e.g., luma table, chroma table) are obtained using Equation (12).

number

[0108] [Table 2]

[0109] In the slice header of the current slice, one or more APS indices (e.g., up to 7 APS indices) may be signaled to specify the luma filter sets that may be used for the current slice. The filtering may be controlled at one or more appropriate levels, such as at the picture level, slice level, or CTB level. In one embodiment, the filtering may be further controlled at the CTB level. A flag may be signaled to indicate whether the ALF is applied to the luma CTB. The luma CTB may select a filter set from among a plurality of fixed filter sets (e.g., 16 fixed filter sets) and a filter set (also referred to as a signaled filter set) that are signaled in the APS. A filter set index may be signaled to the luma CTB to indicate the filter set to be applied (e.g., a filter set among the plurality of fixed filter sets and the signaled filter set). The plurality of fixed filter sets may be predefined and hard coded in the encoder and decoder and may be referred to as a predefined filter set.

[0110] For the chroma component, an APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if an APS has more than one chroma filter set, a filter set index can be signaled for each chroma CTB.

[0111] The filter coefficients may be quantized with a norm equal to 128. To reduce the complexity of multiplications, bitstream conformance may be applied such that the coefficient values ​​of non-center locations are within the range of −27 to 27−1. In one example, center location coefficients are not signaled in the bitstream and may be considered equal to 128.

[0112] In some embodiments, the syntax and semantics of the clipping index and clip values ​​are defined as follows: alf_luma_clip_idx[sfIdx][j] may be used to specify a clipping index of the clip value to use before multiplying the j-th coefficient of the luma filter signaled by sfIdx. Bitstream conformance requirements may include that the value of alf_luma_clip_idx[sfIdx][j] be in the range of 0 to 3, for sfIdx=0 to alf_luma_num_filters_signalled_minus1, and j=0 to 11. The luma filter clip value AlfClipL[adaptation_parameter_set_id] with elements AlfClipL[adaptation_parameter_set_id][filtIdx][j], where filtIdx = 0 to NumAlfFilters-1 and j = 0 to 11, can be derived as specified in Table 2 depending on the bitDepth set equal to BitDepthY and clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j]. alf_chroma_clip_idx[altIdx][j] may be used to specify a clipping index of the clip value to use before multiplying the jth coefficient of the alternative chroma filter with index altIdx. Bitstream conformance requirements may include that the value of alf_chroma_clip_idx[altIdx][j] be in the range 0 to 3, where altIdx=0 to alf_chroma_num_alt_filters_minus1, j=0 to 5. Chroma filter clip values ​​AlfClipC[adaptation_parameter_set_id][altIdx][j] with elements alfClipC[adaptation_parameter_set_id][altIdx][j], altIdx=0 to alf_chroma_num_alt_filters_minus1, j=0 to 5 may be derived as specified in Table 2 depending on bitDepth set equal to BitDepthC and clipIdx set equal to alf_chroma_clip_idx[altIdx][j].

[0113] In one embodiment, the filtering process can be described as follows: On the decoder side, when ALF is enabled for a CTB, samples R(i,j) in a CU (or CB) can be filtered, and filtered sample values ​​R'(i,j) are obtained as shown below using equation (13). In one example, each sample in a CU is filtered.

number

[0114] In a non-linear ALF, multiple sets of clip values ​​can be provided in Table 3. In one example, the luma set includes four clip values ​​{1024, 181, 32, 6} and the chroma set includes four clip values ​​{1024, 161, 25, 4}. The four clip values ​​in the luma set can be selected by approximately equally dividing the full range (e.g., 1024) of the luma block sample values ​​(encoded in 10 bits) in the logarithmic domain. The range can be from 4 to 1024 for the chroma set.

[0115] [Table 3]

[0116] The selected clip value may be encoded in the "alf_data" syntax element as follows: A suitable encoding scheme (e.g., Golomb encoding scheme) may be used to encode the clipping index corresponding to the selected clip value as shown in Table 3. The encoding scheme may be the same encoding scheme used to encode the filter set index.

[0117] In one embodiment, virtual boundary filtering can be used to reduce the line buffer requirements of the ALF. Thus, modified block classification and filtering can be used for samples near CTU boundaries (e.g., horizontal CTU boundaries). The virtual boundary (1130) is the horizontal CTU boundary (1120) as shown in FIG. 11A. samples " can be defined as a line by shifting the sample by N samples can be a positive integer. In one example, N samples is equal to 4 for the luminance component, and N samples is equal to 2 for the chroma component.

[0118] With reference to Figure 11A, modified block classification can be applied to the luma component. In one example, the 1D Laplacian gradient calculation for a 4x4 block (1110) above a virtual boundary (1130) uses only samples above the virtual boundary (1130). Similarly, with reference to Figure 11B, the 1D Laplacian gradient calculation for a 4x4 block (1111) below a virtual boundary (1131) shifted from the CTU boundary (1121) uses only samples below the virtual boundary (1131). Thus, the quantization of the activity value A can be scaled by taking into account the reduction in the number of samples used in the 1D Laplacian gradient calculation.

[0119] For filtering, a symmetric padding operation at the virtual boundary can be used for both luma and chroma components. Figures 12A-12F show an example of such modified ALF filtering for luma components at the virtual boundary. If the sample being filtered is located below the virtual boundary, the adjacent samples located above the virtual boundary can be padded. If the sample being filtered is located above the virtual boundary, the adjacent samples located below the virtual boundary can be padded. With reference to Figure 12A, the adjacent sample C0 can be padded with the sample C2 located below the virtual boundary (1210). With reference to Figure 12B, the adjacent sample C0 can be padded with the sample C2 located above the virtual boundary (1220). With reference to Figure 12C, the adjacent samples C1-C3 can be padded with the samples C5-C7 located below the virtual boundary (1230), respectively. With reference to Figure 12D, the adjacent samples C1-C3 can be padded with the samples C5-C7 located above the virtual boundary (1240), respectively. Referring to Figure 12E, adjacent samples C4-C8 can be padded with samples C10, C11, C12, C11, and C10, respectively, located below the virtual boundary (1250). Referring to Figure 12F, adjacent samples C4-C8 can be padded with samples C10, C11, C12, C11, and C10, respectively, located above the virtual boundary (1260).

[0120] In some examples, the above description may be appropriately adapted when a sample and an adjacent sample are located to the left (or right) and right (or left) of a virtual boundary.

[0121] The cross-component filtering process may apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter may use a luma sample value of a luma component (e.g., luma CB) to refine a chroma component (e.g., chroma CB corresponding to luma CB). In one example, the luma CB and chroma CB are included in a CU.

[0122] FIG. 13 illustrates a cross-component filter (e.g., CC-ALF) used to generate a chroma component, according to one embodiment of the disclosure. In some examples, FIG. 13 illustrates filtering of a first chroma component (e.g., first chroma CB), a second chroma component (e.g., second chroma CB), and a luma component (e.g., luma CB). The luma component may be filtered by a sample adaptive offset (SAO) filter (1310) to generate an SAO filtered luma component (1341). The SAO filtered luma component (1341) may be further filtered by an ALF luma filter (1316) to become a filtered luma CB (1361) (e.g., “Y”).

[0123] The first chroma component may be filtered by an SAO filter (1312) and an ALF chroma filter (1318) to generate a first intermediate component (1352). Furthermore, the SAO filtered luma component (1341) may be filtered by a cross component filter (e.g., CC-ALF) for the first chroma component (1321) to generate a second intermediate component (1342). Subsequently, a filtered first chroma component (1362) (e.g., “Cb”) may be generated based on at least one of the second intermediate component (1342) and the first intermediate component (1352). In one example, the filtered first chroma component (1362) (e.g., “Cb”) may be generated by combining the second intermediate component (1342) and the first intermediate component (1352) with an adder (1322). The cross-component adaptive loop filtering for the first chroma component may include steps performed by a CC-ALF (1321) and steps performed by, for example, an adder (1322).

[0124] The above description can be adapted to the second chroma component. The second chroma component can be filtered by the SAO filter (1314) and the ALF chroma filter (1318) to generate a third intermediate component (1353). Furthermore, the SAO filtered luma component (1341) can be filtered by a cross component filter (e.g., CC-ALF) (1331) for the second chroma component to generate a fourth intermediate component (1343). Subsequently, a filtered second chroma component (1363) (e.g., “Cr”) can be generated based on at least one of the fourth intermediate component (1343) and the third intermediate component (1353). In one example, the filtered second chroma component (1363) (e.g., “Cr”) can be generated by combining the fourth intermediate component (1343) and the third intermediate component (1353) with an adder (1332). In one example, the cross-component adaptive loop filtering for the second chroma component may include steps performed by a CC-ALF (1331) and steps performed, for example, by an adder (1332).

[0125] The cross-component filters (e.g., CC-ALF (1321), CC-ALF (1331)) may operate by applying a linear filter having any suitable filter shape to the luma component (or luma channel) to refine each chroma component (e.g., first chroma component, second chroma component).

[0126] FIG. 14 illustrates an example of a filter (1400) according to an embodiment of the present disclosure. The filter (1400) may include non-zero filter coefficients and zero filter coefficients. The filter (1400) has a diamond shape (1420) formed by the filter coefficients (1410) (shown as solid circles). In one example, the non-zero filter coefficients in the filter (1400) are included in the filter coefficients (1410), and the filter coefficients that are not included in the filter coefficients (1410) are zero. Thus, the non-zero filter coefficients of the filter (1400) are included in the diamond shape (1420), and the filter coefficients that are not included in the diamond shape (1420) are zero. In one example, the number of filter coefficients of the filter (1400) is equal to the number of filter coefficients (1410), which is 18 in the example shown in FIG. 14.

[0127] The CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Referring back to Figure 13, the CC-ALF (1321) and the CC-ALF (1331) may have the same filter shape, such as the diamond shape (1420) shown in Figure 14, and the same number of filter coefficients. In one example, the values ​​of the filter coefficients in the CC-ALF (1321) are different from the values ​​of the filter coefficients in the CC-ALF (1331).

[0128] In general, the filter coefficients (e.g., non-zero filter coefficients) of the CC-ALF can be transmitted, for example, in the APS. In one example, the filter coefficients are coefficients (e.g., 2 10) and may be rounded for a fixed point representation. Application of CC-ALF may be controlled by variable block sizes and signaled by a context coded flag (e.g., a CC-ALF enable flag) received for each block of samples. Context coded flags such as the CC-ALF enable flag may be signaled at any appropriate level, such as the block level. Block sizes along with the CC-ALF enable flags may be received at the slice level for each chroma component. In some examples, block sizes (in chroma samples) 16×16, 32×32, and 64×64 may be supported.

[0129] In general, a luma block may correspond to a chroma block, such as two chroma blocks. The number of samples in each of the chroma block(s) may be less than the number of samples in the luma block. A chroma subsampling format (also referred to as a chroma subsampling format, for example, specified by chroma_format_idc) may indicate a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) between each of the chroma block(s) and the corresponding luma block. In one example, the chroma subsampling format is 4:2:0, and thus the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown in FIG. 15A-FIG. 15B. In one example, the chroma subsampling format is 4:2:2, and thus the chroma horizontal subsampling factor (e.g., SubWidthC) is 2, and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In one example, the chroma subsampling format is 4:4:4, and therefore the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample format (also called the chroma sample position) may indicate a relative position of a chroma sample in a chroma block relative to at least one corresponding luma sample in a luma block.

[0130] 15A-15B illustrate example locations of chroma samples relative to luma samples, according to an embodiment of the present disclosure. With reference to FIG. 15A, luma samples (1501) are arranged in rows (1511)-(1518). The luma samples (1501) shown in FIG. 15A may represent a portion of an image. In one example, a luma block (e.g., Luma CB) includes a luma sample (1501). The luma block may correspond to two chroma blocks having a chroma subsampling format of 4:2:0. In one example, each chroma block includes a chroma sample (1503). Each chroma sample (e.g., chroma sample (1503(1))) corresponds to four luma samples (e.g., luma samples (1501(1)) through (1501(4)). In one example, the four luma samples are the top left sample (1501(1)), the top right sample (1501(2)), the bottom left sample (1501(3)), and the bottom right sample (1501(4)). A chroma sample (e.g., (1503(1))) corresponds to the top left sample (1501(1)) and the bottom left sample (1501(4)). The chroma sample format of a chroma block having chroma sample (1503) located at a left-center position between the top left sample (1501(1)) and the bottom left sample (1501(3)) may be referred to as chroma sample format 0. Chroma sample format 0 indicates relative position 0, which corresponds to a left-center position halfway between the top left sample (1501(1)) and the bottom left sample (1501(3)). The four luma samples (e.g., (1501(1)) through (1501(4))) may be referred to as neighboring luma samples of chroma sample (1503)(1).

[0131] In one example, each chroma block includes a chroma sample (1504). The above description of the chroma sample (1503) may be adapted to the chroma sample (1504), and thus detailed description may be omitted for brevity. Each of the chroma samples (1504) may be located at a central position of four corresponding luma samples, and a chroma sample format of a chroma block having the chroma sample (1504) may be referred to as chroma sample format 1. The chroma sample format 1 indicates a relative position 1 corresponding to the central position of the four luma samples (e.g., (1501(1)) to (1501(4))). For example, one of the chroma samples (1504) may be located at the central portion of the luma samples (1501(1)) to (1501(4)).

[0132] In one example, each chroma block includes a chroma sample (1505). Each of the chroma samples (1505) may be located at a top left position that is the same as the top left sample of the four corresponding luma samples (1501), and the chroma sample format of the chroma block having the chroma samples (1505) may be referred to as a chroma sample format 2. Thus, each of the chroma samples (1505) is located at the same position as the top left sample of the four luma samples (1501) that correspond to the respective chroma sample. The chroma sample format 2 indicates a relative position 2 that corresponds to the top left position of the four luma samples (1501). For example, one of the chroma samples (1505) may be located at the top left position of the luma samples (1501(1))-(1501(4)).

[0133] In one example, each chroma block includes chroma samples (1506). Each of the chroma samples (1506) may be located at an upper center position between a corresponding upper-left sample and a corresponding upper-right sample, and a chroma sample format of a chroma block having the chroma samples (1506) may be referred to as chroma sample format 3. Chroma sample format 3 indicates a relative position 3, which corresponds to a top center position between the upper-left sample (and the upper-right sample). For example, one of the chroma samples (1506) may be located at an upper center position of the luma samples (1501(1))-(1501(4)).

[0134] In one example, each chroma block includes a chroma sample (1507). Each of the chroma samples (1507) may be located at a bottom left position that is co-located with a bottom left sample of the four corresponding luma samples (1501), and the chroma sample format of the chroma block having the chroma sample (1507) may be referred to as chroma sample format 4. Thus, each of the chroma samples (1507) is located at the bottom left sample of the four luma samples (1501) that correspond to the respective chroma sample. The chroma sample format 4 indicates a relative position 4, which corresponds to the bottom left position of the four luma samples (1501). For example, one of the chroma samples (1507) may be located at the bottom left position of the luma samples (1501(1))-(1501(4)).

[0135] In one example, each chroma block includes a chroma sample (1508). Each chroma sample (1508) is located at a bottom center position between a bottom left sample and a bottom right sample, and a chroma sample format of a chroma block having the chroma sample (1508) can be referred to as a chroma sample format 5. The chroma sample format 5 indicates a relative position 5, which corresponds to a bottom center position between a bottom left sample and a bottom right sample of the four luma samples (1501). For example, one of the chroma samples (1508) can be located between a bottom left sample and a bottom right sample of the luma samples (1501(1))-(1501(4)).

[0136] In general, any suitable chroma sample format may be used for the chroma subsampling format. Chroma sample formats 0-5 are examples of chroma sample formats described in chroma subsampling format 4:2:0. Additional chroma sample formats may be used for chroma subsampling format 4:2:0. Additionally, other chroma sample formats and / or variations of chroma sample formats 0-5 may be used for other chroma subsampling formats, such as 4:2:2, 4:4:4, etc. In one example, a chroma sample format that combines chroma samples (1505) and (1507) is used for chroma subsampling format 4:2:2.

[0137] In one example, a luma block may be considered to have alternating rows such as rows (1511)-(1512) each including the top two samples (e.g., (1501(1))-(1501(2))) of the four luma samples (e.g., (1501(1))-(1501(4))) and the bottom two samples (e.g., (1501(3))-(1501(4))) of the four luma samples (e.g., (1501(1))-(1501(4))). Thus, rows (1511), (1513), (1515), and (1517) may be considered to have alternating rows such as rows (1518), (1519), (1520), (1521), (1522), (1523), (1524), (1525), (1526), ​​(1527), (1528), (1529), (1530), (1531), (1532), (1533), (1534), (1535), (1536), (1537), (1538), (1539), (1540), (1541), (1542), (1543), (1544), (1545), (1546), (1547), (1548), (1549), (1550), (1551), (1552), (1553), (1554), (1555), (1556), (1557), (1558), (1559), (1560), (1561), (1562), (1563), (1564), (1565), (1566), (1567), (1568), (1569 ) can be referred to as the current row (also referred to as the top field), and rows (1512), (1514), (1516), and (1518) can be referred to as the next rows (also referred to as the bottom field). The four luminance samples (e.g., (1501(1)) through (1501(4))) are located in the current row (e.g., (1511)) and the next row (e.g., (1512)). Relative positions 2 through 3 are located in the current row, relative positions 0 through 1 are located between each current row and the respective next row, and relative positions 4 through 5 are located in the next row.

[0138] The chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) are arranged in rows (1551)-(1554) within each chroma block. The specific location of the rows (1551)-(1554) may depend on the chroma sample format of the chroma samples. For example, for chroma samples (1503)-(1504) having respective chroma sample formats 0-1, row (1551) is located between rows (1511)-(1512). For chroma samples (1505)-(1506) having respective chroma sample formats 2-3, row (1551) is in the same position as the current row (1511). For chroma samples (1507)-(1508) having respective chroma sample types 4-5, row (1551) is co-located with the next row (1512). The above description can be adapted appropriately for rows (1552)-(1554), and a detailed description will be omitted for the sake of brevity.

[0139] Any suitable scanning method may be used to display, store, and / or transmit the luma blocks and corresponding chroma blocks described above in Figure 15A. In one example, progressive scanning is used.

[0140] Interlace scanning can be used, as shown in Figure 15B. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc equals 1). In one example, the variable chroma location type (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current rows (1511), (1513), (1515), and (1517) and the next rows (1512), (1514), (1516), and (1518) may be scanned separately, for example, the current rows (1511), (1513), (1515), and (1517) may be scanned first, followed by the next rows (1512), (1514), (1516), and (1518). The current row may include luminance samples (1501) and the next row may include luminance samples (1502).

[0141] Similarly, corresponding chroma blocks may be interlaced. Rows (1551) and (1553) containing chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) with no fill may be referred to as the current row (or current chroma row), and rows (1552) and (1554) containing chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) with gray fill may be referred to as the next row (or next chroma row). In one example, during interlacing, rows (1551) and (1553) are scanned first, followed by rows (1552) and (1554).

[0142] The diamond filter shape (1420) of FIG. 14 is designed for a 4:2:0 chroma subsampling format and chroma sample type 0 (e.g., the chroma row is between two luma rows), which may not be efficient with other chroma sample types (e.g., chroma sample types 1 through 5) and other chroma subsampling formats (e.g., 4:2:2 and 4:4:4).

[0143] The coded information of the chroma block or chroma CB (e.g., the first chroma CB or the second chroma CB in FIG. 13) may be decoded from the coded video bit stream. The coded information may indicate that a cross-component filter is applied to the chroma CB. The coded information may further include a chroma subsampling format and a chroma sample format. As described above, the chroma subsampling format may indicate a chroma horizontal subsampling factor and a chroma vertical subsampling factor between the chroma CB and a corresponding luma CB (e.g., the luma CB in FIG. 13). The chroma sample format may indicate a relative position of the chroma sample with respect to at least one corresponding luma sample in the luma CB. In one example, the chroma sample format is signaled in the coded video bit stream. The chroma sample format may be signaled at any appropriate level, such as a sequence parameter set (SPS).

[0144] According to an aspect of the present disclosure, a filter shape of a cross-component filter (e.g., CC-ALF (1321)) in a cross-component filtering process can be determined based on at least one of a chroma subsampling format and a chroma sample format. Furthermore, a first intermediate CB (e.g., intermediate component (1342)) can be generated by applying a cross-component filter having the determined filter shape to a corresponding luma CB (e.g., SAO filtered luma component (1341)). A second intermediate CB (e.g., intermediate component (1352)) can be generated by applying a loop filter (e.g., ALF (1318)) to a chroma CB (e.g., SAO filtered first chroma CB). A filtered chroma CB (e.g., filtered first chroma component (1362) (e.g., “Cb”) of FIG. 13) can be determined based on the first intermediate CB and the second intermediate CB. As described above, the cross-component filter can be CC-ALF, and the loop filter can be ALF.

[0145] The chroma sample format of the chroma block can be indicated in the encoded bitstream when CC-ALF is used. The filter shape of the CC-ALF can depend on the chroma subsampling format (e.g., chroma_format_idc) of the chroma block, the chroma sample format, etc.

[0146] 16 illustrates example cross-component filters (e.g., CC-ALF) (1601)-(1603) with respective filter shapes (1621)-(1623) according to an embodiment of the present disclosure. With reference to FIGS. 14 and 16, filter shapes (1420) and (1621)-(1623) may be used for CC-ALF based on the chroma sample format of the chroma block, for example, when the chroma subsampling format is 4:2:0.

[0147] According to aspects of the present disclosure, the chroma sample type may be one of six chroma sample types 0-5, which indicate six relative positions 0-5, respectively. The six relative positions 0-5 may correspond to the left center position, center position, top left position, top center position, bottom left position, and bottom center position, respectively, of the four luma samples (e.g., (1501(1))-(1501(4))) as shown in FIG. 15A. The filter shape of the cross-component filter may be determined based on the chroma sample type.

[0148] When the chroma sample type of the chroma block is chroma sample type 0, a filter (1400) having the filter shape (1420) of Figure 14 can be used in a cross-component filter (eg, CC-ALF).

[0149] In one example, a square filter shape (e.g., a 4×4 square filter shape, a 2×2 square filter shape) can be used for the cross-component filter (e.g., CC-ALF) when the chroma sample format of the chroma block is chroma sample format 1. The chroma sample being cross-component filtered (e.g., (1504(1)) in FIG. 15A) can be located at the center of the square filter shape because the chroma sample (e.g., (1504(1))) is located at the center of the four corresponding luma samples (e.g., (1501(1)) through (1501(4))).

[0150] In one example, when the chroma sample format of a chroma block is chroma sample format 2, a diamond filter shape can be used in CC-ALF (e.g., a 5x5 diamond filter shape (1621) of filter (1601) or a 3x3 diamond filter shape (1622) of filter (1602)).

[0151] In one example, a diamond filter shape (1623) of the filter (1603) may be used in CC-ALF when the chroma sample format of the chroma block is chroma sample format 3. With reference to Figures 14 and 16, the diamond filter shape (1623) is a geometric transformation (e.g., a 90° rotation) of the filter shape (1420).

[0152] In one example, when the chroma sample format of a chroma block is chroma sample format 4, the filter shape used in CC-ALF may be the same as or similar to that used for chroma sample format 2 (e.g., diamond filter shape (1621) or (1622)). Thus, the filter shape for chroma sample format 4 may be a diamond filter shape, such as a vertically shifted diamond filter shape (1621) or (1622).

[0153] In one example, when the chroma sample format of a chroma block is chroma sample format 5, the filter shape used in CC-ALF may be the same as or similar to that used for chroma sample format 3 (e.g., diamond filter shape (1623)). Thus, the filter shape for chroma sample format 5 may be a diamond filter shape, such as a vertically shifted diamond filter shape (1623).

[0154] In one embodiment, the number of filter coefficients is signaled in the coded video bit stream, such as an APS. With reference to FIG. 14 and FIG. 16, the filter coefficients (1611) may form a diamond shape (1621), with other filter coefficients not included in or otherwise excluded from the filter coefficients (1611) being zero. Thus, the number of filter coefficients of the filter (1601) may refer to the number of filter coefficients of the filter shape (1621) and therefore may be equal to the number of filter coefficients (1611) (e.g., 13). Similarly, the number of filter coefficients of the filter (1602) may refer to the number of filter coefficients of the filter shape (1622) and therefore may be equal to the number of filter coefficients (1612) (e.g., 5). The number of filter coefficients of the filter (1603) may refer to the number of filter coefficients of the filter shape (1623) and therefore may be equal to the number of filter coefficients (1613) (e.g., 18). Similarly, the number of filter coefficients of the filter (1400) may be eighteen.

[0155] Different filter shapes may have different numbers of coefficients. Thus, in some examples, the filter shape may be determined based on the number of filter coefficients. For example, if the number of filter coefficients is 16, the filter shape may be determined to be a 4×4 square filter shape.

[0156] According to aspects of the present disclosure, a number of filter coefficients of the cross-component filter can be signaled in the encoded video bitstream, and a filter shape of the cross-component filter can be determined based on the number of filter coefficients and at least one of a chroma subsampling format and a chroma sample type.

[0157] In one embodiment, a cross-component linear model (CCLM) flag can indicate a chroma sample format. Thus, the filter shape of the CC-ALF can depend on the CCLM flag. In one example, the chroma sample format indicated by the CCLM flag is chroma sample format 0 or 2, and thus the filter shape can be a filter shape (1420) or a diamond filter shape (e.g., a 5×5 diamond filter shape (1621) or a 3×3 diamond filter shape (1622)).

[0158] In one embodiment, a CCLM flag (e.g., sps_cclm_colocated_chroma_flag) is signaled in a coded video bitstream, e.g., SPS. In one example, the CCLM flag (e.g., sps_cclm_colocated_chroma_flag) indicates whether the top-left downsampled luma sample in CCLM intra prediction is co-located with the top-left luma sample. As mentioned above, the chroma sample format indicated by the CCLM flag can be chroma sample format 0 or 2. The filter shape of the CC-ALF may depend on the sps_cclm_colocated_chroma_flag or similar information (e.g., information indicating whether the top-left downsampled luma sample in CCLM intra prediction is co-located with the top-left luma sample) of CCLM signaled in SPS.

[0159] In some examples, the cross-component filter (e.g., CC-ALF) includes a large number (e.g., 18) of multiplications per chroma sample (e.g., Cb chroma sample or Cr chroma sample) and therefore has a high cost, for example, in computational complexity. The number of multiplications is based on the number of filter coefficients in the CC-ALF (e.g., 18 filter coefficients of filters (1400) and (1603), 13 filter coefficients of filter (1601), and 5 filter coefficients of filter (1602)). For example, the number of multiplications is equal to the number of filter coefficients in the CC-ALF. According to aspects of the present disclosure, the number of bits representing the CC-ALF filter coefficients of the CC-ALF can be constrained to be equal to or less than K bits. K can be a positive integer, such as 8. Thus, the CC-ALF filter coefficients are expressed as [-2 K-1 ~2 K-1 −1]. The range of CC-ALF filter coefficients in CC-ALF can be constrained to be K bits or less so that simpler multipliers (e.g., with fewer bits) for CC-ALF can be used.

[0160] In one embodiment, the range of the CC-ALF filter coefficients is −24 From 2 4 Alternatively, the CC-ALF filter coefficients are constrained between -2 5 From 2 5 -1, where K is 6 bits.

[0161] In one example, several different values ​​of the CC-ALF filter coefficients are constrained to a particular number, such as K bits. A lookup table may be used when applying the CC-ALF.

[0162] The CC-ALF filter coefficients of the CC-ALF may be encoded and signaled using fixed-length coding. For example, if the CC-ALF filter coefficients are constrained to K bits, a K-bit fixed-length coding may be used to signal the CC-ALF filter coefficients. If K is relatively small, such as 8 bits, the fixed-length coding may be more efficient than other methods, such as variable-length coding.

[0163] Referring back to FIG. 13, according to an embodiment of the present disclosure, if the dynamic range (or luma bit depth) of the luma sample value (e.g., (1341)) is greater than L bits, the luma sample value of luma CB (e.g., (1341)) may be shifted to have a dynamic range of L bits. L may be a positive integer such as 8. Subsequently, an intermediate component (e.g., (1342) or (1343)) may be generated by applying CC-ALF (e.g., (1321) or (1331)) to the shifted luma sample value.

[0164] Referring back to FIG. 13, if the luma bit depth of the luma sample value (e.g., the luma sample value (1341) of the SAO filtered luma component) is higher than L bits, the cross-component adaptive loop filtering using CC-ALF (e.g., (1321) or (1331)) can be modified as described below. The luma sample value (e.g., (1341)) can first be shifted to a dynamic range of L bits. The shifted luma sample value can be unsigned L bits. In one example, L is 8. The shifted luma sample value can then be used as an input to CC-ALF (e.g., (1321) or (1331)). Thus, the shifted luma sample value and the CC-ALF filter coefficients can be multiplied. As described above, the CC-ALF filter coefficients can be K bits (e.g., 8 bits or [-2 K-1 ~2 K-1 -1]), and an unsigned L-bit by signed K-bit multiplier can be used. In one example, K and L are 8 bits, and a relatively simple and efficient unsigned 8-bit by signed 8-bit multiplier (e.g., a multiplier based on a single instruction, multiple data (SIMD) instruction) can be used to improve filtering efficiency.

[0165] Referring to FIG. 13, according to an embodiment of the present disclosure, the downsampled luma CB can be generated by applying a downsampling filter to the luma CB. Thus, the chroma horizontal subsampling factor and the chroma vertical subsampling factor between the first chroma CB (or the second chroma CB) and the downsampled luma CB are 1. The downsampling filter can be applied at any suitable step before the downsampled luma CB is used as an input to the CC-ALF (e.g., (1321)). In one example, the downsampling filter is applied between the SAO filter (1310) and the CC-ALF (e.g., (1321)), thus the SAO filtered luma component (1341) is first downsampled, and then the downsampled and SAO filtered luma component is sent to the CC-ALF (1321).

[0166] As described above, the filter shape of the CC-ALF can be determined based on the chroma subsampling format and / or the chroma sample format, and thus, in some examples, different filter shapes can be used for different chroma sample formats. Alternatively, because the downsampled luma samples are aligned with chroma samples whose chroma horizontal and vertical subsampling factors are one, the CC-ALF (e.g., (1321)) can use a unified filter shape when the input to the CC-ALF is a downsampled luma CB. The unified filter shape can be independent of the chroma subsampling format and the chroma sample format of the chroma CB. Thus, an intermediate CB (e.g., intermediate component (1342)) can be generated by applying a CC-ALF with a unified filter shape to a downsampled luma CB.

[0167] Referring to FIG. 13, in one example, for a chroma subsampling format that is a YUV (e.g., YCbCr or YCgCo) format, a downsampling filter may be applied to luma samples in the luma CB to derive downsampled luma samples whose positions are aligned with the chroma samples in the first chroma CB, and then a combined filtering shape may be applied to the CC-ALF (e.g., (1321)) to cross-filter the downsampled luma samples to generate intermediate components (1342).

[0168] The down-sampling filter may be any suitable filter. In one example, the down-sampling filter corresponds to the filter applied to the co-located luma samples in CCLM mode. Thus, the luma samples are down-sampled using the same down-sampling filter applied to the co-located luma samples in CCLM mode.

[0169] In one example, the downsampling filter is a {1,2,1;1,2,1} / 8 filter and the chroma subsampling format is 4:2:0. So for a chroma 4:2:0 format, the luma samples are downsampled by applying a {1,2,1;1,2,1} / 8 filter.

[0170] The filter shape (or unified filter shape) of the cross-component filter (e.g., CC-ALF) may have any suitable shape. In one example, the filter shape of the cross-component filter (e.g., CC-ALF) is one of a 7×7 diamond shape, a 7×7 square shape, a 5×5 diamond shape, a 5×5 square shape, a 3×3 diamond shape, and a 3×3 square shape.

[0171] FIG. 17 shows a flow chart outlining a process (1700) according to one embodiment of the present disclosure. The process (1700) may be used to reconstruct a block (e.g., CB) in an image of an encoded video sequence. The process (1700) may be used in the reconstruction of a block to generate a prediction block of the block being reconstructed. The term block may be interpreted as a prediction block, CB, CU, etc. In various embodiments, the process (1700) is performed by a processing circuit of the terminal devices (310), (320), (330), and (340), a processing circuit performing the function of the video encoder (403), a processing circuit performing the function of the video decoder (410), a processing circuit performing the function of the video decoder (510), a processing circuit performing the function of the video encoder (603), etc. In some embodiments, the process (1700) is implemented in software instructions, and thus the processing circuit performs the process (1700) when the processing circuit executes the software instructions. The process starts at (S1701) and proceeds to (S1710). In one example, the block is a chroma block, such as a chroma CB corresponding to a luma CB. In one example, the chroma block and the corresponding luma CB are in a CU.

[0172] At (S1710), coded information for the chroma CB may be decoded from the coded video bit stream. The coded information may indicate that a cross-component filter is applied to the chroma CB and may further indicate a chroma subsampling format and a chroma sample format. The chroma subsampling format may indicate a chroma horizontal subsampling factor and a chroma vertical subsampling factor between the chroma CB and a corresponding luma CB, as described above. The chroma subsampling format may be any suitable format, such as 4:2:0, 4:2:2, 4:4:4, etc. The chroma sample format may indicate a relative position of the chroma sample with respect to at least one corresponding luma sample in the luma CB, as described above. In one example, for a chroma subsampling format of 4:2:0, the chroma sample format may be one of chroma sample formats 0-5 described above with reference to FIG. 15A-FIG. 15B.

[0173] At (S1720), a filter shape of the cross-component filter may be determined based on at least one of the chroma subsampling format and the chroma sample type. In an example, referring to FIG. 13, the cross-component filter is used in cross-component filtering (e.g., CC-ALF filtering), and the cross-component filter may be CC-ALF. The filter shape may be any suitable shape depending on the chroma subsampling format and / or the chroma sample type. For a chroma subsampling format of 4:2:0, the filter shape may be one of the filter shapes (1420) and (1621)-(1623) based on the chroma sample type. The filter shape may be a change (e.g., a geometric transformation such as a rotation or shift) of one of the filter shapes (1420) and (1621)-(1623) based on the chroma sample type.

[0174] At (S1730), a first intermediate CB may be generated by applying a loop filter (eg, ALF) to the chroma CB (eg, the SAO filtered chroma CB).

[0175] In (S1740), for example, when the chroma subsampling format is 4:2:0 and the chroma sample format is chroma sample 0, a second intermediate CB can be generated by applying a cross-component filter (e.g., CC-ALF) having a determined filter shape (e.g., filter shape (1420)) to the corresponding luma CB.

[0176] At (S1750), a filtered chroma CB (e.g., the filtered first chroma component (1362)) may be determined based on the first intermediate CB (e.g., the intermediate component (1342)) and the second intermediate CB (e.g., the intermediate component (1352)). The process (1700) proceeds to (S1799) and ends.

[0177] The process (1700) may be adapted as appropriate. Steps of the process (1700) may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0178] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0179] The techniques described above can be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 18 illustrates a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.

[0180] Computer software can be encoded using any suitable machine code or computer language that can be assembled, compiled, linked, etc. to create code containing instructions that can be executed directly, or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0181] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0182] 18 for the computer system (1800) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (1800).

[0183] The computer system (1800) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), video (2D video, 3D video including stereoscopic video, etc.).

[0184] The input human interface devices may include one or more (only one of each) of a keyboard (1801), a mouse (1802), a trackpad (1803), a touch screen (1810), a data glove (not shown), a joystick (1805), a microphone (1806), a scanner (1807), a camera (1808).

[0185] The computer system (1800) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touch screen (1810), data gloves (not shown), or joystick (1805), although there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (1809), headphones (not shown)), visual output devices (e.g., screens (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without tactile feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional hypervisual output via means such as stereo output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0186] The computer system (1800) may also include human accessible storage devices and their associated media, such as optical media (1821) including CD / DVD ROM / RW (1820) having media such as CDs / DVDs, thumb drives (1822), removable hard drives or solid state drives (1823), legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0187] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0188] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The networks may be, for example, wireless, wired, optical. The networks may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1849) (e.g., a USB port of the computer system (1800)). Others are generally integrated into the core of the computer system (1800) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1800) can communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast TV), transmit only unidirectional (e.g., CANbus to a particular CANbus device), or bidirectional, for example, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used with each of these networks and network interfaces, as described above.

[0189] The aforementioned human interface devices, human access storage devices, and network interfaces may be attached to a core (1840) of the computer system (1800).

[0190] The cores (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1843), hardware accelerators for specific tasks (1844), graphics adapters (1850), and the like. These devices may be connected via a system bus (1848), along with read only memory (ROM) (1845), random access memory (1846), internal mass storage such as an internal non-user accessible hard drive, SSD, and the like (1847). In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, and the like. Peripherals may be attached directly to the core's system bus (1848) or via a peripheral bus (1849). In one example, a display (1810) may be connected to the graphics adapter (1850). Peripheral bus architectures include PCI, USB, etc.

[0191] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) can execute certain instructions that, in combination, may constitute the computer code described above. That computer code may be stored in ROM (1845) or RAM (1846). Transient data may also be stored in RAM (1846), while persistent data may be stored, for example, in internal mass storage (1847). Rapid storage and retrieval in any of the memory devices may be made possible by the use of cache memory, which may be closely associated with one or more of the CPU (1841), GPU (1842), mass storage device (1847), ROM (1845), RAM (1846), etc.

[0192] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0193] By way of example and not limitation, a computer system having the architecture (1800), and in particular the cores (1840), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, as well as media associated with specific storage of the cores (1840) of a non-transitory nature, such as the core internal mass storage (1847) or ROM (1845). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the cores (1840). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the cores (1840), and in particular the processors therein (including a CPU, GPU, FPGA, etc.) to perform certain operations or certain portions of certain operations described herein, including defining data structures stored in RAM (1846) and modifying such data structures according to operations defined by the software. Additionally, or alternatively, the computer system may provide functionality as a result of logic embodied in hardwired or otherwise circuitry (e.g., accelerator (1844)), which may operate in place of or in conjunction with software to perform certain operations or certain portions of certain operations described herein. References to software may encompass logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0194] Appendix A: Acronyms JEM joint exploration model VVC versatile video coding BMS benchmark set MV Motion Vector HEVC High Efficiency Video Coding MPM most probable mode WAIP Wide-Angle Intra Prediction SEI Supplementary Enhancement Information VUI Video Usability Information GOP Groups of Pictures TU Transform Units, PU Prediction Units Prediction Units CTU Coding Tree Units CTBs Coding Tree Blocks PB Prediction Blocks HRD Hypothetical Reference Decoder SDR standard dynamic range SNR Signal Noise Ratio CPU Central Processing Units GPU Graphics Processing Units CRT Cathode Ray Tube LCD Liquid-Crystal Display OLED Organic Light-Emitting Diode CD Compact Disc DVD Digital Video Disc Digital Video Disc ROM Read-Only Memory RAM Random Access Memory ASIC Application-Specific Integrated Circuit PLD Programmable Logic Device LAN Local Area Network GSM Global System for Mobile Communications LTE Long-Term Evolution CANBus Controller Area Network Bus Controller Area Network Bus USB Universal Serial Bus PCI Peripheral Component Interconnect FPGA Field Programmable Gate Areas SSD solid-state drive IC Integrated Circuit CU Coding Unit Encoder PDPC Position Dependent Prediction Combination ISP Intra Sub-Partitions SPS Sequence Parameter Setting

[0195] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]

[0196] 101 Samples 102 Arrow 103 Arrow 104 Square Block 180 Schematic diagram 201 Current Block 202 Ambient Samples 203 Ambient Samples 204 Ambient Samples 205 Ambient Samples 206 Ambient Samples 300 Communication Systems 310 Terminal Equipment 320 Terminal Equipment 330 Terminal Equipment 340 Terminal Equipment 350 Network 400 Communication Systems 401 Video Source 402 Video Image Stream 403 Video Encoder 404 Encoded Video Data 405 Streaming Server 406 Client Subsystem 407 Copy 408 Client Subsystem 409 Copy 410 Video Decoder 411 Video image output stream 412 Display 413 Capture Subsystem 420 Electronic equipment 430 Electronic equipment 501 Channel 510 Video Decoder 512 Rendering Device 515 Buffer Memory 520 Analyzer 521 Symbols 530 Electronic equipment 531 Receiver 551 Scaler / Descaler Unit 552 Intra Image Prediction Unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Image Memory 558 Image Buffer 601 Video Sources 603 Video Encoder 620 Electronic equipment 630 Source Encoder 632 encoding engine 633 Decoder 634 Reference Image Memory 635 Predictor 640 Transmitter 643 encoded video data 645 Entropy Encoder 650 Controller 660 Communication Channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoder 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 873 Residual Decoder 730 Intercoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 874 Reconstruction Module 880 Inter Decoder 910 ALF 911 ALF 920 Elements 921 Elements 923 Elements 924 elements 925 elements 926 elements 927 elements 928 elements 929 Elements 930 Elements 931 elements 932 elements 940 elements 941 elements 942 elements 943 elements 944 elements 945 elements 946 elements 947 elements 948 elements 949 elements 950 elements 951 elements 952 elements 953 elements 954 elements 956 elements 957 elements 958 elements 959 elements 960 elements 961 elements 962 elements 963 Elements 964 elements 1110 Block 1111 Block 1120 Horizontal CTU Boundary 1121 CTU boundary 1130 Virtual Boundary 1131 Virtual Boundary 1210 Virtual Boundary 1220 Virtual Boundary 1230 Virtual Boundary 1240 Virtual Boundary 1250 Virtual Boundary 1260 Virtual Boundary 1310 SAO Brightness 1312 SAO Cb 1314 SAO Cr 1316 ALF Brightness 1318 ALF Saturation 1321 CC-ALF Cb 1322 Adder 1332 Adder 1331 CC-ALF Cr 1341 SAO filtered luminance component 1342 Intermediate Component 1343 Intermediate Component 1352 Intermediate Component 1353 Intermediate Component 1361 Brightness CB 1362 First chroma component 1363 Second chroma component 1400 Filter 1410 Filter Coefficients 1420 Diamond filter shape 1501 Luminance Samples 1501(1) Luminance Samples 1501(2) Luminance Samples 1501(3) Luminance Samples 1501(4) Luminance Samples 1503 saturation samples 1503(1) Saturation Samples 1504 saturation samples 1504(1) Saturation Samples 1505 saturation samples 1506 saturation samples 1507 Saturation Samples 1508 saturation samples Line 1511 1512 lines Line 1513 1514 lines Line 1515 1516 lines 1517 lines 1518 lines 1551 lines 1552 lines 1553 lines 1554 lines 1601 Filter 1602 Filter 1603 Filter 1611 Filter Coefficients 1612 filter coefficients 1613 Filter Coefficients 1621 Diamond filter shape 1622 Diamond filter shape 1623 Diamond filter shape 1700 Processing 1800 Computer Systems 1801 Keyboard 1802 Mouse 1803 Trackpad 1805 Joystick 1806 Mike 1807 Scanner 1808 Camera 1809 Speaker 1810 Touch screen, display 1820 CD / DVD ROM / RW 1821 Optical media 1822 Thumb Drive 1823 Solid State Drive 1840 Core 1841 CPU 1842 GPU 1843 FPGA 1844 ACCL. 1845 ROM 1846 RAM 1847 Internal Mass Storage, Core Internal Mass Storage 1848 System Bus 1849 Peripheral Bus 1850 Graphics Adapter 1854 Network Interface 1855 one or more communication networks

Claims

1. 1. A method for video decoding in a decoder, comprising: decoding coded information of a chroma coding block (CB) from the coded video bit stream, the coded information indicating that a cross-component filter is applied to a luma CB, the coded information further indicating a chroma subsampling format and a chroma sample type indicating a relative position of a chroma sample with respect to at least one luma sample in a corresponding luma CB; determining a filter shape of the cross-component filter to be applied to the luma CB based on at least one of the chroma subsampling format and the chroma sample type; generating a first intermediate CB by applying a loop filter to the chroma CB; generating a second intermediate CB by applying the cross-component filter having the determined filter shape to the corresponding luminance CB; determining a filtered saturation CB based on the first intermediate CB and the second intermediate CB; Including, a number of filter coefficients of the cross-component filter is signaled in the encoded video bit stream; The method, wherein determining the filter shape includes determining the filter shape of the cross-component filter based on the number of filter coefficients and the at least one of the chroma subsampling format and the chroma sample type.

2. The method of claim 1 , wherein the chroma sample format is signaled in the encoded video bit stream.

3. the chroma subsampling format is 4:2:0; the at least one luma sample includes four luma samples, the luma samples being a top left sample, a top right sample, a bottom left sample, and a bottom right sample; the chroma sample format is one of six chroma sample formats 0-5, each of which indicates six relative positions 0-5 of the chroma samples, the six relative positions 0-5 corresponding to a left center position between the top left sample and the bottom left sample, a center position of the four luma samples, a top left position coinciding with the top left sample, a top center position between the top left sample and the top right sample, a bottom left position coinciding with the bottom left sample, and a bottom center position between the bottom left sample and the bottom right sample; determining the filter shape includes determining the filter shape of the cross-component filter based on the chroma sample type. The method according to claim 1 or 2.

4. 4. The method of claim 3, wherein the encoded video bit stream includes a cross-component linear model (CCLM) flag indicating that the chroma sample format is 0 or 2.

5. The method according to any one of claims 1 to 4, wherein the cross-component filter is a cross-component adaptive loop filter (CC-ALF) and the loop filter is an adaptive loop filter (ALF).

6. The method of claim 1 , wherein the range of filter coefficients of the cross-component filter is less than or equal to K bits, where K is a positive integer.

7. The method of claim 6 , wherein the filter coefficients of the cross-component filters are coded using fixed-length coding.

8. Shifting the corresponding luminance CB luminance sample value to have a dynamic range of 8 bits based on a dynamic range of the luminance sample value that is greater than 8 bits, where K is 8 bits; generating the second intermediate CB includes applying the cross-component filter having the determined filter shape to the shifted luminance sample values. The method of claim 6.

9. 1. A method for video decoding in a decoder, comprising: decoding coded information of a chroma coding block (CB) from a coded video bit stream, the coded information indicating that a cross-component filter is applied to the luma coding block (CB) based on a corresponding luma coding block (CB); generating a downsampled luma CB by applying a down-sampling filter to the corresponding luma CB, where a chroma horizontal subsampling factor and a chroma vertical subsampling factor between the chroma CB and the downsampled luma CB are 1; generating a first intermediate CB by applying a loop filter to the chroma CB; generating a second intermediate CB by applying the cross-component filter to the downsampled luma CB, the filter shape of the cross-component filter being independent of the chroma subsampling format and chroma sample type of the chroma CB, the chroma sample type indicating a relative position of a chroma sample with respect to at least one luma sample in the corresponding luma CB; determining a filtered saturation CB based on the first intermediate CB and the second intermediate CB; Including, A method according to claim 1, wherein a number of filter coefficients of the cross-component filters are signaled in the encoded video bit stream.

10. The method of claim 9 , wherein the down-sampling filter corresponds to a filter applied to co-located luma samples in CCLM mode.

11. 10. The method of claim 9, wherein the down-sampling filter is a {1,2,1;1,2,1} / 8 filter and the chroma subsampling format is 4:2:

0.

12. 12. The method of claim 9, wherein the filter shape of the cross-component filter is one of a 7x7 diamond shape, a 7x7 square shape, a 5x5 diamond shape, a 5x5 square shape, a 3x3 diamond shape, and a 3x3 square shape.

13. 9. An apparatus for video decoding, comprising a processing circuit, the processing circuit being configured to perform the method of any one of claims 1 to 8.

14. A program for causing one or more processors to carry out the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cross Component Filter

    JP2019525679A