Method and apparatus for video filtering
Customizable nonlinear mapping-based filters address inefficiencies in representing less likely intra-prediction directions, enhancing video compression efficiency by reducing bit usage and redundancy.
Patent Information
- Application Number
- JP2025182518
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-28
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-23
AI Technical Summary
Existing video coding technologies face inefficiencies in representing intra-prediction directions, particularly those that are statistically less likely, leading to increased bit usage and reduced compression efficiency.
Implementing a nonlinear mapping-based filter with customizable filter shapes, such as cross-component sample offset (CCSO) and local sample offset (LSO) filters, to enhance video encoding and decoding processes, allowing for more efficient representation of intra-prediction directions.
The solution reduces the bit usage for less likely intra-prediction directions, improving video compression efficiency and reducing redundancy in encoded video data.
Smart Images

Figure 2026012316000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 160,560, entitled "FLEXIBLE FILTER SHAPE FOR SAMPLE OFFSET," filed March 12, 2021, which claims the benefit of priority to U.S. Patent Application No. 17 / 449,199, entitled "METHOD AND APPARATUS FOR VIDEO FILTERING," filed September 28, 2021. The entire disclosures of the above applications are incorporated herein by reference in their entirety.
[0002] This disclosure describes embodiments related generally to video coding. [Background technology]
[0003] The background art discussion provided herein is intended to generally present the context for the present disclosure. The inventors' work, to the extent that this work is described in this background art section, and aspects of the description that may not be considered prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.
[0004] Video coding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luma samples and associated chroma samples. The series of pictures can have a fixed or variable picture rate (also informally known as a frame rate), for example, 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920 x 1080 luma sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GBytes of storage space.
[0005] One goal of video coding and decoding is to reduce redundancy in an input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements by more than two orders of magnitude, in some cases. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to techniques that allow an exact copy of the original signal to be reconstructed from a compressed version of the original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough to allow the reconstructed signal to be used for its intended purpose. For video, lossy compression is widely used. The amount of acceptable distortion depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that higher acceptable / tolerable distortion may result in a higher compression ratio.
[0006] Video encoders and decoders may utilize techniques from several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.
[0007] Video codec technology may include a technique known as intra-coding, in which sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, may be used to reset the decoder state and therefore may be used as the first picture in a coded video bitstream and video session or as a still image. Samples of intra-blocks may undergo a transform, and the transform coefficients may be quantized before entropy coding. Intra-prediction may be a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after the transform, the fewer bits are required for a given quantization step size to represent the block after entropy coding.
[0008] Conventional intra-coding, e.g., as known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that rely on surrounding sample data and / or metadata obtained during the encoding / decoding of, for example, spatially neighboring blocks of data or preceding blocks of data in decoding order. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from reference pictures.
[0009] Intra-prediction can take many different forms. If more than one of these techniques can be used in a given video coding technique, the technique in use can be coded as an intra-prediction mode. In certain cases, a mode can have sub-modes and / or parameters, which can be coded separately or included in a mode codeword. The codeword used for a given mode / sub-mode / parameter combination can affect the coding efficiency achieved by intra-prediction, and therefore can affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] A specific mode of intra prediction was introduced in H.264, improved in H.265, and further refined in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark sets (BMS). A predictor block can be formed using neighboring sample values belonging to already available samples. The sample values of the neighboring samples are copied into the predictor block according to the direction. A reference to the direction in use can be coded in the bitstream or can itself be predicted.
[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions known from the 33 possible predictor directions in H.265 (corresponding to the 33 angular modes out of the 35 intra modes). The point where the arrows converge (101) represents the sample being predicted. The arrows represent the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from a sample or samples located to the upper right and at a 45-degree angle from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from a sample or samples located to the lower left and at a 22.5-degree angle from horizontal.
[0012] 1A, a square block (104) of 4x4 samples (depicted by a thick dashed line) is shown in the upper left. The square block (104) includes 16 samples, each of which is labeled with "S," its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Because the block size is 4x4 samples, S44 is located in the lower right. Reference samples, which follow a similar numbering scheme, are also shown. The reference samples are labeled with R, their Y position (e.g., row index), and their X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are neighbors of the block being reconstructed, so negative values do not need to be used.
[0013] Intra-picture prediction may work by copying reference sample values from neighboring samples as assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction consistent with the arrow (102) for this block, i.e., the sample is predicted from a prediction sample or samples located to the upper right at a 45-degree angle from horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Next, sample S44 is predicted from reference sample R08.
[0014] In certain cases, especially when the orientation is not evenly divisible by 45 degrees, the values of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.
[0015] The number of possible directions has increased as video coding technology has evolved. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and as of the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and specific entropy coding techniques are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predicted from nearby directions used in nearby, already decoded blocks.
[0016] Figure 1B shows a schematic diagram (180) showing 65 intra-prediction directions with JEM to illustrate the increasing number of prediction directions over time.
[0017] The mapping of intra-prediction direction bits in a coded video bitstream to represent directions may vary from video coding technique to technique, ranging from simple direct mappings from prediction directions to intra-prediction modes, to codewords, to complex adaptive schemes including most-probable modes, and similar techniques. However, in all cases, there may be certain directions that are statistically less likely than certain other directions in the video content. Because the goal of video compression is to reduce redundancy, these less likely directions are represented with more bits than more likely directions in a well-performing video coding technique.
[0018] Motion compensation may be a lossy compression technique used to predict a newly reconstructed picture or picture portion after blocks of sample data from a previously reconstructed picture or portion thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the latter may indirectly be a temporal dimension).
[0019] In some video compression techniques, the MV applicable to a particular area of sample data can be predicted from other MVs, for example, from an MV associated with another area of sample data that is spatially adjacent to the area being reconstructed and precedes that MV in decoding order. Doing so can substantially reduce the amount of data required to code the MV, thereby eliminating redundancy and increasing compression. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical likelihood that areas larger than the area to which a single MV is applicable will move in a similar direction, and therefore, in some cases, can be predicted using similar motion vectors derived from MVs in nearby areas. This results in the MV found for a given area being similar or identical to the MV predicted from surrounding MVs, and after entropy coding, can be represented using fewer bits than would be used to code the MV directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself can be lossy, for example, due to rounding errors when calculating a predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Here, we will explain a technique called "spatial merging" from among the many MV prediction mechanisms provided by H.265.
[0021] Referring to Figure 2, the current block (201) contains samples found by the encoder during the motion search process to be predictable from a spatially shifted previous block of the same size. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, e.g., the most recent (in decoding order) reference picture, using MVs associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, and B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture as neighboring blocks. Summary of the Invention [Means for solving the problem]
[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit determines an offset value associated with a first filter shape configuration of a nonlinear mapping-based filter based on a signal in a coded video bitstream carrying video. The first filter shape configuration has a number of filter taps less than five. The processing circuit applies the nonlinear mapping-based filter to samples to be filtered using the offset value associated with the first filter shape configuration.
[0023] In some examples, the nonlinear mapping-based filter includes at least one of a cross-component sample offset (CCSO) filter and a local sample offset (LSO) filter.
[0024] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered. In another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.
[0025] In some examples, the first filter shape configuration includes a single filter tap. The processing circuit determines an average sample value within the area and calculates a difference between the reconstructed sample value at the position of the single filter tap and the average sample value within the area. The processing circuit then applies a nonlinear mapping-based filter to the samples to be filtered based on the difference between the reconstructed sample value at the position of the single filter tap and the average sample value within the area.
[0026] In some examples, the processing circuit selects a first filter shape configuration from a group of filter shape configurations of nonlinear mapping-based filters. In one example, the filter shape configurations in the group each have a number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration. In some examples, the processing circuit decodes an index from a coded video bitstream carrying video, the index indicating selection of the first filter shape configuration from the group of nonlinear mapping-based filters. For example, the processing circuit decodes the index from syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.
[0027] In some examples, the processing circuit quantizes delta values of sample values at two filter tap positions to one of a number of possible quantized outputs, where the number of possible quantized outputs is an integer in the range of 1 to 1024, inclusive. For example, the processing circuit decodes an index from syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, where the index indicates the number of possible quantized outputs.
[0028] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video encoding / decoding.
[0029] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief explanation of the drawings]
[0030] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 1 is a diagram of an exemplary intra-prediction direction. [Figure 2] FIG. 1 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 3] 1 is a simplified schematic block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a simplified schematic block diagram of a communication system (400) according to one embodiment. [Figure 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 10 is a block diagram of a decoder according to another embodiment. [Figure 9] 1A and 1B illustrate examples of filter shapes according to embodiments of the present disclosure. [Figure 10A] FIG. 10 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10B] FIG. 10 illustrates another example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10C] FIG. 10 illustrates yet another example of sub-sampled positions used to calculate gradients according to embodiments of the present disclosure. [Figure 10D] FIG. 10 illustrates yet another example of sub-sampled positions used to calculate gradients according to embodiments of the present disclosure. [Figure 11A] FIG. 1 illustrates an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 11B] FIG. 10 illustrates another example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 12A] FIG. 10 illustrates an example of a symmetric padding operation on a virtual boundary according to an embodiment of the present disclosure. [Figure 12B] FIG. 10 illustrates another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12C] FIG. 10 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12D] FIG. 10 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12E] FIG. 10 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12F] FIG. 10 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 13] FIG. 2 illustrates example partitions of a picture according to some embodiments of the present disclosure. [Figure 14] 1A-1C illustrate quadtree division patterns for pictures in some examples. [Figure 15] FIG. 1 illustrates a cross-component filter according to an embodiment of the present disclosure. [Figure 16] FIG. 1 illustrates an example of a filter shape according to an embodiment of the present disclosure. [Figure 17] FIG. 10 illustrates an example syntax for a cross-component filter according to some embodiments of the present disclosure. [Figure 18A] FIG. 2 illustrates an exemplary location of chroma samples relative to luma samples, according to one embodiment of the present disclosure. [Figure 18B] FIG. 10 illustrates an exemplary location of chroma samples relative to luma samples in accordance with another embodiment of the present disclosure. [Figure 19] FIG. 1 illustrates an example of direction finding according to an embodiment of the present disclosure. [Figure 20] FIG. 10 illustrates an example of subspace projection in some examples. [Figure 21] 1 is a table of multiple sample adaptive offset (SAO) types according to one embodiment of the present disclosure. [Figure 22] 10A-10C illustrate example patterns for pixel classification at edge offset in some examples. [Figure 23] 10 is a table of edge offset pixel classification rules in some examples. [Figure 24] FIG. 10 illustrates an example of syntax that may be signaled. [Figure 25] 10A-10C illustrate examples of filter support areas according to some embodiments of the present disclosure. [Figure 26] 10A-10C illustrate examples of another filter support area according to some embodiments of the present disclosure. [Figure 27A] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27B] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27C] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 28] FIG. 10 illustrates eight filter shape configurations of three filter taps in one example. [Figure 29] FIG. 12 illustrates 12 filter shape configurations for three filter taps in one example. [Figure 30] FIG. 1 illustrates an example of two candidate filter shape configurations for a nonlinear mapping-based filter. [Figure 31] 1 is a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 32] 1 is a flowchart outlining a process according to one embodiment of the present disclosure. [Figure 33] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0031] Figure 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes multiple terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.
[0032] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.
[0033] In the example of FIG. 3 , the terminal devices 310, 320, 330, and 340 may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure need not be so limited. Embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. The network 350 represents any number of networks that convey coded video data between the terminal devices 310, 320, 330, and 340, including, for example, wired and / or wireless communication networks. The communication network 350 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network 350 may not be important to the operation of the present disclosure unless otherwise described herein below.
[0034] 4 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter, which may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0035] The streaming system may include a video source (401) and a capture subsystem (413), which may include, for example, a digital camera, that creates an uncompressed video picture stream (402). In one example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402), depicted as a thick line to emphasize its high data volume compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize its lower data volume compared to the video picture stream (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard, such as ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.
[0036] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0037] 5 shows a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used in place of the video decoder (410) in the example of FIG. 4.
[0038] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive encoded video data along with other data, such as coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). The receiver (531) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be external to the video decoder (510) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (510), for example, to combat network jitter, and another buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. When the receiver (531) is receiving data from a store / forward device of sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (515) may not be needed or may be small. For use with best-effort packet networks such as the Internet, a buffer memory (515) may be needed, may be relatively large, preferably adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).
[0039] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device, such as a render device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, context-sensitive or non-context-sensitive arithmetic coding, etc. The parser (520) may extract from the coded video sequence at least one set of subgroup parameters for a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (520) may also extract from the coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0040] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to produce symbols (521).
[0041] The reconstruction of the symbols (521) can involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.
[0042] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:
[0043] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients and control information from the parser (520) as symbols (521), including the transform to be used, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (551) can output blocks containing sample values that can be input to an aggregator (555).
[0044] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (558). The current picture buffer (558), for example, buffers the partially reconstructed and / or fully reconstructed current picture. The aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0045] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (553) can access the reference picture memory (557) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples can be added by the aggregator (555) to the output of the scalar / inverse transform unit (551) to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches the prediction samples can be controlled by a motion vector and are available to the motion-compensated prediction unit (553) in the form of a symbol (521) that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.
[0046] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of the coded picture or coded video sequence, and may also be responsive to previously reconstructed and loop-filtered sample values.
[0047] The output of the loop filter unit (556) may be a sample stream that can be output to the render device (512) as well as stored in a reference picture memory (557) for use in future inter-picture prediction.
[0048] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before beginning reconstruction of the next coded picture.
[0049] The video decoder (510) may perform decoding operations according to a predetermined video compression technique in a standard such as ITU-T Rec. H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to both the syntax and the profile of the video compression technique or standard as documented in the video compression technique or standard. Specifically, the profile may select certain tools from among all tools available in the video compression technique or standard as the only tools available for use under that profile. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by the specification of a hypothetical reference decoder (HRD) and HRD buffer management metadata signaled in the coded video sequence.
[0050] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0051] 6 shows a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.
[0052] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture the video image(s) to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0053] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description focuses on samples.
[0054] According to one embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units as described below. For clarity, coupling is not depicted. Parameters set by the controller (650) may include rate control-related parameters (e.g., picture skip, quantizer, lambda value for rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions for the video encoder (603) optimized for a particular system design.
[0055] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (630) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to that of the (remote) decoder (because in the video compression techniques contemplated by the disclosed subject matter, any compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding of the symbol stream produces bit-exact results regardless of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, e.g., due to channel errors) is also used in some related technologies.
[0056] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510), which is described in detail above in connection with Figure 5. However, with brief reference also to Figure 5, symbols may be available, and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, and the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).
[0057] An observation that can be made in this regard is that decoder techniques other than analysis / entropy decoding present in a decoder must necessarily be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operation. A description of the encoder techniques may be omitted, as they are the opposite of the decoder techniques described comprehensively. Only in certain areas are more detailed descriptions required and are provided below.
[0058] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this way, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference picture(s) that may be selected as predictive reference(s) for the input picture.
[0059] The local video decoder (633) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (630). The operation of the coding engine (632) may preferably be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by a video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).
[0060] The predictor (635) may perform the predictive search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable predictive references for the new picture. The predictor (635) may operate on one sample block per pixel block to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634).
[0061] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0062] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0063] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) and prepare them for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0064] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0065] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs support various types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variants of I-pictures and their respective uses and characteristics.
[0066] A predictive picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0067] A bidirectionally predicted picture (B picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0068] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded with reference to one previously coded reference picture by spatial prediction or temporal prediction. Blocks of a B-picture may be predictively coded with reference to one or two previously coded reference pictures by spatial prediction or temporal prediction.
[0069] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.
[0070] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, and VUI parameter set fragments.
[0071] Video may be captured as a time-sequence of multiple source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0072] In some embodiments, inter-picture prediction may use a bi-prediction technique. According to the bi-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (but may be in the past and future, respectively, in display order). A block in the current picture may be coded with a first motion vector pointing toward a first reference block in the first reference picture and a second motion vector pointing toward a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.
[0073] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0074] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three CTUs: one luma coding tree block (CTB) and two chroma CTBs. Each CTU may be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU may be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) of 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0075] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block of sample values (e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.
[0076] In an HEVC example, the video encoder (703) receives a matrix of sample values for a processing block, such as a predictive block of 8x8 samples. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-predictive mode, e.g., using rate-distortion optimization. When the processing block is coded in intra-mode, the video encoder (703) may use intra-prediction techniques to code the processing block into a coded picture, and when the processing block is coded in inter-mode or bi-predictive mode, the video encoder (703) may use inter-prediction or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.
[0077] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculation unit (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), which are coupled together as shown in Figure 7.
[0078] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundant information through an inter-encoding technique, a motion vector, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.
[0079] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with previously coded blocks in the same picture, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.
[0080] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is intra mode, the general controller (721) controls the switch (726) to select an intra mode result to be used by the residual calculation unit (723) and controls the entropy encoder (725) to select intra prediction information to include in the bitstream. If the mode is inter mode, the general controller (721) controls the switch (726) to select an inter prediction result to be used by the residual calculation unit (723) and controls the entropy encoder (725) to select inter prediction information to include in the bitstream.
[0081] The residual calculation unit (723) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate on the residual data and encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients then undergo a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be suitably used by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and intra-prediction information. In some examples, the decoded blocks are appropriately processed to generate decoded pictures, which may be buffered in a memory circuit (not shown) and used as reference pictures.
[0082] The entropy encoder (725) is configured to format a bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information in accordance with an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that, according to the disclosed subject matter, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, residual information is not present.
[0083] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.
[0084] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), which are coupled together as shown in Figure 8.
[0085] The entropy decoder (871) may be configured to reconstruct, from a coded picture, certain symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra-prediction information or inter-prediction information) that can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, the latter two being merged or separate submodes), certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively, residual information, e.g., in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter-prediction mode or a bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and be provided to the residual decoder (873).
[0086] The inter decoder (880) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0087] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0088] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (data paths not shown may be low-level control information only).
[0089] The reconstruction module (874) is configured to combine, in the spatial domain, the residual as output by the residual decoder (873) and the prediction result (possibly as output by an inter- or intra-prediction module) to form a reconstructed block that may be part of a reconstructed picture, and the reconstructed block may be part of a reconstructed video. It should be noted that other appropriate operations, such as a deblocking operation, may be performed to improve visual quality.
[0090] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.
[0091] Aspects of the present disclosure provide filtering techniques for video coding / decoding.
[0092] An adaptive loop filter (ALF) with block-based filter adaptation can be applied by the encoder / decoder to reduce artifacts. For the luma component, for example, one of multiple filters (e.g., 25 filters) can be selected for a 4x4 luma block based on local gradient direction and activity.
[0093] The ALFs can have any suitable shape and size. Referring to FIG. 9, the ALFs (910)-(911) have diamond shapes, such as a 5x5 diamond shape for the ALF (910) and a 7x7 diamond shape for the ALF (911). In the ALF (910), elements (920)-(932) form the diamond shape and can be used in the filtering process. Seven values (e.g., C0-C6) can be used for the elements (920)-(932). In the ALF (911), elements (940)-(964) form the diamond shape and can be used in the filtering process. Thirteen values (e.g., C0-C12) can be used for the elements (940)-(964).
[0094] Referring to FIG. 9, in some examples, two ALFs (910)-(911) having diamond filter shapes are used. A 5x5 diamond-shaped filter (910) may be applied to a chroma component (e.g., a chroma block, chroma CB), and a 7x7 diamond-shaped filter (911) may be applied to a luma component (e.g., a luma block, luma CB). Other suitable shapes and sizes may be used in the ALFs. For example, a 9x9 diamond-shaped filter may be used.
[0095] The filter coefficients at the locations indicated by the values (e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, if the ALF includes a clipping function, the clip values at those locations may be non-zero.
[0096] For block classification of the luma component, a 4x4 block (or luma block, luma CB) can be categorized or classified as one of multiple (e.g., 25) classes. The classification index C is calculated using Equation (1) by using the quantized value of the directionality parameter D and the activity value A.
number
number
number
number
number
number
number
[0097] To reduce the complexity of the block classification described above, a subsampled 1-D Laplacian calculation can be applied. Figures 10A-10D show the gradient g in the vertical direction (Figure 10A), horizontal direction (Figure 10B), and two diagonal directions d1 (Figure 10C) and d2 (Figure 10D). v , g h , g d1 , and g d2 10A shows an example of the subsampled positions used to compute the vertical gradient g. The same subsampled positions can be used for gradient computations in different directions. In FIG. 10A, the label "V" indicates the vertical gradient g. vIn Figure 10B, the label "H" indicates the subsampled positions for computing the horizontal gradient g h In FIG. 10C, the label "D1" indicates the subsampled positions for computing the d1 diagonal gradient g d1 In Figure 10D, the label "D2" indicates the subsampled positions for computing the d2 diagonal gradient g d2 indicates the subsampled positions for computing
[0098] horizontal g v and vertical g h The maximum value of the gradient of
number
number
number
number
number
number
number
number
number
number
number
[0099] The activity value A can be calculated as follows:
number
number
[0100] For the chroma components in a picture, no block classification is applied, and therefore a single set of ALF coefficients can be applied for each chroma component.
[0101] A geometric transformation can be applied to the filter coefficients and corresponding filter clip values (also called clip values). Before filtering a block (e.g., a 4x4 luma block), for example, the gradient values (e.g., g v , g h , g d1 and / or g d2Depending on the filter coefficients f(k,l), a geometric transformation such as a rotation or a diagonal and vertical flip can be applied to the filter coefficients f(k,l) and the corresponding filter clip values c(k,l). The geometric transformation applied to the filter coefficients f(k,l) and the corresponding filter clip values c(k,l) can be equivalent to applying a geometric transformation to the samples within the region supported by the filter. The geometric transformation can make the different blocks to which the ALF is applied more similar by aligning their respective directionality.
[0102] Three geometric transformations can be performed, including a diagonal flip, a vertical flip, and a rotation, as described by equations (9)-(11), respectively. f D (k,l)=f(l,k),c D (k,l)=c(l,k) Equation (9) f V (k,l)=f(k,Kl-1),c V (k,l)=c(k,Kl-1) Equation (10) f R (k,l)=f(Kl-1,k),c R (k,l)=c(Kl-1,k) Equation (11) Here, K is the size of the ALF or filter, and 0≦k, 1≦K-1 are the coordinates of the coefficients. For example, the filter f or clip value matrix (or clip matrix) c has position (0,0) in the upper left corner and position (K-1,K-1) in the lower right corner. Transforms can be applied to the filter coefficients f(k,l) and clip values c(k,l) depending on the gradient values calculated for the block. An example of the relationship between the transforms and the four gradients is summarized in Table 1.
[0103] [Table 1]
[0104] In some embodiments, the ALF filter parameters are signaled in an adaptive parameter set (APS) of a picture. In the APS, one or more sets (e.g., up to 25 sets) of luma filter coefficients and clip value indices may be signaled. In one example, a set of the one or more sets may include luma filter coefficients and one or more clip value indices. One or more sets (e.g., up to 8 sets) of chroma filter coefficients and clip value indices may be signaled. To reduce signaling overhead, filter coefficients of different classifications (e.g., having different classification indices) of the luma component may be merged. In the slice header, the index of the APS used for the current slice may be signaled.
[0105] In one embodiment, a clip value index (also referred to as a clipping index) can be decoded from the APS. The clip value index may be used to determine a corresponding clip value, for example, based on a relationship between the clip value index and the corresponding clip value. The relationship may be predefined and stored in the decoder. In one example, the relationship is described by a table, such as a luma table (e.g., used for the luma CB) of clip value indexes and corresponding clip values, a chroma table (e.g., used for the chroma CB) of clip value indexes and corresponding clip values, etc. The clip value may depend on the bit depth B. The bit depth B may refer to the internal bit depth, the bit depth of the reconstructed samples in the CB to be filtered, etc. In some examples, the tables (e.g., luma table, chroma table) are obtained using Equation (12).
number
[0106] [Table 2]
[0107] The slice header of the current slice may signal one or more APS indices (e.g., up to seven APS indices) to specify the luma filter sets that can be used for the current slice. The filtering process may be controlled at one or more appropriate levels, such as the picture level, slice level, or CTB level. In one embodiment, the filtering process may be further controlled at the CTB level. A flag may be signaled to indicate whether an ALF is applied to the luma CTB. The luma CTB may select a filter set from multiple fixed filter sets (e.g., 16 fixed filter sets) and a filter set (also referred to as a signaled filter set) signaled in the APS. A filter set index may be signaled to the luma CTB to indicate the filter set to be applied (e.g., a filter set among the multiple fixed filter sets and signaled filter sets). The multiple fixed filter sets may be predefined and hard-coded in the encoder and decoder and may be referred to as predefined filter sets.
[0108] For chroma components, an APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there is more than one chroma filter set in the APS, a filter set index can be signaled for each chroma CTB.
[0109] The filter coefficients may be quantized with a norm equal to 128. To reduce multiplication complexity, bitstream adaptation may be applied so that the coefficient values of non-center positions are within the range of −27 to 27−1, inclusive. In one example, center position coefficients are not signaled in the bitstream and may be considered equal to 128.
[0110] In some embodiments, the syntax and semantics of clipping indexes and clip values are defined as follows: alf_luma_clip_idx[sfIdx][j] may be used to specify the clipping index of the clip value to use before multiplying by the jth coefficient of the signaled luma filter indicated by sfIdx. Bitstream conformance requirements may include that when sfIdx=0 to alf_luma_num_filters_signalled_minus1 and j=0 to 11, the value of alf_luma_clip_idx[sfIdx][j] is in the range of 0 to 3, inclusive. The luma filter clip value AlfClipL[adaptation_parameter_set_id] having element AlfClipL[adaptation_parameter_set_id][filtIdx][j] may be derived as specified in Table 2 depending on bitDepth set equal to BitDepthY and clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j] when filtIdx = 0 to NumAlfFilters-1 and j = 0 to 11. alf_chroma_clip_idx[altIdx][j] can be used to specify the clipping index of the clip value to use before multiplying by the jth coefficient of the alternate chroma filter with index altIdx. Bitstream conformance requirements may include that the value of alf_chroma_clip_idx[altIdx][j] is in the range of 0 to 3, inclusive, when altIdx=0 to alf_chroma_num_alt_filters_minus1 and j=0 to 5. The chroma filter clip value AlfClipC[adaptation_parameter_set_id][altIdx] with element AlfClipC[adaptation_parameter_set_id][altIdx][j] may be derived as specified in Table 2 depending on bitDepth set equal to BitDepthC and clipIdx set equal to alf_chroma_clip_idx[altIdx][j] when altIdx=0 to alf_chroma_num_alt_filters_minus1 and j=0 to 5.
[0111] In one embodiment, the filtering process can be described as follows: On the decoder side, when ALF is enabled for a CTB, samples R(i,j) in a CU (or CB) can be filtered, and filtered sample values R′(i,j) are obtained as shown using the following equation (13): In one example, each sample in a CU is filtered.
number
[0112] In a nonlinear ALF, multiple sets of clip values can be provided in Table 3. In one example, the luma set includes four clip values {1024, 181, 32, 6}, and the chroma set includes four clip values {1024, 161, 25, 4}. The four clip values in the luma set can be selected by approximately equally dividing the full range (e.g., 1024) of the luma block sample values (coded with 10 bits) in the logarithmic domain. This range can be from 4 to 1024 for the chroma set.
[0113] [Table 3]
[0114] The selected clip value can be coded in the "alf_data" syntax element as follows: An appropriate encoding scheme (e.g., Golomb encoding scheme) can be used to encode the clipping index corresponding to the selected clip value as shown in Table 3. The encoding scheme can be the same encoding scheme used to encode the filter set index.
[0115] In one embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of ALF. Therefore, modified block classification and filtering can be used for samples near CTU boundaries (e.g., horizontal CTU boundaries). The virtual boundary (1130) is defined by dividing the horizontal CTU boundary (1120) by "N" as shown in FIG. 11A. samples ” can be defined as a line by shifting the sample by N samples can be a positive integer. In one example, N samples is equal to 4 for the luma component and N samples is equal to 2 for the chroma component.
[0116] Referring to Figure 11A, modified block classification can be applied to the luma component. In one example, a 1D Laplacian gradient calculation for a 4x4 block (1110) above a virtual boundary (1130) uses only samples above the virtual boundary (1130). Similarly, referring to Figure 11B, a 1D Laplacian gradient calculation for a 4x4 block (1111) below a virtual boundary (1131) shifted from the CTU boundary (1121) uses only samples below the virtual boundary (1131). Thus, the quantization of the activity value A can be scaled by taking into account the reduced number of samples used in the 1D Laplacian gradient calculation.
[0117] For the filtering process, symmetric padding operations at the virtual boundary may be used for both the luma and chroma components. Figures 12A-12F show examples of such modified ALF filtering for the luma component at the virtual boundary. If the sample being filtered is located below the virtual boundary, neighboring samples located above the virtual boundary may be padded. If the sample being filtered is located above the virtual boundary, neighboring samples located below the virtual boundary may be padded. Referring to Figure 12A, neighboring sample C0 may be padded with sample C2 located below the virtual boundary (1210). Referring to Figure 12B, neighboring sample C0 may be padded with sample C2 located above the virtual boundary (1220). Referring to Figure 12C, neighboring samples C1-C3 may be padded with samples C5-C7 located below the virtual boundary (1230), respectively. Referring to Figure 12D, neighboring samples C1-C3 may be padded with samples C5-C7 located above the virtual boundary (1240), respectively. Referring to Figure 12E, neighboring samples C4-C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located below the virtual boundary (1250). Referring to Figure 12F, neighboring samples C4-C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located above the virtual boundary (1260).
[0118] In some examples, the above description can be appropriately adapted when the sample(s) and neighboring sample(s) are located to the left (or right) and right (or left) of the virtual boundary.
[0119] According to aspects of the present disclosure, to improve coding efficiency, a picture may be divided based on a filtering process. In some examples, a CTU is also referred to as a largest coding unit (LCU). In one example, a CTU or LCU may have a size of 64x64 pixels. In some embodiments, an LCU-aligned picture quadtree partitioning may be used for filtering-based partitioning. In some examples, a coding unit-synchronized picture quadtree-based adaptive loop filter may be used. For example, a luma picture is divided into multiple multi-level quadtree partitions, and the boundaries of each partition are aligned with the boundaries of the LCUs. Each partition has its own filtering process and is therefore referred to as a filter unit (FU).
[0120] In some examples, a two-pass encoding flow may be used. In the first pass of the two-pass encoding flow, a quadtree partitioning pattern for the picture and a best filter for each FU may be determined. In some embodiments, the determination of the quadtree partitioning pattern for the picture and the determination of the best filter for the FU are based on filtering distortion. The filtering distortion may be estimated by a fast filtering distortion estimation (FFDE) technique during the determination process. The picture is divided using a quadtree partition. The reconstructed picture may be filtered according to the determined quadtree partitioning pattern and the selected filters for all FUs.
[0121] In the second pass of the two-pass encoding flow, the CU-synchronized ALF is turned on / off. According to the ALF on / off result, the first filtered picture is partially restored by the reconstructed picture.
[0122] Specifically, in some examples, a top-down partitioning method is employed to divide a picture into multi-level quadtree partitions using a rate-distortion criterion. Each partition is called a filter unit (FU). The partitioning process aligns the quadtree partitions to LCU boundaries. The encoding order of the FUs follows the z-scan order.
[0123] 13 illustrates an example partition according to some embodiments of the present disclosure. In the example of FIG. 13, a picture (1300) is divided into 10 FUs, and the encoding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, FU9.
[0124] Figure 14 shows a quadtree partitioning pattern (1400) for a picture (1300). In the example of Figure 14, a partition flag is used to indicate the partition pattern for the picture. For example, "1" indicates that quadtree partitioning is performed on the block, and "0" indicates that the block should not be further divided. In some examples, the minimum size FU has the LCU size, and no partition flag is required for the minimum size FU. The partition flag is encoded and transmitted in z-order as shown in Figure 14.
[0125] In some examples, the filter for each FU is selected from two filter sets based on a rate-distortion criterion. The first set has half-symmetric square and diamond filters derived for the current FU. The second set is obtained from a time delay filter buffer, which stores filters previously derived for FUs of the previous picture. The filter with the smallest rate-distortion cost of these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate-distortion costs of the four child FUs are calculated. A picture quadtree division pattern can be determined by recursively comparing the rate-distortion costs of the division case and the non-division case.
[0126] In some examples, the maximum quadtree division level can be used to limit the maximum number of FUs. In one example, if the maximum quadtree division level is 2, the maximum number of FUs is 16. Furthermore, during quadtree division determination, correlation values for deriving Wiener coefficients for the 16 FUs at the lowest quadtree level (smallest FU) can be reused. The remaining FUs can derive their Wiener filters from the correlations of the 16 FUs at the lowest quadtree level. Therefore, in this example, only one frame buffer access is performed to derive filter coefficients for all FUs.
[0127] After the quadtree partitioning pattern is determined, CU-synchronized ALF on / off control can be performed to further reduce filtering distortion. By comparing the filtering distortion and non-filtering distortion at each leaf CU, the leaf CU can explicitly switch ALF on / off in its local region. In some cases, coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off results.
[0128] The cross-component filtering process may apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter may use luma sample values of a luma component (e.g., a luma CB) to refine a chroma component (e.g., a chroma CB corresponding to the luma CB). In one example, the luma CB and the chroma CB are included in a CU.
[0129] FIG. 15 illustrates a cross-component filter (e.g., CC-ALF) used to generate chroma components according to one embodiment of the present disclosure. In some examples, FIG. 15 illustrates a filtering process of a first chroma component (e.g., first chroma CB), a second chroma component (e.g., second chroma CB), and a luma component (e.g., luma CB). The luma component may be filtered by a sample adaptive offset (SAO) filter (1510) to generate an SAO-filtered luma component (1541). The SAO-filtered luma component (1541) may be further filtered by an ALF luma filter (1516) to become a filtered luma CB (1561) (e.g., “Y”).
[0130] The first chroma component may be filtered by an SAO filter (1512) and an ALF chroma filter (1518) to generate a first intermediate component (1552). Further, the SAO-filtered luma component (1541) may be filtered by a cross-component filter (e.g., CC-ALF) for the first chroma component (1521) to generate a second intermediate component (1542). Subsequently, a filtered first chroma component (1562) (e.g., “Cb”) may be generated based on at least one of the second intermediate component (1542) and the first intermediate component (1552). In one example, the filtered first chroma component (1562) (e.g., “Cb”) may be generated by combining the second intermediate component (1542) and the first intermediate component (1552) with an adder (1522). The cross-component adaptive loop filtering process for the first chroma component may include steps performed by the CC-ALF (1521) and steps performed by, for example, an adder (1522).
[0131] The above description can be adapted to the second chroma component. The second chroma component can be filtered by the SAO filter (1514) and the ALF chroma filter (1518) to generate a third intermediate component (1553). Furthermore, the SAO-filtered luma component (1541) can be filtered by a cross-component filter (e.g., CC-ALF) for the second chroma component (1531) to generate a fourth intermediate component (1543). Subsequently, a filtered second chroma component (1563) (e.g., “Cr”) can be generated based on at least one of the fourth intermediate component (1543) and the third intermediate component (1553). In one example, the filtered second chroma component (1563) (e.g., “Cr”) can be generated by combining the fourth intermediate component (1543) and the third intermediate component (1553) with the adder (1532). In one example, the cross-component adaptive loop filtering process for the second chroma component may include steps performed by the CC-ALF (1531) and steps performed by, for example, an adder (1532).
[0132] The cross-component filters (e.g., CC-ALF(1521), CC-ALF(1531)) can operate by applying a linear filter having any suitable filter shape to the luma component (or luma channel) to refine each chroma component (e.g., first chroma component, second chroma component).
[0133] FIG. 16 illustrates an example of a filter (1600) according to one embodiment of the present disclosure. The filter (1600) may include non-zero and zero filter coefficients. The filter (1600) has a diamond shape (1620) formed by filter coefficients (1610) (shown as solid circles). In one example, non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and filter coefficients not included in the filter coefficients (1610) are zero. Thus, non-zero filter coefficients of the filter (1600) are included in the diamond shape (1620), and filter coefficients not included in the diamond shape (1620) are zero. In one example, the number of filter coefficients of the filter (1600) is equal to the number of filter coefficients (1610), which is 18 in the example shown in FIG. 16.
[0134] The CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Referring back to Figure 15, the CC-ALF (1521) and the CC-ALF (1531) may have the same filter shape, such as the diamond shape (1620) shown in Figure 16, and the same number of filter coefficients. In one example, the values of the filter coefficients in the CC-ALF (1521) are different from the values of the filter coefficients in the CC-ALF (1531).
[0135] In general, the filter coefficients (e.g., non-zero filter coefficients) of CC-ALF may be transmitted, for example, in APS. In one example, the filter coefficients may be transmitted as coefficients (e.g., 2 10) and may be rounded for fixed-point representation. The application of CC-ALF is controlled by variable block sizes and may be signaled by a context-coded flag (e.g., a CC-ALF enable flag) received for each block of samples. Context-coded flags, such as the CC-ALF enable flag, may be signaled at any appropriate level, such as the block level. Block sizes along with CC-ALF enable flags may be received at the slice level for each chroma component. In some examples, block sizes (in chroma samples) of 16x16, 32x32, and 64x64 may be supported.
[0136] 17 shows an example syntax of CC-ALF according to some embodiments of the present disclosure. In the example of FIG. 17, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index indicating whether a cross-component Cb filter is used, and if so, the index of the cross-component Cb filter. For example, if alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the cross-component Cb filter is not applied to the block of Cb color component samples at the luma position (xCtb, yCtb). If alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not 0, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the filter to apply. For example, the alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]th cross-component Cb filter is applied to the block of Cb color component samples at luma position (xCtb, yCtb).
[0137] 17, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is used to indicate whether to use a cross-component Cr filter or a cross-component Cr filter index. For example, if alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the cross-component Cr filter is not applied to the block of Cr color component samples at the luma position (xCtb, yCtb). If alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not 0, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the cross-component Cr filter. For example, the alf_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]th cross-component Cr filter may be applied to a block of Cr color component samples at luma position (xCtb, yCtb).
[0138] In some examples, a chroma subsampling technique is used, so that the number of samples in each of the chroma blocks may be less than the number of samples in the luma blocks. The chroma subsampling format (also referred to as a chroma subsampling format, for example, specified by chroma_format_idc) may indicate a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) between each of the chroma blocks and the corresponding luma block. In one example, the chroma subsampling format is 4:2:0, so that the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown in Figures 18A-18B. In one example, the chroma subsampling format is 4:2:2, so that the chroma horizontal subsampling factor (e.g., SubWidthC) is 2 and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In one example, the chroma subsampling format is 4:4:4, and therefore the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample format (also referred to as the chroma sample position) may indicate the relative position of a chroma sample within a chroma block relative to at least one corresponding luma sample within the luma block.
[0139] 18A-18B illustrate exemplary locations of chroma samples relative to luma samples according to an embodiment of the present disclosure. Referring to FIG. 18A, luma samples (1801) are located in rows (1811)-(1818). The luma samples (1801) shown in FIG. 18A may represent a portion of a picture. In one example, a luma block (e.g., luma CB) includes the luma sample (1801). The luma block may correspond to two chroma blocks having a 4:2:0 chroma subsampling format. In one example, each chroma block includes a chroma sample (1803). Each chroma sample (e.g., chroma sample (1803(1))) corresponds to four luma samples (e.g., luma samples (1801(1)) to (1801(4))). In one example, the four luma samples are the top left sample (1801(1)), the top right sample (1801(2)), the bottom left sample (1801(3)), and the bottom right sample (1801(4)). A chroma sample (e.g., (1803(1))) corresponds to the top left sample (1801(1)) and the bottom left sample (18 The chroma sample format of a chroma block having chroma sample (1803) located at the left center position between the top left sample (1801(1)) and the bottom left sample (1801(3)) can be referred to as chroma sample format 0. Chroma sample format 0 indicates relative position 0, which corresponds to the left center position halfway between the top left sample (1801(1)) and the bottom left sample (1801(3)). The four luma samples (e.g., (1801(1)) through (1801(4))) can be referred to as neighboring luma samples of chroma sample (1803)(1).
[0140] In one example, each chroma block includes a chroma sample (1804). The above description of the chroma sample (1803) may be adapted to the chroma sample (1804), and therefore, detailed description may be omitted for brevity. Each of the chroma samples (1804) may be located at a central position of four corresponding luma samples, and the chroma sample format of a chroma block having the chroma sample (1804) may be referred to as chroma sample format 1. Chroma sample format 1 indicates relative position 1, which corresponds to the central position of the four luma samples (e.g., (1801(1)) to (1801(4))). For example, one of the chroma samples (1804) may be located at the central portion of the luma samples (1801(1)) to (1801(4)).
[0141] In one example, each chroma block includes chroma samples (1805). Each of the chroma samples (1805) may be located at the top left position, which is the same position as the top left sample of the four corresponding luma samples (1801). The chroma sample format of a chroma block having chroma samples (1805) may be referred to as chroma sample format 2. Therefore, each of the chroma samples (1805) is located at the same position as the top left sample of the four luma samples (1801) corresponding to the respective chroma sample. Chroma sample format 2 indicates relative position 2, which corresponds to the top left position of the four luma samples (1801). For example, one of the chroma samples (1805) may be located at the top left position of luma samples (1801(1)) to (1801(4)).
[0142] In one example, each chroma block includes chroma samples (1806). Each of the chroma samples (1806) may be located at a top center position between a corresponding top-left sample and a corresponding top-right sample, and the chroma sample format of the chroma block having the chroma samples (1806) may be referred to as chroma sample format 3. Chroma sample format 3 indicates relative position 3, which corresponds to the top center position between the top-left sample and the top-right sample. For example, one of the chroma samples (1806) may be located at a top center position of luma samples (1801(1)) to (1801(4)).
[0143] In one example, each chroma block includes a chroma sample (1807). Each of the chroma samples (1807) may be located at the bottom left position, which is the same position as the bottom left sample of the four corresponding luma samples (1801). The chroma sample format of the chroma block having the chroma sample (1807) may be referred to as chroma sample format 4. Therefore, each of the chroma samples (1807) is located at the same position as the bottom left sample of the four luma samples (1801) corresponding to the respective chroma sample. Chroma sample format 4 indicates relative position 4, which corresponds to the bottom left position of the four luma samples (1801). For example, one of the chroma samples (1807) may be located at the bottom left position of luma samples (1801(1)) to (1801(4)).
[0144] In one example, each chroma block includes chroma samples (1808). Each chroma sample (1808) is located at a bottom center position between the bottom left sample and the bottom right sample, and the chroma sample format of a chroma block having the chroma samples (1808) can be referred to as a chroma sample format of 5. The chroma sample format of 5 indicates a relative position of 5, which corresponds to a bottom center position between the bottom left sample and the bottom right sample of the four luma samples (1801). For example, one of the chroma samples (1808) can be located between the bottom left sample and the bottom right sample of luma samples (1801(1)) to (1801(4)).
[0145] In general, any suitable chroma sample format can be used for the chroma subsampling format. Chroma sample formats 0-5 are examples of chroma sample formats described in chroma subsampling format 4:2:0. Additional chroma sample formats can be used for chroma subsampling format 4:2:0. Furthermore, other chroma sample formats and / or variations of chroma sample formats 0-5 can be used for other chroma subsampling formats, such as 4:2:2, 4:4:4, etc. In one example, a chroma sample format combining chroma samples (1805) and (1807) is used for chroma subsampling format 4:2:2.
[0146] In one example, a luma block may be considered to have alternating rows, such as rows 1811-1812, each containing the top two samples (e.g., 1801(1)-1801(2)) of the four luma samples (e.g., 1801(1)-1801(4))) and the bottom two samples (e.g., 1801(3)-1801(4)) of the four luma samples (e.g., 1801(1)-1801(4)). Thus, rows 1811, 1813, 1815, and 1817 may be considered to have alternating rows, such as rows 1818-1819, 1819-1820, 1810-1821, 1811-1822, 1812-1823, 1813-1824, 1814-1825, 1815-1826, 1816-1827, 1817-1828, 1818-1829, 1819-1900, 1900-2001, 1900-2102, 1900-2110, 1900-2120, 1900-2131, 1900-2142, 1900-2153, 1900-2161, 1900-2172, 1900-2183, 1900-2184, 1900-2195, 1900-2196, 1900-2197, 1900-2198, 1900-2199, 1910-2199, 191 ) can be referred to as the current row (also called the top field), and rows (1812), (1814), (1816), and (1818) can be referred to as the next row (also called the bottom field). Four luma samples (e.g., (1801(1)) to (1801(4))) are located in the current row (e.g., (1811)) and the next row (e.g., (1812)). Relative positions 2 to 3 are located in the current row, relative positions 0 to 1 are located between each current row and the respective next row, and relative positions 4 to 5 are located in the next row.
[0147] Chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 are located in rows 1851-1854 within each chroma block. The specific locations of rows 1851-1854 may depend on the chroma sample format of the chroma samples. For example, for chroma samples 1803-1804 having chroma sample formats 0-1, respectively, row 1851 is located between rows 1811-1812. For chroma samples 1805-1806 having chroma sample formats 2-3, respectively, row 1851 is in the same position as the current row 1811. For chroma samples 1807-1808 having respective chroma sample types 4-5, row 1851 is in the same position as the next row 1812. The above description can be adapted appropriately for rows 1852-1854, and a detailed description will be omitted for the sake of brevity.
[0148] Any suitable scanning method may be used to display, store, and / or transmit the luma blocks and corresponding chroma blocks described above in Figure 18A. In one example, progressive scanning is used.
[0149] Interlaced scanning can be used, as shown in Figure 18B. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc equals 1). In one example, the variable chroma location format (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current rows (1811), (1813), (1815), and (1817) and the next rows (1812), (1814), (1816), and (1818) can be scanned separately, for example, the current rows (1811), (1813), (1815), and (1817) can be scanned first, followed by the next rows (1812), (1814), (1816), and (1818). The current row can include luma sample (1801), and the next row can include luma sample (1802).
[0150] Similarly, corresponding chroma blocks can be interlaced. Rows 1851 and 1853, which contain chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 with no fill, can be referred to as the current row (or current chroma row), and rows 1852 and 1854, which contain chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 with gray fill, can be referred to as the next row (or next chroma row). In one example, during interlaced scanning, rows 1851 and 1853 are scanned first, followed by rows 1852 and 1854.
[0151] In some examples, constrained directional enhancement filtering techniques can be used. The use of an in-loop constrained directional enhancement filter (CDEF) is to remove coding artifacts while preserving image details. In one example (e.g., HEVC), the sample adaptive offset (SAO) algorithm can achieve a similar goal by defining signal offsets for different classes of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, CDEF can be constrained to be easily vectorizable (i.e., executable with single instruction multiple data (SIMD) operations). Note that other nonlinear filters, such as median filters and bilateral filters, cannot be treated similarly.
[0152] In some cases, the amount of ringing artifacts in a coded image tends to be roughly proportional to the quantization step size. Although the level of detail is a property of the input image, the minimum level of detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of the ringing will generally be smaller than the amplitude of the detail.
[0153] The CDEF can be used to identify the orientation of each block and adaptively filter along the identified orientation and along orientations rotated 45 degrees from the identified orientation at smaller angles. In some examples, the encoder can look up the filter strength or signal the filter strength explicitly, allowing for a high degree of control over blurring.
[0154] Specifically, in some examples, the direction search is performed on the reconstructed pixels immediately after the deblocking filter. Because these pixels are available to the decoder, the direction can be searched by the decoder; therefore, in one example, the direction does not require signaling. In some examples, the direction search can operate on a specific block size, such as an 8x8 block, that is small enough to properly handle non-linear edges but large enough to reliably estimate the direction when applied to the quantized image. Also, imposing a certain directionality on the 8x8 region facilitates vectorization of the filter. In some examples, each block (e.g., 8x8) can be compared with a fully directional block to determine the difference. A fully directional block is a block in which all pixels along a line in a certain direction have the same value. In one example, a difference measure, such as the sum of squared differences (SSD) or root mean square (RMS) error, can be calculated for the block and each fully directional block. The fully directional block with the smallest difference (e.g., minimum SSD, minimum RMS, etc.) can then be determined, and the direction of the determined fully directional block can be the direction that best matches the pattern within the block.
[0155] FIG. 19 shows an example of direction search according to an embodiment of the present disclosure. In one example, block (1910) is an 8x8 block that is reconstructed and output from the deblocking filter. In the example of FIG. 19, the direction search can determine a direction from eight directions indicated by (1920) for block (1910). Eight fully directional blocks (1930) are formed corresponding to the eight directions (1920), respectively. A fully directional block corresponding to a certain direction is a block in which pixels along the directional line have the same value. Furthermore, difference measures such as SSD and RMS error can be calculated for block (1910) and each fully directional block (1930). In the example of FIG. 19, the RMS error is indicated by (1940). As indicated by (1943), the RMS error between block (1910) and fully directional block (1933) is the smallest, and therefore direction (1923) is the direction that best matches the pattern of block (1910).
[0156] After the block direction is identified, a nonlinear low-pass directional filter can be determined. For example, the filter taps of the nonlinear low-pass directional filter can be aligned along the identified direction to reduce ringing while preserving the directional edge or pattern. However, in some instances, directional filtering alone is not sufficient to reduce ringing. In one example, additional filter taps are used for pixels that are not aligned with the identified direction. The extra filter taps are treated more conservatively to reduce the risk of blurring. Thus, the CDEF includes first-order and second-order filter taps. In one example, the complete 2D CDEF filter can be expressed as equation (14):
number
[0157] In some examples, in-loop reconstruction schemes are used in post-deblocking video coding, in addition to the deblocking operation, generally to remove noise and improve edge quality. In one example, the in-loop reconstruction schemes are switchable within a frame for tiles of appropriate size. The in-loop reconstruction schemes are based on separable symmetric Wiener filters, dual self-guided filters with subspace projection, and domain-transformed recursive filters. Because content statistics can change substantially within a frame, the in-loop reconstruction schemes are integrated into a switchable framework that can trigger different schemes in different regions of a frame.
[0158] A separable symmetric Wiener filter can be one of the in-loop restoration methods. In some instances, every pixel of the degraded frame can be reconstructed as a non-causal filtered version of the pixels in a w × w window around it, where w = 2r + 1 is odd for integer r. The two-dimensional filter taps are expressed as w in column vector form. 2 If the vector is represented by a 1 × 1 element vector F, then direct LMMSE optimization gives F=H -1 The filter parameters are derived given by M, where H=E[XX T ] is the autocovariance of x and w in a w × w window around the pixel 2 is a column-wise vectorized version of the samples of M=E[YX T ] is the cross-correlation between x and the scalar source sample y to be estimated. In one example, the encoder can estimate H and M from a realization of the deblocked frame and the source, and send the resulting filter F to the decoder. However, doing so would require that w 2Not only does transmitting the taps incur a significant bitrate cost, but non-separable filtering significantly complicates decoding. In some embodiments, several additional constraints are imposed on the nature of F. For the first constraint, F is constrained to be separable, and filtering can be implemented as a separable horizontal and vertical w-tap convolution. For the second constraint, each of the horizontal and vertical filters is constrained to be symmetric. For the third constraint, both the horizontal and vertical filter coefficients are assumed to sum to one.
[0159] Doubly self-guided filtering using subspace projection can be one of the in-loop restoration methods. Guided filtering is an image filtering technique that uses a local linear model shown by equation (15): y=Fx+G Equation (15) The above formula is used to calculate the filtered output y from the unfiltered sample x. Here, F and G are determined based on the statistics of the guidance image in the neighborhood of the degraded image and the filtered pixel. If the guidance image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. In one example, a specific form of self-guided filtering can be used. The specific form of self-guided filtering depends on two parameters: the radius r and the noise parameter e, and is listed as the following steps: 1. The mean μ and variance σ of the pixels in a (2r+1) × (2r+1) window around every pixel 2 This step can be efficiently implemented with box filtering based on integral imaging. 2. Calculate for all pixels: f = σ 2 / (σ 2 +e);g=(1-f)μ 3. Calculate F and G for every pixel as the average of the f and g values in a 3x3 window around the pixel being used.
[0160] The specific form of the self-guided filter is controlled by r and e, with larger r resulting in larger spatial variance and larger e resulting in larger range variance.
[0161] Figure 20 shows an example of subspace projection in some examples. As shown in Figure 20, even if neither of the inexpensive reconstructions X1, X2 is close to the source Y, a suitable multiplier {α, β} can make them quite close to the source Y as long as they are somewhat shifted in the right direction.
[0162] In some examples (e.g., HEVC), a filtering technique called sample adaptive offset (SAO) can be used. In some examples, SAO is applied to the reconstructed signal after the deblocking filter. SAO can use an offset value provided in the slice header. In some examples, for luma samples, the encoder can determine whether to apply (enable) SAO to the slice. When SAO is enabled, the current picture allows for recursive division of the coding unit into four sub-regions, and each sub-region can select an SAO type from multiple SAO types based on features within the sub-region.
[0163] FIG. 21 shows a table (2100) of multiple SAO types according to one embodiment of the present disclosure. Table (2100) shows SAO types 0 to 6. Note that SAO type 0 is used to indicate no SAO application. Furthermore, each SAO type, SAO type 1 to SAO type 6, includes multiple categories. SAO can reduce distortion by classifying reconstructed pixels in subregions into categories and adding offsets to pixels in each category within the subregion. In some examples, edge characteristics can be used to classify pixels in SAO types 1 to 4, and pixel intensity can be used to classify pixels in SAO types 5 and 6.
[0164] Specifically, in one embodiment, such as SAO Types 5-6, band offsets (BOs) can be used to classify all pixels in a subregion into multiple bands. Each band of the multiple bands contains pixels in the same intensity interval. In some examples, the intensity range is equally divided into multiple intervals, such as 32 intervals ranging from 0 to the maximum intensity value (e.g., 255 for 8-bit pixels), with each interval associated with an offset. Furthermore, in one example, the 32 bands are divided into two groups, such as a first group and a second group. The first group contains the middle 16 bands (e.g., 16 intervals in the middle of the intensity range), and the second group contains the remaining 16 bands (e.g., 8 intervals at the low end of the intensity range and 8 intervals at the high end of the intensity range). In one example, only the offset of one of the two groups is transmitted. In some embodiments, when pixel classification operations in BO are used, the five most significant bits of each pixel can be used directly as a band index.
[0165] Additionally, in one embodiment, such as SAO Types 1-4, edge offset (EO) can be used to determine pixel classification and offset. For example, pixel classification can be determined based on a one-dimensional three-pixel pattern taking into account edge direction information.
[0166] Figure 22 shows an example of a three-pixel pattern for pixel classification at edge offset in some examples. In the example of Figure 22, the first pattern (2210) (shown by three gray pixels) is referred to as a 0-degree pattern (the 0-degree pattern is associated with a horizontal direction), the second pattern (2220) (shown by three gray pixels) is referred to as a 90-degree pattern (the 90-degree pattern is associated with a vertical direction), the third pattern (2230) (shown by three gray pixels) is referred to as a 135-degree pattern (the 135-degree pattern is associated with a 135-degree diagonal direction), and the fourth pattern (2240) (shown by three gray pixels) is referred to as a 45-degree pattern (the 45-degree pattern is associated with a 45-degree diagonal direction). In one example, one of the four direction patterns shown in Figure 22 can be selected taking into account edge direction information of the subregion. The selection can be transmitted in the video bitstream, in one example, coded as side information. The pixels within the sub-region can then be classified into multiple categories by comparing each pixel with two neighboring pixels in a direction associated with the directional pattern.
[0167] Figure 23 shows a table (2300) of pixel classification rules for edge offsets in some examples. Specifically, pixel c (shown in each pattern in Figure 22) is compared with two neighboring pixels (shown in gray in each pattern in Figure 22), and pixel c can be classified into one of categories 0 to 4 based on the comparison using the pixel classification rules shown in Figure 23.
[0168] In some embodiments, decoder-side SAO can operate independently of the largest coding unit (LCU) (e.g., CTU) to conserve line buffers. In some examples, the top and bottom pixels within each LCU are not SAO processed when the 90-degree, 135-degree, and 45-degree classification patterns are selected. The leftmost and rightmost columns of pixels within each LCU are not SAO processed when the 0-degree, 135-degree, and 45-degree patterns are selected.
[0169] FIG. 24 shows an example (2400) of syntax that may need to be signaled for a CTU when parameters are not merged from neighboring CTUs. For example, the syntax element sao_type_idx[cldx][rx][ry] can be signaled to indicate the SAO type of the subregion. The SAO type may be BO (Band Offset) or EO (Edge Offset). A value of 0 for sao_type_idx[cldx][rx][ry] indicates that SAO is OFF, values of 1 to 4 indicate that one of four EO categories corresponding to 0°, 90°, 135°, and 45° is used, and a value of 5 indicates that BO is used. In the example of FIG. 24, each of the BO and EO types has four SAO offset values that are signaled (sao_offset[cIdx][rx][ry][0] to sao_offset[cIdx][rx][ry][3]).
[0170] In general, the filtering process can use reconstructed samples of a first color component as an input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component, which can be the same as the first color component or another color component different from the first color component.
[0171] In a related example of cross-component filtering (CCF), filter coefficients are derived based on several mathematical equations. The derived filter coefficients are signaled from the encoder side to the decoder side, and the derived filter coefficients are used to generate offsets using a linear combination. The generated offsets are then added to reconstructed samples as a filtering process. For example, offsets are generated based on a linear combination of the filtering coefficients and luma samples, and the generated offsets are added to reconstructed chroma samples. A related example of CCF is based on the assumption of a linear mapping relationship between reconstructed luma sample values and delta values between the original chroma samples and the reconstructed chroma samples. However, the mapping between the reconstructed luma sample values and delta values between the original chroma samples and the reconstructed chroma samples does not necessarily follow a linear mapping process. Therefore, the coding performance of CCF may be limited under the assumption of a linear mapping relationship.
[0172] In some examples, nonlinear mapping techniques may be used for cross-component filtering and / or same-color component filtering without significant signaling overhead. In one example, nonlinear mapping techniques may be used in cross-component filtering to generate cross-component sample offsets. In another example, nonlinear mapping techniques may be used in same-color component filtering to generate local sample offsets.
[0173] For convenience, a filtering process using a nonlinear mapping technique can be referred to as sample offset with nonlinear mapping (SO-NLM). SO-NLM in the cross-component filtering process can be referred to as cross-component sample offset (CCSO). SO-NLM in same-color component filtering can be referred to as local sample offset (LSO). A filter using a nonlinear mapping technique can be referred to as a nonlinear mapping-based filter. Nonlinear mapping-based filters can include CCSO filters, LSO filters, etc.
[0174] In one example, CCSO and LSO can be used as loop filtering to reduce distortion of reconstructed samples. CCSO and LSO do not rely on the assumption of a linear mapping used in the associated exemplary CCF. For example, CCSO does not rely on the assumption of a linear mapping relationship between luma reconstructed sample values and delta values between original chroma samples and the chroma reconstructed samples. Similarly, LSO does not rely on the assumption of a linear mapping relationship between reconstructed sample values of color components and delta values between original samples of the color components and the reconstructed samples of the color components.
[0175] The following description describes an SO-NLM filtering process that uses reconstructed samples of a first color component (e.g., Y or Cb or Cr, or R or G or B) as input to generate an output, and the output of the filtering process is applied to a second color component. If the second color component is the same color component as the first color component, the description is applicable to LSO, and if the second color component is different from the first color component, the description is applicable to CCSO.
[0176] In SO-NLM, a nonlinear mapping is derived on the encoder side. The nonlinear mapping is between the reconstructed samples of the first color component within the filter support region and the offset added to the second color component within the filter support region. If the second color component is the same as the first color component, the nonlinear mapping is used in LSO; if the second color component is different from the first color component, the nonlinear mapping is used in CCSO. The domain of the nonlinear mapping is determined by the different combinations of reconstructed samples of the processed input (also called possible reconstructed sample value combinations).
[0177] The SO-NLM technique can be illustrated using a specific example. In the specific example, a reconstruction sample from a first color component located in a filter support area (also called a "filter support region") is determined. The filter support area is an area to which a filter can be applied, and the filter support area can have any suitable shape.
[0178] FIG. 25 illustrates an example of a filter support area (2500) according to some embodiments of the present disclosure. The filter support area (2500) includes four reconstructed samples of a first color component: P0, P1, P2, and P3. In the example of FIG. 25, the four reconstructed samples may form a cross shape in the vertical and horizontal directions, and the center of the cross is the location of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is indicated by C. The sample of the second color component at the center position is indicated by F. The second color component may be the same as the first color component P0-P3 or may be different from the first color component P0-P3.
[0179] FIG. 26 shows an example of another filter support area (2600) according to some embodiments of the present disclosure. The filter support area (2600) includes four reconstructed samples P0, P1, P2, and P3 of a first color component that form a square. In the example of FIG. 26, the center position of the square is the position of the sample to be filtered. The sample of the same color component as P0 to P3 at the center position is indicated by C. The sample of the second color component at the center position is indicated by F. The second color component may be the same as the first color component P0 to P3 or may be different from the first color component P0 to P3.
[0180] The reconstructed samples are input to the SO-NLM filter and processed appropriately to form the filter taps. In one example, the positions of the reconstructed samples that are input to the SO-NLM filter are called filter tap positions. In a particular example, the reconstructed samples are processed in two steps:
[0181] In the first step, delta values are calculated between P0 to P3 and C. For example, m0 indicates the delta value between P0 and C, m1 indicates the delta value between P1 and C, m2 indicates the delta value between P2 and C, and m3 indicates the delta value between P3 and C.
[0182] In a second step, the delta values m0-m3 are further quantized, and the quantized values are denoted as d0, d1, d2, and d3. In one example, the quantized value may be one of −1, 0, or 1 based on the quantization process. For example, if m is less than −N (N is a positive value and is called the quantization step size), the value m may be quantized to −1; if m is in the range of [−N,N], the value m may be quantized to 0; and if m is greater than N, the value m may be quantized to 1. In some examples, the quantization step size N may be one of 4, 8, 12, 16, etc.
[0183] In some embodiments, the quantized values d0-d3 are filter taps and may be used to identify one combination within the filter domain. For example, the filter taps d0-d3 may form a combination in the filter domain. Each filter tap may have three quantized values, so if four filter taps are used, the filter domain contains 81 (3 x 3 x 3 x 3) combinations.
[0184] 27A-27C show a table 2700 having 81 combinations according to one embodiment of the present disclosure. The table 2700 includes 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, the first column includes an index of the combination, the second column includes a value of the filter tap d0 for the combination, the third column includes a value of the filter tap d1 for the combination, the fourth column includes a value of the filter tap d2 for the combination, the fifth column includes a value of the filter tap d3 for the combination, and the sixth column includes an offset value associated with the combination for nonlinear mapping. In one example, once the filter taps d0-d3 are determined, an offset value (represented by s) associated with the combination of d0-d3 can be determined according to the table 2700. In one example, the offset values s0-s80 are integers such as 0, 1, -1, 3, -3, 5, -5, and -7.
[0185] In some embodiments, the final filtering process of the SO-NLM may be applied as shown in equation (16): f'=clip(f+s) Equation (16) where f is the reconstructed sample of the second color component to be filtered, and s is an offset value determined according to the filter taps that result from processing the reconstructed sample of the first color component, such as using table (2700). The sum of the reconstructed sample F and the offset value s is further clipped to a range associated with the bit depth to determine the final filtered sample f' of the second color component.
[0186] It should be noted that in the case of LSO, the second color component in the above description is the same as the first color component, and in the case of CCSO, the second color component in the above description may be different from the first color component.
[0187] It should be noted that the above description may be adjusted for other embodiments of the present disclosure.
[0188] In some examples, at the encoder side, the encoding device may derive a mapping between the reconstructed samples of the first color component within the filter support region and an offset to be added to the reconstructed samples of the second color component. The mapping may be any suitable linear or nonlinear mapping. Then, a filtering process may be applied at the encoder side and / or the decoder side based on the mapping. For example, the mapping may be appropriately notified to the decoder (e.g., the mapping may be included in the coded video bitstream transmitted from the encoder side to the decoder side), and then the decoder may perform the filtering process based on the mapping.
[0189] According to some aspects of the present disclosure, the performance of a nonlinear mapping-based filter, such as a CCSO filter or an LSO filter, depends on the filter shape configuration. The filter shape configuration (also referred to as filter shape) of a filter may refer to the characteristics of the pattern formed by the filter tap positions. The pattern may be defined by various parameters such as the number of filter taps, the geometric shape of the filter tap positions, and the distance of the filter tap positions to the center of the pattern. Using a fixed filter shape configuration may limit the performance of the nonlinear mapping-based filter.
[0190] As shown in Figures 24 and 25 and Figures 27A to 27C, some examples use a 5-tap filter design for filter shape configuration of a nonlinear mapping-based filter. The 5-tap filter design may use tap positions P0, P1, P2, P3, and C. The 5-tap filter design for filter shape configuration may result in a look-up table (LUT) with 81 entries, as shown in Figures 27A to 27C. The LUT for sample offsets needs to be signaled from the encoder side to the decoder side, and signaling the LUT may contribute to a large portion of the signaling overhead and affect coding efficiency using a nonlinear mapping-based filter. According to some aspects of the present disclosure, the number of filter taps may be different from 5. In some examples, the number of filter taps may be reduced and still capture information within the filter support area, improving coding efficiency.
[0191] According to one aspect of the present disclosure, the number of filter taps of a nonlinear mapping-based filter, such as a CCSO filter or an LSO filter, can be any integer from 1 to M, where M is an integer. In some examples, the value of M is 1024 or any other suitable number.
[0192] In some examples, the number of filter taps of the nonlinear mapping-based filter is three. In one example, the positions of the three filter taps include a center position. In another example, the positions of the three filter taps exclude a center position. The center position refers to the position of the reconstructed sample to be filtered.
[0193] In some examples, the number of filter taps of the nonlinear mapping-based filter is five. In one example, the five filter tap positions include a center position. In another example, the five filter tap positions exclude a center position. The center position refers to the position of the reconstructed sample to be filtered.
[0194] In some examples, the number of filter taps of the nonlinear mapping-based filter is 1. In one example, the position of one filter tap is a central position. In another example, the position of one filter tap is not a central position. The central position refers to the position of the reconstructed sample to be filtered. In one example, the CCSO filter has a 1-tap filter design, and the delta value of the CCSO filter can be calculated as (p-μ), where p is the reconstructed sample value located at the filter tap and μ is the average sample value within a given area s. s can be a coding block, a CTU / SB, or a picture.
[0195] In some embodiments, the filter shape configuration of the nonlinear mapping-based filter is switchable between groups of filter shape configurations. In some examples, the groups of filter shape configurations may be candidates for the nonlinear mapping-based filter. During encoding / decoding, the filter shape configuration of the nonlinear mapping-based filter may change from one of the filter shape configurations in the group to another of the filter shape configurations in the group.
[0196] According to one aspect of the present disclosure, the filter shape configurations within a group of nonlinear mapping-based filters may have the same number of filter taps.
[0197] In some examples, the filter shape configurations in the group of nonlinear mapping-based filters each have three filter taps.
[0198] FIG. 28 illustrates eight filter shape configurations of three filter taps in one example. Specifically, a first filter shape configuration includes three filter taps at positions labeled "1" and "C," with position "C" being the center position of position "1." A second filter shape configuration includes three filter taps at positions labeled "2" and "C," with position "C" being the center position of position "2." A third filter shape configuration includes three filter taps at positions labeled "3" and "C," with position "C" being the center position of position "3." A fourth filter shape configuration includes three filter taps at positions labeled "4" and "C," with position "C" being the center position of position "4." A fifth filter shape configuration includes three filter taps at positions labeled "5" and "C," with position "C" being the center position of position "5." The sixth filter shape configuration includes three filter taps at positions labeled "6" and position "C," with position "C" being the center position of position "6." The seventh filter shape configuration includes three filter taps at positions labeled "7" and position "C," with position "C" being the center position of position "7." The eighth filter shape configuration includes three filter taps at positions labeled "8" and position "C," with position "C" being the center position of position "8."
[0199] In one example, eight filter shape configurations are candidates for a nonlinear mapping-based filter, which can switch from one of the eight filter shape configurations to another of the eight filter shape configurations during encoding / decoding.
[0200] FIG. 29 illustrates 12 filter shape configurations of three filter taps in one example. Specifically, a first filter shape configuration includes three filter taps at positions labeled "1" and "C," with position "C" being the center position of position "1." A second filter shape configuration includes three filter taps at positions labeled "2" and "C," with position "C" being the center position of position "2." A third filter shape configuration includes three filter taps at positions labeled "3" and "C," with position "C" being the center position of position "3." A fourth filter shape configuration includes three filter taps at positions labeled "4" and "C," with position "C" being the center position of position "4." A fifth filter shape configuration includes three filter taps at positions labeled "5" and "C," with position "C" being the center position of position "5." The sixth filter shape configuration includes three filter taps at positions labeled "6" and position "C," with position "C" being the center position of position "6." The seventh filter shape configuration includes three filter taps at positions labeled "7" and position "C," with position "C" being the center position of position "7." The eighth filter shape configuration includes three filter taps at positions labeled "8" and position "C," with position "C" being the center position of position "8." The ninth filter shape configuration includes three filter taps at positions labeled "9" and "C," with position "C" being the center position of position "9." The tenth filter shape configuration includes three filter taps at positions labeled "10" and position "C," with position "C" being the center position of position "10." The eleventh filter shape configuration includes three filter taps at positions labeled "11" and position "C," with position "C" being the center position of position "11." The twelfth filter configuration includes three filter taps at positions labeled "12" and position "C," with position "C" being the center position of position "12."
[0201] In one example, 12 filter shape configurations are candidates for a nonlinear mapping-based filter, which can switch from one of the 12 filter shape configurations to another of the 12 filter shape configurations during encoding / decoding.
[0202] According to another aspect of the present disclosure, the filter shape configurations within a group of nonlinear mapping-based filters can have different numbers of filter taps.
[0203] 30 shows an example of two candidate filter shape configurations for a nonlinear mapping-based filter. For example, a first of the two candidate filter shape configurations includes three filter taps, and a second of the two candidate filter shape configurations includes five filter taps. Specifically, the first filter shape configuration includes three filter taps at positions labeled "p0," "C," and "p2" (e.g., at the positions of the dashed circles), and the second filter shape configuration includes five filter taps at positions labeled "p0," "p1," "p2," "p3," and "C."
[0204] In one example, two filter shape configurations are candidates for a nonlinear mapping-based filter, which can switch from one of the two filter shape configurations to another of the two filter shape configurations during encoding / decoding.
[0205] It should be noted that the filter shape configuration of the nonlinear mapping-based filter can be switched at various levels. In one example, the filter shape configuration of the nonlinear mapping-based filter can be switched at a sequence level. In another example, the filter shape configuration of the nonlinear mapping-based filter can be switched at a picture level. In another example, the filter shape configuration of the nonlinear mapping-based filter can be switched at a CTU level or a superblock (SB) level. A superblock is the largest coding block in some examples. In another example, the filter shape configuration of the nonlinear mapping-based filter can be switched at a coding block level (e.g., a CU level).
[0206] In some examples, a selection of a filter shape configuration from a group of filter shape configurations is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected filter shape configuration is signaled in a high-level syntax (HLS), such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a frame header.
[0207] It should be noted that the above description uses a nonlinear mapping-based filter to illustrate the technique for switching filter shape configurations, and the technique for switching filter shape configurations may be applied to other loop filters, including, but not limited to, a cross-component adaptive loop filter, an adaptive loop filter, a loop restoration filter, etc. For example, a loop filter may have a group of candidate filter shape configurations. During encoding / decoding, the loop filter may switch its filter shape configuration from one of the candidate filter shape configurations to another of the candidate filter shape configurations. The candidate filter shape configurations within a group may have different numbers of filter taps and / or different relative positions with respect to the samples to be filtered. It should be noted that the filter shape configuration of a nonlinear mapping-based filter may be switched at various levels. In one example, the filter shape configuration of a loop filter may be switched at a sequence level. In another example, the filter shape configuration of a loop filter may be switched at a picture level. In another example, the filter shape configuration of a loop filter may be switched at a CTU level or a superblock (SB) level. In another example, the filter shape configuration of a loop filter may be switched at a coding block level (e.g., a CU level). In some examples, a selection of a filter shape configuration from a group of filter shape configurations is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected filter shape configuration is signaled in a high-level syntax (HLS), such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a frame header.
[0208] According to another aspect of the present disclosure, quantization is used during the filtering process of the nonlinear mapping-based filter. For example, in the second step of SO-NLM, delta values m0-m3 are quantized to determine d0, d1, d2, and d3, and each of d0, d1, d2, and d3 may be one of three quantized outputs (e.g., −1, 0, and 1). It should be noted that in some examples, the number of quantized outputs of the nonlinear mapping-based filter may be any suitable integer. In one example, the number of quantized outputs of the nonlinear mapping-based filter may be a number from 1 to an upper limit, such as 1024. Note that the upper limit is not limited to 1024.
[0209] In some examples, the number of quantized outputs is 3, and the quantized outputs can be represented as -1, 0, and 1.
[0210] In some examples, the number of quantized outputs is 5, and the quantized outputs can be represented as -2, -1, 0, 1, 2.
[0211] According to one aspect of the present disclosure, the number of quantization outputs of a nonlinear mapping-based filter can be switched at various levels (e.g., changed from one integer to another) during encoding / decoding. In one example, the number of quantization outputs of a nonlinear mapping-based filter can be switched at a sequence level. In another example, the number of quantization outputs of a nonlinear mapping-based filter can be switched at a picture level. In another example, the number of quantization outputs of a nonlinear mapping-based filter can be switched at a CTU level or a superblock (SB) level. In another example, the number of quantization outputs of a nonlinear mapping-based filter can be switched at a coding block level (e.g., a CU level). In some examples, the selection of the number of quantization outputs is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected number of quantization outputs is signaled in a high-level syntax (HLS), such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a frame header.
[0212] FIG. 31 shows a flowchart outlining a process (3100) according to one embodiment of the present disclosure. The process (3100) may be used to reconstruct video carried in a coded video bitstream. When the term block is used, the block may be interpreted as a prediction block, a coding unit, a luma block, a chroma block, or the like. In various embodiments, the process (3100) is performed by a processing circuit, such as the processing circuitry of the terminal devices (310), (320), (330), and (340), a processing circuit performing the functions of the video encoder (403), a processing circuit performing the functions of the video decoder (410), a processing circuit performing the functions of the video decoder (510), or a processing circuit performing the functions of the video encoder (603). In some embodiments, the process (3100) is implemented by software instructions, and thus, the processing circuit performs the process (3100) when the processing circuit executes the software instructions. The process starts at (S3101) and proceeds to (S3110).
[0213] At (S3110), an offset value associated with a first filter shape configuration of the nonlinear mapping-based filter is determined based on a signal in a coded video bitstream carrying video, wherein the number of filter taps of the first filter shape configuration is less than five.
[0214] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered. In another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.
[0215] In some examples, the first filter shape configuration includes a single filter tap, and calculates an average sample value in an area and a difference between the reconstructed sample value at the single filter tap position and the average sample value in the area, and a nonlinear mapping-based filter is applied to the samples to be filtered based on the difference between the reconstructed sample value at the single filter tap position and the average sample value in the area.
[0216] In some examples, the first filter shape configuration is selected from a group of filter shape configurations of the nonlinear mapping-based filter. In one example, the filter shape configurations in the group each have the same number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration.
[0217] In some examples, an index is decoded from a coded video bitstream carrying video. The index indicates a selection of a first filter shape configuration from a group of nonlinear mapping-based filters. The index may be decoded from syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.
[0218] At (S3120), a nonlinear mapping-based filter is applied to the samples to be filtered using an offset value associated with the first filter shape configuration.
[0219] In some examples, to apply a nonlinear mapping-based filter, a delta value of sample values at two filter tap positions is quantized to one of a number of possible quantization outputs. The number of possible quantization outputs is an integer ranging from 1 to 1024, inclusive. For example, an index is decoded from syntax signaling of at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, and the index indicates the number of possible quantization outputs.
[0220] The process (3100) proceeds to (S3199) and ends.
[0221] It should be noted that in some examples, the nonlinear mapping-based filter is a cross-component sample offset (CCSO) filter, and in some other examples, the nonlinear mapping-based filter is a local sample offset (LSO) filter.
[0222] The process (3100) may be adapted as appropriate. Step(s) of the process (3100) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.
[0223] FIG. 32 shows a flowchart outlining a process (3200) according to one embodiment of the present disclosure. The process (3200) may be used to encode video in a coded video bitstream. When the term block is used, the block may be interpreted as a prediction block, a coding unit, a luma block, a chroma block, or the like. In various embodiments, the process (3200) is performed by a processing circuit, such as the processing circuitry of the terminal devices (310), (320), (330), and (340), a processing circuit performing the functions of the video encoder (403), or a processing circuit performing the functions of the video encoder (603). In some embodiments, the process (3200) is implemented by software instructions, and thus, the processing circuit performs the process (3200) when the processing circuit executes the software instructions. The process starts at (S3201) and proceeds to (S3210).
[0224] At (S3210), a nonlinear mapping-based filter is applied to samples to be filtered in the video using offset values associated with a first filter shape configuration, the number of filter taps of the first filter shape configuration being less than five.
[0225] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered. In another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.
[0226] In some examples, the first filter shape configuration includes a single filter tap, and calculates an average sample value in an area and a difference between the reconstructed sample value at the single filter tap position and the average sample value in the area, and a nonlinear mapping-based filter is applied to the samples to be filtered based on the difference between the reconstructed sample value at the single filter tap position and the average sample value in the area.
[0227] In some examples, the first filter shape configuration is selected from a group of filter shape configurations of the nonlinear mapping-based filter. In one example, the filter shape configurations in the group each have the same number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration.
[0228] In some examples, an index is encoded in a coded video bitstream carrying video. The index indicates a selection of a first filter shape configuration from a group of nonlinear mapping-based filters. The index may be signaled by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.
[0229] In some examples, to apply a nonlinear mapping-based filter, a delta value of sample values at two filter tap positions is quantized to one of a number of possible quantization outputs. The number of possible quantization outputs is an integer ranging from 1 to 1024, inclusive. For example, an index is encoded in the coded video bitstream by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, where the index indicates the number of possible quantization outputs.
[0230] At (S3220), the offset value is coded in a coded video bitstream that carries the video.
[0231] The process (3200) proceeds to (S3299) and ends.
[0232] It should be noted that in some examples, the nonlinear mapping-based filter is a cross-component sample offset (CCSO) filter, and in some other examples, the nonlinear mapping-based filter is a local sample offset (LSO) filter.
[0233] The process (3200) may be adapted as appropriate. Step(s) of the process (3200) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.
[0234] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0235] The techniques described above may be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 33 illustrates a computer system (3300) suitable for implementing certain embodiments of the disclosed subject matter.
[0236] Computer software can be coded using any suitable machine code or computer language that is capable of undergoing mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. directly, or via interpretation, microcode execution, etc.
[0237] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0238] 33 for computer system (3300) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be construed as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (3300).
[0239] The computer system (3300) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users using, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0240] The input human interface devices may include one or more (only one of each shown) of a keyboard (3301), a mouse (3302), a trackpad (3303), a touch screen (3310), a data glove (not shown), a joystick (3305), a microphone (3306), a scanner (3307), and a camera (3308).
[0241] The computer system (3300) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (3310), data gloves (not shown), or joystick (3305), although some haptic feedback devices may not function as input devices), audio output devices (such as speakers (3309), headphones (not shown)), visual output devices (such as screens (3310) including CRT screens, LCD screens, plasma screens, and OLED screens (each with or without touchscreen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions by means of stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0242] The computer system (3300) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (3320) with CD / DVD or similar media (3321), thumb drives (3322), removable hard drives or solid state drives (3323), legacy magnetic media (not shown) such as tape and floppy disks, and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0243] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0244] The computer system (3300) may also include an interface (3354) to one or more communication networks (3355). The network may be, for example, wireless, wired, or optical. The network may further be local, wide area, metropolitan, vehicular, industrial, real-time, delay-tolerant, or the like. Examples of networks include local area networks such as Ethernet; cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, and the like; television wired or wireless wide area digital networks including cable, satellite, and terrestrial television; and vehicular and industrial networks including CANBus. Certain networks typically require an external network interface adapter connected to a particular general data port or peripheral bus (3349) (e.g., a USB port on the computer system 3300), while others are typically integrated into the core of the computer system (3300) by connection to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3300) can communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to a particular CANbus device), or bidirectional, e.g., to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.
[0245] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to the core (3340) of the computer system (3300).
[0246] The cores (3340) may include one or more central processing units (CPUs) (3341), graphics processing units (GPUs) (3342), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (3343), task-specific hardware accelerators (3344), graphics adapters (3350), etc. Such devices may be connected via a system bus (3348), along with read-only memory (ROM) (3345), random access memory (3346), and internal mass storage (3347), such as an internal non-user-accessible hard drive or SSD. In some computer systems, the system bus (3348) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (3348) or via a peripheral bus (3349). In one example, a display (3310) may be connected to the graphics adapter (3350). Peripheral bus architectures include PCI and USB.
[0247] The CPU (3341), GPU (3342), FPGA (3343), and accelerator (3344) can execute specific instructions that, in combination, may constitute the aforementioned computer code. This computer code may be stored in ROM (3345) or RAM (3346). Temporary data may also be stored in RAM (3346), while permanent data may be stored, for example, in internal mass storage (3347). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (3341), GPU (3342), mass storage (3347), ROM (3345), and RAM (3346), etc.
[0248] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0249] By way of example and not limitation, a computer system (3300) having the architecture, and in particular the core (3340), may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be the user-accessible mass storage introduced above, as well as media associated with specific storage of the core (3340) that is non-transitory, such as the core's internal mass storage (3347) or ROM (3345). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (3340). The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause the core (3340), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM (3346) and modifying such data structures according to software-defined processes. Additionally, or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (3344)) that may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software may encompass logic, and vice versa. Where appropriate, references to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding MPM: Most Probable Mode WAIP: Wide-angle Intra Prediction SEI: Supplemental Extended Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted Block HRD: Hypothetical Reference Decoder SDR: Standard Dynamic Range SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit PDPC: Position-dependent prediction combination ISP: Intra-subpartition SPS: Sequence parameter settings
[0250] While this disclosure has described several exemplary embodiments, there are modifications, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art can devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]
[0251] 101 Samples 102 Arrow 103 Arrow 104 Square Blocks 201 Current Block 202 Samples 203 Samples 204 Samples 205 samples 206 Samples 300 Communication Systems 310 Terminal Devices 320 terminal devices 330 Terminal Devices 340 Terminal Devices 350 Network 400 Communication Systems 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 Encoded Video Data, Encoded Video Bitstream 405 Streaming Server 406 Client Subsystem 407 encoded video data, input copy 408 Client Subsystem 409 Encoded video data, copy 410 Video Decoder 411 Video Picture Output Stream 412 Display 413 Capture Subsystem 420 Electronic Devices 430 Electronic Devices 501 Channel 510 Video Decoder 512 render device 515 buffer memory 520 Parser 521 Symbol 530 Electronic Devices 531 Receiver 551 Scaler / Descaler Unit 552 Intra-picture prediction unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Sources 603 Video Coder, Video Encoder 620 Electronic Devices 630 Source Coder 632 Coding Engine 633 Local Video Decoder 634 Reference Picture Memory, Reference Picture Cache 635 Predictors 640 Transmitter 643 coded video sequence 645 Entropy Coder 650 Controller 660 Communication Channels 703 Video Encoder 721 General Controller 722 Intra Encoder 723 Residual Calculation Unit 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 Interencoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 910 Diamond-shaped filter 911 Diamond-shaped filter 920~932 elements 940~964 elements 1110 Block 1111 Block 1120 Horizontal CTU Boundary 1121 CTU boundary 1130 Virtual Boundary 1131 Virtual Boundary 1210 Virtual Boundary 1220 Virtual Boundary 1230 Virtual Boundary 1240 Virtual Boundary 1250 Virtual Boundary 1260 Virtual Boundary 1300 Pictures 1400 Quadtree Partitioning Pattern 1510 sample adaptive offset filter 1512 SAO filter 1514 SAO filter 1516 LF Luma Filter 1518 ALF Chroma Filter 1522 Adder 1532 Adder 1541 SAO filtered luma component 1542 Second intermediate component 1543 Fourth intermediate component 1552 First intermediate component 1553 Third intermediate component 1561 Filtered Luma CB 1562 filtered first chroma component 1563 filtered second chroma component 1600 filters 1610 filter coefficients 1620 Diamond shape 1801 Luma Sample 1802 Luma Sample 1803 Chroma Samples 1804 Chroma Samples 1805 Chroma Samples 1806 Chroma Samples 1807 Chroma Samples 1808 Chroma Samples Lines 1811~1818 Lines 1851~1854 1910 Block 1920 direction 1923 direction 1930 Fully Oriented Block 1933 Fully directional block 1940 RMS error 2210 First Pattern 2220 Second Pattern 2230 Third Pattern 2240 Fourth Pattern 2500 filter support area 2600 filter support area 3300 Computer Systems 3301 Keyboard 3302 Mouse 3303 Trackpad 3305 Joystick 3306 Microphone 3307 Scanner 3308 Camera 3309 Speaker 3310 Touch Screen, Display 3320 CD / DVD ROM / RW 3321 Medium 3322 thumb drive 3323 Removable Hard Drive or Solid State Drive 3341 Central Processing Unit 3342 Graphics Processing Unit 3343 Field Programmable Gate Area 3344 Accelerator 3345 Read-Only Memory 3346 Random Access Memory 3347 Internal Mass Storage 3348 System Bus 3349 Peripheral bus 3350 graphics adapter 3354 Interface 3355 Communication Networks
Claims
[Claim 1] 1. A method for filtering in video decoding, comprising: determining, by a processor, an offset value associated with a first filter shape configuration of a video filter based on a signal in a coded video bitstream carrying video, wherein the number of filter taps of the first filter shape configuration is less than five and the video filter is based on a non-linear mapping; applying, by the processor, the nonlinear mapping-based filter to the samples to be filtered using the offset value associated with the first filter shape configuration; A method comprising: