Method and device for video filtering

JP2024019652A5Active Publication Date: 2025-05-02TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023215040
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-28
Filing Date
2023-12-20
Publication Date
2025-05-02
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing video coding techniques face inefficiencies in representing intra-prediction directions, particularly those that are statistically less likely, leading to increased bit usage and reduced compression efficiency.

Method used

Implementing a nonlinear mapping-based filter with adaptive filter configurations, such as cross-component sample offset (CCSO) and local sample offset (LSO) filters, to enhance video encoding/decoding by optimizing the representation of intra-prediction directions.

Benefits of technology

Improves video compression efficiency by reducing the number of bits required to represent less likely intra-prediction directions, thereby enhancing data reduction and overall coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and a device for video encoding and decoding.SOLUTION: In some examples, a device for video decoding comprises a processing circuit. The processing circuit determines an off-set value associated with a first filter shape configuration of a video filter according to a signal in a coded video stream that transmits a video. The video filter is based on nonlinear mapping. The number of filter taps of the first filter shape configuration is less than five. The processing circuit applies the video filter to a sample to be filtered using the offset value associated with the first filter shape configuration.SELECTED DRAWING: Figure 31
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 160,560, entitled "FLEXIBLE FILTER SHAPE FOR SAMPLE OFFSET," filed on March 12, 2021, which claims the benefit of priority to U.S. Provisional Application No. 17 / 449,199, entitled "METHOD AND APPARATUS FOR VIDEO FILTERING," filed on September 28, 2021. The entire disclosures of the above applications are hereby incorporated by reference in their entireties.

[0002] SUMMARY This disclosure describes embodiments related generally to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. The inventors' work, to the extent that this work is described in this background section, and aspects of the description that may not be considered prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure.

[0004] Video coding and decoding may be performed using inter-picture prediction with motion compensation. Uncompressed digital video may include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luma samples and associated chroma samples. The sequence of pictures may have a fixed or variable picture rate (also informally known as frame rate) of, for example, 60 pictures per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luma sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 GByte of storage space.

[0005] One objective of video coding and decoding may be the reduction of redundancy in the input video signal through compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations thereof, may be used. Lossless compression refers to techniques where an exact copy of the original signal can be reconstructed from a compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to utilize the reconstructed signal for its intended application. For video, lossy compression is widely used. The amount of acceptable distortion depends on the application, e.g., a user of a particular consumer streaming application may tolerate higher distortion than a user of a television distribution application. The achievable compression ratio may reflect that a higher acceptable / tolerable distortion may result in a higher compression ratio.

[0006] Video encoders and decoders may utilize techniques from a number of broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007] Video codec techniques may include a technique known as intra-coding, in which sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures, and their derivatives, such as independent decoder refresh pictures, may be used to reset the decoder state and thus may be used as the first picture in a coded video bitstream and video session or as still images. Samples of an intra block may undergo a transform, and the transform coefficients may be quantized before entropy coding. Intra prediction may be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value after the transform and the smaller the AC coefficients, the fewer bits are required for a given quantization step size to represent the block after entropy coding.

[0008] Conventional intra-coding, e.g. as known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to predict from surrounding sample data and / or metadata obtained during encoding / decoding of, e.g., spatially neighboring blocks of data or preceding blocks of data in decoding order. Such techniques are hereafter referred to as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed and not from reference pictures.

[0009] There may be many different forms of intra-prediction. If more than one of such techniques may be used in a given video coding technique, the technique in use may be coded in intra-prediction mode. In certain cases, a mode may have sub-modes and / or parameters that may be coded separately or included in the mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination may affect the coding efficiency improvement of intra-prediction, and therefore may affect the entropy coding technique used to convert the codeword into a bitstream.

[0010] A specific mode of intra prediction was introduced in H.264, improved in H.265, and further improved in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark sets (BMS). A predictor block may be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of the neighboring samples are copied to the predictor block according to the direction. The reference to the direction in use may be coded in the bitstream or may itself be predicted.

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to the 33 angular modes out of the 35 intra modes). The point (101) where the arrows converge represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from a sample or samples located to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from a sample or samples located to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] 1A, at the top left is shown a square block (104) of 4×4 samples (indicated by a thick dashed line). The square block (104) contains 16 samples, each of which is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions in the block (104). S44 is at the bottom right because the size of the block is 4×4 samples. Also shown are reference samples that follow a similar numbering scheme. The reference sample is labeled with R, its Y position (e.g., row index), and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are neighbors of the block being reconstructed, therefore negative values ​​do not need to be used.

[0013] Intra-picture prediction may work by copying reference sample values ​​from neighboring samples as assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction that coincides with the arrow (102), i.e., the sample is predicted from a prediction sample or samples that are at an angle of 45 degrees from the horizontal and to the upper right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the orientation is not evenly divisible by 45 degrees, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.

[0015] The number of possible directions is increasing as video coding techniques develop. In H.264 (2003), nine different directions can be represented. This has increased to 33 in H.265 (2013), and at the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been performed to identify the most likely directions, and certain techniques of entropy coding are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Furthermore, the direction itself may be predicted from nearby directions used in nearby already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (180) showing 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.

[0017] The mapping of intra-prediction direction bits in a coded video bitstream that represent the direction may vary from one video coding technique to another, ranging from simple direct mappings from prediction directions to intra-prediction modes, to codewords, to complex adaptation schemes including most probable modes, and similar techniques. In all cases, however, there may be certain directions that are statistically less likely than certain other directions in the video content. Because the goal of video compression is to reduce redundancy, these less likely directions are represented by more bits than more likely directions in a video coding technique that works well.

[0018] Motion compensation may be a lossy compression technique, and may refer to a technique used to predict a newly reconstructed picture or picture portion after blocks of sample data from a previously reconstructed picture or portion thereof (reference picture) are spatially shifted in a direction indicated by a motion vector (hereafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, with the third dimension being an indication of the reference picture in use (the latter may indirectly be a temporal dimension).

[0019] In some video compression techniques, the MV applicable to a particular area of ​​sample data may be predicted from other MVs, e.g., from an MV associated with another area of ​​sample data that is spatially adjacent to the area being reconstructed and precedes that MV in decoding order. Doing so may substantially reduce the amount of data required to code the MV, thereby eliminating redundancy and increasing compression. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction may work effectively because there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in similar directions, and therefore, in some cases, can be predicted using similar motion vectors derived from MVs of nearby areas. This makes the MV found for a given area similar or the same as the MV predicted from the surrounding MVs, and after entropy coding, it can be represented with fewer bits than would be used if coding the MV directly. In some cases, MV prediction may be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lossy, e.g., due to rounding errors when computing a predictor from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Here, we will explain a technique called "spatial merging" among the many MV prediction mechanisms provided by H.265.

[0021] Referring to Figure 2, the current block (201) contains samples found by the encoder during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of coding its MV directly, the MV can be derived from metadata associated with one or more reference pictures, e.g., the most recent (in decoding order) reference picture, using MVs associated with any one of five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture that neighboring blocks are using. Summary of the Invention [Means for solving the problem]

[0022] Aspects of the present disclosure provide a method and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit determines an offset value associated with a first filter shape configuration of a nonlinear mapping-based filter based on a signal in a coded video bitstream carrying the video. The number of filter taps of the first filter shape configuration is less than five. The processing circuit applies the nonlinear mapping-based filter to samples to be filtered using the offset value associated with the first filter shape configuration.

[0023] In some examples, the nonlinear mapping-based filter includes at least one of a cross-component sample offset (CCSO) filter and a local sample offset (LSO) filter.

[0024] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered, hi another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.

[0025] In some examples, the first filter shape configuration includes a single filter tap. The processing circuit determines an average sample value in the area and calculates a difference between the reconstructed sample value at the position of the single filter tap and the average sample value in the area. The processing circuit then applies a nonlinear mapping-based filter to the samples to be filtered based on the difference between the reconstructed sample value at the position of the single filter tap and the average sample value in the area.

[0026] In some examples, the processing circuit selects a first filter shape configuration from a group of filter shape configurations of the nonlinear mapping-based filter. In one example, the filter shape configurations in the group each have a number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration. In some examples, the processing circuit decodes an index from a coded video bitstream carrying the video, the index indicating a selection of the first filter shape configuration from the group of the nonlinear mapping-based filter. For example, the processing circuit decodes the index from syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.

[0027] In some examples, the processing circuit quantizes a delta value of the sample values ​​at the two filter tap positions to one of a number of possible quantization outputs, the number of possible quantization outputs being an integer in the range of 1 to 1024, inclusive. For example, the processing circuit decodes an index from syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, the index indicating the number of possible quantization outputs.

[0028] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video encoding / decoding.

[0029] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0030] [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is a diagram of an example intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Diagram 3] FIG. 3 is a simplified schematic block diagram of a communication system (300) according to one embodiment. [Figure 4] FIG. 4 is a simplified schematic block diagram of a communication system (400) according to one embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] FIG. 4 is a block diagram of an encoder according to another embodiment. [Figure 8] FIG. 4 is a block diagram of a decoder according to another embodiment. [Figure 9] 1A-1C are diagrams illustrating examples of filter shapes according to embodiments of the present disclosure. [Figure 10A] FIG. 13 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10B] FIG. 13 illustrates another example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10C] FIG. 13 illustrates yet another example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10D] FIG. 13 illustrates yet another example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 11A] FIG. 1 illustrates an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 11B] FIG. 13 illustrates another example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 12A] FIG. 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12B] FIG. 13 illustrates another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12C] FIG. 13 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12D] FIG. 13 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12E] FIG. 13 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12F] FIG. 13 illustrates yet another example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 13] FIG. 2 illustrates an example partition of a picture according to some embodiments of the present disclosure. [Figure 14] A diagram showing quadtree division patterns of pictures in some examples. [Figure 15] FIG. 2 illustrates a cross-component filter according to an embodiment of the present disclosure. [Figure 16] FIG. 2 illustrates an example of a filter shape according to an embodiment of the present disclosure. [Figure 17] A diagram illustrating an example syntax for a cross-component filter according to some embodiments of the present disclosure. [Figure 18A] FIG. 2 illustrates an exemplary position of a chroma sample relative to a luma sample, according to one embodiment of the present disclosure. [Figure 18B] FIG. 13 illustrates an exemplary position of a chroma sample relative to a luma sample in accordance with another embodiment of the present disclosure. [Figure 19] FIG. 2 illustrates an example of direction finding according to an embodiment of the present disclosure. [Figure 20] FIG. 1 illustrates an example of a subspace projection in some examples. [Figure 21] 1 is a table of multiple sample adaptive offset (SAO) types according to one embodiment of the present disclosure. [Figure 22] FIG. 13 illustrates example patterns for pixel classification at edge offset in some examples. [Diagram 23] 1 is a table of edge offset pixel classification rules in some examples. [Figure 24] FIG. 1 illustrates an example of syntax that may be signaled. [Diagram 25] 1A-1C are diagrams illustrating examples of filter support areas according to some embodiments of the present disclosure. [Figure 26] 13A-13C are diagrams illustrating examples of another filter support area according to some embodiments of the present disclosure. [Figure 27A] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27B] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27C] 1 is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 28] FIG. 1 illustrates eight filter shape configurations of three filter taps in one example. [Figure 29] FIG. 12 illustrates 12 filter shape configurations of three filter taps in one example. [Diagram 30] FIG. 1 shows an example of two candidate filter shape configurations for a nonlinear mapping-based filter. [Diagram 31] 1 is a flowchart outlining a process according to one embodiment of the present disclosure. [Diagram 32] 1 is a flowchart outlining a process according to one embodiment of the present disclosure. [Diagram 33] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0031] FIG. 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform a one-way transmission of data. For example, the terminal device (310) may code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more coded video bitstreams. The terminal device (320) may receive the coded video data from the network (350), decode the coded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. Unidirectional data transmission may be common, such as in media serving applications.

[0032] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) for bidirectional transmission of coded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to recover the video pictures, and display the video pictures on an accessible display device according to the recovered video data.

[0033] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure may not be so limited. The embodiments of the present disclosure apply to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that convey coded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. The communication network (350) may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of the network (350) may not be important to the operation of the present disclosure unless otherwise described herein below.

[0034] 4 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0035] The streaming system may include a video source (401) and a capture subsystem (413) that may include, for example, a digital camera that creates, for example, a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples taken by a digital camera. The stream of video pictures (402), depicted as a thick line to emphasize the high amount of data compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the lower amount of data compared to the stream of video pictures (402), may be stored on a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410), for example, in an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and creates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard, such as ITU-T Recommendation H.265.In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0036] It should be noted that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).

[0037] 5 shows a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of the video decoder (410) of the example of FIG. 4.

[0038] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510), in the same or another embodiment, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the encoded video data. The receiver (531) may receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be external to the video decoder (510) (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (510), for example to combat network jitter, and another buffer memory (515) internal to the video decoder (510), for example to handle playback timing. When the receiver (531) is receiving data from a store / forward device of sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be needed or may be small. For use with best-effort packet networks such as the Internet, a buffer memory (515) may be needed, may be relatively large, and preferably adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder (510).

[0039] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device such as a render device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530) as shown in FIG. 5. The control information for the rendering device(s) may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context-dependent, etc. The parser (520) may extract from the coded video sequence at least one set of subgroup parameters for a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include Group of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) may also extract from the coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0040] The parser (520) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to produce symbols (521).

[0041] The reconstruction of the symbols (521) can involve a number of different units, depending on the type of coded video picture or part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.

[0042] Beyond the functional blocks already mentioned, the video decoder (510) may be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0043] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients and control information from the parser (520) as symbols (521), including the transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0044] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from a current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) adds, possibly on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0045] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit (553) may access the reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) associated with the block, these samples may be added to the output of the scalar / inverse transform unit (551) by the aggregator (555) to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory (557) from which the motion compensated prediction unit (553) fetches the prediction samples may be controlled by a motion vector and are available to the motion compensated prediction unit (553) in the form of a symbol (521) that may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory (557) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0046] The output samples of the aggregator (555) may be subject to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques that are controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also be responsive to meta-information obtained during decoding of a coded picture or previous (in decoding order) portion of the coded video sequence, and may also be responsive to previously reconstructed and loop filtered sample values.

[0047] The output of the loop filter unit (556) may be a sample stream that can be output to the render device (512) as well as stored in a reference picture memory (557) for use in future inter-picture prediction.

[0048] Once a particular coded picture is fully reconstructed, it may be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before beginning reconstruction of the next coded picture.

[0049] The video decoder (510) may perform decoding operations according to a given video compression technique in a standard such as ITU-T Rec. H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used in the sense that the coded video sequence complies with both the syntax and the profile of the video compression technique or standard as documented in the video compression technique or standard. Specifically, the profile may select a particular tool as the only tool available for use under that profile, among all tools available in the video compression technique or standard. Also, what is required for compliance may be that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a Hypothetical Reference Decoder (HRD) and metadata of HRD buffer management signaled in the coded video sequence.

[0050] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] 6 shows a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) of the example of FIG. 4.

[0052] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture the video image(s) to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0053] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of separate pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0054] According to one embodiment, the video encoder (603) may code and compress pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required by the application. Enforcing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units as described below. For clarity, couplings are not depicted. Parameters set by the controller (650) may include rate control related parameters (picture skip, quantizer, lambda value for rate distortion optimization techniques, etc.), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions for the video encoder (603) optimized for a particular system design.

[0055] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an oversimplified explanation, in one example, the coding loop can include a source coder (630) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and reference picture(s)) and a (local) decoder (633) built into the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a similar manner that the (remote) decoder also does (because in the video compression techniques contemplated in the disclosed subject matter, any compression between the symbols and the coded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). The decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), so the content in the reference picture memory (634) is also bit-exact between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values ​​as the reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, e.g., due to channel errors) is also used in some related technologies.

[0056] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510), which is described in detail above in connection with Figure 5. However, with brief reference also to Figure 5, symbols may be available, and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, and the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0057] An observation that may be made at this point is that decoder techniques other than analysis / entropy decoding present in a decoder must necessarily be present in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the decoder operation. A description of the encoder techniques may be omitted, since they are the inverse of the decoder techniques described generically. Only for certain areas are more detailed descriptions required and are provided below.

[0058] In operation, in some examples, the source coder (630) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from the video sequence designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference picture(s) that may be selected as predictive reference(s) to the input picture.

[0059] The local video decoder (633) may decode the coded video data of pictures that may be designated as reference pictures based on the symbols created by the source coder (630). The operation of the coding engine (632) may preferably be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence may usually be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in a reference picture cache (634). In this way, the video encoder (603) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures obtained by the far-end video decoder (without transmission errors).

[0060] The predictor (635) may perform the prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may serve as suitable prediction references for the new picture. The predictor (635) may operate on one sample block per pixel block to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).

[0061] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0062] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.

[0063] The transmitter (640) may buffer the coded video sequence(s) created by the entropy coder (645) and prepare them for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) may merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0064] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:

[0065] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs support various types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variants of I-pictures and their respective uses and characteristics.

[0066] A predictive picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0067] A bidirectionally predicted picture (B-picture) may be a picture that can be coded and decoded using intra- or inter-prediction that uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predictive picture may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0068] A source picture may generally be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and coded block by block. A block may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, a block of an I picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P picture may be predictively coded with reference to one previously coded reference picture by spatial prediction or by temporal prediction. A block of a B picture may be predictively coded with reference to one or two previously coded reference pictures by spatial prediction or by temporal prediction.

[0069] The video encoder (603) may perform coding operations according to a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0070] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, and VUI parameter set fragments, etc.

[0071] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture may be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0072] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are prior to the decoding order of the current picture in the video (but the display order may be past and future, respectively). A block in the current picture may be coded by a first motion vector pointing toward a first reference block in the first reference picture and a second motion vector pointing toward a second reference block in the second reference picture. A block may be predicted by a combination of the first reference block and the second reference block.

[0073] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0074] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, a picture in a sequence of video pictures is divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three CTBs, which are one luma coding tree block (CTB) and two chroma CTBs. Each CTU may be recursively quadtree-partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels may be divided into one CU of 64×64 pixels, or four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. In general, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values) of 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0075] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and to encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0076] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as a predictive block of 8×8 samples. The video encoder (703) determines whether the processing block is best coded using intra mode, inter mode, or bi-predictive mode, for example using rate-distortion optimization. When the processing block is coded in intra mode, the video encoder (703) may use intra prediction techniques to code the processing block into a coded picture, and when the processing block is coded in inter mode or bi-predictive mode, the video encoder (703) may use inter prediction or bi-predictive techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, the merge mode may be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0077] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), coupled together as shown in FIG.

[0078] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-prediction information (e.g., a description of redundant information by inter-encoding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a prediction block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.

[0079] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with blocks already coded in the same picture, generate quantized coefficients after transformation, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0080] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines a mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is an intra mode, the general controller (721) controls the switch (726) to select an intra mode result to be used by the residual calculator (723) and controls the entropy encoder (725) to select intra prediction information to be included in the bitstream, and if the mode is an inter mode, the general controller (721) controls the switch (726) to select an inter prediction result to be used by the residual calculator (723) and controls the entropy encoder (725) to select inter prediction information to be included in the bitstream.

[0081] The residual calculation unit (723) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate on the residual data and encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data may be suitably used by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and the intra-prediction information. In some examples, the decoded blocks may be appropriately processed to generate decoded pictures, which may be buffered in a memory circuit (not shown) and used as reference pictures.

[0082] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in accordance with an appropriate standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. It should be noted that, according to the disclosed subject matter, when coding a block in a merged sub-mode of either an inter mode or a bi-prediction mode, the residual information is not present.

[0083] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) of the example of FIG. 4.

[0084] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872), coupled together as shown in FIG.

[0085] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols representing syntax elements that make up the coded picture. Such symbols may include, for example, prediction information (e.g., intra-prediction information or inter-prediction information, etc.) that may identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, the latter two being merged submode or separate submode), certain samples or metadata used for prediction by the intra-decoder (872) or inter-decoder (880), respectively, residual information, e.g., in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter-prediction mode or a bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880), and if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and is provided to the residual decoder (873).

[0086] The inter decoder (880) is configured to receive the inter prediction information and to generate inter prediction results based on the inter prediction information.

[0087] The intra decoder (872) is configured to receive the intra prediction information and to generate a prediction result based on the intra prediction information.

[0088] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantized transform coefficients, and to process the inverse quantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (the data path not shown as this may be low-volume control information only).

[0089] The reconstruction module (874) is configured to combine, in the spatial domain, the residual as output by the residual decoder (873) and the prediction result (possibly as output by an inter- or intra-prediction module) to form a reconstructed block that may be part of a reconstructed picture, and the reconstructed block may be part of a reconstructed video. It should be noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.

[0090] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0091] Aspects of the present disclosure provide filtering techniques for video coding / decoding.

[0092] An adaptive loop filter (ALF) with block-based filter adaptation can be applied by the encoder / decoder to reduce artifacts. For the luma component, one of multiple filters (e.g., 25 filters) can be selected for a 4×4 luma block, for example, based on local gradient direction and activity.

[0093] The ALF may have any suitable shape and size. Referring to FIG. 9, the ALFs (910)-(911) have diamond shapes, such as a 5×5 diamond shape for the ALF (910) and a 7×7 diamond shape for the ALF (911). In the ALF (910), the elements (920)-(932) form a diamond shape and may be used in the filtering process. Seven values ​​(e.g., C0-C6) may be used for the elements (920)-(932). In the ALF (911), the elements (940)-(964) form a diamond shape and may be used in the filtering process. Thirteen values ​​(e.g., C0-C12) may be used for the elements (940)-(964).

[0094] Referring to FIG. 9, in some examples, two ALFs (910)-(911) with diamond filter shapes are used. A 5×5 diamond shaped filter (910) may be applied to a chroma component (e.g., a chroma block, chroma CB), and a 7×7 diamond shaped filter (911) may be applied to a luma component (e.g., a luma block, luma CB). Other suitable shapes and sizes may be used in the ALFs. For example, a 9×9 diamond shaped filter may be used.

[0095] The filter coefficients at the locations indicated by the values ​​(e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, if the ALF includes a clipping function, the clip values ​​at those locations may be non-zero.

[0096] For block classification of the luma component, a 4×4 block (or luma block, luma CB) can be categorized or classified as one of multiple (e.g., 25) classes. The classification index C is calculated using Equation (1) by the quantized value of the directionality parameter D and the activity value A:

number

number

number

number

number

number

number

[0097] To reduce the complexity of the block classification described above, a subsampled 1-D Laplacian calculation can be applied. Figures 10A-10D show the gradient g in the vertical direction (Figure 10A), horizontal direction (Figure 10B), and two diagonal directions d1 (Figure 10C) and d2 (Figure 10D). v , g h , g d1 , and g d2 10A shows an example of the subsampled positions used to compute the vertical gradient g vIn FIG. 10B, the label “H” indicates the subsampled positions for computing the horizontal gradient g h In FIG. 10C, the label “D1” indicates the subsampled positions for computing the d1 diagonal gradient g d1 In FIG. 10D, the label “D2” indicates the subsampled positions for computing the d2 diagonal gradient g d2 indicates the subsampled positions for computing

[0098] horizontal g v and the vertical direction g h The maximum value of the gradient of

number

number

number

number

number

number

number

number

number

number

number

[0099] The activity value A can be calculated as follows:

number

number

[0100] For the chroma components in a picture, no block classification is applied and therefore a single set of ALF coefficients can be applied for each chroma component.

[0101] A geometric transformation may be applied to the filter coefficients and the corresponding filter clip values ​​(also called clip values). Before filtering a block (e.g., a 4×4 luma block), for example, the gradient values ​​(e.g., g v , g h , g d1 and / or g d2), a geometric transformation such as a rotation or a diagonal and vertical flip can be applied to the filter coefficients f(k,l) and the corresponding filter clip values ​​c(k,l). The geometric transformation applied to the filter coefficients f(k,l) and the corresponding filter clip values ​​c(k,l) can be equivalent to applying a geometric transformation to the samples in the region supported by the filter. The geometric transformation can make the different blocks to which the ALF is applied more similar by aligning their respective directionality.

[0102] Three geometric transformations can be performed, including a diagonal flip, a vertical flip, and a rotation, as described by equations (9) through (11), respectively. f D (k,l)=f(l,k),c D (k,l)=c(l,k) Equation (9) f V (k,l)=f(k,Kl-1),c V (k,l)=c(k,Kl-1) Equation (10) f R (k,l)=f(Kl-1,k),c R (k,l)=c(Kl-1,k) Equation (11) where K is the size of the ALF or filter, and 0≦k,1≦K-1 are the coordinates of the coefficients. For example, a filter f or clip value matrix (or clip matrix) c has position (0,0) in its upper left corner and position (K-1,K-1) in its lower right corner. Transformations can be applied to the filter coefficients f(k,l) and clip values ​​c(k,l) depending on the gradient values ​​calculated for the block. An example of the relationship between the transformations and the four gradients is summarized in Table 1.

[0103] [Table 1]

[0104] In some embodiments, the ALF filter parameters are signaled in an adaptive parameter set (APS) of a picture. In the APS, one or more sets (e.g., up to 25 sets) of luma filter coefficients and clip value indexes may be signaled. In one example, a set of the one or more sets may include luma filter coefficients and one or more clip value indexes. One or more sets (e.g., up to 8 sets) of chroma filter coefficients and clip value indexes may be signaled. To reduce signaling overhead, filter coefficients of different classifications (e.g., having different classification indexes) of luma components may be merged. In the slice header, an index of the APS used for the current slice may be signaled.

[0105] In one embodiment, a clip value index (also called a clipping index) can be decoded from the APS. The clip value index can be used to determine a corresponding clip value, for example, based on a relationship between the clip value index and the corresponding clip value. The relationship can be predefined and stored in the decoder. In one example, the relationship is described by a table, such as a luma table (e.g., used for luma CB) of clip value index and corresponding clip value, a chroma table (e.g., used for chroma CB) of clip value index and corresponding clip value, etc. The clip value can depend on the bit depth B. The bit depth B can refer to the internal bit depth, the bit depth of the reconstructed samples in the CB to be filtered, etc. In some examples, the tables (e.g., luma table, chroma table) are obtained using Equation (12).

number

[0106] [Table 2]

[0107] In the slice header of the current slice, one or more APS indexes (e.g., up to 7 APS indexes) may be signaled to specify the luma filter sets that can be used for the current slice. The filtering process may be controlled at one or more appropriate levels, such as the picture level, slice level, CTB level, etc. In one embodiment, the filtering process may be further controlled at the CTB level. A flag may be signaled to indicate whether the ALF is applied to the luma CTB. The luma CTB may select a filter set from among multiple fixed filter sets (e.g., 16 fixed filter sets) and a filter set (also referred to as a signaled filter set) signaled in the APS. A filter set index may be signaled to the luma CTB to indicate the filter set to be applied (e.g., a filter set among the multiple fixed filter sets and the signaled filter set). The multiple fixed filter sets may be predefined and hard-coded in the encoder and decoder and may be referred to as predefined filter sets.

[0108] For chroma components, an APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if an APS has more than one chroma filter set, a filter set index can be signaled for each chroma CTB.

[0109] The filter coefficients may be quantized with a norm equal to 128. To reduce multiplication complexity, bitstream conformance may be applied such that the coefficient values ​​of non-center positions are within the range of −27 to 27−1, inclusive. In one example, center position coefficients are not signaled in the bitstream and may be considered equal to 128.

[0110] In some embodiments, the syntax and semantics of the clipping index and clip values ​​are defined as follows: alf_luma_clip_idx[sfIdx][j] may be used to specify the clipping index of the clip value to use before multiplying the jth coefficient of the signaled luma filter indicated by sfIdx. Bitstream conformance requirements may include that when sfIdx=0 to alf_luma_num_filters_signalled_minus1 and j=0 to 11, the value of alf_luma_clip_idx[sfIdx][j] is in the range of 0 to 3, inclusive. The luma filter clip value AlfClipL[adaptation_parameter_set_id] with element AlfClipL[adaptation_parameter_set_id][filtIdx][j] may be derived as specified in Table 2 in response to bitDepth set equal to BitDepthY and clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j] when filtIdx=0 to NumAlfFilters-1 and j=0 to 11. alf_chroma_clip_idx[altIdx][j] may be used to specify a clipping index of the clip value to use before multiplying the jth coefficient of the alternate chroma filter with index altIdx. Bitstream conformance requirements may include that the value of alf_chroma_clip_idx[altIdx][j] is in the range of 0 to 3, inclusive, when altIdx=0 to alf_chroma_num_alt_filters_minus1 and j=0 to 5. The chroma filter clip value AlfClipC[adaptation_parameter_set_id][altIdx] with element AlfClipC[adaptation_parameter_set_id][altIdx][j] may be derived as specified in Table 2 in response to bitDepth set equal to BitDepthC and clipIdx set equal to alf_chroma_clip_idx[altIdx][j] when altIdx=0 to alf_chroma_num_alt_filters_minus1 and j=0 to 5.

[0111] In one embodiment, the filtering process can be described as follows: On the decoder side, when ALF is enabled for CTB, the sample R(i,j) in the CU (or CB) can be filtered, and the filtered sample value R'(i,j) is obtained as shown using the following equation (13). In one example, each sample in the CU is filtered.

number

[0112] In a non-linear ALF, multiple sets of clip values ​​can be provided in Table 3. In one example, the luma set includes four clip values ​​{1024, 181, 32, 6} and the chroma set includes four clip values ​​{1024, 161, 25, 4}. The four clip values ​​in the luma set can be selected by approximately equally dividing the full range (e.g., 1024) of the luma block sample values ​​(coded in 10 bits) in the logarithmic domain. The range can be from 4 to 1024 for the chroma set.

[0113] [Table 3]

[0114] The selected clip value may be coded in the "alf_data" syntax element as follows: A suitable encoding scheme (e.g., Golomb encoding scheme) may be used to encode the clipping index corresponding to the selected clip value as shown in Table 3. The encoding scheme may be the same encoding scheme used to encode the filter set index.

[0115] In one embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of the ALF. Thus, modified block classification and filtering can be used for samples near a CTU boundary (e.g., a horizontal CTU boundary). The virtual boundary (1130) is defined as the horizontal CTU boundary (1120) as "N samples " can be defined as a line by shifting the sample by N samples can be a positive integer. In one example, N samples is equal to 4 for the luma component, and N samples is equal to 2 for the chroma components.

[0116] Referring to Figure 11A, modified block classification can be applied to the luma component. In one example, the 1D Laplacian gradient calculation for a 4x4 block (1110) above the virtual boundary (1130) uses only samples above the virtual boundary (1130). Similarly, referring to Figure 11B, the 1D Laplacian gradient calculation for a 4x4 block (1111) below the virtual boundary (1131) shifted from the CTU boundary (1121) uses only samples below the virtual boundary (1131). Thus, the quantization of the activity value A can be scaled by taking into account the reduction in the number of samples used in the 1D Laplacian gradient calculation.

[0117] For the filtering process, a symmetric padding operation at the virtual boundary may be used for both the luma and chroma components. Figures 12A-12F show an example of such modified ALF filtering for the luma component at the virtual boundary. If the sample being filtered is located below the virtual boundary, the neighboring samples located above the virtual boundary may be padded. If the sample being filtered is located above the virtual boundary, the neighboring samples located below the virtual boundary may be padded. With reference to Figure 12A, the neighboring sample C0 may be padded with the sample C2 located below the virtual boundary (1210). With reference to Figure 12B, the neighboring sample C0 may be padded with the sample C2 located above the virtual boundary (1220). With reference to Figure 12C, the neighboring samples C1-C3 may be padded with the samples C5-C7 located below the virtual boundary (1230), respectively. With reference to Figure 12D, the neighboring samples C1-C3 may be padded with the samples C5-C7 located above the virtual boundary (1240), respectively. Referring to Figure 12E, the neighboring samples C4-C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located below the virtual boundary (1250). Referring to Figure 12F, the neighboring samples C4-C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located above the virtual boundary (1260).

[0118] In some examples, the above description may be appropriately adapted when the sample(s) and neighboring sample(s) are located to the left (or right) and right (or left) of a virtual boundary.

[0119] According to aspects of the present disclosure, to improve coding efficiency, a picture may be divided based on a filtering process. In some examples, a CTU is also referred to as a largest coding unit (LCU). In one example, a CTU or an LCU may have a size of 64x64 pixels. In some embodiments, an LCU-aligned picture quadtree partition may be used for filtering-based partitioning. In some examples, a coding unit-synchronized picture quadtree-based adaptive loop filter may be used. For example, a luma picture is divided into multiple multi-level quadtree partitions, and the boundaries of each partition are aligned to the boundaries of an LCU. Each partition has its own filtering process and is therefore referred to as a filter unit (FU).

[0120] In some examples, a two-pass encoding flow may be used. In the first pass of the two-pass encoding flow, a quadtree partitioning pattern of the picture and a best filter for each FU may be determined. In some embodiments, the determination of the quadtree partitioning pattern of the picture and the determination of the best filter for the FU are based on the filtering distortion. The filtering distortion may be estimated by a fast filtering distortion estimation (FFDE) technique during the determination process. The picture is partitioned using a quadtree partition. The reconstructed picture may be filtered according to the determined quadtree partitioning pattern and the selected filters of all the FUs.

[0121] In the second pass of the two-pass encoding flow, the CU-synchronous ALF on / off control is performed. According to the ALF on / off result, the first filtered picture is partially restored by the reconstructed picture.

[0122] Specifically, in some examples, a top-down partitioning method is adopted to divide a picture into multi-level quadtree partitions using a rate-distortion criterion. Each partition is called a filter unit (FU). The partitioning process aligns the quadtree partitions to the boundaries of LCUs. The encoding order of the FUs follows the z-scan order.

[0123] Figure 13 illustrates an example partition according to some embodiments of the present disclosure. In the example of Figure 13, a picture (1300) is partitioned into 10 FUs, and the encoding order is FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, FU9.

[0124] FIG. 14 illustrates a quadtree partitioning pattern (1400) for a picture (1300). In the example of FIG. 14, a partitioning flag is used to indicate the partitioning pattern for the picture. For example, a "1" indicates that quadtree partitioning is performed for the block, and a "0" indicates that the block should not be further divided. In some examples, the minimum size FU has an LCU size, and no partitioning flag is required for the minimum size FU. The partitioning flag is encoded and transmitted in z-order as shown in FIG. 14.

[0125] In some examples, the filter for each FU is selected from two filter sets based on a rate-distortion criterion. The first set has 1 / 2 symmetric square and diamond filters derived for the current FU. The second set is obtained from a time-delay filter buffer, which stores filters previously derived for the FU of the previous picture. The filter with the smallest rate-distortion cost of these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate-distortion costs of the four child FUs are calculated. A picture quadtree division pattern can be determined by recursively comparing the rate-distortion costs of the division case and the non-division case.

[0126] In some examples, the maximum quadtree division level can be used to limit the maximum number of FUs. In one example, if the maximum quadtree division level is 2, the maximum number of FUs is 16. Furthermore, during the quadtree division decision, correlation values ​​for deriving Wiener coefficients for the 16 FUs at the lowest quadtree level (smallest FU) can be reused. The remaining FUs can derive their Wiener filters from the correlations of the 16 FUs at the lowest quadtree level. Thus, in this example, only one frame buffer access is made to derive the filter coefficients for all FUs.

[0127] After the quadtree division pattern is determined, CU-synchronous ALF on / off control can be performed to further reduce the filtering distortion. By comparing the filtering distortion and the non-filtering distortion in each leaf CU, the leaf CU can explicitly switch ALF on / off in its local region. In some examples, the coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off results.

[0128] The cross-component filtering process may apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter may use a luma sample value of a luma component (e.g., luma CB) to refine a chroma component (e.g., chroma CB corresponding to luma CB). In one example, the luma CB and the chroma CB are included in a CU.

[0129] FIG. 15 illustrates a cross-component filter (e.g., CC-ALF) used to generate chroma components according to one embodiment of the present disclosure. In some examples, FIG. 15 illustrates a filtering process of a first chroma component (e.g., first chroma CB), a second chroma component (e.g., second chroma CB), and a luma component (e.g., luma CB). The luma component may be filtered by a sample adaptive offset (SAO) filter (1510) to generate a SAO filtered luma component (1541). The SAO filtered luma component (1541) may be further filtered by an ALF luma filter (1516) to become a filtered luma CB (1561) (e.g., “Y”).

[0130] The first chroma component may be filtered by an SAO filter (1512) and an ALF chroma filter (1518) to generate a first intermediate component (1552). Furthermore, the SAO filtered luma component (1541) may be filtered by a cross-component filter (e.g., CC-ALF) for the first chroma component (1521) to generate a second intermediate component (1542). Subsequently, a filtered first chroma component (1562) (e.g., “Cb”) may be generated based on at least one of the second intermediate component (1542) and the first intermediate component (1552). In one example, the filtered first chroma component (1562) (e.g., “Cb”) may be generated by combining the second intermediate component (1542) and the first intermediate component (1552) with an adder (1522). The cross-component adaptive loop filtering process for the first chroma component may include steps performed by a CC-ALF (1521) and steps performed, for example, by an adder (1522).

[0131] The above description can be adapted to the second chroma component. The second chroma component can be filtered by the SAO filter (1514) and the ALF chroma filter (1518) to generate a third intermediate component (1553). Furthermore, the SAO filtered luma component (1541) can be filtered by a cross-component filter (e.g., CC-ALF) (1531) for the second chroma component to generate a fourth intermediate component (1543). Subsequently, a filtered second chroma component (1563) (e.g., “Cr”) can be generated based on at least one of the fourth intermediate component (1543) and the third intermediate component (1553). In one example, the filtered second chroma component (1563) (e.g., “Cr”) can be generated by combining the fourth intermediate component (1543) and the third intermediate component (1553) with the adder (1532). In one example, the cross-component adaptive loop filtering process for the second chroma component may include steps performed by a CC-ALF (1531) and steps performed, for example, by an adder (1532).

[0132] The cross-component filters (e.g., CC-ALF (1521), CC-ALF (1531)) may operate by applying a linear filter having any suitable filter shape to the luma component (or luma channel) to refine each chroma component (e.g., first chroma component, second chroma component).

[0133] FIG. 16 illustrates an example of a filter (1600) according to an embodiment of the present disclosure. The filter (1600) may include non-zero filter coefficients and zero filter coefficients. The filter (1600) has a diamond shape (1620) formed by the filter coefficients (1610) (shown as solid circles). In one example, the non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and the filter coefficients that are not included in the filter coefficients (1610) are zero. Thus, the non-zero filter coefficients of the filter (1600) are included in the diamond shape (1620), and the filter coefficients that are not included in the diamond shape (1620) are zero. In one example, the number of filter coefficients of the filter (1600) is equal to the number of filter coefficients (1610), which is 18 in the example shown in FIG. 16.

[0134] The CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Referring back to Figure 15, the CC-ALF (1521) and the CC-ALF (1531) may have the same filter shape, such as the diamond shape (1620) shown in Figure 16, and the same number of filter coefficients. In one example, the values ​​of the filter coefficients in the CC-ALF (1521) are different from the values ​​of the filter coefficients in the CC-ALF (1531).

[0135] In general, the filter coefficients (e.g., non-zero filter coefficients) of the CC-ALF may be transmitted, for example, in the APS. In one example, the filter coefficients may be transmitted as coefficients (e.g., 2 10) and may be rounded for fixed-point representation. Application of CC-ALF may be controlled by variable block sizes and signaled by a context coded flag (e.g., a CC-ALF enable flag) received for each block of samples. Context coded flags such as the CC-ALF enable flag may be signaled at any appropriate level, such as at the block level. Block sizes along with CC-ALF enable flags may be received at the slice level for each chroma component. In some examples, block sizes (in chroma samples) 16×16, 32×32, and 64×64 may be supported.

[0136] 17 illustrates an example syntax of CC-ALF according to some embodiments of the present disclosure. In the example of FIG. 17, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index indicating whether to use a cross-component Cb filter, and if so, the index of the cross-component Cb filter. For example, if alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, then the cross-component Cb filter is not applied to the block of Cb color component samples at luma position (xCtb, yCtb). If alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not 0, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the filter to apply. For example, the alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]th cross-component Cb filter is applied to the block of Cb color component samples at luma position (xCtb, yCtb).

[0137] 17, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is used to indicate whether to use a cross-component Cr filter or a cross-component Cr filter index. For example, if alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, then the cross-component Cr filter is not applied to the block of Cr color component samples at the luma position (xCtb, yCtb). alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the cross-component Cr filter, if it is not 0. For example, the alf_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]-th cross-component Cr filter may be applied to a block of Cr color component samples at luma position (xCtb, yCtb).

[0138] In some examples, a chroma subsampling technique is used, and thus the number of samples in each of the chroma block(s) may be less than the number of samples in the luma block. The chroma subsampling format (also referred to as chroma subsampling format, for example, specified by chroma_format_idc) may indicate a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) between each of the chroma block(s) and the corresponding luma block. In one example, the chroma subsampling format is 4:2:0, and thus the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown in FIG. 18A-FIG. 18B. In one example, the chroma subsampling format is 4:2:2, and thus the chroma horizontal subsampling factor (e.g., SubWidthC) is 2, and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In one example, the chroma subsampling format is 4:4:4, and therefore the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample format (also referred to as the chroma sample position) may indicate a relative position of a chroma sample in a chroma block with respect to at least one corresponding luma sample in the luma block.

[0139] 18A-18B illustrate example locations of chroma samples relative to luma samples according to an embodiment of the present disclosure. With reference to FIG. 18A, luma samples (1801) are located in rows (1811)-(1818). The luma samples (1801) illustrated in FIG. 18A may represent a portion of a picture. In one example, a luma block (e.g., luma CB) includes the luma sample (1801). The luma block may correspond to two chroma blocks having a chroma subsampling format of 4:2:0. In one example, each chroma block includes a chroma sample (1803). Each chroma sample (e.g., chroma sample (1803(1))) corresponds to four luma samples (e.g., luma samples (1801(1)) through (1801(4)). In one example, the four luma samples are the top left sample (1801(1)), the top right sample (1801(2)), the bottom left sample (1801(3)), and the bottom right sample (1801(4)). A chroma sample (e.g., (1803(1))) corresponds to the top left sample (1801(1)) and the bottom left sample (18 The chroma sample format of a chroma block having chroma sample (1803) located at a left center position between the top left sample (1801(1)) and the bottom left sample (1801(3)) may be referred to as chroma sample format 0. Chroma sample format 0 indicates relative position 0, which corresponds to a left center position halfway between the top left sample (1801(1)) and the bottom left sample (1801(3)). The four luma samples (e.g., (1801(1)) through (1801(4))) may be referred to as neighboring luma samples of chroma sample (1803)(1).

[0140] In one example, each chroma block includes a chroma sample (1804). The above description regarding the chroma sample (1803) may be adapted to the chroma sample (1804), and thus detailed description may be omitted for brevity. Each of the chroma samples (1804) may be located at a central position of the four corresponding luma samples, and the chroma sample format of the chroma block having the chroma sample (1804) may be referred to as chroma sample format 1. The chroma sample format 1 indicates a relative position 1, which corresponds to the central position of the four luma samples (e.g., (1801(1)) to (1801(4))). For example, one of the chroma samples (1804) may be located at the central portion of the luma samples (1801(1)) to (1801(4)).

[0141] In one example, each chroma block includes a chroma sample (1805). Each of the chroma samples (1805) may be located at a top left position that is the same as the top left sample of the four corresponding luma samples (1801), and the chroma sample format of the chroma block having the chroma samples (1805) may be referred to as a chroma sample format 2. Thus, each of the chroma samples (1805) is located at the same position as the top left sample of the four luma samples (1801) corresponding to the respective chroma sample. The chroma sample format 2 indicates a relative position 2, which corresponds to the top left position of the four luma samples (1801). For example, one of the chroma samples (1805) may be located at the top left position of the luma samples (1801(1)) to (1801(4)).

[0142] In one example, each chroma block includes a chroma sample (1806). Each of the chroma samples (1806) may be located at a top center position between a corresponding top left sample and a corresponding top right sample, and a chroma sample format of a chroma block having the chroma sample (1806) may be referred to as a chroma sample format 3. The chroma sample format 3 indicates a relative position 3, which corresponds to a top center position between a top left sample and a top right sample. For example, one of the chroma samples (1806) may be located at a top center position of the luma samples (1801(1))-(1801(4)).

[0143] In one example, each chroma block includes a chroma sample (1807). Each of the chroma samples (1807) may be located at a bottom left position that is co-located with the bottom left sample of the four corresponding luma samples (1801), and the chroma sample format of the chroma block having the chroma sample (1807) may be referred to as chroma sample format 4. Thus, each of the chroma samples (1807) is co-located with the bottom left sample of the four luma samples (1801) corresponding to the respective chroma sample. The chroma sample format 4 indicates a relative position 4, which corresponds to the bottom left position of the four luma samples (1801). For example, one of the chroma samples (1807) may be located at the bottom left position of the luma samples (1801(1))-(1801(4)).

[0144] In one example, each chroma block includes a chroma sample (1808). Each chroma sample (1808) is located at a bottom center position between a bottom left sample and a bottom right sample, and a chroma sample format of a chroma block having the chroma sample (1808) can be referred to as a chroma sample format 5. The chroma sample format 5 indicates a relative position 5, which corresponds to a bottom center position between the bottom left sample and the bottom right sample of the four luma samples (1801). For example, one of the chroma samples (1808) can be located between the bottom left sample and the bottom right sample of the luma samples (1801(1)) to (1801(4)).

[0145] In general, any suitable chroma sample format may be used for the chroma subsampling format. Chroma sample formats 0-5 are examples of chroma sample formats described in chroma subsampling format 4:2:0. Additional chroma sample formats may be used for chroma subsampling format 4:2:0. Additionally, other chroma sample formats and / or variations of chroma sample formats 0-5 may be used for other chroma subsampling formats, such as 4:2:2, 4:4:4, etc. In one example, a chroma sample format that combines chroma samples (1805) and (1807) is used for chroma subsampling format 4:2:2.

[0146] In one example, a luma block may be considered to have alternating rows, such as rows (1811)-(1812), each including the top two samples (e.g., (1801(1))-(1801(2))) of the four luma samples (e.g., (1801(1))-(1801(4))) and the bottom two samples (e.g., (1801(3))-(1801(4))) of the four luma samples (e.g., (1801(1))-(1801(4))). Thus, rows (1811), (1813), (1815), and (1817) may be considered to have alternating rows, such as rows (1818), (1819), (1820), (1821), (1822), (1823), (1824), (1825), (1826), (1827), (1828), (1829), (1830), (1831), (1832), (1833), (1834), (1835), (1836), (1837), (1838), (1839), (1840), (1841), (1842), (1843), (1844), (1845), (1846), (1847), (1848), (1849), (1850), (1851), (1852), (1853), (1854), (1855), (1856), (1857), (1858), (1859), (1860), (1861), (1862), (1863), (1864), (1865), (1865), (1866), (1867), (18 ) can be referred to as the current row (also referred to as the top field), and rows (1812), (1814), (1816), and (1818) can be referred to as the next rows (also referred to as the bottom field). The four luma samples (e.g., (1801(1)) through (1801(4))) are located in the current row (e.g., (1811)) and the next row (e.g., (1812)). Relative positions 2 through 3 are located in the current row, relative positions 0 through 1 are located between each current row and the respective next row, and relative positions 4 through 5 are located in the next row.

[0147] Chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) are located in rows (1851)-(1854) within each chroma block. The specific location of rows (1851)-(1854) may depend on the chroma sample format of the chroma samples. For example, for chroma samples (1803)-(1804) having respective chroma sample formats 0-1, row (1851) is located between rows (1811)-(1812). For chroma samples (1805)-(1806) having respective chroma sample formats 2-3, row (1851) is in the same position as the current row (1811). For chroma samples (1807)-(1808) having respective chroma sample types 4-5, row (1851) is in the same position as the next row (1812). The above description can be adapted appropriately for rows (1852)-(1854), and a detailed description will be omitted for the sake of brevity.

[0148] Any suitable scanning method may be used to display, store, and / or transmit the luma blocks and corresponding chroma blocks described above in Figure 18A. In one example, progressive scanning is used.

[0149] Interlace scanning can be used, as shown in Figure 18B. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc equals 1). In one example, the variable chroma location format (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current rows (1811), (1813), (1815), and (1817) and the next rows (1812), (1814), (1816), and (1818) may be scanned separately, for example, the current rows (1811), (1813), (1815), and (1817) may be scanned first, followed by the next rows (1812), (1814), (1816), and (1818). The current row may include luma samples (1801) and the next row may include luma samples (1802).

[0150] Similarly, corresponding chroma blocks can be interlaced. Rows (1851) and (1853) containing chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) with no fill can be referred to as the current row (or current chroma row), and rows (1852) and (1854) containing chroma samples (1803), (1804), (1805), (1806), (1807), or (1808) with gray fill can be referred to as the next row (or next chroma row). In one example, during interlacing, rows (1851) and (1853) are scanned first, followed by rows (1852) and (1854).

[0151] In some examples, constrained directional enhancement filtering techniques can be used. The use of an in-loop constrained directional enhancement filter (CDEF) is to remove coding artifacts while preserving image details. In one example (e.g., HEVC), the sample adaptive offset (SAO) algorithm can achieve a similar objective by defining signal offsets for different classes of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, CDEF can be constrained to be easily vectorizable (i.e., executable with single instruction multiple data (SIMD) operations). Note that other nonlinear filters such as median filters, bilateral filters, etc. cannot be treated similarly.

[0152] In some cases, the amount of ringing artifacts in a coded image tends to be roughly proportional to the quantization step size. Although the level of detail is a property of the input image, the minimum level of detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of the ringing will generally be smaller than the amplitude of the detail.

[0153] The CDEF can be used to identify the orientation of each block and adaptively filter along the identified orientation and along orientations rotated 45 degrees from the identified orientation at smaller angles. In some examples, the encoder can look up the filter strength or can signal the filter strength explicitly, allowing a high degree of control over blurring.

[0154] Specifically, in some examples, the direction search is performed on the reconstructed pixels immediately after the deblocking filter. These pixels are available to the decoder, so the direction can be searched by the decoder, and thus, in one example, the direction does not require signaling. In some examples, the direction search can operate on a particular block size, such as an 8×8 block, that is small enough to properly handle non-linear edges, but large enough to reliably estimate the direction when applied to the quantized image. Also, having a certain directionality in the 8×8 region makes it easier to vectorize the filter. In some examples, each block (e.g., 8×8) can be compared to a fully directional block to determine the difference. A fully directional block is a block in which all pixels along a line in a direction have the same value. In one example, a difference measure, such as sum of squared differences (SSD), root mean square (RMS) error, etc., for the block and each fully directional block can be calculated. The fully directional block with the smallest difference (e.g., smallest SSD, smallest RMS, etc.) can then be determined, and the direction of the determined fully directional block can be the direction that best matches the pattern in the block.

[0155] FIG. 19 illustrates an example of a direction search according to an embodiment of the present disclosure. In one example, block (1910) is an 8×8 block that is reconstructed and output from a deblocking filter. In the example of FIG. 19, the direction search can determine a direction from eight directions, indicated by (1920), for block (1910). Eight fully directional blocks (1930) are formed corresponding to the eight directions (1920), respectively. A fully directional block corresponding to a direction is a block whose pixels along the line of the direction have the same value. In addition, difference measures such as SSD, RMS error, etc., of block (1910) and each fully directional block (1930) can be calculated. In the example of FIG. 19, the RMS error is indicated by (1940). As indicated by (1943), the RMS error between block (1910) and fully directional block (1933) is the smallest, and thus direction (1923) is the direction that best matches the pattern of block (1910).

[0156] After the block orientation is identified, a nonlinear low-pass directional filter can be determined. For example, the filter taps of the nonlinear low-pass directional filter can be aligned along the identified orientation to reduce ringing while preserving the directional edge or pattern. However, in some examples, directional filtering alone cannot sufficiently reduce ringing. In one example, additional filter taps are also used for pixels that are not aligned with the identified orientation. The extra filter taps are treated more conservatively to reduce the risk of blurring. Thus, the CDEF includes first and second order filter taps. In one example, the complete 2D CDEF filter can be expressed as Equation (14):

number

[0157] In some examples, in-loop restoration schemes are used in post-deblocking video coding besides the deblocking operation, generally to remove noise and improve edge quality. In one example, the in-loop restoration schemes are switchable within a frame for tiles of appropriate size. The in-loop restoration schemes are based on a separable symmetric Wiener filter, a dual self-guided filter with subspace projection, and a domain transform recursive filter. Because content statistics can change substantially within a frame, the in-loop restoration schemes are integrated into a switchable framework that can trigger different schemes in different regions of a frame.

[0158] A separable symmetric Wiener filter can be one of the in-loop restoration schemes. In some examples, every pixel of the degraded frame can be reconstructed as a non-causal filtered version of the pixels in a w × w window around it, where w = 2r + 1 is odd for integer r. The 2-D filter taps are expressed as w in column vectorized form. 2 If F is represented by a 1 × 1 element vector, then direct LMMSE optimization gives F = H -1 The filter parameters are derived given by M, where H=E[XX T ] is the autocovariance of x and w in a w × w window around the pixel. 2 is a column-wise vectorized version of the samples of T ] is the cross-correlation between x and the scalar source sample y to be estimated. In one example, the encoder can estimate H and M from a realization of the deblocked frame and the source, and send the resulting filter F to the decoder. However, doing so would result in a large error in the cross-correlation between x and the scalar source sample y. 2Not only does it incur a significant bit rate cost in transmitting the taps, but non-separable filtering makes decoding very complicated. In some embodiments, some additional constraints are imposed on the nature of F. For the first constraint, F is constrained to be separable, and filtering can be implemented as a separable horizontal and vertical w-tap convolution. For the second constraint, each of the horizontal and vertical filters is constrained to be symmetric. For the third constraint, it is assumed that both the horizontal and vertical filter coefficients sum to one.

[0159] Doubly self-guided filtering using subspace projection can be one of the in-loop restoration methods. Guided filtering is an image filtering technique that uses a local linear model shown by equation (15): y=Fx+G Equation (15) The above formula is used to calculate the filtered output y from the unfiltered sample x. Here, F and G are determined based on the statistics of the degraded image and the guidance image in the neighborhood of the filtered pixel. If the guide image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. In one example, a specific form of self-guided filtering can be used. The specific form of self-guided filtering depends on two parameters, namely the radius r and the noise parameter e, and is listed as the following steps: 1. The mean μ and variance σ of the pixels in a (2r+1)×(2r+1) window around every pixel 2 This step can be efficiently implemented with box filtering based on integral imaging. 2. Calculate for every pixel: f = σ 2 / (σ 2 +e);g=(1-f)μ 3. Calculate F and G for every pixel as the average of the f and g values ​​in a 3x3 window around the pixel being used.

[0160] The particular form of the self-guided filter is controlled by r and e, with larger r resulting in larger spatial variance and larger e resulting in larger range variance.

[0161] Figure 20 shows an example of subspace projection in some examples. As shown in Figure 20, even if none of the inexpensive reconstructions X1, X2 are close to the source Y, a suitable multiplier {α, β} can make them quite close to the source Y as long as they are somewhat moved in the right direction.

[0162] In some examples (e.g., HEVC), a filtering technique called sample adaptive offset (SAO) can be used. In some examples, SAO is applied to the reconstructed signal after the deblocking filter. SAO can use an offset value given in the slice header. In some examples, for luma samples, the encoder can determine whether to apply (enable) SAO to the slice. When SAO is enabled, the current picture allows a recursive division of the coding unit into four sub-regions, and each sub-region can select an SAO type from multiple SAO types based on features within the sub-region.

[0163] FIG. 21 illustrates a table (2100) of multiple SAO types according to an embodiment of the present disclosure. In the table (2100), SAO types 0 to 6 are shown. Note that SAO type 0 is used to indicate no SAO application. Furthermore, each SAO type, SAO type 1 to SAO type 6, includes multiple categories. SAO can reduce distortion by classifying reconstructed pixels in a sub-region into categories and adding an offset to pixels of each category in the sub-region. In some examples, edge characteristics can be used for pixel classification in SAO types 1 to 4, and pixel intensity can be used for pixel classification in SAO types 5 to 6.

[0164] Specifically, in one embodiment, such as SAO types 5-6, a band offset (BO) may be used to classify all pixels of a sub-region into multiple bands. Each band of the multiple bands includes pixels of the same intensity interval. In some examples, the intensity range is divided equally into multiple intervals, such as 32 intervals from 0 to a maximum intensity value (e.g., 255 for 8-bit pixels), and each interval is associated with an offset. Further, in one example, the 32 bands are divided into two groups, such as a first group and a second group. The first group includes the middle 16 bands (e.g., 16 intervals in the middle of the intensity range), and the second group includes the remaining 16 bands (e.g., 8 intervals on the low side of the intensity range, and 8 intervals on the high side of the intensity range). In one example, only one offset of the two groups is transmitted. In some embodiments, when pixel classification operations in BO are used, the five most significant bits of each pixel may be used directly as a band index.

[0165] Additionally, in one embodiment, such as SAO types 1-4, edge offset (EO) can be used to determine pixel classification and offset. For example, pixel classification can be determined based on a one-dimensional three-pixel pattern taking into account edge direction information.

[0166] FIG. 22 shows an example of a three-pixel pattern for pixel classification in edge offset in some examples. In the example of FIG. 22, the first pattern (2210) (shown by three gray pixels) is called a 0 degree pattern (the 0 degree pattern is associated with a horizontal direction), the second pattern (2220) (shown by three gray pixels) is called a 90 degree pattern (the 90 degree pattern is associated with a vertical direction), the third pattern (2230) (shown by three gray pixels) is called a 135 degree pattern (the 135 degree pattern is associated with a 135 degree diagonal direction), and the fourth pattern (2240) (shown by three gray pixels) is called a 45 degree pattern (the 45 degree pattern is associated with a 45 degree diagonal direction). In one example, one of the four direction patterns shown in FIG. 22 can be selected taking into account edge direction information of the sub-region. The selection can be transmitted in the video bitstream coded as side information in one example. The pixels within the sub-region can then be classified into multiple categories by comparing each pixel to its two neighboring pixels in a direction associated with the directional pattern.

[0167] 23 shows a table (2300) of pixel classification rules for edge offsets in some examples. Specifically, pixel c (shown in each pattern in FIG. 22) is compared with two neighboring pixels (shown in gray in each pattern in FIG. 22), and pixel c can be classified into one of categories 0 to 4 based on the comparison according to the pixel classification rules shown in FIG. 23.

[0168] In some embodiments, the decoder-side SAO can operate independently of the largest coding unit (LCU) (e.g., CTU) so that line buffers can be saved. In some examples, the top and bottom pixels in each LCU are not SAO processed when the 90 degree, 135 degree, and 45 degree classification patterns are selected. The leftmost and rightmost columns of pixels in each LCU are not SAO processed when the 0 degree, 135 degree, and 45 degree patterns are selected.

[0169] FIG. 24 shows an example of syntax (2400) that may need to be signaled for a CTU if parameters are not merged from neighboring CTUs. For example, the syntax element sao_type_idx[cldx][rx][ry] may be signaled to indicate the SAO type of the sub-region. The SAO type may be BO (band offset) or EO (edge ​​offset). A value of 0 for sao_type_idx[cldx][rx][ry] indicates that SAO is OFF, values ​​from 1 to 4 indicate that one of the four EO categories corresponding to 0°, 90°, 135°, and 45° is used, and a value of 5 indicates that BO is used. In the example of FIG. 24, each of the BO and EO types has four SAO offset values ​​to be signaled (sao_offset[cIdx][rx][ry][0] to sao_offset[cIdx][rx][ry][3]).

[0170] In general, the filtering process may use reconstructed samples of a first color component as an input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component, which may be the same as the first color component or may be another color component different from the first color component.

[0171] In a related example of cross-component filtering (CCF), filter coefficients are derived based on some mathematical equations. The derived filter coefficients are signaled from the encoder side to the decoder side, and the derived filter coefficients are used to generate an offset using a linear combination. The generated offset is then added to the reconstructed sample as a filtering process. For example, an offset is generated based on a linear combination of the filtering coefficients and the luma sample, and the generated offset is added to the reconstructed chroma sample. A related example of CCF is based on the assumption of a linear mapping relationship between the reconstructed luma sample value and the delta value between the original chroma sample and the reconstructed chroma sample. However, the mapping between the reconstructed luma sample value and the delta value between the original chroma sample and the reconstructed chroma sample does not necessarily follow a linear mapping process, and therefore the coding performance of CCF may be limited under the assumption of a linear mapping relationship.

[0172] In some examples, the nonlinear mapping technique may be used for cross-component filtering and / or same-color component filtering without significant signaling overhead. In one example, the nonlinear mapping technique may be used in cross-component filtering to generate cross-component sample offsets. In another example, the nonlinear mapping technique may be used in same-color component filtering to generate local sample offsets.

[0173] For convenience, the filtering process using the nonlinear mapping technique can be referred to as sample offset with nonlinear mapping (SO-NLM). The SO-NLM in the cross-component filtering process can be referred to as cross-component sample offset (CCSO). The SO-NLM in the same-color component filtering can be referred to as local sample offset (LSO). The filter using the nonlinear mapping technique can be referred to as a nonlinear mapping-based filter. The nonlinear mapping-based filter can include a CCSO filter, an LSO filter, etc.

[0174] In one example, CCSO and LSO can be used as loop filtering to reduce distortion of the reconstructed samples. CCSO and LSO do not rely on the assumption of linear mapping used in the associated exemplary CCF. For example, CCSO does not rely on the assumption of a linear mapping relationship between luma reconstructed sample values ​​and delta values ​​between original chroma samples and chroma reconstructed samples. Similarly, LSO does not rely on the assumption of a linear mapping relationship between reconstructed sample values ​​of color components and delta values ​​between original samples of color components and reconstructed samples of color components.

[0175] In the following description, a SO-NLM filtering process is described that uses reconstructed samples of a first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component. If the second color component is the same color component as the first color component, the description is applicable to LSO, and if the second color component is different from the first color component, the description is applicable to CCSO.

[0176] In SO-NLM, a nonlinear mapping is derived on the encoder side. The nonlinear mapping is between the reconstructed samples of the first color component in the filter support region and an offset to be added to the second color component in the filter support region. If the second color component is the same as the first color component, the nonlinear mapping is used in LSO, and if the second color component is different from the first color component, the nonlinear mapping is used in CCSO. The domain of the nonlinear mapping is determined by the different combinations of reconstructed samples of the processed input (also called possible reconstructed sample value combinations).

[0177] The technique of SO-NLM can be illustrated using a specific example. In the specific example, a reconstruction sample from a first color component located in a filter support area (also called a "filter support region") is determined. The filter support area is an area to which a filter can be applied, and the filter support area can have any suitable shape.

[0178] FIG. 25 illustrates an example of a filter support area (2500) according to some embodiments of the present disclosure. The filter support area (2500) includes four reconstructed samples of a first color component: P0, P1, P2, and P3. In the example of FIG. 25, the four reconstructed samples may form a cross in the vertical and horizontal directions, and the center position of the cross is the position of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is indicated by C. The sample of the second color component at the center position is indicated by F. The second color component may be the same as the first color component of P0-P3 or may be different from the first color component of P0-P3.

[0179] FIG. 26 illustrates an example of another filter support area (2600) according to some embodiments of the present disclosure. The filter support area (2600) includes four reconstructed samples P0, P1, P2, and P3 of a first color component forming a square. In the example of FIG. 26, the center position of the square is the position of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is indicated by C. The sample of the second color component at the center position is indicated by F. The second color component may be the same as the first color component of P0-P3 or may be different from the first color component of P0-P3.

[0180] The reconstructed samples are input to the SO-NLM filter and processed appropriately to form the filter taps. In one example, the positions of the reconstructed samples that are input to the SO-NLM filter are called filter tap positions. In a particular example, the reconstructed samples are processed in two steps:

[0181] In the first step, delta values ​​are calculated between P0 to P3 and C. For example, m0 indicates the delta value between P0 and C, m1 indicates the delta value between P1 and C, m2 indicates the delta value between P2 and C, and m3 indicates the delta value between P3 and C.

[0182] In a second step, the delta values ​​m0-m3 are further quantized, and the quantized values ​​are denoted as d0, d1, d2, d3. In one example, the quantized value can be one of -1, 0, 1 based on the quantization process. For example, if m is less than -N (N is a positive value and is called the quantization step size), the value m can be quantized to -1, if m is in the range of [-N,N], the value m can be quantized to 0, and if m is greater than N, the value m can be quantized to 1. In some examples, the quantization step size N can be one of 4, 8, 12, 16, etc.

[0183] In some embodiments, the quantized values ​​d0-d3 are filter taps and may be used to identify one combination in the filter domain. For example, the filter taps d0-d3 may form a combination in the filter domain. Each filter tap may have three quantized values, so if four filter taps are used, the filter domain contains 81 (3×3×3×3) combinations.

[0184] 27A-27C show a table (2700) having 81 combinations according to one embodiment of the present disclosure. The table (2700) includes 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, the first column includes an index of the combination, the second column includes a value of filter tap d0 for the combination, the third column includes a value of filter tap d1 for the combination, the fourth column includes a value of filter tap d2 for the combination, the fifth column includes a value of filter tap d3 for the combination, and the sixth column includes an offset value associated with the combination for the non-linear mapping. In one example, once the filter taps d0-d3 are determined, an offset value (represented by s) associated with the combination of d0-d3 can be determined according to the table (2700). In one example, the offset values ​​s0-s80 are integers such as 0, 1, -1, 3, -3, 5, -5, -7, etc.

[0185] In some embodiments, the final filtering process of the SO-NLM may be applied as shown in equation (16): f'=clip(f+s) Equation (16) where f is the reconstructed sample of the second color component to be filtered and s is an offset value determined according to the filter taps that are the result of processing the reconstructed sample of the first color component, such as using table (2700). The sum of the reconstructed sample F and the offset value s is further clipped to a range associated with the bit depth to determine the final filtered sample f' of the second color component.

[0186] It should be noted that in the case of LSO, the second color component in the above description is the same as the first color component, and in the case of CCSO, the second color component in the above description may be different from the first color component.

[0187] It should be noted that the above description may be adjusted for other embodiments of the present disclosure.

[0188] In some examples, at the encoder side, the encoding device may derive a mapping between the reconstructed samples of the first color component within the filter support region and an offset to be added to the reconstructed samples of the second color component. The mapping may be any suitable linear or nonlinear mapping. Then, the filtering process may be applied at the encoder side and / or the decoder side based on the mapping. For example, the mapping is appropriately notified to the decoder (e.g., the mapping is included in the coded video bitstream transmitted from the encoder side to the decoder side), and then the decoder may perform the filtering process based on the mapping.

[0189] According to some aspects of the present disclosure, the performance of a nonlinear mapping-based filter, such as a CCSO filter, an LSO filter, etc., depends on the filter shape configuration. The filter shape configuration (also called filter shape) of a filter can refer to the characteristics of the pattern formed by the filter tap positions. The pattern can be defined by various parameters such as the number of filter taps, the geometric shape of the filter tap positions, the distance of the filter tap positions to the center of the pattern, etc. Using a fixed filter shape configuration may limit the performance of the nonlinear mapping-based filter.

[0190] As shown in FIG. 24 and FIG. 25 and FIG. 27A-27C, some examples use a 5-tap filter design for filter shape configuration of the nonlinear mapping based filter. The 5-tap filter design can use tap positions of P0, P1, P2, P3, and C. The 5-tap filter design for filter shape configuration can result in a look-up table (LUT) with 81 entries, as shown in FIG. 27A-27C. The LUT of sample offset needs to be signaled from the encoder side to the decoder side, and the signaling of the LUT may contribute to a large portion of the signaling overhead and affect the coding efficiency of using the nonlinear mapping based filter. According to some aspects of the present disclosure, the number of filter taps may be different from 5. In some examples, the number of filter taps can be reduced and still capture information within the filter support area, improving the coding efficiency.

[0191] According to one aspect of the present disclosure, the number of filter taps of a nonlinear mapping based filter, such as a CCSO filter, an LSO filter, etc., can be any integer from 1 to M, where M is an integer. In some examples, the value of M is 1024 or any other suitable number.

[0192] In some examples, the number of filter taps of the nonlinear mapping based filter is three. In one example, the positions of the three filter taps include a central position. In another example, the positions of the three filter taps exclude a central position. The central position refers to a position of the reconstructed sample to be filtered.

[0193] In some examples, the number of filter taps of the nonlinear mapping based filter is five. In one example, the five filter tap locations include a central position. In another example, the five filter tap locations exclude a central position. The central position refers to a position of the reconstructed sample to be filtered.

[0194] In some examples, the number of filter taps of the nonlinear mapping based filter is 1. In one example, the location of one filter tap is a center position. In another example, the location of one filter tap is not a center position. The center position refers to the location of the reconstructed sample to be filtered. In one example, the CCSO filter has a 1-tap filter design, and the delta value of the CCSO filter can be calculated as (p-μ), where p is the reconstructed sample value located at the filter tap, and μ is the average sample value within a given area s. s can be a coding block, or a CTU / SB, or a picture.

[0195] In some embodiments, the filter shape configuration of the nonlinear mapping based filter is switchable between groups of filter shape configurations. In some examples, the groups of filter shape configurations may be candidates for the nonlinear mapping based filter. During encoding / decoding, the filter shape configuration of the nonlinear mapping based filter may change from one of the filter shape configurations in the group to another one of the filter shape configurations in the group.

[0196] According to one aspect of the present disclosure, the filter shape configurations within a group of nonlinear mapping based filters may have the same number of filter taps.

[0197] In some examples, the filter shape configurations in the group of nonlinear mapping based filters each have three filter taps.

[0198] FIG. 28 illustrates eight filter shape configurations of three filter taps in one example. Specifically, a first filter shape configuration includes three filter taps at positions labeled "1" and "C", with position "C" being the center position of position "1". A second filter shape configuration includes three filter taps at positions labeled "2" and position "C", with position "C" being the center position of position "2". A third filter shape configuration includes three filter taps at positions labeled "3" and position "C", with position "C" being the center position of position "3". A fourth filter shape configuration includes three filter taps at positions labeled "4" and position "C", with position "C" being the center position of position "4". A fifth filter shape configuration includes three filter taps at positions labeled "5" and "C", with position "C" being the center position of position "5". The sixth filter shape configuration includes three filter taps at positions labeled "6" and position "C", where position "C" is the center position of position "6". The seventh filter shape configuration includes three filter taps at positions labeled "7" and position "C", where position "C" is the center position of position "7". The eighth filter shape configuration includes three filter taps at positions labeled "8" and position "C", where position "C" is the center position of position "8".

[0199] In one example, eight filter shape configurations are candidates for a nonlinear mapping-based filter, which can switch from one of the eight filter shape configurations to another of the eight filter shape configurations during encoding / decoding.

[0200] FIG. 29 illustrates twelve filter shape configurations of three filter taps in one example. Specifically, a first filter shape configuration includes three filter taps at positions labeled "1" and "C", with position "C" being the center position of position "1". A second filter shape configuration includes three filter taps at positions labeled "2" and position "C", with position "C" being the center position of position "2". A third filter shape configuration includes three filter taps at positions labeled "3" and position "C", with position "C" being the center position of position "3". A fourth filter shape configuration includes three filter taps at positions labeled "4" and position "C", with position "C" being the center position of position "4". A fifth filter shape configuration includes three filter taps at positions labeled "5" and "C", with position "C" being the center position of position "5". The sixth filter shape configuration includes three filter taps at positions labeled "6" and position "C", with position "C" being the center position of position "6". The seventh filter shape configuration includes three filter taps at positions labeled "7" and position "C", with position "C" being the center position of position "7". The eighth filter shape configuration includes three filter taps at positions labeled "8" and position "C", with position "C" being the center position of position "8". The ninth filter shape configuration includes three filter taps at positions labeled "9" and "C", with position "C" being the center position of position "9". The tenth filter shape configuration includes three filter taps at positions labeled "10" and position "C", with position "C" being the center position of position "10". The eleventh filter shape configuration includes three filter taps at positions labeled "11" and position "C", with position "C" being the center position of position "11". The twelfth filter shape configuration includes three filter taps at positions labeled "12" and position "C," with position "C" being the center position of position "12."

[0201] In one example, 12 filter shape configurations are candidates for a nonlinear mapping-based filter. During encoding / decoding, the nonlinear mapping-based filter can switch from one of the 12 filter shape configurations to another of the 12 filter shape configurations.

[0202] According to another aspect of the disclosure, the filter shape configurations within a group of nonlinear mapping based filters can have different numbers of filter taps.

[0203] FIG. 30 illustrates an example of two candidate filter shape configurations for a nonlinear mapping-based filter. For example, a first filter shape configuration of the two candidate filter shape configurations includes three filter taps, and a second filter shape configuration of the two candidate filter shape configurations includes five filter taps. Specifically, the first filter shape configuration includes three filter taps at positions labeled "p0", "C", and "p2" (e.g., at the positions of the dashed circle), and the second filter shape configuration includes five filter taps at positions labeled "p0", "p1", "p2", "p3", and "C".

[0204] In one example, the two filter shape configurations are candidates for a nonlinear mapping-based filter. During encoding / decoding, the nonlinear mapping-based filter can switch from one of the two filter shape configurations to another of the two filter shape configurations.

[0205] It should be noted that the filter shape configuration of the nonlinear mapping based filter can be switched at various levels. In one example, the filter shape configuration of the nonlinear mapping based filter can be switched at a sequence level. In another example, the filter shape configuration of the nonlinear mapping based filter can be switched at a picture level. In another example, the filter shape configuration of the nonlinear mapping based filter can be switched at a CTU level or a superblock (SB) level. A superblock is the largest coding block in some examples. In another example, the filter shape configuration of the nonlinear mapping based filter can be switched at a coding block level (e.g., a CU level).

[0206] In some examples, a selection of a filter shape configuration from a group of filter shape configurations is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected filter shape configuration is signaled in a high level syntax (HLS), such as a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, frame header, etc.

[0207] It should be noted that the above description uses a nonlinear mapping-based filter to illustrate the technique of switching filter shape configurations, and the technique of switching filter shape configurations may be applied to other loop filters, including but not limited to cross-component adaptive loop filters, adaptive loop filters, loop restoration filters, etc. For example, a loop filter may have a group of candidate filter shape configurations. During encoding / decoding, the loop filter may switch its filter shape configuration from one of the candidate filter shape configurations to another one of the candidate filter shape configurations. The candidate filter shape configurations in a group may have different numbers of filter taps and / or different relative positions with respect to the sample to be filtered. It should be noted that the filter shape configuration of the nonlinear mapping-based filter may be switched at various levels. In one example, the filter shape configuration of the loop filter may be switched at a sequence level. In another example, the filter shape configuration of the loop filter may be switched at a picture level. In another example, the filter shape configuration of the loop filter may be switched at a CTU level or a superblock (SB) level. In another example, the filter shape configuration of the loop filter may be switched at a coding block level (e.g., a CU level). In some examples, a selection of a filter shape configuration from a group of filter shape configurations is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected filter shape configuration is signaled in a high level syntax (HLS), such as a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, frame header, etc.

[0208] According to another aspect of the present disclosure, quantization is used during the filtering process of the nonlinear mapping-based filter. For example, in the second step of SO-NLM, the delta values ​​m0-m3 are quantized to determine d0, d1, d2, d3, and each of d0, d1, d2, d3 may be one of three quantized outputs (e.g., −1, 0, and 1). It should be noted that in some examples, the number of quantized outputs of the nonlinear mapping-based filter may be any suitable integer. In one example, the number of quantized outputs of the nonlinear mapping-based filter may be a number from 1 to an upper limit, such as 1024. Note that the upper limit is not limited to 1024.

[0209] In some examples, the number of quantized outputs is 3, and the quantized outputs can be represented as -1, 0, and 1.

[0210] In some examples, the number of quantized outputs is 5, and the quantized outputs can be represented as -2, -1, 0, 1, 2.

[0211] According to one aspect of the present disclosure, the number of quantization outputs of the nonlinear mapping based filter can be switched (e.g., changed from one integer to another) at various levels during encoding / decoding. In one example, the number of quantization outputs of the nonlinear mapping based filter can be switched at a sequence level. In another example, the number of quantization outputs of the nonlinear mapping based filter can be switched at a picture level. In another example, the number of quantization outputs of the nonlinear mapping based filter can be switched at a CTU level or a superblock (SB) level. In another example, the number of quantization outputs of the nonlinear mapping based filter can be switched at a coding block level (e.g., a CU level). In some examples, the selection of the number of quantization outputs is signaled in a coding bitstream carrying the video. In one example, an index indicating the selected number of quantization outputs is signaled in a high level syntax (HLS), such as a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a frame header, etc.

[0212] FIG. 31 shows a flow chart outlining the process (3100) according to one embodiment of the present disclosure. The process (3100) may be used to reconstruct video carried in a coded video bitstream. When the term block is used, the block may be interpreted as a prediction block, a coding unit, a luma block, a chroma block, and the like. In various embodiments, the process (3100) is performed by processing circuits such as the processing circuits of the terminal devices (310), (320), (330), and (340), the processing circuits performing the functions of the video encoder (403), the processing circuits performing the functions of the video decoder (410), the processing circuits performing the functions of the video decoder (510), the processing circuits performing the functions of the video encoder (603), and the like. In some embodiments, the process (3100) is implemented in software instructions, and thus, the processing circuits perform the process (3100) when the processing circuits execute the software instructions. The process starts at (S3101) and proceeds to (S3110).

[0213] At (S3110), an offset value associated with a first filter shape configuration of the nonlinear mapping based filter is determined based on a signal in a coded video bitstream carrying video. The number of filter taps of the first filter shape configuration is less than five.

[0214] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered, hi another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.

[0215] In some examples, the first filter shape configuration includes a single filter tap, and calculates an average sample value in an area and a difference between the reconstructed sample value at the single filter tap position and the average sample value in the area, and a nonlinear mapping-based filter is applied to the samples to be filtered based on the difference between the reconstructed sample value at the single filter tap position and the average sample value in the area.

[0216] In some examples, the first filter shape configuration is selected from a group of filter shape configurations of the nonlinear mapping-based filter. In one example, the filter shape configurations in the group each have the same number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration.

[0217] In some examples, an index is decoded from a coded video bitstream carrying a video. The index indicates a selection of a first filter shape configuration from a group of nonlinear mapping-based filters. The index may be decoded from syntax signaling of at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.

[0218] At (S3120), a nonlinear mapping based filter is applied to the samples to be filtered using an offset value associated with the first filter shape configuration.

[0219] In some examples, to apply a nonlinear mapping-based filter, a delta value of sample values ​​at two filter tap positions is quantized to one of a number of possible quantization outputs. The number of possible quantization outputs is an integer ranging from 1 to 1024, inclusive. For example, an index is decoded from syntax signaling of at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, and the index indicates the number of possible quantization outputs.

[0220] The process (3100) proceeds to (S3199) and ends.

[0221] It should be noted that in some examples, the nonlinear mapping-based filter is a cross-component sample offset (CCSO) filter, and in some other examples, the nonlinear mapping-based filter is a local sample offset (LSO) filter.

[0222] The process (3100) may be adapted as appropriate. Step(s) of the process (3100) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0223] FIG. 32 shows a flow chart outlining the process (3200) according to one embodiment of the present disclosure. The process (3200) may be used to encode video in a coded video bitstream. When the term block is used, the block may be interpreted as a prediction block, a coding unit, a luma block, a chroma block, etc. In various embodiments, the process (3200) is performed by a processing circuit, such as a processing circuit of the terminal devices (310), (320), (330) and (340), a processing circuit performing the function of a video encoder (403), a processing circuit performing the function of a video encoder (603), etc. In some embodiments, the process (3200) is implemented in software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (3200). The process starts at (S3201) and proceeds to (S3210).

[0224] At (S3210), a nonlinear mapping based filter is applied to samples to be filtered in the video using offset values ​​associated with a first filter shape configuration, the number of filter taps of the first filter shape configuration being less than five.

[0225] In one example, the filter tap positions of the first filter shape configuration include the positions of the samples to be filtered, hi another example, the filter tap positions of the first filter shape configuration exclude the positions of the samples to be filtered.

[0226] In some examples, the first filter shape configuration includes a single filter tap, and calculates an average sample value in an area and a difference between the reconstructed sample value at the single filter tap position and the average sample value in the area, and a nonlinear mapping-based filter is applied to the samples to be filtered based on the difference between the reconstructed sample value at the single filter tap position and the average sample value in the area.

[0227] In some examples, the first filter shape configuration is selected from a group of filter shape configurations of the nonlinear mapping-based filter. In one example, the filter shape configurations in the group each have the same number of filter taps. In another example, one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration.

[0228] In some examples, an index is encoded in a coded video bitstream carrying a video. The index indicates a selection of a first filter shape configuration from a group of nonlinear mapping-based filters. The index may be signaled by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header.

[0229] In some examples, to apply a nonlinear mapping-based filter, a delta value of sample values ​​at two filter tap positions is quantized to one of a number of possible quantization outputs. The number of possible quantization outputs is an integer ranging from 1 to 1024, inclusive. For example, an index is encoded in the coded video bitstream by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, and the index indicates the number of possible quantization outputs.

[0230] At (S3220), the offset value is coded in a coded video bitstream that carries the video.

[0231] The process (3200) proceeds to (S3299) and ends.

[0232] It should be noted that in some examples, the nonlinear mapping-based filter is a cross-component sample offset (CCSO) filter, and in some other examples, the nonlinear mapping-based filter is a local sample offset (LSO) filter.

[0233] The process (3200) may be adapted as appropriate. Step(s) of the process (3200) may be modified and / or omitted. Additional step(s) may be added. Any suitable order of performance may be used.

[0234] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0235] The techniques described above may be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 33 illustrates a computer system (3300) suitable for implementing certain embodiments of the disclosed subject matter.

[0236] Computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. directly, or via interpretation, microcode execution, etc.

[0237] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0238] The components illustrated in Figure 33 for the computer system (3300) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of the computer system (3300).

[0239] The computer system (3300) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users using, for example, tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (voice, music, ambient sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0240] The input human interface devices may include one or more (only one of each shown) of a keyboard (3301), a mouse (3302), a trackpad (3303), a touch screen (3310), a data glove (not shown), a joystick (3305), a microphone (3306), a scanner (3307), a camera (3308).

[0241] The computer system (3300) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, by haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (3310), data gloves (not shown), or joystick (3305), although there may also be haptic feedback devices that do not function as input devices), audio output devices (such as speakers (3309), headphones (not shown)), visual output devices (such as screens (3310) including CRT screens, LCD screens, plasma screens, OLED screens (each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output, or output in more than three dimensions by means of stereographic output, etc.), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0242] The computer system (3300) may also include human accessible storage devices and associated media, such as optical media, including CD / DVD ROM / RW (3320) with CD / DVD or similar media (3321), thumb drives (3322), removable hard drives or solid state drives (3323), legacy magnetic media such as tape and floppy disks (not shown), and dedicated ROM / ASIC / PLD based devices such as security dongles (not shown).

[0243] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0244] The computer system (3300) may also include an interface (3354) to one or more communication networks (3355). The network may be, for example, wireless, wired, optical. The network may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television and terrestrial television, vehicular and industrial including CANBus, etc. Certain networks typically require an external network interface adapter that connects to a specific general data port or peripheral bus (3349) (e.g., a USB port on the computer system 3300), while others are typically integrated into the core of the computer system (3300) by connection to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3300) may communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast TV), unidirectional transmit only (e.g., CANbus to a particular CANbus device), or bidirectional, for example to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces, as described above.

[0245] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be connected to the core (3340) of the computer system (3300).

[0246] The cores (3340) may include one or more central processing units (CPUs) (3341), graphics processing units (GPUs) (3342), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (3343), hardware accelerators for specific tasks (3344), graphics adapters (3350), and the like. Such devices may be connected via a system bus (3348), along with read-only memory (ROM) (3345), random access memory (3346), internal mass storage (3347), such as internal non-user accessible hard drives, SSDs, and the like. In some computer systems, the system bus (3348) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs and GPUs, and the like. Peripheral devices may be connected directly to the core's system bus (3348) or may be connected via a peripheral bus (3349). In one example, a display (3310) may be connected to the graphics adapter (3350). Peripheral bus architectures include PCI and USB, among others.

[0247] The CPU (3341), GPU (3342), FPGA (3343), and accelerator (3344) can execute certain instructions that in combination may constitute the aforementioned computer code. This computer code may be stored in ROM (3345) or RAM (3346). Temporary data may also be stored in RAM (3346), while permanent data may be stored, for example, in internal mass storage (3347). Rapid storage and retrieval from any of the memory devices may be made possible by the use of cache memory, which may be closely associated with one or more of the CPU (3341), GPU (3342), mass storage (3347), ROM (3345), and RAM (3346), etc.

[0248] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the type well known and available to those skilled in the computer software arts.

[0249] By way of example and not by way of limitation, a computer system (3300) having an architecture, and in particular a core (3340), may provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with the user-accessible mass storage introduced above, as well as the core's internal mass storage (3347) or specific storage of the core (3340) that is of a non-transitory nature, such as ROM (3345). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (3340). The computer-readable media may include one or more memory devices or chips, depending on the particular needs. The software may cause the core (3340), and in particular the processors (including CPU, GPU, FPGA, etc.) therein, to perform certain processes or certain parts of certain processes described herein, including defining data structures stored in RAM (3346) and modifying such data structures according to the software-defined processes. Additionally, or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (3344)) that may operate in place of or in conjunction with software to perform certain processes or certain portions of certain processes described herein. Where appropriate, reference to software may encompass logic, and vice versa. Where appropriate, reference to computer-readable medium may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any appropriate combination of hardware and software. Appendix A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding MPM: Most Probable Mode WAIP: Wide-angle Intra Prediction SEI: Supplemental Extended Information VUI: Video Usability Information GOP: Group of Pictures TU: conversion unit PU: Prediction Unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted block HRD: Hypothetical Reference Decoder SDR: Standard Dynamic Range SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD:Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit CU: coding unit PDPC: Position-dependent prediction combination ISP: Intra Subpartition SPS: Sequence Parameter Settings

[0250] While this disclosure has described several exemplary embodiments, there are modifications, substitutions, and various substitute equivalents which are within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods which, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]

[0251] 101 Samples 102 Arrow 103 Arrow 104 Square Block 201 Current Block 202 Sample 203 Sample 204 Sample 205 Samples 206 Samples 300 Communication Systems 310 Terminal Devices 320 Terminal Devices 330 Terminal Devices 340 Terminal Devices 350 Network 400 Communication Systems 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 encoded video data, encoded video bitstream 405 Streaming Server 406 Client Subsystem 407 encoded video data, input copy 408 Client Subsystem 409 Encoded video data, copy 410 Video Decoder 411 Video Picture Output Stream 412 Display 413 Capture Subsystem 420 Electronic Devices 430 Electronic Devices 501 Channel 510 Video Decoder 512 Render Device 515 Buffer Memory 520 Parser 521 Symbols 530 Electronic Devices 531 Receiver 551 Scaler / Descaler Unit 552 Intra-picture prediction unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 601 Video Sources 603 Video Coder, Video Encoder 620 Electronic Devices 630 Source Coder 632 Coding Engine 633 Local Video Decoder 634 Reference Picture Memory, Reference Picture Cache 635 Predictors 640 Transmitter 643 coded video sequence 645 Entropy Coder 650 Controller 660 Communication Channels 703 Video Encoder 721 General Controller 722 Intra Encoder 723 Residual Calculation Unit 724 Residual Encoder 725 Entropy Encoder 726 Switch 728 Residual Decoder 730 InterEncoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Reconstruction Module 880 Interdecoder 910 Diamond shaped filter 911 Diamond shaped filter 920~932 elements 940~964 elements 1110 Block 1111 Block 1120 Horizontal CTU Boundary 1121 CTU boundary 1130 Virtual Boundary 1131 Virtual Boundary 1210 Virtual Boundary 1220 Virtual Boundary 1230 Virtual Boundary 1240 Virtual Boundary 1250 Virtual Boundary 1260 Virtual Boundary 1300 Pictures 1400 Quadtree division pattern 1510 sample adaptive offset filter 1512 SAO Filter 1514 SAO Filter 1516 LF Luma Filter 1518 ALF Chroma Filter 1522 Adder 1532 Adder 1541 SAO filtered luma component 1542 Second intermediate component 1543 4th intermediate component 1552 First intermediate component 1553 Third intermediate component 1561 Filtered Luma CB 1562 filtered first chroma component 1563 Filtered second chroma component 1600 Filter 1610 Filter Coefficients 1620 Diamond shape 1801 Luma Sample 1802 Luma Sample 1803 Chroma Sample 1804 Chroma Sample 1805 Chroma Sample 1806 Chroma Samples 1807 Chroma Sample 1808 Chroma Samples Lines 1811~1818 Lines 1851~1854 1910 Block 1920 direction 1923 direction 1930 Fully directional block 1933 Fully directional block 1940 RMS error 2210 First Pattern 2220 Second Pattern 2230 Third Pattern 2240 Fourth Pattern 2500 Filter Support Area 2600 Filter Support Area 3300 Computer Systems 3301 Keyboard 3302 Mouse 3303 Trackpad 3305 Joystick 3306 Microphone 3307 Scanner 3308 Camera 3309 Speaker 3310 Touch screen, display 3320 CD / DVD ROM / RW 3321 Medium 3322 Thumb Drive 3323 Removable Hard Drive or Solid State Drive 3341 Central Processing Unit 3342 Graphics Processing Unit 3343 Field Programmable Gate Area 3344 Accelerator 3345 Read-Only Memory 3346 Random Access Memory 3347 Internal Mass Storage 3348 System Bus 3349 Surrounding bus 3350 Graphics Adapter 3354 Interface 3355 Communication Networks

Claims

1. 1. A method for filtering in video encoding, comprising: applying, by a processor, a nonlinear mapping based filter to samples in the video to be filtered using offset values ​​associated with a first filter shape configuration of a video filter, the first filter shape configuration having a number of filter taps less than five; coding, by the processor, the offset value in a coded video bitstream carrying the video; A method comprising:

2. The method of claim 1 , wherein the video filter comprises at least one of a cross-component sample offset (CCSO) filter and a local sample offset (LSO) filter.

3. The method of claim 1 , wherein the filter tap locations of the first filter shape configuration comprise the sample locations.

4. The method of claim 1 , wherein the filter tap positions of the first filter shape configuration are exclusive of the sample positions.

5. The first filter configuration includes a single filter tap, and the method further comprises: determining an average sample value within an area; calculating the difference between the reconstructed sample value at the single filter tap position and the average sample value within the area; applying the video filter to the sample to be filtered based on the difference between the reconstructed sample value at the position of the single filter tap and the average sample value within the area; The method of claim 1 , comprising:

6. selecting the first filter shape configuration from a group of filter shape configurations of the video filter; The method of claim 1 further comprising:

7. The method of claim 6 , wherein the filter shape configurations in the group each have the number of filter taps.

8. The method of claim 6 , wherein one or more filter shape configurations in the group have a different number of filter taps than the first filter shape configuration.

9. encoding an index in the coded video bitstream carrying the video, the index indicating a selection of the first filter shape configuration from the group of video filters; The method of claim 6 further comprising:

10. encoding said index by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header. The method of claim 9 further comprising:

11. quantizing a delta value of the sample values ​​at the two filter tap positions to one of a number of possible quantized outputs, said number of possible quantized outputs being integers in the range of 1 to 1024, inclusive; The method of claim 1 further comprising:

12. encoding an index in the coded video bitstream by syntax signaling at least one of a block level, a coding tree unit (CTU) level, a superblock (SB) level, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a tile header, and a frame header, the index indicating the number of possible quantization outputs; The method of claim 11 further comprising:

13. An apparatus for video decoding configured to perform a method according to any one of claims 1 to 12.

14. A computer program comprising computer instructions configured, when executed by a processor, to cause the processor to perform a method according to any one of claims 1 to 12.