Method and apparatus for video filtering
Nonlinear mapping-based filters in video coding technologies address inefficiencies in intra-prediction and motion compensation, enhancing compression ratios and video quality in streaming and conferencing applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2026-02-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing video coding technologies face inefficiencies in reducing redundancy and distortion, particularly in intra-prediction and motion compensation, leading to suboptimal compression ratios and quality in video streaming and conferencing applications.
Implementing nonlinear mapping-based filters, such as inter-component sample offset (CCSO) and local sample offset (LSO) filters, to enhance video filtering in loop filter chains, improving the accuracy and efficiency of video encoding and decoding processes.
Enhances video compression by reducing redundancy and distortion, leading to improved compression ratios and video quality in streaming and conferencing applications.
Smart Images

Figure 2026090331000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to U.S. Patent Application No. 17 / 449,126, “Method and Apparatus for Video Filtering,” filed on 28 September 2021, which in turn claims priority to U.S. Provisional Application No. 63 / 164,478, “Flexible Filter Location for Sample Offset,” filed on 22 March 2021. The entire disclosure of the prior applications is incorporated herein by reference.
[0002] This disclosure generally describes embodiments related to video coding. [Background technology]
[0003] The background art descriptions provided herein are intended to provide a general context for this disclosure. The work of the inventors named herein, to the extent described in this background art section, and any aspects of the descriptions that may otherwise not qualify as prior art at the time of filing, are not expressly or implicitly recognized as prior art to this disclosure.
[0004] Video coding and decoding can be performed using interpicture prediction with motion compensation. Uncompressed digital video can contain a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated saturation samples. The series of pictures can have a fixed or variable picture rate (also informally known as frame rate), for example, 60 pictures per second or 60Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a 60Hz frame rate) requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires more than 600 GB of storage space.
[0005] One purpose of video coding and decoding may be to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements by two orders of magnitude or more, in some cases. Both lossless and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to a technique that allows an exact copy of the original signal to be restored from the compressed original signal. With lossy compression, the restored signal may not be identical to the original signal, but the distortion between the original and restored signals is small enough to make the restored signal useful for its intended purpose. For video, lossy compression is widely adopted. The amount of distortion that can be tolerated depends on the application; for example, users of a particular consumer streaming application may tolerate higher distortion than users of a television distribution application. The feasible compression ratio can reflect that the greater the tolerance / tolerance of distortion, the higher the compression ratio can be.
[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transformation, quantization, and entropy coding.
[0007] Video codec techniques may include a technique known as intra coding. In intra coding, sample values are represented without referencing samples or other data from a previously restored reference picture. In some video codecs, the picture is spatially subdivided into blocks of samples. A picture can be an intra picture when all blocks of samples are coded in intra mode. Intra pictures, and their derivatives such as independent decoder refresh pictures, can be used to reset the decoder state and can therefore be used as the first picture in a coded video bitstream and video session, or as a still image. Samples in intra blocks may be subjected to transformations, and the transformation coefficients can be quantized before entropy coding. Intra prediction may be a technique to minimize the sample values in the pre-transformation region. In some cases, the smaller the post-transformation DC value and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the post-entropy-coded block.
[0008] Traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include, for example, techniques that try from surrounding sample data and / or metadata acquired during the encoding / decoding of spatially adjacent and preceding blocks of data in the decoding order. Such techniques will hereafter be referred to as “intra-prediction” techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being restored, and not from the reference picture.
[0009] Intra-prediction can take many different forms. When two or more of these techniques can be used in a given video coding technique, the techniques in use can be coded in intra-prediction mode. In some examples, a mode can have submodes and / or parameters, which can be coded individually or included in a mode codeword. The choice of codeword for a given mode / submode / parameter combination can affect the improvement in coding efficiency by intra-prediction, and thus can affect the entropy coding technique used to convert the codeword into a bitstream.
[0010] Certain modes of intra-prediction were introduced in H.264, improved in H.265, and further refined in newer coding techniques such as Joint Search Models (JEM), Versatile Video Coding (VVC), and Benchmark Sets (BMS). Predictor blocks can be formed using neighboring sample values belonging to already available samples. The sample values of neighboring samples are copied to the predictor block according to direction. References to the direction in use may be coded within the bitstream or predicted themselves.
[0011] Referring to Figure 1A, the lower right shows a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to 33 of the 35 intra-modes). The point where the arrows converge (101) represents the predicted sample. The arrows indicate the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples located to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples located to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.
[0012] Referring further to FIG. 1A, in the upper left, a square block (104) of 4×4 samples (indicated by the thick dashed line) is depicted. The square block (104) contains 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within the block (104). Since the block is 4×4 samples in size, S44 is in the lower right. Further reference samples following a similar numbering scheme are shown. The reference samples are labeled with R, its Y position (e.g., row index), and X position (column index) with respect to the block (104). In both H.264 and H.265, since the predicted samples are adjacent to the block being restored, there is no need to use negative values.
[0013] Intrapicture prediction can function by copying the reference sample values from adjacent samples as assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating a prediction direction that matches the arrow (102) for this block, i.e., one or more predicted samples where the samples are at a 45-degree angle from horizontal and in the upper right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.
[0014] In certain cases, especially when the direction is not evenly divisible by 45 degrees, the values of multiple reference samples may be combined, for example, by interpolation, to calculate the reference sample.
[0015] The number of possible directions has been increasing as video coding technology has evolved. In H.264 (2003), nine different directions could be represented. That increased to 33 in H.265 (2013), and at the time of the present disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques of entropy coding are used to represent those likely directions with a small number of bits, accepting a certain penalty for less likely directions. Further, the direction itself can sometimes be predicted from the neighboring directions used in neighboring already decoded blocks.
[0016] FIG. 1B shows a schematic diagram (180) depicting 65 intra prediction directions by JEM to show an increasing number of prediction directions over time.
[0017] The mapping of intra prediction direction bits within the coded video bitstream representing the direction may vary for each video coding technology, and can range, for example, from a simple direct mapping of the prediction direction to the intra prediction mode, to complex adaptive schemes including codewords, the most likely modes, and similar techniques. However, in all cases, there can be certain directions that occur statistically less frequently within the video content than certain other directions. Since the purpose of video compression is reduction of redundancy, those less likely directions are represented by a greater number of bits than the more likely directions in a well-functioning video coding technology.
[0018] Motion compensation can be a lossy compression technique that can be used to predict a newly restored picture or part of a picture after blocks of sample data from a previously restored picture or part of a picture (reference picture) have been spatially shifted in the direction indicated by a motion vector (MV). In some cases, the reference picture may be the same as the picture currently being restored. The MV can have two dimensions, X and Y, or three dimensions, the third of which is the reference picture being used (the latter of which may indirectly be the time dimension).
[0019] In some video compression techniques, the motion vector (MV) applicable to a particular region of sample data can be predicted from other MVs, for example, from MVs related to another region of sample data that is spatially adjacent to the region being reconstructed and precedes that MV in the decoding order. Doing so significantly reduces the amount of data required to code the MV, thereby eliminating redundancy and increasing the compression ratio. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical probability that regions larger than the region to which a single MV is applicable will move in a similar direction, and therefore, in some cases, can be predicted using similar motion vectors derived from the MVs of adjacent regions. As a result, the detected MV for a given region is similar to or identical to the MV predicted from the surrounding MVs and can be represented with fewer bits than would be used when coding the MV directly after entropy coding. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating the predictor from some surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, "High Efficiency Video Coding," December 2016). Of the many MV prediction mechanisms provided by H.265, the technique described herein is hereafter referred to as "spatial merging."
[0021] Referring to Figure 2, the current block (201) contains samples that the encoder found to be predictable from a previous block of the same size, but spatially shifted, during the motion search process. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent reference picture (in decoding order), using the MV associated with one of the five surrounding samples, denoted as A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, the MV prediction can use predictors from the same reference pictures used by adjacent blocks. [Overview of the Initiative] [Means for solving the problem]
[0022] Aspects of this disclosure provide methods and apparatus for filtering in video coding / decoding. In some examples, the apparatus for video filtering includes a processing circuit. The processing circuit determines a first offset value for applying a nonlinear mapping-based filter based on a first reconstructed sample at a first node in a loop filter chain. The processing circuit then adds the first offset value to an intermediate reconstructed sample at a second node along the loop filter chain to generate a second reconstructed sample.
[0023] In some examples, the nonlinear mapping-based filter is an inter-component sample offset (CCSO) filter, where the intermediate reconstructed sample and the first reconstructed sample are samples of different color components.
[0024] In some examples, the nonlinear mapping-based filter is a local sample offset (LSO) filter, where the intermediate reconstructed sample and the first reconstructed sample are samples of the same color component.
[0025] In some embodiments, the first and second nodes are the same node, such as an input node, an output node, or an intermediate node in a loop filter chain. In one example, the first restored sample is generated by the processing module before the deblocking filter. In another example, the first restored sample is generated by the deblocking filter. In yet another example, the first restored sample is generated by the constrained directional expansion filter. In yet another example, the first restored sample is generated by the loop restored filter.
[0026] In some examples, the first and second nodes belong to different nodes. In one example, the first reconstructed sample is generated by the processing module before the deblocking filter, and the intermediate reconstructed sample is generated by at least one of the following: the deblocking filter, the constrained directional expansion filter, or the loop reconstructed filter.
[0027] In some examples, the first reconstructed sample is generated by a deblocking filter, and the intermediate reconstructed sample is generated by at least one of a constrained directional expansion filter or a loop reconstructed filter.
[0028] In some examples, the first reconstructed sample is generated by a constrained directional extension filter, and the intermediate reconstructed samples are generated by a loop reconstructed filter.
[0029] A part of the present disclosure also provides a non-temporary computer-readable medium that, when executed by a computer, stores instructions causing the computer to perform one of the methods for video encoding / decoding.
[0030] Further features, properties, and various advantages of the disclosed subject matter will become clearer from the detailed description and accompanying drawings below. [Brief explanation of the drawing]
[0031] [Figure 1A] This is a schematic diagram of an exemplary subset of intra-predictive modes. [Figure 1B] This is an example diagram of the intra-prediction direction. [Figure 2] This is a schematic diagram of a current block and its surrounding spatial merge candidates in one example. [Figure 3] This is a schematic diagram of a simplified block diagram of a communication system (300) according to one embodiment. [Figure 4] This is a schematic diagram of a simplified block diagram of a communication system (400) according to one embodiment. [Figure 5] This is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 6] This is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Figure 7] This is a block diagram of an encoder according to another embodiment. [Figure 8] This is a block diagram of a decoder according to another embodiment. [Figure 9] This figure shows an example of a filter shape according to the embodiments of this disclosure. [Figure 10A] This figure shows an example of subsampled locations used to calculate the gradient according to an embodiment of the present disclosure. [Figure 10B] This figure shows an example of subsampled locations used to calculate the gradient according to an embodiment of the present disclosure. [Figure 10C] This figure shows an example of subsampled locations used to calculate the gradient according to an embodiment of the present disclosure. [Figure 10D] This figure shows an example of subsampled locations used to calculate the gradient according to an embodiment of the present disclosure. [Figure 11A] This figure shows an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 11B] This figure shows an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 12A] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 12B] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 12C] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 12D] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 12E] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 12F] This figure shows an example of symmetric padding behavior at a virtual boundary according to an embodiment of the present disclosure. [Figure 13] This figure shows examples of picture partitions according to some embodiments of the present disclosure. [Figure 14] This figure shows quadtree partitioning patterns for pictures in several examples. [Figure 15] This figure shows an inter-component filter according to one embodiment of the present disclosure. [Figure 16] This figure shows an example of a filter shape according to one embodiment of the present disclosure. [Figure 17] This figure shows example syntax for inter-component filtering according to some embodiments of the present disclosure. [Figure 18A] This figure shows an exemplary position of a chrominance sample relative to a luminance sample according to an embodiment of the present disclosure. [Figure 18B] This figure shows an exemplary position of a chrominance sample relative to a luminance sample according to an embodiment of the present disclosure. [Figure 19] This figure shows an example of a direction search according to one embodiment of the present disclosure. [Figure 20] This figure shows an example of subspace projection in several cases. [Figure 21] This is a table of multiple sample adaptation offset (SAO) types according to one embodiment of the present disclosure. [Figure 22] This figure shows examples of patterns for pixel classification at edge offsets in several cases. [Figure 23] This is a table for pixel classification rules for edge offsets in several examples. [Figure 24] This figure shows an example of a syntax that can be signaled. [Figure 25] This figure shows an example of a filter support area according to some embodiments of the present disclosure. [Figure 26] This figure shows an example of another filter support area according to some embodiments of the present disclosure. [Figure 27A] This is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27B] This is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 27C] This is a table having 81 combinations according to one embodiment of the present disclosure. [Figure 28] This figure shows the seven filter shape configurations for three filter taps in one example. [Figure 29] Here are some example block diagrams of loop filter chains. [Figure 30A] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 30B] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 30C] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 30D]This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 31A] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 31B] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 32A] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 32B] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 32C] This figure shows an example of a loop filter chain that includes nonlinear mapping-based filters at different positions within the loop filter chain. [Figure 33] This figure shows another example of a loop filter chain that includes a nonlinear mapping-based filter in several examples. [Figure 34] This figure shows an example of a loop filter chain that includes multiple nonlinear mapping-based filters in several examples. [Figure 35] This figure shows another example of a loop filter chain that includes multiple nonlinear mapping-based filters in several examples. [Figure 36] This is a flowchart outlining the process according to one embodiment of the present disclosure. [Figure 37] This is a schematic diagram of a computer system according to one embodiment. [Modes for carrying out the invention]
[0032] Figure 3 shows a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes, for example, a plurality of terminal devices that can communicate with each other over a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected over the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, terminal device (310) can encode video data (e.g., a stream of video pictures captured by terminal device (310)) for transmission to another terminal device (320) over the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. Terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media serving applications, for example.
[0033] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of coded video data, which may occur, for example, during a video conference. In the case of bidirectional transmission of data, in one example, each terminal device of terminal devices (330) and (340) can encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other terminal device of terminal devices (330) and (340) via the network (350). Each terminal device of terminal devices (330) and (340) can also receive coded video data transmitted by the other terminal device of terminal devices (330) and (340), decode the coded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.
[0034] In the example in Figure 3, terminal devices (310), (320), (330), and (340) may be represented as a server, a personal computer, and a smartphone, but the principles of this disclosure are not limited thereto. Embodiments of this disclosure find applications using laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit coded video data between terminal devices (310), (320), (330), and (340), including, for example, wired and / or wireless communication networks. Communication networks (350) can exchange data over circuit-switched channels and / or packet-switched channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of network (350) may not be important to the operation of this disclosure unless described below herein.
[0035] Figure 4 shows an example of the arrangement of a video encoder and video decoder in a streaming environment as an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, such as video conferencing, digital television, and storing compressed video on digital media including CDs, DVDs, and memory sticks.
[0036] The streaming system may include, for example, a capture subsystem (413) which may include a video source (401), such as a digital camera, that creates a stream (402) of uncompressed video pictures. In one example, the stream (402) of video pictures includes a sample captured by the digital camera. The stream (402) of video pictures, depicted as a thick line to emphasize the large amount of data compared to encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize the small amount of data compared to the stream (402), can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) in Figure 4, can access a streaming server (405) to retrieve copies (407) and (409) of encoded video data (404). Client subsystem (406) may include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes the input copy (407) of the encoded video data and creates an output stream (411) of a video picture that can be rendered on a display (412) (e.g., a display screen) or other rendering device (without rendering). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to specific video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265.For example, a video coding standard under development is informally known as Multipurpose Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0037] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may also include a video encoder (not shown).
[0038] Figure 5 shows a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of Figure 4.
[0039] The receiver (531) can receive one or more coded video sequences to be decoded by the video decoder (510), in the same or different embodiments, one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. Coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores coded video data. The receiver (531) can receive coded video data together with other data that may be transferred to their respective usage entities (not described), e.g., coded audio data and / or auxiliary data streams. The receiver (531) can isolate coded video sequences from other data. To counteract network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter, "Parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, it may be located outside the video decoder (510) (not described). In yet other applications, for example, to counteract network jitter, a buffer memory (not described) may exist outside the video decoder (510), and in addition, for example, to handle playout timing, another buffer memory (515) may exist inside the video decoder (510). When the receiver (531) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be required or may be small. For use in best-effort packet networks such as the Internet, the buffer memory (515) may be required, may be relatively large, may be advantageously adaptive in size, and may be at least partially implemented in the operating system or a similar element (not described) outside the video decoder (510).
[0040] The video decoder (510) may include a parser (520) to recover symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510), and potentially information for controlling rendering devices, such as a rendering device (512) (e.g., a display screen), which is not an integral part of the electronic device (530) but can be coupled to the electronic device (530), as shown in Figure 5. The control information for rendering devices may be in the form of supplemental extension information (SEI messages) or parameter set fragments (not depicted) of video usability information (VUI). The parser (520) can parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding, etc., with or without context sensitivity. The parser (520) can extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include picture groups (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), predictive units (PU), and so on. The parser (520) can also extract information from the coded video sequence such as transform coefficients, quantizer parameter values, and motion vectors.
[0041] The parser (520) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) in order to create a symbol (521).
[0042] The reconstruction of symbol (521) may involve multiple different units, depending on the type of coded video picture or part thereof (such as between and within pictures, between and within blocks), as well as other factors. How each unit is involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the following multiple units is not depicted for clarity.
[0043] In addition to the functional blocks already described, the video decoder (510) can be conceptually subdivided into several functional units, as described below. In actual implementations operating under commercial constraints, many of these units can interact closely with each other and integrate with each other at least partially. However, the following conceptual subdivision into functional units is appropriate for describing the disclosed subject matter.
[0044] The first unit is the scaler / inverse unit (551). The scaler / inverse unit (551) receives control information from the parser (520) as symbols (521), including the quantization conversion coefficients, which conversion to use, block size, quantization coefficients, and quantization scaling matrix. The scaler / inverse unit (551) can output a block containing sample values that can be input to the aggregator (555).
[0045] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use predictive information from previously restored pictures but can use predictive information from previously restored portions of the current picture. Such predictive information can be provided by the intra-picture predictive unit (552). In some cases, the intra-picture predictive unit (552) generates a block of the same size and shape as the block being restored, using the surrounding already restored information fetched from the current picture buffer (558). The current picture buffer (558) buffers, for example, partially restored current pictures and / or fully restored current pictures. The aggregator (555) may, in some cases, add the predictive information generated by the intra-predictive unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a sample-by-sample basis.
[0046] In other cases, the output samples of the scaler / inverse unit (551) may be associated with an intercoded and potentially motion-compensated block. In such cases, the motion-compensated prediction unit (553) can access the reference picture memory (557) to fetch samples to be used for prediction. After motion-compensating the fetched samples according to the symbols (521) associated with the block, these samples can be added to the output of the scaler / inverse unit (551) by the aggregator (555) to generate output sample information (in this case, called residual samples or residual signals). The address in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches the prediction samples can be controlled by motion vectors available to the motion-compensated prediction unit (553) in the form of symbols (521), which may have X, Y, and reference picture components, for example. Motion compensation may also include interpolation of sample values fetched from the reference picture memory (557) when the exact motion vectors of the subsamples are used, motion vector prediction mechanisms, etc.
[0047] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technique may include in-loop filtering techniques that are controlled by parameters contained in the coded video sequence (also called coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to previously restored and loop-filtered sample values, as well as metadata obtained during decoding of earlier parts (in decoding order) of the coded picture or coded video sequence.
[0048] The output of the loop filter unit (556) can be a sample stream that can be output to the rendering device (512) as well as stored in reference picture memory (557) for use in future interpicture prediction.
[0049] A particular coded picture, once fully restored, can be used as a reference picture for future predictions. For example, once the coded picture corresponding to the current picture is fully restored and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and any unused current picture buffer can be reallocated before starting the restoration of the next coded picture.
[0050] The video decoder (510) can perform decoding operations according to a specified video compression technique in a standard such as ITU-T Rec.H.265. The coded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence adheres to both the syntax of the video compression technique or standard and the documented profile within the video compression technique. Specifically, a profile can select several tools from all the tools available in the video compression technique or standard as the only tools available for use under that profile. Furthermore, compliance may require that the complexity of the coded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum recovery sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further restricted by the virtual reference decoder (HRD) specification and the metadata for HRD buffer management signaled within the coded video sequence.
[0051] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately restore the original video data. The additional data may take the form of, for example, extension layers of time, space, or signal-to-noise ratio (SNR), redundant slices, redundant pictures, forward error correction codes, etc.
[0052] Figure 6 shows a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of Figure 4.
[0053] The video encoder (603) can capture video images that are encoded by the video encoder (603) and can receive video samples from a video source (601) which is not part of the electronic device (620) in the example in Figure 6. In another example, the video source (601) is part of the electronic device (620).
[0054] The video source (601) can provide a source video sequence encoded by a video encoder (603) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of separate pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following description will focus on samples.
[0055] According to one embodiment, the video encoder (603) can encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For simplicity, the coupling is not depicted. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.
[0056] In some embodiments, the video encoder (603) is configured to operate in a coding loop. For the most simplified explanation, in one example, the coding loop may include a source coder (630) (which is involved in creating symbols, such as a symbol stream, based on, for example, an input picture to be coded and a reference picture), as well as a (local) decoder (633) built into the video encoder (603). The decoder (633) deconstructs the symbols to create sample data in a similar manner to how a (remote) decoder also creates them (since any compression between the symbols and the coded video bitstream is reversible in the video compression techniques considered in the disclosed subject). The deconstructed sample stream (sample data) is fed into the reference picture memory (634). Since decoding the symbol stream leads to bit-accurate results regardless of the decoder's location (local or remote), the contents in the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the predictive part of the encoder "sees" the exact same sample values as the reference picture samples that the decoder "sees" when using the predictions during decoding. This fundamental principle of reference picture synchronization (and the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is also used in several related techniques.
[0057] The operation of the “local” decoder (633) may be the same as that of a “remote” decoder, such as the video decoder (510), which has already been described in detail above with reference to Figure 5. However, referring again briefly to Figure 5, since symbols are available and the encoding / decoding of symbols to the coded video sequence by the entropy coder (645) and parser (520) may be reversible, the entropy decoding portion of the video decoder (510), including the buffer memory (515), and the parser (520) may not be fully implemented in the local decoder (633).
[0058] An observation that can be made at this point is that any decoder techniques other than parsing / entropy decoding present in the decoder must also exist in substantially the same functional form within the corresponding encoder. Therefore, the subject matter disclosed will focus on the operation of the decoder. A description of encoder techniques can be omitted, as it is the inverse of a comprehensive description of decoder techniques. More detailed explanations are only necessary in specific areas, and are provided below.
[0059] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding, predictively coding the input picture by referencing one or more previously coded pictures from a video sequence designated as “reference pictures”. In this way, the coding engine (632) codes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture that may be selected as a predictive reference to the input picture.
[0060] The local video decoder (633) can decode the coded video data of a picture that may be designated as a reference picture based on symbols created by the source coder (630). The operation of the coding engine (632) may, advantageously, be a lossy process. When coded video data can be decoded by a video decoder (not shown in Figure 6), the restored video sequence may be a replica of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that may be performed by the video decoder on the reference picture so that the restored reference picture is stored in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the restored reference picture that has common content as the restored reference picture obtained by the far-end video decoder (without transmission errors).
[0061] The predictor (635) can perform predictive searches for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or reference picture motion vectors, block shapes, and other specific metadata that can serve as appropriate predictive references for the new picture. The predictor (635) can operate on sample blocks pixel by pixel to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).
[0062] The controller (650) can manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.
[0063] The outputs of all the aforementioned functional units can undergo entropy coding within the entropy coder (645). The entropy coder (645) converts the symbols generated by the various functional units into coded video sequences by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.
[0064] The transmitter (640) can buffer the coded video sequence created by the entropy coder (645) and prepare it for transmission over the communication channel (660), which may be a hardware / software link to a storage device that stores coded video data. The transmitter (640) can merge the coded video data from the video coder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0065] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) can assign a specific coded picture type to each coded picture, which may affect the coding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types:
[0066] An intra-picture (I-picture) can be a picture that can be coded and decoded without using any other pictures in the sequence as a source of prediction. Some video codecs enable various types of intra-pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are familiar with their variations of I-pictures, as well as their respective uses and characteristics.
[0067] A predictive picture (P-picture) can be a picture that can be coded and decoded using intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0068] A bidirectional predictive picture (B-picture) can be a picture that can be coded and decoded using intra-prediction or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use three or more reference pictures and associated metadata to reconstruct a single block.
[0069] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block is coded. Blocks may be predictively coded by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks in picture I may be non-predictively coded, or they may be predictively coded by referencing already coded blocks in the same picture (spatial prediction or intra-prediction). Pixel blocks in picture P may be predictively coded via spatial prediction or temporal prediction by referencing one previously coded reference picture. Blocks in picture B may be predictively coded via spatial prediction or temporal prediction by referencing one or two previously coded reference pictures.
[0070] The video encoder (603) can perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec.H.265. In this operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technique or standard being used.
[0071] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include time / space / SNR extension layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0072] Video may be captured as multiple source pictures (video pictures) in a time series. Intra-picture prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (time or other) correlations between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block in the current picture is analogous to a reference block in a reference picture that has been previously coded and is still buffered within the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension to identify the reference picture if multiple reference pictures are used.
[0073] In some embodiments, a dual prediction technique can be used in interpicture prediction. According to the dual prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which are earlier in the decoding order than the current picture in the video (but may be past and future in the display order, respectively). A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture, and a second motion vector pointing to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.
[0074] Furthermore, to improve coding efficiency, merge mode techniques can be used in interpicture prediction.
[0075] According to some embodiments of this disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one luminance CTB and two saturation CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU, or four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as inter-prediction type or intra-prediction type. The CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes one luminance prediction block (PB) and two saturation PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values) such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0076] Figure 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values in the current video picture within a sequence of video pictures and to encode the processing block into a coded picture which is part of a coded video sequence. In one example, the video encoder (703) is used instead of the video encoder (403) in the example of Figure 4.
[0077] In the HEVC example, the video encoder (703) receives a matrix of sample values for a processing block, such as an 8x8 sample prediction block. The video encoder (703) determines whether the processing block is best coded using intra-mode, inter-mode, or bi-prediction mode, for example, using rate-distortion optimization. When the processing block is coded in intra-mode, the video encoder (703) can encode the processing block into a coded picture using the intra-prediction technique; when the processing block is coded in inter-mode or bi-prediction mode, the video encoder (703) can encode the processing block into a coded picture using the inter-prediction technique or the bi-prediction technique, respectively. In certain video coding techniques, the merge mode may be an inter-picture prediction submode where the motion vectors are derived from one or more motion vector predictors, without the advantage of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components such as a mode determination module (not shown) to determine the mode of the processing block.
[0078] In the example shown in Figure 7, the video encoder (703) includes an interencoder (730), an intraencoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general-purpose controller (721), and an entropy encoder (725), all coupled together as shown in Figure 7.
[0079] The interencoder (730) is configured to receive a sample of the current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in the previous and subsequent pictures), generate interprediction information (e.g., a description of redundant information, motion vectors, and merge mode information by the intercoding technique), and compute an interprediction result (e.g., a predicted block) based on the interprediction information using any appropriate technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.
[0080] The intra encoder (722) is configured to receive a sample of the current block (e.g., a processing block), and optionally compare the block to an already coded block in the same picture, generate quantization coefficients after the transformation, and optionally also generate intra prediction information (e.g., intra prediction direction information by one or more intra coding techniques). In one example, the intra encoder (722) also computes an intra prediction result (e.g., a prediction block) based on the intra prediction information and reference block in the same picture.
[0081] The general-purpose controller (721) is configured to determine general-purpose control data and to control other components of the video encoder (703) based on the general-purpose control data. For example, the general-purpose controller (721) determines the mode of a block and provides control signals to the switch (726) based on the mode. For example, when the mode is intra-mode, the general-purpose controller (721) controls the switch (726) to select intra-mode results for use by the residual calculator (723), controls the entropy encoder (725) to select intra-prediction information, and includes the intra-prediction information in the bitstream. When the mode is inter-mode, the general-purpose controller (721) controls the switch (726) to select inter-prediction results for use by the residual calculator (723), controls the entropy encoder (725) to select inter-prediction information, and includes the inter-prediction information in the bitstream.
[0082] The residual calculator (723) is configured to calculate the difference (residual data) between the receiving block and the prediction result selected from the intra-encoder (722) or inter-encoder (730). The residual encoder (724) is configured to operate on the residual data to encode the residual data and generate conversion coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate conversion coefficients. The conversion coefficients are then quantized to obtain quantized conversion coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (722) and inter-encoder (730). For example, an interencoder (730) can generate a decoded block based on decoded residual data and interprediction information, and an intraencoder (722) can generate a decoded block based on decoded residual data and intraprediction information. The decoded block is appropriately processed to generate a decoded picture, which is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.
[0083] The entropy encoder (725) is configured to format the bitstream to include the encoded blocks. The entropy encoder (725) is configured to include various information according to an appropriate standard such as the HEVC standard. For example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Note that, according to the disclosed subject, residual information is not present when coding blocks in either inter-mode or bi-prediction mode merge submodes.
[0084] Figure 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture which is part of an encoded video sequence, and to decode the encoded picture to produce a restored picture. In one example, the video decoder (810) is used instead of the video decoder (410) in the example of Figure 4.
[0085] In the example shown in Figure 8, the video decoder (810) includes an entropy decoder (871), an interdecoder (880), a residual decoder (873), a restoration module (874), and an intradecoder (872) coupled together as shown in Figure 8.
[0086] The entropy decoder (871) can be configured to recover specific symbols from the coded picture that represent the syntactic elements that make up the coded picture. Such symbols can identify the mode in which the block is coded (e.g., intra-mode, inter-mode, bi-prediction mode, merge sub-mode, or the latter two of another sub-mode), each of which can identify specific samples or metadata used for prediction by the intra decoder (872) or inter-decoder (880), and may include prediction information (e.g., intra-prediction information or inter-prediction information), such as residual information in the form of quantization transformation coefficients. For example, when the prediction mode is inter-mode or bi-prediction mode, inter-prediction information is provided to the inter-decoder (880), and when the prediction type is intra-prediction type, intra-prediction information is provided to the intra-decoder (872). The residual information can undergo inverse quantization and be provided to the residual decoder (873).
[0087] The interdecoder (880) is configured to receive interprediction information and generate interprediction results based on the interprediction information.
[0088] The intra decoder (872) is configured to receive intra prediction information and generate prediction results based on the intra prediction information.
[0089] The residual decoder (873) is configured to perform inverse quantization to extract inverse quantization conversion coefficients, process these coefficients, and convert the residuals from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (QP)), which may be provided by the entropy decoder (871) (the data path is not depicted as this may only contain a small amount of control information).
[0090] The restoration module (874) combines the residual output by the residual decoder (873) and the prediction results (possibly output by the inter-prediction module or intra-prediction module) in the spatial domain to form a restored block that may be part of the restored picture, and similarly, the restored picture may be part of the restored video. Note that other appropriate operations, such as deblocking, can be performed to improve the appearance.
[0091] It should be noted that the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using any suitable technique. In one embodiment, the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (603), as well as the video decoders (410), (510), and (810), can be implemented using one or more processors that execute software instructions.
[0092] Aspects of this disclosure provide filtering techniques for video coding / decoding.
[0093] Adaptive loop filters (ALFs) with block-based filter adaptation can be applied by encoders / decoders to reduce artifacts. For luminance components, for example, one of several filters (e.g., 25 filters) can be selected for a 4x4 luminance block based on the direction and activity of local gradients.
[0094] ALF can have any suitable shape and size. Referring to Figure 9, ALF(910)~(911) have rhombus shapes such as a 5x5 rhombus shape for ALF(910) and a 7x7 rhombus shape for ALF(911). In ALF(910), elements (920)~(932) form a rhombus shape and can be used in the filtering process. Seven values (e.g., C0~C6) can be used for elements (920)~(932). In ALF(911), elements (940)~(964) form a rhombus shape and can be used in the filtering process. Thirteen values (e.g., C0~C12) can be used for elements (940)~(964).
[0095] Referring to Figure 9, in some examples, two ALFs (910) and (911) with rhombic filter shapes are used. The 5x5 rhombic filter (910) can be applied to the saturation component (e.g., saturation block, saturation CB), and the 7x7 rhombic filter (911) can be applied to the luminance component (e.g., luminance block, luminance CB). Other suitable shapes and sizes can be used in the ALF. For example, a 9x9 rhombic filter can be used.
[0096] The filter coefficients at the positions indicated by the values (e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, if the ALF includes a clipping function, the clipping values at those positions may be non-zero.
[0097] In the case of block classification of luminance components, a 4x4 block (or luminance block, luminance CB) can be categorized or classified as one of several (e.g., 25) classes. The classification index C is calculated using equation (1) based on the quantized values of the directional parameter D and the activation value A.
number
number
number
number
[0098] To reduce the complexity of the block classification described above, a subsampled 1-D Laplacian calculation can be applied. Figures 10A to 10D show the gradients g in the vertical direction (Figure 10A), the horizontal direction (Figure 10B), and the two diagonal directions d1 (Figure 10C) and d2 (Figure 10D), respectively. v , g h , g d1 , and g d2 An example of a subsampled location used to calculate the gradient is shown. The same subsampled location can be used to calculate the gradient in a different direction. In Figure 10A, the label "V" is the vertical gradient g.v shows the subsampled positions for calculating. In FIG. 10B, the label "H" is the horizontal gradient g h shows the subsampled positions for calculating. In FIG. 10C, the label "D1" is the diagonal gradient g of d1 d1 shows the subsampled positions for calculating. In FIG. 10D, the label "D2" is the diagonal gradient g of d2 d2 shows the subsampled positions for calculating.
[0099] The horizontal gradient g h and the vertical gradient g v maximum value
Number
Number
Number
Number
Number
Number
Number
Number
number
number
number
[0100] Activity value A is,
number
number
[0101] In the case of saturation components within a picture, block classification is not applied, and therefore, a single set of ALF coefficients can be applied to each saturation component.
[0102] Geometric transformations can be applied to filter coefficients and corresponding filter clipping values (also called clipping values). Before filtering a block (e.g., a 4x4 luminance block), for example, a gradient value calculated for the block (e.g., g) can be applied. v , g h , g d1 , and / or g d2Depending on the context, geometric transformations such as rotation, diagonal flipping, and vertical flipping can be applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l). The geometric transformation applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l) may be equivalent to applying a geometric transformation to the samples within the region supported by the filter. Geometric transformations can make different blocks to which ALF is applied more similar by aligning their respective orientations.
[0103] Three geometric transformations can be performed, including diagonal flip, vertical flip, and rotation, as described by equations (9) to (11). f D (k,l)=f(l,k), c D (k,l)=c(l,k) Equation (9) f V (k,l)=f(k,Kl-1), c V (k,l)=c(k,Kl-1) Equation (10) f R (k,l)=f(Kl-1,k), c R (k,l)=c(Kl-1,k) Equation (11) Here, K is the size of the ALF or filter, and 0 ≤ k, l ≤ K-1 are the coordinates of the coefficients. For example, the upper left corner of the filter f or the clipping value matrix (or clipping matrix) c is (0,0) and the lower right corner is (K-1,K-1). The transformation can be applied to the filter coefficients f(k,l) and clipping value c(k,l) depending on the gradient values calculated for the block. An example of the relationship between the transformation and the four gradients is summarized in Table 1.
[0104] [Table 1]
[0105] In some embodiments, ALF filter parameters are signaled within an Adaptive Parameter Set (APS) for the picture. The APS can signal one or more sets (e.g., up to 25 sets) of luminance filter coefficients and clipping value indices. In one example, one of these sets may include luminance filter coefficients and one or more clipping value indices. One or more sets (e.g., up to 8 sets) of saturation filter coefficients and clipping value indices can also be signaled. To reduce signaling overhead, filter coefficients for different classifications (e.g., with different classification indices) for the luminance component can be merged. The slice header can signal the index of the APS used for the current slice.
[0106] In one embodiment, a clipping value index (also called a clipping index) can be decoded from the APS. The clipping value index can be used, for example, to determine a corresponding clipping value based on the relationship between the clipping value index and the corresponding clipping value. The relationship can be predefined and stored in the decoder. In one example, the relationship is described by tables such as a luminance table (used for, e.g., a luminance CB) and a saturation table (used for, e.g., a saturation CB) for the clipping value index and the corresponding clipping value. The clipping value may depend on the bit depth B, which can refer to the internal bit depth, the bit depth of the recovered sample in the filtered CB, etc. In some examples, the tables (e.g., the luminance table, the saturation table) are obtained using equation (12).
number
[0107] [Table 2]
[0108] The slice header for the current slice can signal one or more APS indices (e.g., up to seven APS indices) to specify the luminance filter set that can be used for the current slice. The filtering process can be controlled at one or more appropriate levels, such as the picture level, slice level, or CTB level. In one embodiment, the filtering process can be further controlled at the CTB level. A flag can be signaled to indicate whether the ALF is applied to the luminance CTB. The luminance CTB can select a filter set from among multiple fixed filter sets (e.g., 16 fixed filter sets) and filter sets (also called signaled filter sets) that are signaled within the APS. A filter set index can be signaled to the luminance CTB to indicate the filter set being applied (e.g., a filter set among multiple fixed filter sets and signaled filter sets). Multiple fixed filter sets can be predefined and hardcoded in the encoder and decoder and can be called predefined filter sets.
[0109] For saturation components, the APS index can be signaled within the slice header to indicate the saturation filter set currently used for slicing. At the CTB level, if there are two or more saturation filter sets within the APS, the filter set index can be signaled for each saturation CTB.
[0110] The filter coefficients can be quantized to a norm equal to 128. To reduce the complexity of the multiplication, bitstream fit can be applied so that non-center position coefficient values can be within the range of -27 to 27-1, including both ends. In one example, the center position coefficients are not signaled in the bitstream and can be considered equal to 128.
[0111] In some embodiments, the syntax and semantics of the clipping index and clipping value are defined as follows: alf_luma_clip_idx[sfIdx][j] can be used to specify the clipping index of the clipping value to use before multiplying by the j-th coefficient of the signaled luminance filter indicated by sfIdx. Bitstream compatibility requirements may include the requirement that the value of alf_luma_clip_idx[sfIdx][j], where sfIdx=0~alf_luma_num_filters_signalled_minus1 and j=0~11, should be within the range of 0~3, including both ends. A luminance filter clipping value AlfClipL[adaptation_parameter_set_id] having element AlfClipL[adaptation_parameter_set_id][filtIdx][j], where filtIdx=0~NumAlfFilters-1 and j=0~11, can be derived as specified in Table 2, depending on a bitDepth set equal to BitDepthY and a clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j]. alf_chroma_clip_idx[altIdx][j] can be used to specify the clipping index of the clipping value to use before multiplying by the j-th coefficient of the alternative chroma filter having index altIdx. Bitstream compatibility requirements may include the requirement that the value of alf_chroma_clip_idx[altIdx][j], where altIdx=0~alf_chroma_num_alt_filters_minus1 and j=0~5, should be in the range of 0~3, including both ends. A chroma filter clipping value AlfClipC[adaptation_parameter_set_id][altIdx] having element AlfClipC[adaptation_parameter_set_id][altIdx][j], where altIdx=0~alf_chroma_num_alt_filters_minus1 and j=0~5, can be derived as specified in Table 2, depending on a bitDepth set equal to BitDepthC and clipIdx set equal to alf_chroma_clip_idx[altIdx][j].
[0112] In one embodiment, the filtering process can be described as follows: On the decoder side, when ALF is enabled for CTB, the sample R(i,j) in CU (or CB) can be filtered, and the filtered sample value R'(i,j) is obtained using equation (13) as shown below. In one example, each sample in CU is filtered.
number
[0113] In a nonlinear ALF, multiple sets of clipping values can be provided in Table 3. For example, the luminance set contains four clipping values {1024, 181, 32, 6}, and the saturation set contains four clipping values {1024, 161, 25, 4}. The four clipping values in the luminance set can be selected by dividing the entire range of sample values (e.g., 1024) for the luminance block (encoded in 10 bits) into approximately equal parts within the logarithmic domain. The range can be from 4 to 1024 for the saturation set.
[0114] [Table 3]
[0115] The selected clipping values can be coded within the “alf_data” syntax element as follows: To encode the clipping index corresponding to the selected clipping values shown in Table 3, an appropriate encoding scheme (e.g., Golomb encoding scheme) can be used. The encoding scheme may be the same one used to encode the filter set index.
[0116] In one embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of the ALF. Thus, modified block classification and filtering can be employed for samples near the CTU boundary (e.g., horizontal CTU boundary). The virtual boundary (1130) is defined as "N" as shown in Figure 11A, with respect to the horizontal CTU boundary (1120). samples By shifting only the samples, it can be defined as a line, N samples n can be a positive integer. For example, N samples For the luminance component, it is equal to 4, N samples For the saturation component, it is equal to 2.
[0117] Referring to Figure 11A, the modified block classification can be applied to the luminance component. For example, in the 1D Laplacian gradient calculation of a 4x4 block (1110) above a virtual boundary (1130), only samples above the virtual boundary (1130) are used. Similarly, referring to Figure 11B, in the 1D Laplacian gradient calculation of a 4x4 block (1111) below a virtual boundary (1131) shifted from the CTU boundary (1121), only samples below the virtual boundary (1131) are used. The quantization of the activation value A can be scaled accordingly by taking into account the reduction in the number of samples used in the 1D Laplacian gradient calculation.
[0118] For the filtering process, symmetric padding operations at the virtual boundary can be used for both luminance and chroma components. Figures 12A to 12F show examples of such modified ALF filtering for luminance components at the virtual boundary. When the sample being filtered is located below the virtual boundary, adjacent samples located above the virtual boundary can be padded. When the sample being filtered is located above the virtual boundary, adjacent samples located below the virtual boundary can be padded. Referring to Figure 12A, adjacent sample C0 can be padded by sample C2, which is located below the virtual boundary (1210). Referring to Figure 12B, adjacent sample C0 can be padded by sample C2, which is located above the virtual boundary (1220). Referring to Figure 12C, adjacent samples C1 to C3 can be padded by samples C5 to C7, which are located below the virtual boundary (1230), respectively. Referring to Figure 12D, adjacent samples C1 to C3 can be padded by samples C5 to C7, which are located above the virtual boundary (1240), respectively. Referring to Figure 12E, adjacent samples C4-C8 can be padded by samples C10, C11, C12, C11, and C10, respectively, which are located below the virtual boundary (1250). Referring to Figure 12F, adjacent samples C4-C8 can be padded by samples C10, C11, C12, C11, and C10, respectively, which are located above the virtual boundary (1260).
[0119] In some examples, the above explanation can be appropriately applied when a sample and its neighbors are located to the left (or right) and right (or left) of a virtual boundary.
[0120] According to aspects of this disclosure, a picture can be partitioned based on a filtering process to improve coding efficiency. In some examples, the CTU is also called the Maximum Coding Unit (LCU). In one example, the CTU or LCU may have a size of 64 × 64 pixels. In some embodiments, an LCU-aligned picture quadtree partition can be used for filtering-based partitioning. In some examples, a coding unit-synchronous picture quadtree-based adaptive loop filter can be used. For example, a luminance picture can be partitioned into several multilevel quadtree partitions, the boundary of each partition aligned with the boundary of the LCU. Each partition has its own filtering process and may therefore be called a Filter Unit (FU).
[0121] In some examples, a two-pass coding flow can be used. In the first pass of the two-pass coding flow, the quadtree partitioning pattern of the picture and the best filter for each FU can be determined. In some embodiments, the determination of the quadtree partitioning pattern of the picture and the best filter for each FU is based on filtering distortion. Filtering distortion can be estimated during the determination process by a Fast Filtering Distortion Estimation (FFDE) technique. The picture is partitioned using quadtree partitions. The recovered picture can be filtered according to the determined quadtree partitioning pattern and the selected filters for all FUs.
[0122] In the second pass of the two-pass coding flow, the on / off control of the CU-synchronous ALF is performed. Depending on the ALF on / off result, the initially filtered picture is partially recovered by the restored picture.
[0123] Specifically, in some examples, a top-down partitioning procedure is employed to divide an image into multilevel quadtree partitions using a rate-distortion criterion. Each partition is called a filter unit (FU). The partitioning process aligns the quadtree partitions to the LCU boundaries. The encoding order of the FUs follows the z-scan order.
[0124] Figure 13 shows examples of partitions according to some embodiments of the present disclosure. In the example in Figure 13, picture(1300) is divided into 10 FUs, with the encoding order being FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.
[0125] Figure 14 shows a quadtree partitioning pattern (1400) for a picture (1300). In the example in Figure 14, partitioning flags are used to indicate the partitioning pattern of the picture. For example, "1" indicates that a quadtree partition will be performed on the block, and "0" indicates that it will not be partitioned further. In some examples, the smallest FU has an LCU size, and partitioning flags are not required for the smallest FU. The partitioning flags are encoded and transmitted in z order as shown in Figure 14.
[0126] In some examples, the filter for each FU is selected from two filter sets based on a rate distortion criterion. The first set contains 1 / 2 symmetric square and rhombus filters derived for the current FU. The second set comes from a time-delay filter buffer, which stores filters previously derived for the FU of previous pictures. The filter with the minimum rate distortion cost from these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four child FUs, the rate distortion costs of the four child FUs are calculated. By recursively comparing the rate distortion costs with and without division, the picture quadtree partitioning pattern can be determined.
[0127] In some examples, the maximum quadtree partitioning level may be used to limit the maximum number of FUs. In one example, when the maximum quadtree partitioning level is 2, the maximum number of FUs is 16. Furthermore, during the quadtree partitioning decision, the correlation values can be reused to derive the Wiener coefficients of the 16 FUs (minimum FUs) at the lowest quadtree level. The remaining FUs can have their Wiener filters derived from the correlations of the 16 FUs at the lowest quadtree level. Thus, in that example, only one framebuffer access is performed to derive the filter coefficients for all FUs.
[0128] After the quadtree partitioning pattern is determined, CU-synchronous ALF on / off control can be implemented to further reduce filtering distortion. By comparing the filtering distortion and unfiltering distortion at each leaf CU, the leaf CU can explicitly switch the ALF on or off in its local region. In some examples, coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off results.
[0129] The intercomponent filtering process can apply intercomponent filters, such as intercomponent adaptive loop filters (CC-ALF). These intercomponent filters can use luminance sample values from the luminance component (e.g., luminance CB) to refine the saturation component (e.g., saturation CB corresponding to the luminance CB). In one example, the luminance CB and saturation CB are contained within the CU.
[0130] Figure 15 shows an inter-component filter (e.g., CC-ALF) used to generate a chroma component according to one embodiment of the present disclosure. In some examples, Figure 15 shows filtering processes for a first chroma component (e.g., first chroma CB), a second chroma component (e.g., second chroma CB), and a luminance component (e.g., luminance CB). The luminance component can be filtered by a sample-adaptive offset (SAO) filter (1510) to produce an SAO-filtered luminance component (1541). The SAO-filtered luminance component (1541) can be further filtered by an ALF luminance filter (1516) to become a filtered luminance CB (1561) (e.g., "Y").
[0131] The first saturation component can be filtered by an SAO filter (1512) and an ALF saturation filter (1518) to generate a first intermediate component (1552). Furthermore, the SAO-filtered luminance component (1541) can be filtered by an inter-component filter (e.g., CC-ALF) (1521) for the first saturation component to generate a second intermediate component (1542). Subsequently, a filtered first saturation component (1562) (e.g., "Cb") can be generated based on at least one of the second intermediate component (1542) and the first intermediate component (1552). In one example, the filtered first saturation component (1562) (e.g., "Cb") can be generated by combining the second intermediate component (1542) and the first intermediate component (1552) with an adder (1522). The intercomponent adaptive loop filtering process for the first saturation component may include steps performed by a CC-ALF (1521) and steps performed by, for example, an adder (1522).
[0132] The above description can be adapted to a second saturation component. The second saturation component can be filtered by an SAO filter (1514) and an ALF saturation filter (1518) to generate a third intermediate component (1553). Furthermore, the SAO-filtered luminance component (1541) can be filtered by an inter-component filter (e.g., CC-ALF) (1531) for the second saturation component to generate a fourth intermediate component (1543). Subsequently, a filtered second saturation component (1563) (e.g., "Cr") can be generated based on at least one of the fourth intermediate component (1543) and the third intermediate component (1553). In one example, the filtered second saturation component (1563) (e.g., "Cr") can be generated by combining the fourth intermediate component (1543) and the third intermediate component (1553) with an adder (1532). In one example, the inter-component adaptive loop filtering process for the second saturation component may include steps performed by a CC-ALF (1531) and steps performed by, for example, an adder (1532).
[0133] Inter-component filters (e.g., CC-ALF(1521), CC-ALF(1531)) can operate by applying a linear filter with any appropriate filter shape to the luminance component (or luminance channel) to improve each chroma component (e.g., the first chroma component, the second chroma component).
[0134] Figure 16 shows an example of a filter (1600) according to one embodiment of the present disclosure. The filter (1600) may include non-zero filter coefficients and zero filter coefficients. The filter (1600) has a rhombus shape (1620) formed by the filter coefficients (1610) (indicated by the black-filled circles). In one example, the non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and the filter coefficients not included in the filter coefficients (1610) are zero. Thus, the non-zero filter coefficients in the filter (1600) are included in the rhombus shape (1620), and the filter coefficients not included in the rhombus shape (1620) are zero. In one example, the number of filter coefficients in the filter (1600) is equal to the number of filter coefficients (1610), which is 18 in the example shown in Figure 16.
[0135] A CC-ALF can contain any suitable filter coefficients (also called CC-ALF filter coefficients). Referring back to Figure 15, CC-ALF(1521) and CC-ALF(1531) can have the same filter shape, such as the rhombus shape (1620) shown in Figure 16, and the same number of filter coefficients. In one example, the values of the filter coefficients in CC-ALF(1521) are different from the values of the filter coefficients in CC-ALF(1531).
[0136] In general, filter coefficients within CC-ALF (e.g., non-zero filter coefficients) can be transmitted within APS, for example. In one example, the filter coefficients are multiples (e.g., 2 10It can be scaled by and rounded for fixed-point representation. The application of CC-ALF is controlled for variable block sizes and can be signaled by context coding flags (e.g., CC-ALF enable flags) received for each block of samples. Context coding flags, such as the CC-ALF enable flag, can be signaled at any appropriate level, such as the block level. The block size, along with the CC-ALF enable flag, can be received at the slice level for each saturation component. In some examples, block sizes of 16x16, 32x32, and 64x64 (within saturation samples) can be supported.
[0137] Figure 17 shows syntax examples for CC-ALF according to several embodiments of the present disclosure. In the example in Figure 17, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index indicating whether an inter-component Cb filter is used, and if so, the index of the inter-component Cb filter. For example, when alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the intercomponent Cb filter is not applied to the block of Cb color component samples at the luminance position (xCtb, yCtb), and when alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not equal to 0, alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index for the filter to be applied. For example, the alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]-th intercomponent Cb filter is applied to a block of Cb color component samples at the luminance position (xCtb, yCtb).
[0138] Furthermore, in the example in Figure 17, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is used to indicate whether an intercomponent Cr filter is used, and if so, it is the index of the intercomponent Cr filter. For example, when alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, the intercomponent Cr filter is not applied to the block of Cr color component samples at the luminance position (xCtb, yCtb), and when alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not equal to 0, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is the index of the intercomponent Cr filter. For example, the alf_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY]-th intercomponent Cr filter can be applied to a block of Cr color component samples at the luminance position (xCtb, yCtb).
[0139] In some examples, a chroma subsampling technique is used, so that the number of samples in each chroma block can be less than the number of samples in the luminance block. The chroma subsampling format (also called the chroma subsampling format, specified, for example, by chroma_format_idc) can specify the horizontal chroma subsampling coefficient (e.g., SubWidthC) and the vertical chroma subsampling coefficient (e.g., SubHeightC) between each chroma block and the corresponding luminance block. In one example, the chroma subsampling format is 4:2:0, so that the horizontal chroma subsampling coefficient (e.g., SubWidthC) and the vertical chroma subsampling coefficient (e.g., SubHeightC) are 2, as shown in Figures 18A and 18B. In another example, the chroma subsampling format is 4:2:2, so that the horizontal chroma subsampling coefficient (e.g., SubWidthC) is 2 and the vertical chroma subsampling coefficient (e.g., SubHeightC) is 1. In one example, the saturation subsampling format is 4:4:4, and therefore the horizontal saturation subsampling coefficient (e.g., SubWidthC) and the vertical saturation subsampling coefficient (e.g., SubHeightC) are 1. The saturation sample type (also called the saturation sample position) can indicate the relative position of a saturation sample in a saturation block to at least one corresponding luminance sample in a luminance block.
[0140] Figures 18A and 18B show exemplary positions of chroma samples relative to luminance samples according to embodiments of the present disclosure. Referring to Figure 18A, luminance sample (1801) is located in rows (1811) to (1818). The luminance sample (1801) shown in Figure 18A may represent a portion of a picture. In one example, a luminance block (e.g., luminance CB) contains luminance sample (1801). A luminance block may correspond to two chroma blocks having a 4:2:0 chroma subsampling format. In one example, each chroma block contains chroma sample (1803). Each saturation sample (for example, saturation sample (1803(1))) corresponds to four luminance samples (for example, luminance samples (1801(1)) to (1801(4))). In one example, the four luminance samples are the top-left sample (1801(1)), the top-right sample (1801(2)), the bottom-left sample (1801(3)), and the bottom-right sample (1801(4)). The saturation sample (for example, (1803(1))) corresponds to the top-left sample (1801(1)) and the bottom-left sample (18 A chroma sample type of a chroma block located at the left-center position between 01(3)) and having the chroma sample (1803) can be called chroma sample type 0. Chroma sample type 0 indicates relative position 0, which corresponds to the left-center position midway between the top-left sample (1801(1)) and the bottom-left sample (1801(3)). The four luminance samples (e.g., (1801(1)) to (1801(4))) can be called adjacent luminance samples of chroma sample (1803)(1).
[0141] In one example, each saturation block contains a saturation sample (1804). The above description referring to saturation sample (1803) can be adapted to saturation sample (1804), and therefore, for brevity, detailed descriptions can be omitted. Each of the saturation samples (1804) can be located at the center of four corresponding luminance samples, and the saturation sample type of a saturation block having saturation samples (1804) can be called saturation sample type 1. Saturation sample type 1 indicates relative position 1 corresponding to the center of four luminance samples (e.g., (1801(1)) to (1801(4))). For example, one of the saturation samples (1804) can be located in the central part of luminance samples (1801(1)) to (1801(4)).
[0142] In one example, each saturation block contains a saturation sample (1805). Each saturation sample (1805) can be located in the top-left position, corresponding to the top-left sample of the four corresponding luminance samples (1801). The saturation sample type of a saturation block containing a saturation sample (1805) can be called saturation sample type 2. Thus, each saturation sample (1805) is located in the same position as the top-left sample of the four luminance samples (1801) corresponding to its respective saturation sample. Saturation sample type 2 indicates relative position 2, corresponding to the top-left position of the four luminance samples (1801). For example, one of the saturation samples (1805) can be located in the top-left position of luminance samples (1801(1)) to (1801(4)).
[0143] In one example, each saturation block contains a saturation sample (1806). Each of the saturation samples (1806) can be located at the upper center position between the corresponding upper-left sample and the corresponding upper-right sample, and the saturation sample type of a saturation block containing saturation samples (1806) can be called saturation sample type 3. Saturation sample type 3 indicates relative position 3, which corresponds to the upper center position between the upper-left sample and the upper-right sample. For example, one of the saturation samples (1806) can be located at the upper center position of the luminance samples (1801(1)) to (1801(4)).
[0144] In one example, each saturation block contains a saturation sample (1807). Each saturation sample (1807) can be located in the lower left position, corresponding to the lower left sample of the four corresponding luminance samples (1801). The saturation sample type of a saturation block containing saturation samples (1807) can be called saturation sample type 4. Thus, each saturation sample (1807) is located in the same position as the lower left sample of the four luminance samples (1801) corresponding to each saturation sample. Saturation sample type 4 indicates relative position 4, corresponding to the lower left position of the four luminance samples (1801). For example, one of the saturation samples (1807) can be located in the lower left position of luminance samples (1801(1)) to (1801(4)).
[0145] In one example, each saturation block contains a saturation sample (1808). Each of the saturation samples (1808) is located at the lower center position between the lower left and lower right samples, and the saturation sample type of a saturation block containing saturation samples (1808) can be called saturation sample type 5. Saturation sample type 5 indicates a relative position 5 corresponding to the lower center position between the lower left and lower right samples of the four luminance samples (1801). For example, one of the saturation samples (1808) may be located between the lower left and lower right samples of the luminance samples (1801(1)) to (1801(4)).
[0146] In general, any suitable saturation sample type can be used in a saturation subsampling format. Saturation sample types 0-5 are exemplary saturation sample types described in the saturation subsampling format 4:2:0. Additional saturation sample types can be used in the saturation subsampling format 4:2:0. Furthermore, other saturation sample types and / or variations of saturation sample types 0-5 can be used in other saturation subsampling formats such as 4:2:2 and 4:4:4. For example, a saturation sample type combining saturation samples (1805) and (1807) is used in the saturation subsampling format 4:2:2.
[0147] In one example, each luminance block is considered to have alternating rows such as (1811)~(1812) containing the top two samples (e.g., (1801(1))~(1801(4))) and the bottom two samples (e.g., (1801(3))~(1801(4))) of the four luminance samples (e.g., (1801(1))~(1801(4))). Thus, rows (1811), (1813), (1815), and (181 7) can be called the current row (also called the top field), and rows (1812), (1814), (1816), and (1818) can be called the next row (also called the bottom field). Four luminance samples (e.g., (1801(1)) to (1801(4))) are located in the current row (e.g., (1811)) and the next row (e.g., (1812)). Relative positions 2 to 3 are located in the current row, relative positions 0 to 1 are located between each current row and their respective next row, and relative positions 4 to 5 are located in the next row.
[0148] Saturation samples (1803), (1804), (1805), (1806), (1807), or (1808) are located in rows (1851) to (1854) within their respective saturation blocks. The specific position in rows (1851) to (1854) can depend on the saturation sample type of the saturation sample. For example, for saturation samples (1803) to (1804), each having saturation sample types 0 to 1, row (1851) is located between rows (1811) and (1812). For saturation samples (1805) to (1806), each having saturation sample types 2 to 3, row (1851) is in the same position as the current row (1811). For saturation samples (1807)-(1808) with saturation sample types 4-5, row (1851) is in the same position as the next row (1812). The above explanation can be appropriately adapted to rows (1852)-(1854), and for brevity, a detailed explanation is omitted.
[0149] Any suitable scanning method can be used to display, store, and / or transmit the luminance blocks and corresponding saturation blocks described above in Figure 18A. In one example, sequential scanning is used.
[0150] As shown in Figure 18B, skip scanning can be used. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc is equal to 1). In one example, the variable chroma location type (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current rows (1811), (1813), (1815), and (1817), as well as the next rows (1812), (1814), (1816), and (1818), can be scanned separately. For example, the current rows (1811), (1813), (1815), and (1817) can be scanned first, followed by the next rows (1812), (1814), (1816), and (1818). The current row may contain a luminance sample (1801), and the next row may contain a luminance sample (1802).
[0151] Similarly, corresponding saturation blocks can be scanned by skipping them. Rows (1851) and (1853) containing saturation samples (1803), (1804), (1805), (1806), (1807), or (1808) without fills can be called the current row (or current saturation row), and rows (1852) and (1854) containing saturation samples (1803), (1804), (1805), (1806), (1807), or (1808) with gray fills can be called the next row (or next saturation row). In one example, during a skip scan, rows (1851) and (1853) are scanned first, followed by rows (1852) and (1854).
[0152] In some cases, constrained directional extension filtering techniques can be used. The use of an in-loop constrained directional extension filter (CDEF) can remove coding artifacts while preserving image detail. In one example (e.g., HEVC), the sample-adaptive offset (SAO) algorithm can achieve a similar objective by defining signal offsets for different classes of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some cases, CDEF can be constrained to be easily vectorizable (i.e., implementable with single-instruction multiple-data (SIMD) operations). Note that other nonlinear filters, such as median filters and bilateral filters, cannot be treated similarly.
[0153] In some cases, the amount of ringing artifacts in the coded image tends to be approximately proportional to the quantization step size. While the amount of detail is a characteristic of the input image, the smallest detail retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of ringing is generally smaller than the amplitude of detail.
[0154] CDEF can be used to identify the orientation of each block, then adaptively filter along the identified orientation, and filter less along the orientation rotated 45 degrees from the identified orientation. In some examples, the encoder can retrieve the filter intensity, which can be explicitly signaled, thereby allowing for greater control over blurring.
[0155] Specifically, in some examples, orientation lookup is performed on the recovered pixels immediately after the deblocking filter. Since these pixels are available to the decoder, the decoder can then look up the orientation, and therefore, in one example, the orientation does not require signaling. In some examples, orientation lookup can operate on a specific block size, such as an 8x8 block, which is small enough to handle nonlinear edges well, but large enough to reliably estimate orientation when applied to a quantized image. Having a consistent orientation on an 8x8 area also facilitates the vectorization of the filter. In some examples, each block (e.g., 8x8) can be compared to a perfectly directional block to determine the difference. A perfectly directional block is one in which all pixels along a line in a certain direction have the same value. In one example, a difference measure for each block and perfectly directional block can be calculated, such as the sum of squared differences (SSD), root mean square (RMS) error, etc. Then, a perfectly directional block with the smallest difference (e.g., smallest SSD, smallest RMS, etc.) can be determined, and the orientation of the determined perfectly directional block may be the orientation that best matches the pattern within the block.
[0156] Figure 19 shows an example of direction finding according to one embodiment of the present disclosure. In this example, block (1910) is an 8x8 block that has been restored and output from a deblocking filter. In the example in Figure 19, direction finding can determine the direction of block (1910) from the eight directions indicated by (1920). Eight perfectly directional blocks (1930) are formed corresponding to each of the eight directions (1920). A perfectly directional block corresponding to a direction is a block in which the pixels along the line of that direction have the same value. Furthermore, difference measures such as SSD and RMS error can be calculated for each of block (1910) and the perfectly directional block (1930). In the example in Figure 19, the RMS error is indicated by (1940). As indicated by (1943), the RMS errors of block (1910) and the perfectly directional block (1933) are smallest, and therefore direction (1923) is the direction that best matches the pattern in block (1910).
[0157] After the orientation of the block is identified, a nonlinear low-pass directional filter can be determined. For example, the filter taps of a nonlinear low-pass directional filter can be aligned along the identified orientation to reduce ringing while maintaining directional edges or directional patterns. However, in some examples, directional filtering alone may not sufficiently reduce ringing. In one example, additional filter taps are used on pixels that are not aligned along the identified orientation. To reduce the risk of blurring, the extra filter taps are treated more sparingly. For this reason, the CDEF includes first-order and second-order filter taps. In one example, a complete 2-D CDEF filter is given by equation (14):
number
[0158] In some cases, in addition to deblocking operations, in-loop restoration techniques are used in post-deblocking video coding to generally remove noise and improve edge quality. In one example, the in-loop restoration technique is switchable within a frame for each appropriately sized tile. The in-loop restoration technique is based on a separable symmetric Wiener filter, a dual self-inductive filter with subspace projection, and a region-transform recursive filter. Because content statistics can change significantly within a frame, the in-loop restoration technique is integrated into a switchable framework that can trigger different techniques in different regions of the frame.
[0159] A separable symmetric Wiener filter can be one in-loop reconstruction method. In some examples, every pixel in a degraded frame can be reconstructed as a non-causally filtered version of the pixels in a w×w window around it, where w=2r+1 is odd for an integer r. 2D filter taps are in column vectorized form w 2 When expressed as a single-element vector F, the direct LMMSE optimization is given by F = H -1 This leads to the filter parameters given by M, where H = E[XX] T ] is the autocovariance of x, and w in the w×w window around the pixel. 2 This is a column vectorized version of the individual samples, where M=E[YX T] is the cross-correlation between x and the scalar source sample y that should be estimated. In one example, the encoder can estimate H and M from the realized values in the deblocked frame and source and send the resulting filter F to the decoder. However, if we do that, w 2 Not only does transmitting individual taps incur a considerable bitrate cost, but the inseparable filtering makes decoding excessively complex. In some embodiments, several additional constraints are imposed on the properties of F. The first constraint is that F is constrained to be separable, and as a result, filtering can be performed as separable horizontal and vertical w-tap convolutions. The second constraint is that each of the horizontal and vertical filters is constrained to be symmetric. The third constraint is that it is assumed that the sum of both the horizontal and vertical filter coefficients is 1.
[0160] Dual self-inductive filtering using subspace projection can be one of the in-loop reconstruction methods. Inductive filtering is an image filtering technique in which a local linear model, shown by equation (15), is used to compute a filtered output y from an unfiltered sample x. Equation (15): y = Fx + G Here, F and G are determined based on statistics of the degraded image and the guidance image of the neighboring pixels of the filtered image. If the guide image is the same as the degraded image, the resulting so-called self-inductive filtering has the effect of edge-preserving smoothing. In one example, a specific form of self-inductive filtering can be used. The specific form of self-inductive filtering depends on two parameters: radius r and noise parameter e, and is enumerated as follows: 1. Mean μ and variance σ of pixels within a (2r+1)×(2r+1) window around all pixels. 2 This step can be efficiently performed using box filtering based on integral imaging. 2. For all pixels, f = σ2 / (σ 2 Calculate g = (1-f)μ (+e). 3. Calculate F and G for all pixels as the average of the f and g values within a 3x3 window around the pixel being used.
[0161] The specific form of self-inductive filtering is controlled by r and e, where a larger r results in greater spatial variance, and a larger e results in greater range variance.
[0162] Figure 20 shows an example illustrating subspace projection in several cases. As shown in Figure 20, even when neither the reconstructed X1 nor X2 is close to the source Y, a suitable multiplier {α,β} can bring them fairly close to the source, as long as they are moving somewhat in the right direction.
[0163] In some cases (e.g., HEVC), a filtering technique called Sample Adaptive Offset (SAO) can be used. In some cases, SAO is applied to the reconstructed signal after a deblocking filter. SAO can use an offset value given in the slice header. In some cases, for luminance samples, the encoder can decide whether to apply (enable) SAO to the slice. When SAO is enabled, the current picture allows for recursive division into four sub-regions of the coding unit, and each sub-region can select an SAO type from several SAO types based on the features within the sub-region.
[0164] Figure 21 shows Table (2100) of several SAO types according to one embodiment of the present disclosure. Table (2100) shows SAO types 0 to 6. Note that SAO type 0 is used to indicate that SAO is not applied. Furthermore, each SAO type from SAO type 1 to SAO type 6 includes multiple categories. SAO can reduce distortion by classifying the restored pixels of a subregion into categories and adding an offset to the pixels of each category within the subregion. In some examples, edge characteristics can be used for pixel classification in SAO types 1 to 4, and pixel intensity can be used for pixel classification in SAO types 5 to 6.
[0165] Specifically, in one embodiment, such as SAO types 5-6, a band offset (BO) can be used to classify all pixels in a subregion into multiple bands. Each of the multiple bands contains pixels within the same intensity interval. In some examples, the intensity range is divided equally into multiple intervals, such as 32 intervals from 0 to the maximum intensity value (e.g., 255 for an 8-bit pixel), and each interval is associated with an offset. Furthermore, in one example, the 32 bands are divided into two groups, such as a first group and a second group. The first group contains the central 16 bands (e.g., 16 intervals in the middle of the intensity range), and the second group contains the remaining 16 bands (e.g., 8 intervals on the lower side of the intensity range and 8 intervals on the higher side of the intensity range). In one example, only one offset from the two groups is transmitted. In some embodiments, when pixel classification operations are used in BO, the five most significant bits of each pixel can be used directly as the band index.
[0166] Furthermore, in one embodiment, such as SAO types 1-4, edge offsets (EOs) can be used for pixel classification and offset determination. For example, pixel classification can be determined based on a one-dimensional three-pixel pattern, taking edge directional information into consideration.
[0167] Figure 22 shows examples of 3-pixel patterns for pixel classification at edge offsets in several cases. In the examples in Figure 22, the first pattern (2210) (shown by 3 gray pixels) is called the 0-degree pattern (the 0-degree pattern is associated with the horizontal direction), the second pattern (2220) (shown by 3 gray pixels) is called the 90-degree pattern (the 90-degree pattern is associated with the vertical direction), the third pattern (2230) (shown by 3 gray pixels) is called the 135-degree pattern (the 135-degree pattern is associated with the 135-degree diagonal direction), and the fourth pattern (2240) (shown by 3 gray pixels) is called the 45-degree pattern (the 45-degree pattern is associated with the 45-degree diagonal direction). In one example, one of the four directional patterns shown in Figure 22 can be selected, taking into account the edge directional information of a subregion. The selection can be transmitted in the coded video bitstream as side information in one example. Next, pixels within a sub-region can be classified into several categories by comparing each pixel with its two adjacent pixels in the direction associated with the directional pattern.
[0168] Figure 23 is a table (2300) for pixel classification rules for edge offsets in several examples. Specifically, pixel c (also shown in each pattern in Figure 22) is compared with two adjacent pixels (also shown in gray in each pattern in Figure 22), and pixel c can be classified into one of categories 0 to 4 based on the comparison using the pixel classification rules shown in Figure 23.
[0169] In some embodiments, the decoder-side SAO can operate independently of the maximum coding unit (LCU) (e.g., CTU) to conserve line buffers. In some examples, pixels in the topmost and bottommost rows within each LCU are not SAO-processed when the 90-degree, 135-degree, and 45-degree classification patterns are selected, and pixels in the leftmost and rightmost columns within each LCU are not SAO-processed when the 0-degree, 135-degree, and 45-degree patterns are selected.
[0170] Figure 24 shows an example of syntax (2400) that may need to be signaled for a CTU when the parameter is not merged from an adjacent CTU. For example, the syntax element sao_type_idx[cldx][rx][ry] can be signaled to indicate the SAO type of a subregion. The SAO type may be BO (band offset) or EO (edge offset). When sao_type_idx[cldx][rx][ry] has a value of 0, it indicates that the SAO is OFF; values from 1 to 4 indicate that one of the four EO categories corresponding to 0°, 90°, 135°, and 45° is used; and a value of 5 indicates that BO is used. In the example in Figure 24, each of the BO type and EO type has four SAO offset values to be signaled (sao_offset[cIdx][rx][ry][0]~sao_offset[cIdx][rx][ry][3]).
[0171] Generally, a filtering process can use a restored sample of a first color component as input (e.g., Y, Cb, or Cr, or R, G, or B) to generate an output, and the output of the filtering process is applied to a second color component that may be the same color component as the first color component, or a different color component from the first color component.
[0172] In a relevant example of intercomponent filtering (CCF), filter coefficients are derived based on several mathematical equations. The derived filter coefficients are signaled from the encoder side to the decoder side, and the derived filter coefficients are used to generate an offset using a linear combination. The generated offset is then added to the restored sample as a filtering process. For example, the offset is generated based on a linear combination of the filtering coefficients with the luminance sample, and the generated offset is added to the restored chroma sample. Relevant examples of CCF are based on the assumption of a linear mapping relationship between the restored luminance sample value and the delta value between the original chroma sample and the restored chroma sample. However, the mapping between the restored luminance sample value and the delta value between the original chroma sample and the restored chroma sample does not necessarily follow a linear mapping process, and therefore, the coding performance of CCF may be limited under the assumption of a linear mapping relationship.
[0173] In some cases, nonlinear mapping techniques can be used in inter-component filtering and / or homochromatic component filtering without significant signaling overhead. In one example, nonlinear mapping techniques can be used in inter-component filtering to generate inter-component sample offsets. In another example, nonlinear mapping techniques can be used in homochromatic component filtering to generate local sample offsets.
[0174] For convenience, a filtering process that uses nonlinear mapping techniques can be called a nonlinear mapping sample offset (SO-NLM). SO-NLM in intercomponent filtering processes can be called intercomponent sample offset (CCSO). SO-NLM in homochromatic component filtering can be called a local sample offset (LSO). A filter that uses nonlinear mapping techniques can be called a nonlinear mapping-based filter. Nonlinear mapping-based filters can include CCSO filters, LSO filters, and the like.
[0175] In one example, CCSO and LSO can be used as loop filtering to reduce distortion in the restored sample. CCSO and LSO do not depend on the linear mapping assumptions used in the relevant exemplary CCF. For example, CCSO does not depend on the assumption of a linear mapping relationship between the luminance restored sample value and the delta value between the original saturation sample and the saturation restored sample. Similarly, LSO does not depend on the assumption of a linear mapping relationship between the color component restored sample value and the delta value between the original sample of the color component and the color component restored sample.
[0176] The following description outlines an SO-NLM filtering process that uses a restored sample of a first color component as input (e.g., Y, Cb, or Cr, or R, G, or B) to generate an output, and the output of the filtering process is applied to a second color component. When the second color component is the same as the first color component, the description is applicable to LSO; when the second color component is different from the first color component, the description is applicable to CCSO.
[0177] In SO-NLM, the nonlinear mapping is derived on the encoder side. The nonlinear mapping lies between the restored sample of the first color component within the filter support region and the offset applied to the second color component within the filter support region. When the second color component is the same as the first color component, the nonlinear mapping is used in LSO; when the second color component is different from the first color component, the nonlinear mapping is used in CCSO. The region of the nonlinear mapping is determined by different combinations of processed input restored samples (also called combinations of possible restored sample values).
[0178] The SO-NLM technique can be illustrated using a specific example. In this example, a reconstructed sample is determined from a first color component located within a filter support area (also called the "filter support region"). The filter support area is the area to which the filter can be applied, and it can have any suitable shape.
[0179] Figure 25 shows an example of a filter support area (2500) according to some embodiments of the present disclosure. The filter support area (2500) includes four restored samples of a first color component: P0, P1, P2, and P3. In the example of Figure 25, the four restored samples can form a cross shape vertically and horizontally, with the center of the cross shape being the position for the sample to be filtered. A sample located at the center and having the same color component as P0-P3 is denoted by C. A sample located at the center and having a second color component is denoted by F. The second color component may be the same as the first color component of P0-P3, or it may be different from the first color component of P0-P3.
[0180] Figure 26 shows an example of another filter support area (2600) according to some embodiments of the present disclosure. The filter support area (2600) includes four restored samples P0, P1, P2, and P3 of a first color component that form a square. In the example of Figure 26, the center position of the square is the position of the sample to be filtered. A sample at the center position that has the same color component as P0-P3 is denoted by C. A sample at the center position that has a second color component is denoted by F. The second color component may be the same as the first color component of P0-P3, or it may be different from the first color component of P0-P3.
[0181] The reconstructed sample is fed into the SO-NLM filter and processed appropriately to form filter taps. In one example, the position of the reconstructed sample that is input to the SO-NLM filter is called the filter tap position. In a specific example, the reconstructed sample is processed in the following two steps:
[0182] In the first step, the delta values between P0, P3, and C are calculated. For example, m0 represents the delta value between P0 and C, m1 represents the delta value between P1 and C, m2 represents the delta value between P2 and C, and m3 represents the delta value between P3 and C.
[0183] In the second step, the delta values m0 to m3 are further quantized, and the quantized values are denoted as d0, d1, d2, and d3. In one example, the quantized value can be one of -1, 0, or 1 based on the quantization process. For example, when m is less than -N (where N is a positive value and is called the quantization step size), the value m can be quantized to -1; when m is within the range [-N, N], the value m can be quantized to 0; and when m is greater than N, the value m can be quantized to 1. In some examples, the quantization step size N can be one of 4, 8, 12, 16, etc.
[0184] In some embodiments, the quantization values d0 to d3 are filter taps and can be used to identify a single combination within the filter region. For example, filter taps d0 to d3 can form a combination within the filter region. Each filter tap can have three quantization values, and therefore, when four filter taps are used, the filter region contains 81 (3 × 3 × 3 × 3) combinations.
[0185] Figures 27A to 27C show a table (2700) having 81 combinations according to one embodiment of the present disclosure. Table (2700) contains 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, the first column contains the index of the combination, the second column contains the value of filter tap d0 for the combination, the third column contains the value of filter tap d1 for the combination, the fourth column contains the value of filter tap d2 for the combination, the fifth column contains the value of filter tap d3 for the combination, and the sixth column contains the offset value associated with the combination for nonlinear mapping. In one example, once filter taps d0 to d3 are determined, the offset values (denoted by s) associated with the combinations d0 to d3 can be determined according to table (2700). In one example, the offset values s0 to s80 are integers such as 0, 1, -1, 3, -3, 5, -5, -7, etc.
[0186] In some embodiments, the final filtering process of SO-NLM can be applied as shown in equation (16), f'=clip(f+s) Equation (16) Here, f is the restored sample of the second color component to be filtered, and s is the offset value determined according to the filter tap, which is the result of processing the restored sample of the first color component, using a table (2700), etc. The sum of the restored sample F and the offset value s is further clipped to a range associated with the bit depth to determine the final filtered sample f' of the second color component.
[0187] Note that in the case of LSO, the second color component in the above explanation is the same as the first color component, while in the case of CCSO, the second color component in the above explanation may be different from the first color component.
[0188] It should be noted that the above description may be adapted for other embodiments of this disclosure.
[0189] In some examples, on the encoder side, the encoding device can derive a mapping between the restored sample of the first color component within the filter support region and the offset applied to the restored sample of the second color component. The mapping can be any suitable linear or nonlinear mapping. The filtering process can then be applied on the encoder side and / or decoder side based on the mapping. For example, the mapping is appropriately communicated to the decoder (e.g., the mapping is included in the coded video bitstream sent from the encoder side to the decoder side), and the decoder can then perform the filtering process based on the mapping.
[0190] The performance of nonlinear mapping-based filters, such as CCSO filters and LSO filters, depends on the filter geometry. The filter geometry (also called the filter shape) refers to the characteristics of the pattern formed by the filter tap locations. The pattern can be defined by various parameters, including the number of filter taps, the geometric shape of the filter tap locations, and the distance of the filter tap locations to the center of the pattern. Using a fixed filter geometry may limit the performance of nonlinear mapping-based filters.
[0191] As shown in Figures 24 and 25 and Figures 27A–27C, some examples use a 5-tap filter design for the filter geometry configuration of a nonlinear mapping-based filter. The 5-tap filter design can use tap positions P0, P1, P2, P3, and C. The 5-tap filter design for the filter geometry configuration can result in a reference table (LUT) with 81 entries, as shown in Figures 27A–27C. The sample offset LUT needs to be signaled from the encoder side to the decoder side, and the signaling of the LUT contributes to a large portion of the signaling overhead, which can affect the coding efficiency when using a nonlinear mapping-based filter. According to some aspects of this disclosure, the number of filter taps may differ from 5. In some examples, the number of filter taps can be reduced, allowing for more information to be captured within the filter support area and improving coding efficiency.
[0192] In some examples, the filter geometry configurations within a group for nonlinear mapping-based filters each have three filter taps.
[0193] Figure 28 shows seven filter shape configurations for three filter taps in one example. Specifically, the first filter shape configuration includes three filter taps at positions labeled "1" and "c", where position "c" is the center position of position "1", the second filter shape configuration includes three filter taps at positions labeled "2" and "c", where position "c" is the center position of position "2", the third filter shape configuration includes three filter taps at positions labeled "3" and "c", where position "c" is the center position of position "3", and the fourth filter shape configuration includes three filter taps at positions labeled "4" and "c" The fifth filter shape configuration includes three filter taps at positions labeled "5" and "c", with position "c" being the center position of position "4". The sixth filter shape configuration includes three filter taps at positions labeled "6" and "c", with position "c" being the center position of position "6". The seventh filter shape configuration includes three filter taps at positions labeled "7" and "c", with position "c" being the center position of position "7".
[0194] According to some aspects of this disclosure, a nonlinear mapping-based filter can be used in conjunction with other in-loop filters in a loop filter chain. The position of the nonlinear mapping-based filter may affect the coding efficiency of the nonlinear mapping-based filter.
[0195] Figure 29 shows block diagrams of loop filter chains (2900) in several examples. A loop filter chain (2900) includes multiple filters connected in series within a filter chain. In one example, a loop filter chain (2900) can be used as a loop filter unit (556). A loop filter chain (2900) can be used in an encoding loop or decoding loop before storing the restored picture in a decoding picture buffer such as a reference picture memory (557). The loop filter chain (2900) receives an input restored sample from a previous processing module and applies filters to the restored sample to generate an output restored sample.
[0196] The loop filter chain (2900) can include any suitable filters. In the example in Figure 29, the loop filter chain (2900) includes a deblocking filter (labeled Deblocking), a constrained directional extension filter (labeled CDEF), and an in-loop reconstruction filter (labeled LR) connected in the chain. The loop filter chain (2900) has an input node (2901), an output node (2909), and several intermediate nodes (2902)~(2903). The input node (2901) of the loop filter chain (2900) receives the input reconstruction sample from the previous processing module, and the input reconstruction sample is provided to the deblocking filter. The intermediate node (2902) receives the reconstruction sample from the deblocking filter (after being processed by the deblocking filter) and provides the reconstruction sample to the CDEF for further filtering. The intermediate node (2903) receives the reconstructed sample from the CDEF (after processing by the CDEF) and provides the reconstructed sample to the LR filter for further filtering. The output node (2909) receives the output reconstructed sample from the LR filter (after processing by the LR filter). The output reconstructed sample can be provided to other processing modules, such as post-processing modules, for further processing.
[0197] Note that the following description illustrates a technique using a nonlinear mapping-based filter based on a loop filter chain (2900). This technique of using a nonlinear mapping-based filter can be used in other suitable loop filter chains.
[0198] According to one aspect of the present disclosure, a nonlinear mapping-based filter can be coupled in series with other filters in a loop filter chain, the inputs and outputs of the nonlinear mapping-based filter can be located at the same node in the loop filter chain, and no other filters exist between the inputs and outputs of the nonlinear mapping-based filter. For example, the nonlinear mapping-based filter receives a reconstructed sample at a node in the loop filter chain, determines a sample offset based on the reconstructed sample at the node in the loop filter chain, and then applies the sample offset to the reconstructed sample at the node in the loop filter chain.
[0199] Figures 30A to 30D show examples of loop filter chains that include nonlinear mapping-based filters coupled in series with other filters within the loop filter chain.
[0200] FIG. 30A shows an example of a loop filter chain (3000A) in one example. The loop filter chain (3000A) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3000A), a non-linear mapping-based filter (labeled as SO-NLM) is applied at the input node. Specifically, the input of the non-linear mapping-based filter (also called the first restored sample) is the input restored sample to the loop filter chain (3000A) at the first node (3011A). Based on the input, the non-linear mapping-based filter generates a sample offset (SO). The sample offset is combined with an intermediate restored sample at the second node (3012A) to generate an output (also called the second restored sample) at the third node (3013A), and the output is provided to a deblocking filter for further filtering. In the example of FIG. 30A, the first node (3011A) and the second node (3012A) are the same node, and the intermediate restored sample is the first restored sample.
[0201] Figure 30B shows an example of a loop filter chain (3000B). The loop filter chain (3000B) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3000B), a nonlinear mapping-based filter (labeled SO-NLM) is applied at the intermediate nodes of the loop filter chain (3000B). Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the reconstructed sample (generated by the deblocking filter) at the first node (3011B). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample at the second node (3012B) to produce an output (also called the second reconstructed sample) at the third node (3013B), which is provided to the CDEF for further filtering. In the example in Figure 30B, the first node (3011B) and the second node (3012B) are the same node, and the intermediate reconstructed sample is the first reconstructed sample.
[0202] FIG. 30C shows an example of a loop filter chain (3000C) in one example. The loop filter chain (3000C) can be used instead of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3000C), a non-linear mapping based filter (labeled as SO-NLM) is applied at an intermediate node of the loop filter chain (3000C). Specifically, the input (also called the first restored sample) of the non-linear mapping based filter is the restored sample (generated by CDEF) at the first node (3011C). Based on the input, the non-linear mapping based filter generates a sample offset (SO). The sample offset is combined with an intermediate restored sample at the second node (3012C) to generate an output (also called the second restored sample) at the third node (3013C), and the output is provided to an LR filter for further filtering. In the example of FIG. 30C, the first node (3011C) and the second node (3012C) are the same node, and the intermediate restored sample is the first restored sample.
[0203] Figure 30D shows an example of a loop filter chain (3000D). The loop filter chain (3000D) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3000D), a nonlinear mapping-based filter (labeled SO-NLM) is applied at the output node. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the reconstructed sample (generated by the LR filter) at the first node (3011D). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample at the second node (3012D) to generate the output (also called the second reconstructed sample) at the third node (3013D), which is the output of the loop filter chain (3000D). In the example in Figure 30D, the first node (3011D) and the second node (3012D) are the same node, and the intermediate reconstructed sample is the first reconstructed sample.
[0204] According to one aspect of the present disclosure, a nonlinear mapping-based filter can be coupled in parallel with one or more filters in a loop filter chain, with at least one filter between the input and output of the nonlinear mapping-based filter. For example, the nonlinear mapping-based filter receives a reconstructed sample at a first node of the loop filter chain, determines a sample offset based on the reconstructed sample at the first node of the loop filter chain, and then applies the sample offset to the reconstructed sample at a second node of the loop filter chain. The reconstructed sample at the second node of the loop filter chain can be obtained by applying one or more filters to the reconstructed sample at the first node of the loop filter chain.
[0205] In some examples, the input to a nonlinear mapping-based filter is the reconstructed sample located after the deblocking filter and before the CDEF, and the output of the nonlinear mapping-based filter is applied to the reconstructed sample located after the CDEF and before the LR filter, or after the LR filter.
[0206] Figure 31A shows an example of a loop filter chain (3100A) that includes a nonlinear mapping-based filter coupled in parallel with a CDEF. The loop filter chain (3100A) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3100A), a nonlinear mapping-based filter (labeled SO-NLM) is applied between two intermediate nodes. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the reconstructed sample (generated by the deblocking filter) at the first node (3111A). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the CDEF) at the second node (3112A) to produce an output (also called the second reconstructed sample) at the third node (3113A), which is then provided to an LR filter for further filtering. In the example in Figure 31A, the CDEF is located between the first node (3111A) and the second node (3112A). The intermediate reconstruction sample is the output of the CDEF.
[0207] Figure 31B shows an example of a loop filter chain (3100B) that includes a nonlinear mapping-based filter coupled in parallel with CDEF and LR filters. The loop filter chain (3100B) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3100B), a nonlinear mapping-based filter (labeled SO-NLM) is applied between the intermediate node and the output node. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the reconstructed sample (generated by the deblocking filter) at the first node (3111B). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the LR filter) at the second node (3112B) to produce the output (also called the second reconstructed sample) at the third node (3113B), which is the output of the loop filter chain (3100B). In the example in Figure 31B, the CDEF and LR filters are located between the first node (3111B) and the second node (3112B). The intermediate reconstruction sample is the output of the LR filter.
[0208] In some examples, the input to a nonlinear mapping-based filter is the reconstructed sample located before the deblocking filter, and the output of the nonlinear mapping-based filter is applied to the reconstructed sample after the deblocking filter and before the CDEF, after the CDEF and before the LR filter, or after the LR filter.
[0209] Figure 32A shows an example of a loop filter chain (3200A) that includes a nonlinear mapping-based filter coupled in parallel with a deblocking filter. The loop filter chain (3200A) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3200A), a nonlinear mapping-based filter (labeled SO-NLM) is applied between the input and intermediate nodes of the loop filter chain. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the input reconstructed sample (the input to the loop filter chain (3200A)) at the first node (3211A). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the deblocking filter) at the second node (3212A) to produce an output (also called the second reconstructed sample) at the third node (3213A), which is provided to the CDEF for further filtering. In the example in Figure 32A, the deblocking filter is located between the first node (3211A) and the second node (3212A). The intermediate reconstruction sample is the output of the deblocking filter.
[0210] Figure 32B shows an example of a loop filter chain (3200B) that includes a nonlinear mapping-based filter coupled in parallel with a deblocking filter and a CDEF. The loop filter chain (3200B) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3200B), a nonlinear mapping-based filter (labeled SO-NLM) is applied between the input and intermediate nodes of the loop filter chain. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the input reconstructed sample at the first node (3211B) (the input to the loop filter chain (3200B)). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the CDEF) at the second node (3212B) to produce an output (also called the second reconstructed sample) at the third node (3213B), which is provided to the LR for further filtering. In the example in Figure 32B, the deblocking filter and CDEF are located between the first node (3211B) and the second node (3212B). The intermediate reconstruction sample is the output of the CDEF.
[0211] Figure 32C shows an example of a loop filter chain (3200C) that includes a nonlinear mapping-based filter coupled in parallel with a deblocking filter, a CDEF, and an LR filter. The loop filter chain (3200C) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3200C), a nonlinear mapping-based filter (labeled SO-NLM) is applied between the input and output nodes of the loop filter chain. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the input reconstructed sample at the first node (3211C) (the input to the loop filter chain (3200B)). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the LR filter) at the second node (3212C) to produce an output (also called the second reconstructed sample) at the third node (3213C), which is the output of the loop filter chain (3200C). In the example in Figure 32C, the deblocking filter, CDEF, and LR filter are located between the first node (3211C) and the second node (3212C). The intermediate reconstructed sample is the output of the LR filter.
[0212] In some examples, the input to the nonlinear mapping-based filter is the reconstructed sample located after the CDEF and before the LR filter, and the output of the nonlinear mapping-based filter is applied to the reconstructed sample after the LR filter.
[0213] Figure 33 shows an example of a loop filter chain (3300) that includes a nonlinear mapping-based filter coupled in parallel with an LR filter. The loop filter chain (3300) can be used in place of the loop filter chain (2900) in an encoding or decoding device. In the loop filter chain (3300), a nonlinear mapping-based filter (labeled SO-NLM) is applied between the intermediate and output nodes of the loop filter chain. Specifically, the input to the nonlinear mapping-based filter (also called the first reconstructed sample) is the reconstructed sample (generated by the CDEF) at the first node (3311). Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the intermediate reconstructed sample (generated by the LR filter) at the second node (3312) to produce the output (also called the second reconstructed sample) at the third node (3313), which is the output of the loop filter chain (3300). In the example in Figure 33, the LR filter is located between the first node (3311) and the second node (3312). The intermediate reconstruction sample is the output of the LR filter.
[0214] According to another aspect of this disclosure, multiple nonlinear mapping-based filters can be applied simultaneously at multiple locations within a loop filter chain. Each of the multiple nonlinear mapping-based filters can be configured as one of the examples shown in Figures 30A–30D, 31A–31B, 32A–32C, and 33.
[0215] Figure 34 shows an example of a loop filter chain (3400) including a first nonlinear mapping-based filter shown by SO-NLM1 and a second nonlinear mapping-based filter shown by SO-NLM2. The loop filter chain (3400) can be used in place of a loop filter chain (2900) in an encoding or decoding device. The first nonlinear mapping-based filter is coupled in parallel with the CDEF in a configuration similar to the example shown in Figure 31A. The second nonlinear mapping-based filter is coupled in series with the other filters in the loop filter chain (3400) in a configuration similar to the example shown in Figure 30D. In one example, the first and second nonlinear mapping-based filters are filters of the same type, such as CCSO or LSO. In another example, the first and second nonlinear mapping-based filters are filters of different types, such as one being a CCSO and the other an LSO.
[0216] Figure 35 shows an example of a loop filter chain (3500) including a first nonlinear mapping-based filter shown by SO-NLM1 and a second nonlinear mapping-based filter shown by SO-NLM2. The loop filter chain (3500) can be used in place of a loop filter chain (2900) in an encoding or decoding device. The first nonlinear mapping-based filter is coupled in parallel with an LR filter in a configuration similar to the example shown in Figure 33. The second nonlinear mapping-based filter is coupled in series with other filters in the loop filter chain (3500) in a configuration similar to the example shown in Figure 30B. In one example, the first and second nonlinear mapping-based filters are filters of the same type, such as CCSO or LSO. In another example, the first and second nonlinear mapping-based filters are filters of different types, such as one being a CCSO and the other an LSO.
[0217] Figure 36 shows a flowchart illustrating an overview of process (3600) according to one embodiment of the present disclosure. Process (3600) can be used for video filtering. When the term block is used, a block may be interpreted as a prediction block, a coding unit, a luminance block, a saturation block, and so on. In various embodiments, process (3600) is performed by processing circuits such as processing circuits in terminal devices (310), (320), (330), and (340), a processing circuit that performs the function of a video encoder (403), a processing circuit that performs the function of a video decoder (410), a processing circuit that performs the function of a video decoder (510), and a processing circuit that performs the function of a video encoder (603). In some embodiments, process (3600) is implemented in software instructions, and so when a processing circuit executes a software instruction, the processing circuit executes process (3600). The process begins at (S3601) and proceeds to (S3610).
[0218] In (S3610), the first offset value for applying the nonlinear mapping-based filter is based on the first restored sample at the first node along the loop filter chain.
[0219] In (S3620), the first offset value is applied to the intermediate reconstructed sample at the second node along the loop filter chain to generate the second reconstructed sample at the third node along the loop filter chain.
[0220] In one example, the nonlinear mapping-based filter is an inter-component sample offset (CCSO) filter, where the intermediate reconstructed sample and the first reconstructed sample are samples of different color components.
[0221] In another example, the nonlinear mapping-based filter is a local sample offset (LSO) filter, where the intermediate reconstructed sample and the first reconstructed sample are samples of the same color component.
[0222] In some embodiments, the first node and the second node can be of the same node, which can be an input node of a loop filter chain, or an output node of a loop filter chain, or an intermediate node of a loop filter chain.
[0223] In one example, the first restored sample is generated by a processing module before the deblocking filter. In another example, the first restored sample is generated by the deblocking filter. In another example, the first restored sample is generated by a constrained directional expansion filter. In another example, the first restored sample is generated by a loop restoration filter.
[0224] In some embodiments, the first node and the second node are of different nodes. In some examples, the first restored sample is generated by a processing module before the deblocking filter, and the intermediate restored sample is generated by at least one of the deblocking filter, the constrained directional expansion filter, or the loop restoration filter. In some examples, the first restored sample is generated by the deblocking filter, and the intermediate restored sample is generated by at least one of the constrained directional expansion filter or the loop restoration filter. In some examples, the first restored sample is generated by the constrained directional expansion filter, and the intermediate restored sample is generated by the loop restoration filter.
[0225] The process (3600) proceeds to (S3699) and ends.
[0226] Note that in some examples, the non-linear mapping-based filter is a component-wise sample offset (CCSO) filter, and in some other examples, the non-linear mapping-based filter is a local sample offset (LSO) filter.
[0227] Process (3600) can be appropriately adapted. The steps of Process (3600) can be modified and / or omitted. Additional steps can be added. Any appropriate order of implementation can be used.
[0228] The embodiments of this disclosure may be used separately or combined in any order. Furthermore, each of the methods (or embodiments), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium.
[0229] The techniques described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 37 shows a computer system (3700) suitable for implementing some embodiments of the disclosed subject matter.
[0230] Computer software can be coded using any suitable machine language or computer language that can undergo assembly, compilation, linking, or similar mechanisms to create code that contains instructions that can be executed directly or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0231] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.
[0232] The components shown in Figure 37 with respect to the computer system (3700) are essentially illustrative and do not imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependence or requirement on any one or combination of components shown in the exemplary embodiment of the computer system (3700).
[0233] The computer system (3700) may include certain human interface input devices. Such human interface input devices can respond to input from one or more human users, for example, via tactile input (such as keystrokes, swipes, or data glove movements), audio input (such as voices or clapping), visual input (such as gestures), or olfactory input (not depicted). Human interface devices can also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as voices, music, or ambient sounds), images (such as scanned images or photographic images taken from a still camera), or video (such as two-dimensional video or three-dimensional video, including stereoscopic video).
[0234] Input human interface devices may include one or more of the following: keyboard (3701), mouse (3702), trackpad (3703), touchscreen (3710), data glove (not shown), joystick (3705), microphone (3706), scanner (3707), and camera (3708) (only one of each is depicted).
[0235] The computer system (3700) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (3710), data glove (not shown), or joystick (3705), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (3709), headphones (not depicted)), visual output devices (such as screens (3710), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input capabilities, each with or without tactile feedback capabilities, some of which may be capable of outputting two-dimensional or three-dimensional or more-dimensional output via means such as stereographic output, virtual reality glasses (not depicted), holographic displays, and smoke tanks (not depicted)), and printers (not depicted).
[0236] A computer system (3700) may also include human-accessible storage devices and associated media such as optical media including CD / DVD ROM / RW (3720) with CD / DVD or similar media (3721), thumb drives (3722), removable hard drives or solid-state drives (3723), legacy magnetic media such as tapes and floppy disks (not described), and specialized ROM / ASIC / PLD-based devices such as security dongles (not described).
[0237] Those skilled in the art should also understand that the term “computer-readable medium” as used in relation to the subject matter currently disclosed does not include transmission media, carrier waves, or other transient signals.
[0238] A computer system (3700) may also include an interface (3754) to one or more communication networks (3755). These networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, automotive, and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet and wireless LANs; cellular networks such as GSM, 3G, 4G, 5G, and LTE; wired or wireless wide-area digital networks for television, including cable TV, satellite TV, and terrestrial broadcast TV; and automotive and industrial networks such as CANBus. Certain networks typically require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (3749) (e.g., a USB port on the computer system (3700)), while other networks are typically integrated into the core of the computer system (3700) by being attached to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3700) can communicate with other entities. Such communication may be unidirectional (e.g., broadcast TV), unidirectional (e.g., CANbus to a specific CANbus device), or bidirectional (e.g., with other computer systems using local or wide-area digital networks). Specific protocols and protocol stacks may be used with each of these networks and network interfaces described above.
[0239] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (3740) of the computer system (3700).
[0240] The core (3740) may include one or more specialized programmable processing units, such as one or more central processing units (CPUs) (3741), graphics processing units (GPUs) (3742), field-programmable gate areas (FPGAs) (3743), hardware accelerators for specific tasks (3744), and graphics adapters (3750). These devices may be connected via a system bus (3748) along with read-only memory (ROM) (3745), random access memory (3746), and internal mass storage such as hard drives and SSDs (3747) that are not accessible to the internal user. In some computer systems, the system bus (3748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripherals may be connected directly to the core's system bus (3748) or via a peripheral bus (3749). For example, a display (3710) may be connected to a graphics adapter (3750). Peripheral bus architectures include PCI, USB, and others.
[0241] The CPU (3741), GPU (3742), FPGA (3743), and accelerator (3744) can, in combination, execute certain instructions that constitute the aforementioned computer code. This computer code can be stored in ROM (3745) or RAM (3746). Transition data can also be stored in RAM (3746), while persistent data can be stored, for example, in internal mass storage (3747). High-speed storage and retrieval of any of the memory devices can be enabled using cache memory, which can be closely associated with one or more CPUs (3741), GPUs (3742), mass storage (3747), ROM (3745), RAM (3746), etc.
[0242] Computer-readable media may contain computer code for performing various computer implementation operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to persons skilled in computer software technology.
[0243] As an example, and not as an limitation, a computer system having an architecture (3700), specifically a core (3740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage as described above, as well as media associated with specific storage of the core (3740) of a non-transient nature, such as core internal mass storage (3747) or ROM (3745). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (3740). The computer-readable media may include one or more memory devices or chips, depending on the specific needs. The software can cause the core (3740), and specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (3746) and modifying such data structures according to processes defined by the software. In addition, or as an alternative, a computer system may provide functionality as a result of logic wired or otherwise embodied within a circuit (e.g., an accelerator (3744)) which can operate in place of or in conjunction with software to perform a particular process or a particular part of a particular process described herein. If necessary, a reference to software may encompass logic and vice versa. If necessary, a reference to a computer-readable medium may encompass a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that embodies logic for execution, or both. This disclosure encompasses any suitable combination of hardware and software. Note A: Acronym JEM: Collaborative Search Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High-Efficiency Video Coding MPM: Most Probable Mode WAIP: Wide-angle intra-prediction SEI: Supplementary and Extended Information VUI: Video Usability Information GOP: Picture Group TU: Conversion Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Virtual Reference Decoder SDR: Standard Dynamic Range SNR: Signal-to-noise ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-only memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logical Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral component interconnection FPGA: Field-Programmable Gate Area SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit PDPC: Location-dependent predictive coupling ISP: Intra Subpartition SPS: Sequence Parameter Settings
[0244] While this disclosure describes several exemplary embodiments, there are many variations, substitutions, and alternative equivalents that fall within the scope of this disclosure. Those skilled in the art will therefore understand that numerous systems and methods embodying the principles of this disclosure, and thus falling within its spirit and scope, can be devised, although these are not expressly illustrated or described herein. [Explanation of symbols]
[0245] 101 samples 102 Arrow 103 Arrow 104 square blocks 180 Schematic Diagram 201 Currently blocked 300 Communication Systems 310 Terminal devices 320 terminal devices 330 terminal devices 340 terminal devices 350 Communication Networks 400 Communication Systems 401 Video Source 402 Video Picture Stream 403 Video Encoder 404 Encoded video data, Encoded video bitstream 405 Streaming Server 406 Client Subsystem 407 Encoded video data, input copy 408 Client Subsystem 409 Encoded video data, copy 410 Video Decoder 411 Video picture output stream 412 displays 413 Capture Subsystem 420 Electronic Devices 430 Electronic Devices 501 Channel 510 Video Decoder 512 Rendering Devices 515 buffer memory 520 Parser 521 Symbols 530 Electronic Devices 531 Receiver 551 Scaler / Inverse Unit 552 Intrapicture Prediction Units 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current picture buffer 601 Video Sources 603 Video encoder, video coder 620 Electronic Devices 630 Source Coder 632 Coding Engine 633 Local Video Decoder 634 Reference picture cache, reference picture memory 635 Predictor 640 Transmitter 643 coded video sequences 645 Entropy Coder 650 Controller 660 communication channels 703 Video Encoder 721 General-purpose controller 722 Intra Encoders 723 Residual Calculator 724 Residual Encoder 725 Entropy Encoder 726 switches 728 Residual Decoder 730 Interencoder 810 Video Decoder 871 Entropy Decoder 872 Intra Decoder 873 Residual Decoder 874 Recovery Module 880 Interdecoder 910 Diamond-shaped filter 911 Rhombus-shaped filter 920-932 elements 940-964 elements 1110 4x4 block 1111 4x4 block 1120 Horizontal CTU Boundary 1121 CTU boundary 1130 Virtual Boundary 1131 Virtual Boundary 1210 Virtual Boundary 1220 Virtual Boundary 1230 Virtual Boundary 1240 virtual boundary 1250 virtual boundary 1260 virtual boundary 1300 Pictures 1400 quadtree partitioning patterns 1510 Sample Adaptive Offset (SAO) Filter 1512 SAO filter 1514 SAO filter 1516 ALF Brightness Filter 1518 ALF Saturation Filter 1521 CC-ALF 1522 Adder 1531 CC-ALF 1532 Adder 1541 SAO-filtered luminance component 1542 Second intermediate component 1543 Fourth intermediate component 1552 First intermediate component 1553 Third intermediate component 1561 Filtered luminance CB 1562 Filtered first saturation component 1563 Filtered second saturation component 1600 filters 1610 filter coefficients 1620 Diamond shape 1801 Brightness Sample 1801(1) Top left sample, brightness sample 1801(2) Upper right sample, brightness sample 1801(3) Lower left sample, brightness sample 1801(4) Lower right sample, brightness sample 1803 Saturation Sample 1803(1) Saturation Sample 1804 Saturation Sample 1805 Saturation Sample 1806 Saturation Sample 1807 Saturation Sample 1808 Saturation Sample 1910 Block 1920 direction 1923 direction 1930 Fully Directional Block 1933 Fully Directional Block Table of 2100 SAO type 2210 First pattern 2220 Second pattern 2230 Third Pattern 2240 The fourth pattern Table for 2300-pixel classification rules 2400 syntax examples 2500 filter support area 2600 filter support area 2700 table 2900 Loop Filter Chain 2901 Input Node 2902 Intermediate Node 2903 Intermediate Node 2909 Output Node 3000A Loop Filter Chain 3000B Loop Filter Chain 3000C Loop Filter Chain 3000D Loop Filter Chain 3011A First node 3012A Second node 3013A Third node 3011B First node 3012B Second node 3013B Third node 3011C First Node 3012C Second Node 3013C Third Node 3011D First node 3012D Second node 3013D Third Node 3100A Loop Filter Chain 3100B Loop Filter Chain 3111A First node 3112A Second node 3113A Third node 3111B First node 3112B Second node 3113B Third node 3200A Loop Filter Chain 3200B Loop Filter Chain 3200C Loop Filter Chain 3211A First node 3212A Second node 3213A Third node 3211B First node 3212B Second node 3213B Third node 3211C First node 3212C Second Node 3213C Third Node 3300 Loop Filter Chain 3311 First node 3312 Second node 3313 Third Node 3400 Loop Filter Chain 3500 Loop Filter Chain 3600 processes 3700 Computer Systems 3701 Keyboard 3702 Mouse 3703 Trackpad 3705 Joystick 3706 Microphone 3707 Scanner 3708 Camera 3709 Speaker 3710 Touchscreen 3720 CD / DVD ROM / RW 3721 CD / DVD or similar media 3722 Thumb Drive 3723 Removable hard drive or solid state drive 3740 cores 3741 Central Processing Unit (CPU) 3742 Graphics Processing Unit (GPU) 3743 Field-Programmable Gate Area (FPGA) 3744 Hardware Accelerator 3745 Read-only memory (ROM) 3746 random access memory 3747 Internal large-capacity storage 3748 System Bus 3749 Local buses 3750 Graphics Adapter 3754 Interface 3755 Communication Network
Claims
1. A method for filtering in video coding, The processor determines a first offset value for applying a nonlinear mapping-based filter based on a first restored sample at a first node along a loop filter chain containing multiple video filters. The processor performs the steps of adding the first offset value to an intermediate reconstructed sample at a second node along the loop filter chain in order to generate a second reconstructed sample. Methods that include...
2. The method according to claim 1, wherein the nonlinear mapping-based filter is an inter-component sample offset (CCSO) filter, and the intermediate reconstructed sample and the first reconstructed sample are samples of different color components.
3. The method according to claim 1, wherein the nonlinear mapping-based filter is a local sample offset (LSO) filter, and the intermediate reconstructed sample and the first reconstructed sample are samples of the same color component.
4. The method according to claim 1, wherein the first node and the second node are the same node.
5. The first node and the second node, The input node of the aforementioned loop filter chain, The output node of the aforementioned loop filter chain, or Intermediate node of the aforementioned loop filter chain The method according to claim 4, which is one of the methods.
6. The first reconstructed sample described above is Processing modules prior to the deblocking filter, The deblocking filter, Constrained directional extension filter, or Loop Restoration Filter The method according to claim 4, which is produced by at least one of the following.
7. The method according to claim 1, wherein the first node and the second node are different nodes.
8. The first reconstructed sample is generated by the processing module before the deblocking filter, and the intermediate reconstructed sample is, The deblocking filter, Constrained directional extension filter, or Loop Restoration Filter The method according to claim 7, which is produced by at least one of the following.
9. The first reconstructed sample is generated by a deblocking filter, and the intermediate reconstructed sample is, Constrained directional extension filter, or Loop Restoration Filter The method according to claim 7, which is produced by at least one of the following.
10. The method according to claim 7, wherein the first reconstructed sample is generated by a constrained directional expansion filter, and the intermediate reconstructed sample is generated by a loop reconstructed filter.
11. Based on a first restored sample at a first node along a loop filter chain containing multiple video filters, a first offset value is determined for applying a nonlinear mapping-based filter. To generate a second reconstructed sample, the first offset value is added to the intermediate reconstructed sample at the second node along the loop filter chain. A device for video coding, comprising a processing circuit configured as such.
12. The apparatus according to claim 11, wherein the nonlinear mapping-based filter is an inter-component sample offset (CCSO) filter, and the intermediate reconstructed sample and the first reconstructed sample are samples of different color components.
13. The apparatus according to claim 11, wherein the nonlinear mapping-based filter is a local sample offset (LSO) filter, and the intermediate reconstructed sample and the first reconstructed sample are samples with the same color component.
14. The apparatus according to claim 11, wherein the first node and the second node are the same node.
15. The first node and the second node, The input node of the aforementioned loop filter chain, The output node of the aforementioned loop filter chain, or Intermediate node of the aforementioned loop filter chain The apparatus according to claim 14, which is one of the above.
16. The first reconstructed sample described above is Processing modules prior to the deblocking filter, The deblocking filter, Constrained directional extension filter, or Loop Restoration Filter The apparatus according to claim 14, which is produced by at least one of the following.
17. The apparatus according to claim 11, wherein the first node and the second node are different nodes.
18. The first reconstructed sample is generated by the processing module before the deblocking filter, and the intermediate reconstructed sample is, The deblocking filter, Constrained directional extension filter, or Loop Restoration Filter The apparatus according to claim 17, which is produced by at least one of the following.
19. The first reconstructed sample is generated by a deblocking filter, and the intermediate reconstructed sample is, Constrained directional extension filter, or Loop Restoration Filter The apparatus according to claim 17, which is produced by at least one of the following.
20. The apparatus according to claim 17, wherein the first reconstructed sample is generated by a constrained directional expansion filter, and the intermediate reconstructed sample is generated by a loop reconstructed filter.