Method, device, program for boundary processing in video encoding
Patent Information
- Application Number
- JP2024099381
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-22
- Filing Date
- 2024-06-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-24
AI Technical Summary
Existing video encoding techniques face challenges in efficiently handling boundary processing during intra-prediction and motion compensation, leading to suboptimal compression ratios and increased data requirements due to inefficiencies in filtering methods.
Implementing a nonlinear mapping-based filter, such as a cross-component sample offset (CCSO) or local sample offset (LSO) filter, in the loop filter chain to process boundary pixel values, combined with loop restoration filters to enhance video encoding efficiency.
Improves video encoding efficiency by reducing redundancy and enhancing compression performance, particularly at frame borders, thereby minimizing data requirements and maintaining image quality.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] [Incorporated by reference] This application claims the benefit of priority to U.S. Patent Application No. 17 / 448,469, entitled "METHOD AND APPARATUS FOR BOUNDARY HANDLING IN VIDEO CODING," filed September 22, 2021, which claims the benefit of priority to U.S. Provisional Application No. 63 / 187,213, entitled "ON LOOP RESTORATION BOUNDARY HANDLING," filed May 11, 2021. The entire disclosure of the prior application is hereby incorporated by reference in its entirety.
[0002] [Technical field] This disclosure describes embodiments generally related to video coding. [Background technology]
[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. Work of the inventors named in this application, to the extent that their work is described in this background section, and aspects of this description that may not otherwise qualify as prior art at the time of filing, are not admitted, either explicitly or implicitly, as prior art to the present disclosure.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a sequence of pictures, each having spatial dimensions of, for example, 1920x1080 luminance samples and associated chrominance samples. The sequence of pictures can have a fixed or variable picture rate (also informally known as frame rate), for example, 60 pictures per second or a picture rate of 60 Hz. Uncompressed video has specific bitrate requirements. For example, 1080p60 4:2:0 video (1920x1080 luminance sample resolution at a frame rate of 60 Hz) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.
[0005] One goal of video encoding and decoding may be the reduction of redundancy in the input video signal through compression. Compression may help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than one order of magnitude. Both lossless and lossy compression, as well as combinations thereof, may be used. Lossless compression refers to techniques whereby an exact copy of the original signal can be reconstructed from the compressed original signal. When lossy compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. For video, lossy compression is widely used. The amount of acceptable distortion depends on the application, e.g., users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio may reflect that a higher tolerable / acceptable distortion can result in a higher compression ratio.
[0006] Video encoders and decoders can utilize techniques from a number of broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.
[0007] Video codec techniques can include a technique known as intra-coding. In intra-coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are coded in intra mode, the picture may be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. Samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized before entropy coding. Intra prediction can be a technique that minimizes sample values in the pre-transform domain. In some cases, the smaller the DC value and the smaller the AC coefficients after the transform, the fewer bits are needed for a given quantization step size to represent the block after entropy coding.
[0008] Traditional intra-coding, for example as known from MPEG-2 generation encoding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to do so from surrounding sample data and / or metadata obtained during encoding / decoding of blocks of data that are spatially nearby and preceding in decoding order. Such techniques are referred to below as "intra-prediction" techniques. It should be noted that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, and not from reference pictures.
[0009] There may be various forms of intra prediction. If more than one such technique is available for a given video coding technique, the technique used may be coded in intra prediction mode. In some cases, a mode may have sub-modes and / or parameters, which may be coded separately or may be included in the mode codeword. Which codeword to use for a given mode / sub-mode / parameter combination may affect the coding efficiency gain through intra prediction, as well as the entropy coding technique used to convert the codeword into a bitstream.
[0010] A mode of intra prediction was introduced in H.264, refined in H.265, and further refined in newer coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using neighboring sample values belonging to already available samples. The sample values of the neighboring samples are copied to the predictor block according to a certain direction. The reference to the direction used can be coded in the bitstream or it may be predicted itself.
[0011] Referring to FIG. 1A, at the bottom right, a subset of 9 known predictor directions from the 33 possible predictor directions (corresponding to the 33 angle modes of the 35 intra modes) of H.265 is depicted. The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample(s) to the top right, at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample(s) to the bottom left of sample (101), at an angle of 22.5 degrees from the horizontal.
[0012] 1A, at the top left, a square block (104) of 4×4 samples is depicted (indicated by a thick dashed line). The square block (104) contains 16 samples, each labeled with an “S” and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Since the block is 4×4 samples in size, S44 is at the bottom right. Additionally, a reference sample is shown that follows a similar numbering scheme. The reference sample is labeled R and its Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the predicted samples are neighbors of the block being reconstructed, so there is no need to use negative values.
[0013] Intra-picture prediction can work by copying reference sample values from neighboring samples filled by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling for this block indicating a prediction direction that is consistent with the arrow (102). That is, the sample is predicted from the top right prediction sample or samples at an angle of 45 degrees from the horizontal. Then samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then sample S44 is predicted from reference sample R08.
[0014] In some cases, particularly when the orientation is not divisible by 45 degrees, the values of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample.
[0015] As video coding technology develops, the number of possible directions has increased. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and JEM / VVC / BMS at the time of this disclosure can support up to 65 directions. Experiments are performed to identify the most likely directions, and certain techniques in entropy coding are used to represent those likely directions with a small number of bits while accepting some penalty for the less likely directions. Furthermore, the direction itself may be predictable from nearby directions used in nearby already decoded blocks.
[0016] FIG. 1B shows a schematic diagram (180) depicting 65 intra prediction directions with JEM to illustrate the increasing number of prediction directions over time.
[0017] The mapping of intra-prediction direction bits in the coded video bitstream representing the direction can vary from one video coding technique to another, for example ranging from a simple direct mapping of prediction directions to intra-prediction modes to complex adaptation schemes involving codewords, most probable modes, and similar techniques. In any case, however, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-performing video coding technique, these less likely ways are represented by a larger number of bits than the more probable directions.
[0018] Motion compensation may be a lossy compression technique and may refer to a technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are used for prediction of a newly reconstructed picture or part thereof after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension being an indication of the reference picture used (which may indirectly be the temporal dimension).
[0019] In some video compression techniques, the MV applicable to a region of sample data can be predicted from other MVs, e.g., from an MV associated with another region of sample data that is spatially adjacent to the region being reconstructed and precedes it in decoding order. Doing so can significantly reduce the amount of data required to encode the MV, thereby removing redundancy and increasing compression. MV prediction works well because, for example, when encoding an input video signal derived from a camera (known as natural video), there is a statistical probability that regions larger than the region to which a single MV is applicable move in a similar direction and can therefore, in some cases, be predicted using similar motion vectors derived from the MVs of nearby regions. As a result, the MV found for a given region will be similar or identical to the MV predicted from the surrounding MVs, which, after entropy encoding, can be represented with fewer bits than would be used if encoding the MVs directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., a sample stream). In other cases, the MV prediction itself may be lossy, for example due to rounding errors in computing the predictor from several surrounding MVs.
[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, a technique hereafter referred to as "spatial merge" is described herein.
[0021] Referring to Figure 2, a current block (201) contains samples that the encoder found during the motion search process to be predictable from a previous block of the same size but spatially shifted. Instead of directly encoding its MV, the MV can be derived from metadata associated with one or more reference pictures, for example from the most recent reference picture (in decoding order) using MVs associated with any of the five surrounding samples, denoted A0, A1, and B0, B1, B2 (202-206, respectively). In H.265, MV prediction can use predictors from the same reference picture that neighboring blocks use. Summary of the Invention
[0022] Aspects of the present disclosure provide methods and apparatus for filtering in video encoding / decoding. In some examples, the apparatus for video coding includes a processing circuit. The processing circuit buffers a first boundary pixel value of a first reconstructed sample at a first node along a loop filter chain. The first node is associated with a nonlinear mapping-based filter applied in the loop filter chain before the loop restoration filter. In one example, the first boundary pixel value is a value of a pixel at a frame boundary. The processing circuit applies the loop restoration filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value.
[0023] In some examples, the nonlinear mapping based filter is a cross-component sample offset (CCSO) filter. In some examples, the nonlinear mapping based filter is a local sample offset (LSO) filter.
[0024] In some examples, the first reconstructed sample at the first node is an input of the nonlinear mapping-based filter. In some examples, the processing circuit buffers a second boundary pixel value of a second reconstructed sample at a second node along the loop filter chain. The second reconstructed sample at the second node is generated after application of the sample offset generated by the nonlinear mapping-based filter. The second boundary pixel value is a value of a pixel at a frame boundary. The processing circuit can then apply a loop restoration filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.
[0025] In one example, the processing circuit combines the sample offsets generated by the nonlinear mapping-based filter with an output of a constrained directional enhancement filter to generate a second reconstructed sample. In another example, the processing circuit combines the sample offsets generated by the nonlinear mapping-based filter with the first reconstructed sample to generate an intermediate reconstructed sample and applies the constrained directional enhancement filter to the intermediate reconstructed sample to generate a second reconstructed sample.
[0026] In some examples, the processing circuit buffers a second boundary pixel value of a second reconstructed sample at a second node along the loop filter, and combines the second reconstructed sample at the second node with the sample offset generated by the nonlinear mapping-based filter to generate a reconstructed sample to be filtered. The second boundary pixel value is a value of a pixel at a frame boundary. The processing circuit then applies a loop restoration filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.
[0027] In some examples, the processing circuit buffers second boundary pixel values from the second reconstructed samples generated by the deblocking filter and applies a constrained directional enhancement filter to the second reconstructed samples to generate intermediate reconstructed samples. The second boundary pixel values are values of pixels at frame boundaries. The processing circuit combines the intermediate reconstructed samples with the sample offsets generated by the nonlinear mapping-based filter to generate the first reconstructed samples. The loop restoration filter can then be applied based on the buffered first boundary pixel values and the buffered second boundary pixel values.
[0028] In some examples, the processing circuit clips the filtered reconstructed samples prior to application of the loop reconstruction filter. In some examples, the processing circuit clips the intermediate reconstructed samples prior to combination with the sample offsets produced by the nonlinear mapping-based filter.
[0029] Aspects of the present disclosure also provide a non-transitory computer-readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform any of the methods for video encoding / decoding. [Brief description of the drawings]
[0030] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1A] FIG. 2 is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B] FIG. 2 is an illustration of an exemplary intra-prediction direction. [Diagram 2] FIG. 2 is a schematic diagram of a current block and its surrounding spatial merging candidates in one example. [Diagram 3] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment. [Figure 4] FIG. 4 is a schematic diagram of a simplified block diagram of a communication system (400) according to an embodiment. [Diagram 5] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment. [Figure 6] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment. [Figure 7] 4 shows a block diagram of an encoder according to another embodiment; [Figure 8] 4 shows a block diagram of a decoder according to another embodiment; [Figure 9] 4 illustrates example filter shapes according to embodiments of the present disclosure. [Figure 10A] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10B] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10C] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 10D] 4 illustrates an example of sub-sampled positions used to calculate gradients according to an embodiment of the present disclosure. [Figure 11A] 1 illustrates an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 11B]1 illustrates an example of a virtual boundary filtering process according to an embodiment of the present disclosure. [Figure 12A] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12B] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12C] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12D] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12E] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 12F] 1 illustrates an example of a symmetric padding operation at a virtual boundary according to an embodiment of the present disclosure. [Figure 13] 1 illustrates an example of picture division according to some embodiments of the present disclosure. [Figure 14] 4 shows quadtree division patterns of pictures in some examples. [Figure 15] 1 illustrates a cross-component filter according to an embodiment of the present disclosure. [Figure 16] 4 illustrates example filter shapes according to certain embodiments of the present disclosure. [Figure 17] 1 illustrates an example syntax for a cross-component filter according to some embodiments of the present disclosure. [Figure 18A] 1 illustrates an example position of a chroma sample relative to a luma sample, according to an embodiment of the present disclosure. [Figure 18B] 1 illustrates an example position of a chroma sample relative to a luma sample, according to an embodiment of the present disclosure. [Figure 19] 1 illustrates an example of direction finding according to an embodiment of the present disclosure. [Figure 20] 13 shows examples illustrating subspace projection in some examples. [Figure 21] 1 illustrates a table of multiple sample adaptive offset (SAO) types, according to an embodiment of the present disclosure. [Figure 22] 13 shows example patterns for pixel classification at edge offset in some examples. [Figure 23] 13 shows a table for pixel classification rules for edge offsets in some examples. [Figure 24] An example of the syntax that can be signaled is shown below. [Diagram 25] 1 illustrates an example of a filter support region, according to some embodiments of the present disclosure. [Figure 26] 13 illustrates another example of a filter support region according to some embodiments of the present disclosure. [Figure 27A] 1 shows a portion of a table having 81 combinations according to an embodiment of the present disclosure. [Figure 27B] 1 shows a portion of a table having 81 combinations according to an embodiment of the present disclosure. [Figure 27C] 1 shows a portion of a table having 81 combinations according to an embodiment of the present disclosure. [Figure 28] 7 shows seven filter shape configurations of three filter taps in one example. [Figure 29] 1 illustrates a block diagram of a loop filter chain in some examples. [Diagram 30] 1 illustrates an example of a loop filter chain including a nonlinear mapping based filter. [Diagram 31] 2 illustrates another example of a loop filter chain that includes a nonlinear mapping based filter. [Diagram 32] 2 illustrates another example of a loop filter chain that includes a nonlinear mapping based filter. [Diagram 33] 2 illustrates another example of a loop filter chain that includes a nonlinear mapping based filter. [Diagram 34] 2 illustrates another example of a loop filter chain that includes a nonlinear mapping based filter. [Diagram 35] 2 illustrates another example of a loop filter chain that includes a nonlinear mapping based filter. [Diagram 36]1 shows a flowchart outlining a process according to an embodiment of the present disclosure. [Figure 37] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0031] FIG. 3 illustrates a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform a unidirectional transmission of data. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to reconstruct the video pictures, and display the video pictures according to the reconstructed video data. One-way data transmission may be common, such as in media service applications.
[0032] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of encoded video data, such as may occur during a video conference. For the bidirectional transmission of data, in one example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) may receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to reconstruct the video pictures, and display the video pictures on an accessible display device in accordance with the reconstructed video data.
[0033] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) may be depicted as a server, a personal computer, and a smartphone, although the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. The network (350) represents any number of networks that convey encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) may exchange data in circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of discussion herein, the architecture and topology of the network (350) may not be important to the operation of the present disclosure, unless otherwise described below.
[0034] 4 shows an arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0035] The streaming system may include a video source (401), e.g., a digital camera, and may include a capture subsystem (413) that generates, e.g., a stream of uncompressed video pictures (402). In one example, the stream of video pictures (402) includes samples captured by a digital camera. The stream of video pictures (402), depicted as a thick line to emphasize its high data volume compared to the encoded video data (404) (or encoded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize its lower data volume compared to the stream of video pictures (402), may be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and generates an output stream (411) of video pictures that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., a video bitstream) can be encoded according to some video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the developing video coding standard is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.
[0036] It is noted that electronic devices 420 and 430 may include other components (not shown). For example, electronic device 420 may include a video decoder (not shown), and electronic device 430 may also include a video encoder (not shown).
[0037] 5 shows a block diagram of a video decoder (510) according to an embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (310) in the example of FIG. 4.
[0038] The receiver (531) may receive one or more coded video sequences to be decoded by the video decoder (510), one coded video sequence at a time, in the same or another embodiment, and the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (531) may receive the coded video data together with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to respective usage entities (not shown). The receiver (531) may separate the coded video sequences from the other data. To combat network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter the "parser"). In some applications, the buffer memory (515) is part of the video decoder (510). In other applications, it may be external to the video decoder (510) (not shown). In yet other applications, there may be a buffer memory (not shown) external to the video decoder (510), e.g., to combat network jitter, and another buffer memory (515) internal to the video decoder (510), e.g., to handle playback timing. If the receiver (531) is receiving data from a store / forward device of sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be required or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (515) may be required and may be relatively large, advantageously of adaptive size, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (510).
[0039] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device such as a render device (512) (e.g., a display screen). The rendering device may not be an integral part of the electronic device (530) as shown in FIG. 5, but may be coupled to the electronic device (530). The control information for the rendering device(s) may be in the form of Supplementary Enhancement Information (SEI messages) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (520) can also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0040] The parser (520) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515) to generate symbols (521).
[0041] The reconstruction of symbols (521) can involve several different units, depending on the type of coded video picture or portions thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are involved and how can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.
[0042] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually divided into a number of functional units, as described below. In a practical implementation working within commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0043] The first unit is a scalar / inverse transform unit (551). The scalar / inverse transform unit (551) receives quantized transform coefficients and control information as symbol(s) (521) from the parser (520). The control information includes which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scalar / inverse transform unit (551) can output blocks containing sample values that can be input to an aggregator (555).
[0044] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information taken from a current picture buffer (558). The current picture buffer (558) may, for example, buffer a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).
[0045] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensation prediction unit (553) may access a reference picture memory (557) to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols (521) for the block, these samples may be added by an aggregator (555) to the output of the scalar / inverse transform unit (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (557) from which the motion compensation unit (553) fetches prediction samples may be controlled by motion vectors available to the motion compensation unit (553) in the form of symbols (521). The symbols may have, for example, X, Y, and reference picture components. Motion compensation may include interpolation of sample values fetched from the reference picture memory (557) when subsample accurate motion vectors are used, motion vector prediction mechanisms, and the like.
[0046] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). Video compression techniques may include in-loop filter techniques that are controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop filtered sample values.
[0047] The output of the loop filter unit (556) can be a sample stream, which can be output to a render device (512) or stored in a reference picture memory (557) for use in future inter-picture prediction.
[0048] Once a coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of a subsequent coded picture.
[0049] The video decoder (510) may perform decoding operations according to a given video compression technique in a standard such as ITU-T Recommendation H.265. The encoded video sequence may be compliant with the syntax prescribed by the video compression technique or standard being used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. In particular, the profile may select certain tools from all tools available in the video compression technique or standard as the only tools available for use under that profile. Compliance may also require that the complexity of the encoded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in units of megasamples per second), maximum reference picture size, etc. The limits set by the level may be further constrained through Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management, possibly signaled in the encoded video sequence.
[0050] In some embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0051] 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.
[0052] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture video images to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).
[0053] The video source (601) may provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures image information locally as a video sequence. The video data may be provided as a number of individual pictures that impart motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.
[0054] According to an embodiment, the video encoder (603) can encode and compress pictures of a source video sequence into an encoded video sequence (643) in real-time or under any other time constraint required by an application. Enforcing an appropriate encoding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is operatively coupled to other functional units, such as those described below. Such couplings are not depicted for clarity. Parameters set by the controller (650) can include parameters related to rate control (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured with other suitable functions for the video encoder (603) optimized for a certain system design.
[0055] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As a simplistic description, in one example, the encoding loop can include a source encoder (630) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture(s)) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in a similar manner that a (remote) decoder would also generate (in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces bit-accurate results regardless of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local and remote encoders. In other words, the predictor of the encoder "sees" exactly the same sample values as the reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, e.g., due to channel errors) is also used in several related technologies.
[0056] The operation of the "local" decoder (633) may be the same as the operation of a "remote" decoder, e.g., the video decoder (410), already described in detail above in connection with Figure 5. However, referring also briefly to Figure 5, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (645) and parser (420) may be lossless, the entropy decoding portion of the video decoder (410), including the buffer memory (415) and the parser (420), may not be fully implemented in the local decoder (633).
[0057] An observation that can be made at this point is that any decoder technique, other than parsing / entropy decoding, that exists in the decoder must exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the decoder operation. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques that are described generically. Only in certain areas is more detailed explanation necessary, which is provided below.
[0058] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence, designated as “reference pictures.” In this manner, the encoding engine (632) encodes differences between pixel blocks of the input picture and pixel blocks of a reference picture(s) that may be selected as predictive references for the input picture.
[0059] The local video decoder (633) can decode the encoded video data of the pictures that may be designated as reference pictures based on the symbols generated by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded in a video decoder (not shown in FIG. 6), the reconstructed video sequence can be a copy of the source video sequence, typically with some errors. The local video decoder (633) can replicate the decoding process that may be performed on the reference pictures by the video decoder and store the reconstructed reference pictures in the reference picture cache (634). In this way, the video encoder (603) can locally store copies of reconstructed reference pictures that have common content (in the absence of transmission errors) as the reconstructed reference pictures that would be obtained by the far-end video decoder.
[0060] The predictor (635) may perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (635) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).
[0061] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0062] The output of all of the above functional units may undergo entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0063] The transmitter (640) can buffer the encoded video sequence generated by the entropy encoder (645) and prepare it for transmission over a communication channel (660), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (630) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).
[0064] A controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a coding picture type to each coded picture. The coding picture type may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0065] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art will recognize these variations of I-pictures and their respective uses and characteristics.
[0066] A predictive picture (P-picture) may be one that can be encoded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0067] Bidirectionally predicted pictures (B-pictures) may be those that can be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0068] A source picture is usually spatially divided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture of the block. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0069] The video encoder (603) may perform encoding operations according to a given video encoding technique or standard, such as ITU-T Recommendation H.265. In its operations, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technique or standard used.
[0070] In some embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0071] A video may be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture to be coded / decoded, called a current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, that block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block in a reference picture, and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0072] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may also be past and future, respectively, in display order) in the video. A block in the current picture can be coded by a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.
[0073] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.
[0074] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels can be partitioned into one CU of 64×64 pixels, or four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In one example, each CU is analyzed to determine a prediction type for that CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. In general, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Taking a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels, such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0075] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block (e.g., a predictive block) of sample values in a current video picture in a sequence of video pictures and to encode the processed block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.
[0076] In an HEVC example, the video encoder (703) receives a matrix of sample values for a processing block, such as a predictive block, such as 8×8 samples. The video encoder (703) determines whether the processing block is best coded using intra mode, inter mode, or bi-predictive mode, e.g., using rate-distortion optimization. If the processing block is coded in intra mode, the video encoder (703) may use intra prediction techniques to code the processing block into a coded picture. If the processing block is coded in inter mode or bi-predictive mode, the video encoder (703) may use inter prediction techniques or bi-predictive techniques, respectively, to code the processing block into a coded picture. In certain video coding techniques, a merge mode may be an inter-picture prediction sub-mode in which motion vectors are derived from one or more motion vector predictors, but without the benefit of coded motion vector components outside the predictors. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.
[0077] In the example of FIG. 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy coder (725), coupled together as shown in FIG.
[0078] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures), generate inter-prediction information (e.g., a description of redundant information due to inter-coding techniques, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the coded video information.
[0079] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), possibly compare the block with blocks already encoded in the same picture, generate transformed and quantized coefficients, and possibly also generate intra prediction information (e.g., intra prediction direction information according to one or more intra encoding techniques). In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0080] The general controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is an intra mode, the general controller (721) controls the switch (726) to select the result of the intra mode for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and include the intra prediction information in the bitstream. If the mode is an inter mode, the general controller (721) controls the switch (726) to select the result of inter prediction for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and include the inter prediction information in the bitstream.
[0081] The residual calculator (723) is configured to calculate a difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to encode the residual data to generate transform coefficients based on the residual data. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform to generate decoded residual data. The decoded residual data can be suitably used by the intra-encoder (722) and the inter-encoder (730). For example, the inter-encoder (730) can generate decoded blocks based on the decoded residual data and the inter-prediction information, and the intra-encoder (722) can generate decoded blocks based on the decoded residual data and the intra-prediction information. The decoded blocks are suitably processed to generate decoded pictures, which can be buffered in a memory circuit (not shown) and used as reference pictures in some examples.
[0082] The entropy encoder (725) is configured to format a bitstream to include the encoded block. The entropy encoder (725) is configured to include various information according to a suitable standard, such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other suitable information in the bitstream. Note that in accordance with the disclosed subject matter, when encoding a block in a merged sub-mode of either the inter mode or the bi-prediction mode, the residual information is not present.
[0083] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive encoded pictures that are part of an encoded video sequence and decode the encoded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.
[0084] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in FIG. 8.
[0085] The entropy decoder (871) may be configured to reconstruct from the coded picture certain symbols representing the syntax elements of which the coded picture is composed. Such symbols may include, for example, prediction information (e.g., intra- or inter-prediction information, etc.) that may identify the mode in which the block is coded (e.g., intra- or inter-prediction mode, merged submode or the latter two in another submode), certain samples or metadata used for prediction by the intra-decoder (872) or the inter-decoder (880), respectively, residual information, for example in the form of quantized transform coefficients, etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880). If the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra-decoder (872). The residual information may undergo inverse quantization and is provided to the residual decoder (873).
[0086] The inter decoder (880) is configured to receive inter prediction information and to generate inter prediction results based on the inter prediction information.
[0087] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0088] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and to process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantizer parameters (QP)), which may be provided by the entropy decoder (871) (data path not depicted since this is only low volume control information).
[0089] The reconstruction module (874) is configured to combine, in the spatial domain, the residual output by the residual decoder (873) and the prediction result (output by the intra- or inter-prediction module, as the case may be) to form a reconstructed block, which may be part of a reconstructed picture, which may be part of a reconstructed video. It is noted that other suitable operations, such as a deblocking operation, may be performed to improve visual quality.
[0090] It should be noted that the video encoders (403), (603), (703) and the video decoders (410), (510), (810) may be implemented using any suitable techniques. In some embodiments, the video encoders (403), (603), (703) and the video decoders (410), (510), (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), (703) and the video decoders (410), (510), (810) may be implemented using one or more processors executing software instructions.
[0091] Aspects of this disclosure provide filtering techniques for video encoding / decoding.
[0092] To reduce artifacts, an adaptive loop filter (ALF) with block-based filter adaptation can be applied by the encoder / decoder. For the luma component, one of multiple filters (e.g., 25 filters) can be selected for a 4x4 luma block, for example, based on local gradient direction and activity.
[0093] The ALF may have any suitable shape and size. Referring to FIG. 9, the ALFs (910)-(911) have a diamond shape, such as a 5×5 diamond shape for the ALF (910) and a 7×7 diamond shape for the ALF (911). In the ALF (910), elements (920)-(932) form a diamond shape and can be used in the filtering process. For elements (920)-(932), seven values (e.g., C0-C6) are available. In the ALF (911), elements (940)-(964) form a diamond shape and can be used in the filtering process. For elements (940)-(964), thirteen values (e.g., C0-C12) are available.
[0094] Referring to FIG. 9, in some examples, two ALFs (910)-(911) with diamond filter shapes are used. A 5×5 diamond shaped filter (910) can be applied to the chroma components (chroma blocks, chroma CB, etc.) and a 7×7 diamond shaped filter (911) can be applied to the luma components (luma blocks, luma CB, etc.). Other suitable shapes and sizes of the ALFs can be used. For example, a 9×9 diamond shaped filter can be used.
[0095] The filter coefficients at the locations indicated by those values (e.g., C0-C6 in (910) or C0-C12 in (920)) may be non-zero. Furthermore, if the ALF includes a clipping function, the clipping values at those locations may be non-zero.
[0096] For block classification of the luma component, a 4×4 block (or luma block, luma CB) can be categorized or classified as one of several (e.g., 25) classes. The classification index C is a function of the directional parameters and the quantized values of the activation values A.
number
number
number
number
[0097] To reduce the complexity of the block classification described above, a subsampled 1-D Laplacian computation can be applied. Figures 10A-10D show the gradients g in the vertical (Figure 10A), horizontal (Figure 10B), and two diagonal directions d1 (Figure 10C) and d2 (Figure 10D). v , g h , g d1 and g d2 10A shows an example of the subsampled positions used to compute the vertical gradient gv In FIG. 10B, the label “H” indicates the subsampling position for computing the horizontal gradient g h In FIG. 10C, the label “D1” indicates the subsampling position for computing the diagonal gradient g d1 In FIG. 10D, the label “D2” indicates the subsampling position for computing the diagonal gradient g d2 indicates the subsampling positions for calculating
[0098] Horizontal and vertical gradients g v and g h The maximum value of g h,v max and the minimum value g h,v min can be set as follows:
number
number
[0099] The activation value A can be calculated as follows:
number
number
[0100] Since no block classification is applied for the chroma components in a picture, a single set of ALF coefficients can be applied for each chroma component.
[0101] A geometric transformation can be applied to the filter coefficients and the corresponding filter clipping values (also referred to as clipping values). Before filtering a block (e.g., a 4×4 luma block), for example, the gradient value (e.g., g v , g h , g d1 , and / or g d2), a geometric transformation such as a rotation or a diagonal and vertical flip can be applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l). The geometric transformation applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l) can be equivalent to applying the geometric transformation to the samples in the region supported by the filter. The geometric transformation can make the different blocks to which the ALF is applied more similar by aligning their respective directionality.
[0102] Three geometric transformations including diagonal flip, vertical flip, and rotation can be performed as described in Equations (9)-(11), respectively.
number
[0103] In some embodiments, the ALF filter parameters are signaled in an adaptive parameter set (APS) for a picture. In the APS, one or more sets (up to 25 sets) of luma filter coefficients and clipping value indexes can be signaled. In one example, a set of the one or more sets can include luma filter coefficients and one or more clipping value indexes. One or more sets (up to 8 sets) of chroma filter coefficients and clipping value indexes can be signaled. To reduce signaling overhead, filter coefficients of different classifications (e.g., with different classification indexes) for the luma component can be merged. In the slice header, an index of the APS used for the current slice can be signaled.
[0104] In an embodiment, a clipping value index (also referred to as a clipping index) can be decoded from the APS. The clipping value index can be used to determine a corresponding clipping value, for example, based on a relationship between the clipping value index and the corresponding clipping value. This relationship can be predefined and stored in the decoder. In one example, the relationship is described by a table, such as a luma table (e.g., used for luma CB) of clipping value index and corresponding clipping value, a chroma table (e.g., used for chroma CB) of clipping value index and corresponding clipping value, etc. The clipping value can depend on the bit depth B. The bit depth B can refer to the internal bit depth, the bit depth of the reconstructed samples in the CB to be filtered, etc. In some examples, the tables (e.g., luma table, chroma table) are obtained using Equation (12).
number
[0105] One or more APS indices (up to 7 APS indices) may be signaled in the slice header for the current slice to specify the luma filter sets that can be used for the current slice. The filtering process may be controlled at one or more appropriate levels, such as picture level, slice level, CTB level, and / or others. In an embodiment, the filtering process may be further controlled at the CTB level. A flag is signaled indicating whether the ALF is applied to the luma CTB. The luma CTB may select a filter set from among a number of fixed filter sets (e.g., 16 fixed filter sets) and a filter set signaled in the APS (also referred to as a signaled filter set). A filter set index is signaled for the luma CTB to indicate the filter set to be applied (e.g., a filter set among the number of fixed filter sets and the signaled filter set). The number of fixed filter sets may be predefined and hard-coded in the encoder and decoder, and may be referred to as a predefined filter set.
[0106] For chroma components, an APS index can be signaled in the slice header to indicate the chroma filter set used for the current slice. At the CTB level, if there are multiple chroma filter sets in an APS, a filter set index can be signaled for each chroma CTB.
[0107] The filter coefficients can be quantized with a norm equal to 128. To reduce the complexity of multiplications, bitstream conformance can be applied such that coefficient values for non-center positions are in the range of -27 to 27-1. In one example, coefficients for center positions are not signaled in the bitstream and can be considered equal to 128.
[0108] In some embodiments, the syntax and semantics of clipping indexes and clipping values are defined as follows: alf_luma_clip_idx[sfIdx][j] may be used to specify a clipping index of the clipping value to use before multiplying the jth coefficient of the luma filter signaled by sfIdx. Bitstream conformance requirements may include that the value of alf_luma_clip_idx[sfIdx][j], for sfIdx=0 to alf_luma_num_filters_signalled_minus1 and j=0 to 11, is in the range 0 to 3, inclusive. The luma filter clipping value AlfClipL[adaptation_parameter_set_id] of element AlfClipL[adaptation_parameter_set_id][filtIdx][j], for filtIdx=0 to NumAlfFilters-1 and j=0 to 11, can be derived as specified in Table 2 depending on bitDepth set equal to BitDepthY and clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx][filtIdx]][j]. alf_chroma_clip_idx[altIdx][j] can be used to specify a clipping index for the clipping value to use before multiplying the jth coefficient of the alternate chroma filter by index altIdx. Bitstream conformance requirements may include that the value of alf_chroma_clip_idx[altIdx][j], for altIdx=0 to alf_chroma_num_alt_filters_minus1, j=0 to 5, be in the range 0 to 3, inclusive. The chroma filter clipping value AlfClipC[adaptation_parameter_set_id][altIdx] with element AlfClipC[adaptation_parameter_set_id][altIdx][j], with altIdx=0 to alf_chroma_num_alt_filters_minus1, j=0 to 5, can be derived as specified in Table 2 depending on bitDepth being set equal to BitDepthC and clipIdx being set equal to alf_chroma_clip_idx[altIdx][j].
[0109] In one embodiment, the filtering process can be described as follows: At the decoder side, when ALF is enabled for CTB, samples R(i,j) in a CU (or CB) can be filtered, resulting in a filtered sample value R'(i,j) as shown below using equation (13). In one example, each sample in a CU is filtered.
number
[0110] In non-linear ALF, multiple sets of clipping values can be provided in Table 3. In one example, the luma set includes four clipping values {1024, 181, 32, 6} and the chroma set includes four clipping values {1024, 161, 25, 4}. The four clipping values in the luma set can be selected by approximately evenly dividing the full range (e.g., 1024) of sample values (encoded in 10 bits) for a luma block in the logarithmic domain. This range can be between 4 and 1024 for the chroma set. [Table 3]
[0111] The selected clipping value may be encoded in the "alf_data" syntax element as follows: An appropriate encoding scheme (e.g., Golomb encoding scheme) may be used to encode the clipping index corresponding to the selected clipping value as shown in Table 3. The encoding scheme may be the same encoding scheme used to encode the filter set index.
[0112] In one embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of the ALF. Thus, for samples near a CTU boundary (e.g., a horizontal CTU boundary), modified block classification and filtering can be used. The virtual boundary (1130) is defined as the horizontal CTU boundary (1120) as "N samples " samples, where N samples can be a positive integer. In one example, N samples is equal to 4 for the luma component, and N samples is equal to 2 for the chroma component.
[0113] With reference to Figure 11A, modified block classification can be applied for the luma component. In one example, for a 1D Laplacian gradient calculation of a 4x4 block (1110) above a virtual boundary (1130), only samples on the virtual boundary (1130) are used. Similarly, with reference to Figure 11B, for a 1D Laplacian gradient calculation of a 4x4 block (1111) below a virtual boundary (1131) shifted from the CTU boundary (1121), only samples below the virtual boundary (1131) are used. By taking into account the reduction in the number of samples used in the 1D Laplacian gradient calculation, the quantization of the activity value A can be scaled accordingly.
[0114] For the filtering process, a symmetric padding operation at the virtual boundary can be used for both the luma and chroma components. FIGS. 12A-12F show an example of such modified ALF filtering for the luma component at the virtual boundary. If the sample to be filtered is located below the virtual boundary, the neighboring samples located above the virtual boundary can be padded. If the sample to be filtered is located above the virtual boundary, the neighboring samples located below the virtual boundary can be padded. Referring to FIG. 12A, the neighboring sample C0 can be padded with a sample C2 located below the virtual boundary (1210). Referring to FIG. 12B, the neighboring sample C0 can be padded with a sample C2 located above the virtual boundary (1220). Referring to FIG. 12C, the neighboring samples C1-C3 can be padded with samples C5-C7 located below the virtual boundary (1230), respectively. Referring to FIG. 12D, the neighboring samples C1-C3 can be padded with samples C5-C7 located above the virtual boundary (1240), respectively. Referring to Figure 12E, the neighboring samples C4 to C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located below the virtual boundary (1250). Referring to Figure 12F, the neighboring samples C4 to C8 may be padded with samples C10, C11, C12, C11, and C10, respectively, located above the virtual boundary (1260).
[0115] In some instances, the above description may be appropriately adapted when a sample and a neighboring sample are located to the left (or right) and right (or left) of a virtual boundary.
[0116] According to an aspect of the present disclosure, a picture can be divided based on a filtering process to improve coding efficiency. In some examples, a CTU is also referred to as a largest coding unit (LCU). In one example, a CTU or an LCU can have a size of 64×64 pixels. In some embodiments, an LCU-aligned picture quadtree partition can be used for filtering-based partitioning. In some examples, a coding unit-synchronous picture quadtree-based adaptive loop filter can be used. For example, a luma picture is divided into several multi-level quadtree partitions, and each partition boundary is aligned to the boundary of an LCU. Each partition has its own filtering process, and is therefore called a filter unit (FU).
[0117] In some examples, a two-pass encoding flow may be used. In a first pass of the two-pass encoding flow, a quadtree partitioning pattern of the picture and a best filter for each FU may be determined. In some embodiments, the determination of the quadtree partitioning pattern of the picture and the determination of the best filter for the FU are based on the filtering distortion. The filtering distortion may be estimated by a fast filtering distortion estimation (FFDE) technique in the decision process. The picture is partitioned using a quadtree partition. The reconstructed picture may be filtered according to the determined quadtree partitioning pattern and the selected filters of all the FUs.
[0118] In the second pass of the two-pass encoding flow, CU-synchronous ALF on / off control is performed. According to the ALF on / off result, the original filtered picture is partially restored by the reconstructed picture.
[0119] Specifically, in some examples, a top-down partitioning strategy is adopted to divide a picture into multi-level quadtree partitions by using a rate-distortion criterion. Each partition is called a filter unit (FU). The partitioning process aligns the quadtree partitions to LCU boundaries. The coding order of the FUs follows the z-scan order.
[0120] Figure 13 shows an example of partitioning according to some embodiments of the present disclosure. In the example of Figure 13, a picture (1300) is partitioned into 10 FUs, with the coding order being FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.
[0121] Figure 14 shows a quadtree partitioning pattern (1400) for a picture (1300). In the example of Figure 14, a split flag is used to indicate the picture partitioning pattern. For example, a "1" indicates that a quadtree partition is performed on the block, and a "0" indicates that the block is not further partitioned. In some examples, the smallest size FU has an LCU size, and no split flag is needed for the smallest size FU. The split flags are coded and transmitted in z-order as shown in Figure 14.
[0122] In some examples, the filter of each FU is selected from two filter sets based on a rate-distortion criterion. The first set has 1 / 2 symmetric square and diamond filters derived for the current FU. The second set comes from a time-delay filter buffer, which stores filters previously derived for FUs of previous pictures. The filter with the minimum rate-distortion cost of these two sets can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further split into four child FUs, the rate-distortion costs of the four child FUs are calculated. By recursively comparing the rate-distortion costs of the split case and the non-split case, a picture quadtree splitting pattern can be determined.
[0123] In some examples, a maximum quadtree division level can be used to limit the maximum number of FUs. In one example, when the maximum quadtree division level is 2, the maximum number of FUs is 16. Furthermore, during the quadtree division decision, correlation values for deriving Wiener coefficients of the 16 FUs at the lowest quadtree level (minimum FUs) can be reused. The remaining FUs can derive Wiener filters from the correlations of the 16 FUs at the lowest quadtree level. Thus, in this example, only one frame buffer access is performed to derive filter coefficients for all FUs.
[0124] After the quadtree division pattern is determined, CU-synchronous ALF on / off control can be performed to further reduce the filtering distortion. By comparing the filtering distortion and the non-filtering distortion in each leaf CU, the leaf CU can explicitly switch ALF on / off in its local region. In some examples, the coding efficiency can be further improved by redesigning the filter coefficients according to the ALF on / off result.
[0125] The cross-component filtering process may apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter may use a luma sample value of a luma component (e.g., luma CB) to refine a chroma component (e.g., chroma CB corresponding to luma CB). In one example, the luma CB and the chroma CB are included in a CU.
[0126] FIG. 15 illustrates a cross-component filter (e.g., CC-ALF) used to generate chroma components according to certain embodiments of the present disclosure. In some examples, FIG. 15 illustrates a filtering process of a first chroma component (e.g., first chroma CB), a second chroma component (e.g., second chroma CB), and a luma component (e.g., luma CB). The luma component may be filtered by a sample adaptive offset (SAO) filter (1510) to generate an SAO filtered luma component (1541). The SAO filtered luma component (1541) may be further filtered by an ALF luma filter (1516) to become a filtered luma CB (1561) (e.g., “Y”).
[0127] The first chroma component may be filtered by an SAO filter (1512) and an ALF chroma filter (1518) to generate a first intermediate component (1552). Furthermore, the SAO filtered luma component (1541) may be filtered by a cross-component filter (e.g., CC-ALF) (1521) for the first chroma component to generate a second intermediate component (1542). A filtered first chroma component (1562) (e.g., “Cb”) may then be generated based on at least one of the second intermediate component (1542) and the first intermediate component (1552). In one example, the filtered first chroma component (1562) (e.g., “Cb”) may be generated by combining the second intermediate component (1542) and the first intermediate component (1552) using an adder (1522). The cross-component adaptive loop filtering process for the first chroma component may include steps performed by a CC-ALF (1521) and steps performed, for example, by an adder (1522).
[0128] The above description can be applied to the second chroma component. The second chroma component can be filtered by the SAO filter (1514) and the ALF chroma filter (1518) to generate a third intermediate component (1553). Furthermore, the SAO filtered luma component (1541) can be filtered by a cross-component filter (e.g., CC-ALF) (1531) for the second chroma component to generate a fourth intermediate component (1543). Then, a filtered second chroma component (1563) (e.g., “Cr”) can be generated based on at least one of the fourth intermediate component (1543) and the third intermediate component (1553). In one example, the filtered second chroma component (1563) (e.g., “Cr”) can be generated by combining the fourth intermediate component (1543) and the third intermediate component (1553) using an adder (1532). In one example, the cross-component adaptive loop filtering process for the second chroma component may include steps performed by a CC-ALF (1531) and steps performed, for example, by an adder (1532).
[0129] The cross-component filters (e.g., CC-ALF (1521), CC-ALF (1531)) may operate by applying a linear filter having any suitable filter shape to the luma component (or luma channel) in order to refine each chroma component (e.g., first chroma component, second chroma component).
[0130] FIG. 16 illustrates an example of a filter (1600) according to an embodiment of the present disclosure. The filter (1600) may include non-zero filter coefficients and zero filter coefficients. The filter (1600) has a diamond shape (1620) (shown as a filled black circle) formed by the filter coefficients (1610). In one example, the non-zero filter coefficients in the filter (1600) are included in the filter coefficients (1610), and the filter coefficients that are not included in the filter coefficients (1610) are zero. Thus, the non-zero filter coefficients in the filter (1600) are included in the diamond shape (1620), and the filter coefficients that are not included in the diamond shape (1620) are zero. In one example, the number of filter coefficients in the filter (1600) is equal to the number of filter coefficients (1610), which is 18 in the embodiment illustrated in FIG. 16.
[0131] The CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Referring back to Figure 15, the CC-ALF (1521) and the CC-ALF (1531) may have the same filter shape, such as the diamond shape (1620) shown in Figure 16, and the same number of filter coefficients. In one example, the values of the filter coefficients in the CC-ALF (1521) are different from the values of the filter coefficients in the CC-ALF (1531).
[0132] In general, filter coefficients in CC-ALF (e.g., non-zero filter coefficients) may be transmitted, for example, in APS. In one example, the filter coefficients may be a certain coefficient (e.g., 2 10) and may be rounded for fixed-point representation. Application of CC-ALF may be controlled with variable block sizes and may be signaled by a context-coded flag (e.g., a CC-ALF enable flag) received for each block of samples. Context-coded flags such as the CC-ALF enable flag may be signaled at any appropriate level, such as the block level. Block sizes may be received at the slice level for each chroma component along with the CC-ALF enable flag. In some examples, block sizes (in chroma samples) of 16×16, 32×32, 64×64 may be supported.
[0133] 17 illustrates an example syntax for CC-ALF according to some embodiments of the present disclosure. In the example of FIG. 17, alf_ctb_cross_component_cross_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index that indicates whether a cross-component Cb filter is used, and if so, the index of that cross-component Cb filter. For example, if alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, then a cross-component Cb filter is not applied to a block of Cb color component samples at luma position (xCtb, yCtb). alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index for the filter to be applied, if it is not equal to 0. For example, the alf_ctb_cross_component_cb_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] cross component Cb filter is applied to the block of Cb color component samples at luma position (xCtb, yCtb).
[0134] Further, in the example of Figure 17, alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is used to indicate whether a cross-component Cr filter is used and is the index of the cross-component Cr filter to be used. For example, if alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is equal to 0, then the cross-component Cr filter is not applied to the block of Cr color component samples at luma position (xCtb, yCtb). alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is an index of a cross-component Cr filter, if alf_ctb_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] is not equal to 0. For example, the alf_cross_component_cr_idc[xCtb>>CtbLog2SizeY][yCtb>>CtbLog2SizeY] cross-component Cr filter may be applied to a block of Cr color component samples at luma position (xCtb, yCtb).
[0135] In some examples, a chroma subsampling technique is used, so that the number of samples in each of the chroma blocks can be less than the number of samples in the luma blocks. The chroma subsampling format (also referred to as chroma subsampling format, for example, specified by chroma_format_idc) can indicate a chroma horizontal subsampling factor (e.g., SubWidthC) and a chroma vertical subsampling factor (e.g., SubHeightC) between each of the chroma blocks and the corresponding luma block. In one example, the chroma subsampling format is 4:2:0, so that the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as shown in Figures 18A-18B. In one example, the chroma subsampling format is 4:2:2, so that the chroma horizontal subsampling factor (e.g., SubWidthC) is 2 and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In one example, the chroma subsampling format is 4:4:4, and thus the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. The chroma sample type (also called the chroma sample position) may indicate the relative position of a chroma sample in a chroma block with respect to at least one corresponding luma sample in the luma block.
[0136] 18A-18B illustrate exemplary locations of chroma samples relative to luma samples according to an embodiment of the present disclosure. With reference to FIG. 18A, luma samples (1801) are located in rows (1811)-(1818). The luma samples (1801) illustrated in FIG. 18A may represent a portion of a picture. In one example, a luma block (e.g., luma CB) includes luma sample (1801). The luma block may correspond to two chroma blocks with a chroma subsampling format of 4:2:0. In one example, each chroma block includes chroma sample (1803). Each chroma sample (e.g., chroma sample (1803(1))) corresponds to four luma samples (e.g., luma samples (1801(1))-(1801(4))). In one example, the four luma samples are the top left sample (1801(1)), the top right sample (1801(2)), the bottom left sample (1801(3)), and the bottom right sample (1801(4)). A chroma sample (e.g., (1803(1))) is located at a left center position between the top left sample (1801(1)) and the bottom left sample (1801(3)), and the chroma sample type of the chroma block having the chroma sample (1803) may be referred to as chroma sample type 0. The chroma sample type 0 indicates a relative position 0, which corresponds to a left center position halfway between the top left sample (1801(1)) and the bottom left sample (1801(3)). The four luma samples (e.g., (1801(1)) to (1801(4))) may be referred to as neighboring luma samples of the chroma sample (1803)(1).
[0137] In one example, each chroma block includes a chroma sample (1804). The above description of the chroma sample (1803) may be applied to the chroma sample (1804), and thus a detailed description may be omitted for brevity. Each of the chroma samples (1804) may be located at a central position of the corresponding four luma samples, and a chroma sample type of a chroma block having the chroma sample (1804) may be referred to as chroma sample type 1. The chroma sample type 1 indicates a relative position 1 corresponding to the central position of the four luma samples (e.g., (1801(1)) to (1801(4))). For example, one of the chroma samples (1804) may be located at the central portion of the luma samples (1801(1)) to (1801(4)).
[0138] In one example, each chroma block includes a chroma sample (1805). Each of the chroma samples (1805) may be located at a top left position that is co-located with the top left sample of the corresponding four luma samples (1801), and the chroma sample type of the chroma block having the chroma sample (1805) may be referred to as a chroma sample type 2. Thus, each chroma sample (1805) is co-located with the top left sample of the four luma samples (1801) corresponding to the respective chroma sample. The chroma sample type 2 indicates a relative position 2, which corresponds to the top left position of the four luma samples (1801). For example, one of the chroma samples (1805) may be located at the top left position of the luma samples (1801(1)) to (1801(4)).
[0139] In one example, each chroma block includes a chroma sample (1806). Each of the chroma samples (1806) may be located at a top center position between a corresponding top-left sample and a corresponding top-right sample, and a chroma sample type of the chroma block having the chroma sample (1806) may be referred to as a chroma sample type 3. The chroma sample type 3 indicates a relative position 3, which corresponds to a top center position between the top-left sample (and the top-right sample). For example, one of the chroma samples (1806) may be located at a top center position of the luma samples (1801(1))-(1801(4)).
[0140] In one example, each chroma block includes a chroma sample (1807). Each of the chroma samples (1807) may be located at a bottom left position that is co-located with the bottom left sample of the corresponding four luma samples (1801), and the chroma sample type of the chroma block having the chroma sample (1807) may be referred to as chroma sample type 4. Thus, each chroma sample (1807) is co-located with the bottom left sample of the four luma samples (1801) corresponding to the respective chroma sample. The chroma sample type 4 indicates a relative position 4, which corresponds to the bottom left position of the four luma samples (1801). For example, one of the chroma samples (1807) may be located at the bottom left position of the luma samples (1801(1)) to (1801(4)).
[0141] In one example, each chroma block includes a chroma sample (1808). Each of the chroma samples (1808) is located at a bottom center position between a bottom left sample and a bottom right sample, and a chroma sample type of the chroma block having the chroma sample (1808) may be referred to as a chroma sample type 5. The chroma sample type 5 indicates a relative position 5, which corresponds to a bottom center position between the bottom left sample and the bottom right sample of the four luma samples (1801). For example, one of the chroma samples (1808) may be located between the bottom left sample and the bottom right sample of the luma samples (1801(1))-(1801(4)).
[0142] In general, any suitable chroma sample type can be used for the chroma subsampling format. Chroma sample types 0-5 are exemplary chroma sample types described for chroma subsampling format 4:2:0. For chroma subsampling format 4:2:0, additional chroma sample types may be used. Furthermore, other chroma sample types and / or variations of chroma sample types 0-5 can be used for other chroma subsampling formats such as 4:2:2, 4:4:4, etc. In one example, a chroma sample type combining chroma samples (1805) and (1807) is used for chroma subsampling format 4:2:2.
[0143] In one example, a luma block may be considered to have alternating rows such as rows (1811)-(1812) including the top two samples (e.g., (1801(1))-(1801(2))) and the bottom two samples (e.g., (1801(3))-(1801(4))) of four luma samples (e.g., (1801(1))-(1801(4))), respectively. Thus, rows (1811), (1813), (1815), and (1817) may be referred to as current rows (also referred to as top field), and rows (1812), (1814), (1816), and (1818) may be referred to as next rows (also referred to as bottom field). The four luma samples (e.g. (1801(1)) to (1801(4))) are located on the current row (e.g. (1811)) and the next row (e.g. (1812)). Relative positions 2 to 3 are located on the current row, relative positions 0 to 1 are located between each current row and the respective next row, and relative positions 4 to 5 are located on the next row.
[0144] Chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 are located in rows 1851-1854 in each chroma block. The specific locations of rows 1851-1854 may depend on the type of the chroma samples. For example, for chroma samples 1803-1804 having respective chroma sample types 0-1, row 1851 is located between rows 1811-1812. For chroma samples 1805-1806 having respective chroma sample types 2-3, row 1851 is co-located with the current row 1811. For chroma samples 1807-1808 having respective chroma sample types 4-5, row 1851 is co-located with the next row 1812. The above description can be applied appropriately to rows 1852-1854, and a detailed description will be omitted for brevity.
[0145] Any suitable scanning method may be used to display, store, and / or transmit the luma blocks and corresponding chroma blocks described above in Figure 18A. In one example, progressive scanning is used.
[0146] As shown in Figure 18B, interlaced scanning can be used. As mentioned above, the chroma subsampling format is 4:2:0 (e.g., chroma_format_idc is equal to 1). In one example, the variable chroma location type (e.g., ChromaLocType) indicates the current row (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next row (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current rows (1811), (1813), (1815), and (1817) and the next rows (1812), (1814), (1816), and (1818) can be scanned separately. For example, current rows (1811), (1813), (1815), and (1817) are scanned first, followed by next rows (1812), (1814), (1816), and (1818). The current row may contain luma sample (1801) and the next row may contain luma sample (1802).
[0147] Similarly, corresponding chroma blocks may be interlaced scanned. Rows 1851 and 1853 containing unfilled chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 may be referred to as the current row (or current chroma row), and rows 1852 and 1854 containing gray filled chroma samples 1803, 1804, 1805, 1806, 1807, or 1808 may be referred to as the next row (or next chroma row). In one example, during interlaced scanning, rows 1851 and 1853 are scanned first, followed by rows 1852 and 1854.
[0148] In some examples, a constrained directional enhancement filtering technique can be used. The use of a constrained directional enhancement filter (CDEF) in the loop can remove coding artifacts while preserving image details. In one example (e.g., HEVC), a sample adaptive offset (SAO) algorithm can achieve a similar goal by defining signal offsets for different classes of pixels. Unlike SAO, CDEF is a nonlinear spatial filter. In some examples, CDEF can be constrained to be easily vectorizable (i.e., implementable with single instruction multiple data (SIMD) operations). Note that other nonlinear filters such as median filters, bilateral filters, etc. cannot be handled in the same way.
[0149] In some cases, the amount of ringing artifacts in an encoded image tends to be roughly proportional to the quantization step size. Although the amount of detail is a property of the input image, the smallest detail that is retained in the quantized image also tends to be proportional to the quantization step size. For a given quantization step size, the amplitude of the ringing is generally smaller than the amplitude of the detail.
[0150] The CDEF can be used to identify the orientation of each block and then adaptively filter along the identified orientation and to a lesser extent along orientations rotated 45 degrees from the identified orientation. In some examples, the encoder can search for filter strengths, which can be explicitly signaled, allowing a high degree of control over blurring.
[0151] Specifically, in some examples, the direction search is performed on the reconstructed pixels immediately after the deblocking filter. These pixels are available to the decoder, so the direction can be searched by the decoder, and thus the direction does not need to be signaled in one example. In some examples, the direction search can operate on a block size, such as an 8×8 block, that is small enough to properly handle non-straight edges, but large enough to reliably estimate the direction when applied to a quantized image. Also, having a constant direction over an 8×8 region makes it easier to vectorize the filter. In some examples, each block (e.g., 8×8) can be compared to a perfectly oriented block to determine the difference. A perfectly oriented block is one in which all of the pixels along a line of one direction have the same value. In one example, a measure of the difference between the block and the perfectly oriented block, such as the sum of squared differences (SSD), root mean square (RMS) error, can be calculated. A perfectly oriented block having the smallest difference (e.g., smallest SSD, smallest RMS, etc.) can then be determined, and the orientation of the determined perfectly oriented block can be the orientation that best matches the pattern within that block.
[0152] FIG. 19 illustrates an example of a direction search according to an embodiment of the present disclosure. In one example, block (1910) is an 8×8 block that is reconstructed and output from a deblocking filter. In the example of FIG. 19, the direction search can determine a direction from eight directions indicated by (1920) for block (1910). Eight perfectly oriented blocks (1930) are formed corresponding to the eight directions (1920), respectively. A perfectly oriented block corresponding to a direction is a block whose pixels along the line of that direction have the same value. In addition, a difference metric such as SSD, RMS error, etc. between block (1910) and each of the perfectly oriented blocks (1930) can be calculated. In the example of FIG. 19, the RMS error is indicated by (1940). As indicated by (1943), the RMS error between block (1910) and the perfectly oriented block (1933) is the smallest, and thus direction (1923) is the direction that best matches the pattern of block (1910).
[0153] After the direction of the block is identified, a nonlinear low-pass directional filter can be determined. For example, the filter taps of the nonlinear low-pass directional filter can be aligned along the identified direction to reduce ringing while preserving directional edges or patterns. However, in some examples, directional filtering alone may not be sufficient to reduce ringing. In one example, additional filter taps are also used for pixels that are not along the identified direction. The additional filter taps are treated more conservatively to reduce the risk of blurring. Thus, the CDEF includes a first-order filter tap and a second-order filter tap. In one example, the complete 2-D CDEF filter can be expressed as Equation (14).
number
[0154] In some examples, in-loop restoration schemes are used in video coding after deblocking to generally remove noise and improve edge quality beyond the deblocking operation. In one example, the in-loop restoration schemes are switchable within a frame for each tile of appropriate size. The in-loop restoration schemes are based on a separable symmetric Wiener filter, a dual self-guided filter with subspace projection, and a domain transform recursive filter. Since content statistics can change substantially within a frame, the in-loop restoration schemes are integrated into a switchable framework where different schemes can be triggered in different regions of a frame.
[0155] According to one aspect of the disclosure, in-loop restoration (LR) (also referred to as LR filter) can use neighboring pixels during LR filtering, such as a window of pixels around the pixel to be filtered. To allow for application of the LR filter at the frame edge, in some examples, some boundary pixel values of pixels at the frame boundary are copied to a boundary buffer before the LR filtering is applied. In some examples, two copies of the boundary pixels are stored in the boundary buffer. In one example, a first copy step is applied after the deblocking filter and before the CDEF, and the first copy step copy of the boundary pixel values of pixels at the frame boundary is denoted as COPY0. A second copy step is applied after the CDEF and before the LR filter, and the second copy step copy of the boundary pixel values of pixels at the frame boundary is denoted as COPY1. The boundary buffer is then used for padding the boundary pixels during the LR filtering process, such as when applying the LR filter to the frame edge. The process of copying the boundary pixel values to the boundary buffer is denoted as boundary processing.
[0156] A separable symmetric Wiener filter can be one of the in-loop restoration schemes. In some examples, every pixel in the degraded frame can be reconstructed as a non-causal filtered version of the pixels in a w × w window around it, where w = 2r + 1 is odd with respect to the integer r. The 2D filter taps are expressed as w in a column vectorized form. 2 If F is a 1×1 vector, then the filter parameters can be calculated by direct LMMSE optimization: F=H -1 M, where H=E[XX T ] is the autocovariance of x, w in a w × w window around the pixel 2 is a column vectorized version of the samples, M=E[YX T] is the cross-correlation of x with the scalar source sample y, which is to be estimated. In one example, the encoder can estimate H and M from realizations in the deblocked frame and the source, and send the resulting filter F to the decoder. However, this is not possible with w 2 Not only does it incur a substantial bit rate cost in transmitting the taps, but non-separable filtering also makes decoding prohibitively complex. In some embodiments, some additional constraints are placed on the nature of F. For the first constraint, F is constrained to be separable so that the filtering can be implemented as a separable horizontal and vertical w-tap convolution. For the second constraint, each of the horizontal and vertical filters is constrained to be symmetric. For the third constraint, both the horizontal and vertical filter coefficients are assumed to sum to one.
[0157] Dual self-guided filtering with subspace projection can be one of the in-loop restoration schemes. Guided filtering is an image filtering technique in which a locally linear model, shown by the following equation (15), is used to calculate the filtered output y from the unfiltered samples x. y=Fx+G Equation (15) Here, F and G are determined based on the statistics of the degraded image and the guide image in the neighborhood of the filtered pixel. If the guide image is the same as the degraded image, the resulting so-called self-guided filtering has the effect of edge-preserving smoothing. In one example, a specific form of self-guided filtering can be used. The specific form of self-guided filtering depends on two parameters, the radius r and the noise parameter e, and is enumerated as the following steps: 1. The mean μ and variance σ of the pixels in a (2r+1) × (2r+1) window around each pixel 2 This step can be implemented efficiently using box filtering based on integral imaging. 2. For every pixel, f=σ2 / (σ 2 +e);g = (1-f)μ. 3. Calculate F and G for each pixel as the average of the f and g values in a 3x3 window around the pixel being used.
[0158] The particular shape of the self-guiding filter is controlled by r and e, where a larger r means a larger spatial variance and a larger e means a larger range variance.
[0159] Figure 20 shows an example showing subspace projection in some cases. As shown in Figure 20, neither of the reconstructions X1, X2 is close to the source Y, but appropriate multipliers {α, β} can bring them much closer to the source Y as long as they are moving in the right direction.
[0160] In some examples, a technique called frame super-resolution (FSR) is used to improve the perceptual quality of the decoded picture. The FSR process is generally applied at a low bit rate and includes four steps. In the first step, the source video is downscaled at the encoder side as a non-prescriptive procedure. In the second step, the downscaled video is encoded, followed by the deblocking filter and CDEF filtering processes. In the third step, a linear upscaling process is applied as a prescriptive procedure to restore the encoded video to the original spatial resolution. In the fourth step, a loop restoration filter is applied to resolve some of the lost high frequencies. In one example, the last two steps can be collectively referred to as the super-resolution process. Similarly, at the decoder side, the decoding, deblocking filter, and CDEF processes can be applied at a lower spatial resolution. The frame is then passed through the super-resolution process. In some examples, the upscaling and downscaling processes are applied only in the horizontal dimension to reduce the line buffer overhead for hardware implementation.
[0161] In some examples (e.g., HEVC), a filtering technique called sample adaptive offset (SAO) can be used. In some examples, SAO is applied to the reconstructed signal after the deblocking filter. SAO can use an offset value given in the slice header. In some examples, for luma samples, the encoder can determine whether to apply (enable) SAO for a slice. When SAO is enabled, the current picture allows recursively splitting the coding unit into four sub-regions, and each sub-region can select an SAO type from multiple SAO types based on features within that sub-region.
[0162] FIG. 21 shows a table (2100) of multiple SAO types according to an embodiment of the present disclosure. In the table (2100), SAO types 0 to 6 are shown. Note that SAO type 0 is used to indicate no SAO application. Also, each SAO type, SAO type 1 to SAO type 6, includes multiple categories. SAO can reduce distortion by classifying reconstructed pixels of a subregion into categories and adding an offset to pixels of each category in the subregion. In some examples, edge characteristics can be used for pixel classification in SAO types 1 to 4, and pixel intensity can be used for pixel classification in SAO types 5 to 6.
[0163] Specifically, in some embodiments, such as SAO types 5-6, a band offset (BO) can be used to classify all pixels of a subregion into multiple bands. Each band of the multiple bands includes pixels in the same intensity interval. In some examples, the intensity range is evenly divided into multiple intervals, such as 32 intervals from zero to a maximum intensity value (e.g., 255 for 8-bit pixels), and each interval is associated with an offset. Furthermore, in one example, the 32 bands are divided into two groups, such as a first group and a second group. The first group includes the middle 16 bands (e.g., the 16 intervals in the middle of the intensity range), and the second group includes the remaining 16 bands (e.g., 8 intervals at the low side of the intensity range and 8 intervals at the high side of the intensity range). In one example, only the offset of one of the two groups is transmitted. In some embodiments, when pixel classification operation in BO is used, the 5 most significant bits of each pixel can be directly used as a band index.
[0164] Additionally, in some embodiments, such as SAO types 1-4, edge offset (EO) can be used for pixel classification and offset determination. For example, pixel classification can be determined based on one-dimensional 3-pixel patterns, taking into account edge direction information.
[0165] FIG. 22 shows an example of a three-pixel pattern for pixel classification in edge offset in some examples. In the example of FIG. 22, a first pattern (2210) (shown with three gray pixels) is referred to as a 0 degree pattern (horizontal direction is associated with a 0 degree pattern), a second pattern (2220) (shown with three gray pixels) is referred to as a 90 degree pattern (vertical direction is associated with a 90 degree pattern), a third pattern (2230) (shown with three gray pixels) is referred to as a 135 degree pattern (135 degree diagonal direction is associated with a 135 degree pattern), and a fourth pattern (2240) (shown with three gray pixels) is referred to as a 45 degree pattern (45 degree diagonal direction is associated with a 45 degree pattern). In one example, one of the four direction patterns shown in FIG. 22 can be selected taking into account edge direction information for the sub-region. The selection can be sent in the encoded video bitstream, in one example as side information. The pixels within the sub-region can then be classified into multiple categories by comparing each pixel with its two neighboring pixels on a direction associated with the directional pattern.
[0166] Fig. 23 shows a table (2300) for pixel classification rules for edge offsets in some examples. Specifically, pixel c (also shown in each pattern of Fig. 22) is compared with two neighboring pixels (shown in gray in each pattern of Fig. 22), and pixel c can be classified into one of categories 0 to 4 based on the comparison according to the pixel classification rules shown in Fig. 23.
[0167] In some embodiments, the decoder-side SAO can operate independently of the largest coding unit (LCU) (e.g., CTU) to conserve line buffers. In some examples, when a 90-degree, 135-degree, 45-degree classification pattern is selected, the pixels in the top and bottom rows in each LCU are not SAO processed. When a 0-degree, 135-degree, 45-degree pattern is selected, the pixels in the leftmost and rightmost columns in each LCU are not SAO processed.
[0168] FIG. 24 shows an example of syntax (2400) that may need to be signaled for a CTU when parameters from neighboring CTUs are not merged. For example, syntax element sao_type_idx[cldx][rx][ry] can be signaled to indicate the SAO type of the sub-region. The SAO type can be BO (band offset) or EO (edge offset). When the value of sao_type_idx[cldx][rx][ry] is 0, it indicates that SAO is off, values of 1 to 4 indicate that one of the four EO categories corresponding to 0°, 90°, 135°, and 45° is used, and a value of 5 indicates that BO is used. In the example of FIG. 24, each of the BO and EO types has four SAO offset values (sao_offset[cIdx][rx][ry][0] to sao_offset[cIdx][rx][ry][3]) that are signaled.
[0169] In general, the filtering process can use a reconstructed sample of a first color component (e.g., Y or Cb or Cr, or R or G or B) as an input to generate an output, and the output of the filtering process is applied to a second color component, which may be the same as the first color component or may be another color component different from the first color component.
[0170] In a related example of cross-component filtering (CCF), filter coefficients are derived based on some mathematical formulas. The derived filter coefficients are signaled from the encoder side to the decoder side, and the derived filter coefficients are used to generate an offset using a linear combination. The generated offset is then added to the reconstructed sample as a filtering process. For example, an offset is generated based on a linear combination of the luma sample and the filtering coefficient, and the generated offset is added to the reconstructed chroma sample. This related example of CCF is based on the assumption of a linear mapping relationship between the reconstructed luma sample value and the delta value between the original chroma sample and the reconstructed chroma sample. However, the mapping between the reconstructed luma sample value and the delta value between the original chroma sample and the reconstructed chroma sample does not necessarily follow a linear mapping process, and thus the coding performance of CCF may be limited under the assumption of a linear mapping relationship.
[0171] In some examples, nonlinear mapping techniques can be used in cross-component filtering and / or same-color component filtering without significant signaling overhead. In one example, nonlinear mapping techniques can be used in cross-component filtering to generate cross-component sample offsets. In another example, nonlinear mapping techniques can be used in same-color component filtering to generate local sample offsets.
[0172] For convenience, a filtering process using a nonlinear mapping technique may be referred to as sample offset by non linear mapping (SO-NLM). SO-NLM in a cross-component filtering process may be referred to as cross-component sample offset (CCSO). SO-NLM in the same color component filtering may be referred to as local sample offset (LSO). A filter using a nonlinear mapping technique may be referred to as a nonlinear mapping-based filter. The nonlinear mapping-based filter may include a CCSO filter, an LSO filter, etc.
[0173] In one example, CCSO and LSO can be used as loop filtering to reduce distortion of reconstructed samples. CCSO and LSO do not rely on the assumption of linear mapping used in the related exemplary CCF. For example, CCSO does not rely on the assumption of a linear mapping relationship between luma reconstructed sample values and delta values between original chroma samples and chroma reconstructed samples. Similarly, LSO does not rely on the assumption of a linear mapping relationship between reconstructed sample values of a color component and delta values between original samples of the color component and reconstructed samples of the color component.
[0174] In the following description, a SO-NLM filtering process is described, which uses reconstructed samples of a first color component as input (e.g., Y or Cb or Cr, or R or G or B) to generate an output, and the output of the filtering process is applied to a second color component. If the second color component is the same color component as the first color component, this description is applicable to LSO. If the second color component is different from the first color component, this description is applicable to CCSO.
[0175] In SO-NLM, a nonlinear mapping is derived at the encoder side. The nonlinear mapping is between the reconstructed samples of the first color component in the filter support region and the offset added to the second color component in the filter support region. If the second color component is the same as the first color component, the nonlinear mapping is used in LSO. If the second color component is different from the first color component, the nonlinear mapping is used in CCSO. The domain of the nonlinear mapping is determined by the different combinations of processed input reconstructed samples (also called possible reconstructed sample value combinations).
[0176] The technique of SO-NLM can be explained by a specific example. In the specific example, a reconstructed sample from a first color component located in a filter support area (also called a "filter support region") is determined. The filter support region is a region within which a filter can be applied, and the filter support region can have any suitable shape.
[0177] FIG. 25 illustrates an example of a filter support region (2500) according to some embodiments of the present disclosure. The filter support region (2500) includes four reconstructed samples of a first color component, P0, P1, P2, and P3. In the example of FIG. 25, the four reconstructed samples may form a cross shape in the vertical and horizontal directions, and the center position of the cross shape is the position of the sample to be filtered. The sample at the center position and of the same color component as P0-P3 is represented as C. The sample at the center position and of the second color component is represented as F. The second color component may be the same as the first color component P0-P3 or may be different from the first color component P0-P3.
[0178] FIG. 26 illustrates another example of a filter support region (2600) according to some embodiments of the present disclosure. The filter support region (2600) includes four reconstructed samples P0, P1, P2, and P3 of a first color component forming a square shape. In the example of FIG. 26, the center position of the square shape is the position of the sample to be filtered. The sample of the same color component as P0-P3 at the center position is represented by C. The sample of the second color component at the center position is represented by F. The second color component may be the same as the first color component P0-P3 or may be different from the first color component P0-P3.
[0179] The reconstructed samples are input to the SO-NLM filter and processed appropriately to form the filter taps. In one example, the positions of the reconstructed samples that are input to the SO-NLM filter are called filter tap positions. In a particular example, the reconstructed samples are processed in two steps:
[0180] In the first step, a delta value is calculated between each of P0-P3 and C. For example, m0 represents the delta value between P0 and C, m1 represents the delta value between P1 and C, m2 represents the delta value between P2 and C, and m3 represents the delta value between P3 and C.
[0181] In a second step, the delta values m0-m3 are further quantized, and the quantized values are represented as d0, d1, d2, d3. In one example, the quantized values may be either -1, 0, or 1 based on the quantization process. For example, if m is less than -N (N is a positive value and is referred to as the quantization step size), the value m may be quantized to -1, if m is in the range of [-N,N], the value m may be quantized to 0, and if m is greater than N, the value m may be quantized to 1. In some examples, the quantization step size N may be one of 4, 8, 12, 16, etc.
[0182] In some embodiments, the quantized values d0-d3 are filter taps and can be used to identify a combination in the filter domain. For example, the filter taps d0-d3 can form a combination in the filter domain. Each filter tap can have three quantized values, so if four filter taps are used, the filter domain contains 81 (3×3×3×3) combinations.
[0183] 27A-27C show a table (2700) having 81 combinations according to an embodiment of the present disclosure. The table (2700) includes 81 rows corresponding to the 81 combinations. In each row corresponding to a combination, the first column includes an index of the combination, the second column includes a value of filter tap d0 for that combination, the third column includes a value of filter tap d1 for that combination, the fourth column includes a value of filter tap d2 for that combination, the fifth column includes a value of filter tap d3 for that combination, and the sixth column includes an offset value associated with that combination for the nonlinear mapping. In one example, when filter taps d0-d3 are determined, an offset value (represented by s) associated with the combination of d0-d3 can be determined according to the table (2700). In one example, the offset values s0-s80 are integers such as 0, 1, -1, 3, -3, 5, -5, -7, etc.
[0184] In some embodiments, a final filtering process of SO-NLM can be applied as shown in equation (16). f'=clip(f+s) Equation (16) where f is the reconstructed sample of the second color component to be filtered and s is an offset value determined according to the filter taps resulting from processing the reconstructed sample of the first color component, such as by using table (2700). The sum of the reconstructed sample F and the offset value s is further clipped to within a range associated with the bit depth to determine the final filtered sample f' of the second color component.
[0185] Note that in the case of LSO, the second color component in the above description is the same as the first color component, and in the case of CCSO, the second color component in the above description may be different from the first color component.
[0186] It should be noted that the above description may be adjusted for other embodiments of the present disclosure.
[0187] In some examples, at the encoder side, the encoding device may derive a mapping between the reconstructed samples of the first color component within the filter support region and an offset to be added to the reconstructed samples of the second color component. The mapping may be any suitable linear or non-linear mapping. The filtering process may then be applied at the encoder side and / or the decoder side based on the mapping. For example, the mapping may be appropriately signaled to the decoder (e.g., the mapping is included in the encoded video bitstream transmitted from the encoder side to the decoder side), and the decoder may then perform the filtering process based on the mapping.
[0188] The implementation of nonlinear mapping based filters such as CCSO filters, LSO filters, etc. depends on the filter shape configuration. The filter shape configuration (also called filter shape) of a filter can refer to the properties of the pattern formed by the filter tap locations. The pattern can be defined by various parameters such as the number of filter taps, the geometric shape of the filter tap locations, the distance of the filter tap locations to the center of the pattern, etc. Using a fixed filter shape configuration can limit the performance of a nonlinear mapping based filter.
[0189] As shown in FIG. 24 and FIG. 25 and FIG. 27A-27C, some examples use a 5-tap filter design for filter shape configuration of the nonlinear mapping based filter. The 5-tap filter design can use tap positions at P0, P1, P2, P3, and C. The 5-tap filter design for filter shape configuration can result in a look-up table (LUT) with 81 entries, as shown in FIG. 27A-27C. The LUT of sample offsets needs to be signaled from the encoder side to the decoder side, and the signaling of the LUT may contribute to a large portion of the signaling overhead and affect the coding efficiency using the nonlinear mapping based filter. According to some aspects of the present disclosure, the number of filter taps may be different from 5. In some examples, the number of filter taps can be reduced and information within the filter support region can still be captured and coding efficiency can be improved.
[0190] In some examples, the filter shape configurations within a group for the nonlinear mapping based filters have three filter taps each.
[0191] FIG. 28 illustrates seven filter shape configurations of three filter taps in one example. Specifically, a first filter shape configuration includes three filter taps at positions labeled as "1" and "c", with position "c" being the center position of position "1". A second filter shape configuration includes three filter taps at positions labeled as "2" and position "c", with position "c" being the center position of position "2". A third filter shape configuration includes three filter taps at positions labeled as "3" and position "c", with position "c" being the center position of position "3". A fourth filter shape configuration includes three filter taps at positions labeled as "4" and position "c", with position "c" being the center position of position "4". A fifth filter shape configuration includes three filter taps at positions labeled as "5" and "c", with position "c" being the center position of position "5". The sixth filter shape configuration includes three filter taps at a position labeled as "6" and at position "c", which is the center position of position "6". The seventh filter shape configuration includes three filter taps at a position labeled as "7" and at position "c", which is the center position of position "7".
[0192] Aspects of the present disclosure provide techniques for integrating video processing tools such as filtering, boundary processing, clipping tools, etc. In some examples, a nonlinear mapping-based filter (e.g., CCSO, LSO) and other tools (e.g., clipping module) are after the deblocking filter and before the LR filter, and a boundary processing process that stores two copies of boundary pixels is also applied after the deblocking filter and before the LR filter. The present disclosure provides various configurations for including a nonlinear mapping-based filter, a boundary processing module, and / or a clipping module after the deblocking filter and before the LR filter.
[0193] FIG. 29 shows a block diagram of a loop filter chain (2900) in some examples. The loop filter chain (2900) includes multiple filters connected in series in a filter chain. In one example, the loop filter chain (2900) can be used as a loop filtering unit, such as the loop filter unit (556). The loop filter chain (2900) can be used in an encoding or decoding loop before storing a reconstructed picture in a decoded picture buffer, such as the reference picture memory (557). The loop filter chain (2900) receives input reconstructed samples from a pre-processing module and applies filters to the reconstructed samples to generate output reconstructed samples.
[0194] The loop filter chain (2900) may include any suitable filter. In the example of FIG. 29, the loop filter chain (2900) includes a deblocking filter (labeled DEBLOCKING), a constrained directional enhancement filter (labeled CDEF), and an in-loop reconstruction filter (labeled LR) connected in a chain. The loop filter chain (2900) has an input node (2901), an output node (2909), and multiple intermediate nodes (2902)-(2903). The input node (2901) of the loop filter chain (2900) receives input reconstructed samples from the pre-processing module, which are provided to the deblocking filter. The intermediate node (2902) receives reconstructed samples (after being processed by the deblocking filter) from the deblocking filter and provides the reconstructed samples to the CDEF for further filtering. The intermediate node (2903) receives reconstructed samples (after being processed by the CDEF) from the CDEF and provides the reconstructed samples to the LR filter for further filtering. The output node (2909) receives the output reconstructed samples from the LR filter (after they have been processed by the LR filter). The output reconstructed samples can be provided to other processing modules, such as a post-processing module, for further processing.
[0195] Note that the following description illustrates a technique for using a nonlinear mapping-based filter in the loop filter chain (2900). The technique for using a nonlinear mapping-based filter can also be used in other suitable loop filter chains.
[0196] According to some aspects of the disclosure, when a nonlinear mapping-based filter is used in the loop filter chain, at least one of the copies (COPY0 and COPY1) of the boundary pixel values used by the boundary processing in the LR filter is associated with the nonlinear mapping-based filter. In some examples, the reconstructed samples from which the boundary pixel values are obtained and buffered may be inputs of the nonlinear mapping-based filter. In some examples, the reconstructed samples from which the boundary pixel values are obtained and buffered may be results of applying the nonlinear mapping-based filter. In some examples, the reconstructed samples from which the boundary pixel values are obtained and buffered may be combined with a sample offset generated by the nonlinear mapping-based filter.
[0197] In some examples, pixels after the deblocking filter and before the CDEF or nonlinear mapping based filter are used for COPY0, and pixels after application of the CDEF or nonlinear mapping based filter and before the LR filter are used for COPY1.
[0198] FIG. 30 shows an example of a loop filter chain (3000) including a nonlinear mapping-based filter and a CDEF between the input and output of the nonlinear mapping-based filter. The loop filter chain (3000) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3000), a deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3011), which is input to a CDEF and a nonlinear mapping-based filter (labeled SO-NLM). The CDEF is applied to the first intermediate reconstructed sample to generate a second intermediate reconstructed sample at a second intermediate node (3012). The nonlinear mapping-based filter generates a sample offset SO based on the first intermediate reconstructed sample. The sample offset SO is combined with the second intermediate reconstructed sample to generate a third intermediate reconstructed sample at a third intermediate node (3013). An LR filter is then applied to the third intermediate reconstructed sample to generate an output of the loop filter chain (3000). Further, the first intermediate reconstructed sample at the first intermediate node (3011) is used to obtain a first copy COPY0 of boundary pixels for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third intermediate node (3013) is used to obtain a second copy COPY1 of boundary pixels for boundary processing in the LR filter.
[0199] In some examples, pixels after the deblocking filter and before application of the CDEF or nonlinear mapping based filter are used for COPY0, and pixels after application of the CDEF and before application of the nonlinear mapping based filter are used for COPY1.
[0200] FIG. 31 shows an example of a loop filter chain (3100) including a nonlinear mapping-based filter and a CDEF between the input and output of the nonlinear mapping-based filter. The loop filter chain (3100) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3100), a deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3111), which is input to a CDEF and a nonlinear mapping-based filter (labeled SO-NLM). The CDEF is applied to the first intermediate reconstructed sample to generate a second intermediate reconstructed sample at a second intermediate node (3112). The nonlinear mapping-based filter generates a sample offset SO based on the first intermediate reconstructed sample. The sample offset SO is combined with the second intermediate reconstructed sample to generate a third intermediate reconstructed sample at a third intermediate node (3113). An LR filter is then applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3100). Further, the first intermediate reconstructed sample at the first intermediate node (3111) is used to obtain a first copy COPY0 of boundary pixels for boundary processing in the LR filter, and the second intermediate reconstructed sample at the second intermediate node (3012) is used to obtain a second copy COPY1 of boundary pixels for boundary processing in the LR filter.
[0201] In some examples, the nonlinear mapping based filter is connected in series with other filters in the loop filter chain, and there are no other filters between the input and output of the nonlinear mapping based filter. In one example, the nonlinear mapping based filter is applied after CDEF. The pixels after the deblocking filter and before CDEF are used for COPY0, and the pixels after the nonlinear mapping based filter and before the LR filter are used for COPY1.
[0202] FIG. 32 shows an example of a loop filter chain (3200) in one example. The loop filter chain (3200) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3200), a CDEF is between a first intermediate node (3211) and a second intermediate node (3212), and a nonlinear mapping-based filter (labeled SO-NLM) is applied at the second intermediate node (3212) of the loop filter chain (3200). Specifically, a deblocking filter generates a first intermediate reconstructed sample at the first intermediate node (3211), and a CDEF is applied to the first intermediate reconstructed sample to generate a second intermediate reconstructed sample at the second intermediate node (3212). The second intermediate reconstructed sample is an input of the nonlinear mapping-based filter. Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the second intermediate reconstructed sample at the second intermediate node (3212) to generate a third intermediate reconstructed sample at the third intermediate node (3213). An LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3200). Furthermore, the first intermediate reconstructed sample at the first intermediate node (3211) is used to obtain a first copy COPY0 of boundary pixels for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third node (3213) is used to obtain a second copy COPY1 of boundary pixels for boundary processing in the LR filter.
[0203] In one example, a nonlinear mapping based filter is applied before CDEF. The pixels after the deblocking filter and before the nonlinear mapping based filter are used for COPY0, and the pixels after CDEF and before the LR filter are used for COPY1.
[0204] FIG. 33 shows an example of a loop filter chain (3300) in one example. The loop filter chain (3300) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3300), a nonlinear mapping-based filter (labeled SO-NLM) is applied at a first intermediate node (3311) of the loop filter chain (3300), and a CDEF is between a second intermediate node (3312) and a third intermediate node (3313). Specifically, the deblocking filter generates a first intermediate reconstructed sample at the first intermediate node (3311). The first intermediate reconstructed sample is an input of the nonlinear mapping-based filter. Based on the input, the nonlinear mapping-based filter generates a sample offset (SO). The sample offset is combined with the first intermediate reconstructed sample at the first intermediate node (3311) to generate a second intermediate reconstructed sample at the second intermediate node (3312). The CDEF is applied to the second intermediate reconstructed sample to generate a third intermediate reconstructed sample at a third intermediate node (3313). An LR filter is applied to the third intermediate reconstructed sample to generate the output of the loop filter chain (3300). Furthermore, the first intermediate reconstructed sample at the first intermediate node (3311) is used to obtain a first copy COPY0 of boundary pixels for boundary processing in the LR filter, and the third intermediate reconstructed sample at the third intermediate node (3313) is used to obtain a second copy COPY1 of boundary pixels for boundary processing in the LR filter.
[0205] According to an aspect of the present disclosure, enabling / disabling of boundary processing of the LR filter (e.g., obtaining and using COPY0 and COPY1 in the LR filter) depends on enabling / disabling of other filtering tools, such as a nonlinear mapping-based filter, CDEF and / or FSR, in the loop filtering chain.
[0206] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the nonlinear mapping-based filter and / or the CDEF and / or the FSR. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of application of the nonlinear mapping-based filter. The flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of application of the CDEF. The flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of application of the FSR. Then, in one example, the flag En-Boundary is a logical combination of the flags En-SO-NLM, En-CDEF, and En-FSR. It should be noted that any suitable logical operator may be used, such as AND, OR, NOT, etc.
[0207] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the nonlinear mapping-based filter and / or the CDEF. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of the application of the nonlinear mapping-based filter. The flag En-CDEF is used to indicate the enabling (e.g., the flag En-CDEF has a binary 1) or disabling (e.g., the flag En-CDEF has a binary 0) of the application of the CDEF. Then, in one example, the flag En-Boundary is a logical combination of the flag En-SO-NLM and the flag En-CDEF. It should be noted that any suitable logical operator such as AND, OR, NOT, etc. may be used.
[0208] In some examples, the enabling / disabling of boundary processing depends on the enabling / disabling of the nonlinear mapping-based filter and / or the FSR. For example, the flag En-Boundary is used to indicate the enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-SO-NLM is used to indicate the enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of application of the nonlinear mapping-based filter. The flag En-FSR is used to indicate the enabling (e.g., the flag En-FSR has a binary 1) or disabling (e.g., the flag En-FSR has a binary 0) of application of the FSR. Then, in one example, the flag En-Boundary is a logical combination of the flag En-SO-NLM and the flag En-FSR. It should be noted that any suitable logical operator such as AND, OR, NOT, etc. may be used.
[0209] In some examples, enabling / disabling of boundary processing depends on the nonlinear mapping-based filter. For example, a flag En-Boundary is used to indicate enabling (e.g., the flag En-Boundary has a binary 1) or disabling (e.g., the flag En-Boundary has a binary 0) of boundary processing. A flag En-SO-NLM is used to indicate enabling (e.g., the flag En-SO-NLM has a binary 1) or disabling (e.g., the flag En-SO-NLM has a binary 0) of application of the nonlinear mapping-based filter. In that case, in one example, the flag En-Boundary can be the flag En-SO-NLM or can be the logical NOT of the flag En-SO-NLM.
[0210] In one embodiment, enabling / disabling of boundary processing depends on enabling / disabling of CDEF and / or FSR.
[0211] In some examples, the enablement / disablement of boundary processing depends on the enablement / disablement of CDEF and / or FSR. For example, the flag En-Boundary is used to indicate the enablement (e.g., the flag En-Boundary has a binary 1) or disablement (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-CDEF is used to indicate the enablement (e.g., the flag En-CDEF has a binary 1) or disablement (e.g., the flag En-CDEF has a binary 0) of application of CDEF. The flag En-FSR is used to indicate the enablement (e.g., the flag En-FSR has a binary 1) or disablement (e.g., the flag En-FSR has a binary 0) of application of FSR. Then, in one example, the flag En-Boundary is a logical combination of the flag En-CDEF and the flag En-FSR. It should be noted that any suitable logical operator such as AND, OR, NOT, etc. may be used.
[0212] In some examples, the enablement / disablement of boundary processing depends on the enablement / disablement of CDEF. For example, the flag En-Boundary is used to indicate the enablement (e.g., the flag En-Boundary has a binary 1) or disablement (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-CDEF is used to indicate the enablement (e.g., the flag En-CDEF has a binary 1) or disablement (e.g., the flag En-CDEF has a binary 0) of application of a nonlinear mapping-based filter. Then, in one example, the flag En-Boundary can be the flag En-CDEF or can be the logical NOT of the flag En-CDEF.
[0213] In one embodiment, enabling / disabling of boundary processing depends on enabling / disabling of FSR.
[0214] In some examples, the enablement / disablement of boundary processing depends on the enablement / disablement of FSR. For example, the flag En-Boundary is used to indicate the enablement (e.g., the flag En-Boundary has a binary 1) or disablement (e.g., the flag En-Boundary has a binary 0) of boundary processing. The flag En-FSR is used to indicate the enablement (e.g., the flag En-FSR has a binary 1) or disablement (e.g., the flag En-FSR has a binary 0) of application of a nonlinear mapping-based filter. Then, in one example, the flag En-Boundary can be the flag En-FSR or can be the logical NOT of the flag En-FSR.
[0215] According to some aspects of the present disclosure, pixel values are clipped before the LR filter. In some examples, the pixels after applying CDEF are clipped first and are denoted as CLIP0. Then, the pixels after applying the nonlinear mapping based filter are clipped and are denoted as CLIP1. Note that in some examples, pixel values are clipped to an appropriate range that is meaningful for the pixel value. In one example, if pixel values are represented by 8 bits, pixel values may be clipped to the range of [0,255].
[0216] It should be noted that in this description, a nonlinear mapping-based filter is used as an example to illustrate a technique for buffering boundary pixel values at different nodes along the loop filter chain and / or a technique for clipping pixel values at different nodes along the loop filter chain. This technique can be used when other coding tools, such as a cross component sample adaptive offset (CCSAO) tool, are applied after the deblocking filter and before the LR filter.
[0217] FIG. 34 shows an example of a loop filter chain (3400) including a nonlinear mapping-based filter with a CDEF between the input and output of the nonlinear mapping-based filter. The loop filter chain (3400) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3400), a nonlinear mapping-based filter (labeled SO-NLM) is applied between two intermediate nodes. A deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3411). The first intermediate reconstructed sample is an input of both the CDEF and the nonlinear mapping-based filter. Based on the first intermediate reconstructed sample, a CDEF is applied to generate a second intermediate reconstructed sample at a second intermediate node (3412). Also, based on the first intermediate reconstructed sample, the nonlinear mapping-based filter generates a sample offset (SO). The second intermediate reconstructed sample is clipped to generate a CLIP0. CLIP0 is combined with a sample offset to generate a third intermediate reconstructed sample at a third intermediate node (3413), which is clipped to generate CLIP1, which is provided to an LR filter for further filtering.
[0218] FIG. 35 shows an example of a loop filter chain (3500) including a nonlinear mapping-based filter with a CDEF between the input and output of the nonlinear mapping-based filter. The loop filter chain (3500) can be used in place of the loop filter chain (2900) in an encoding device or a decoding device. In the loop filter chain (3500), a nonlinear mapping-based filter (labeled SO-NLM) is applied between two intermediate nodes. A deblocking filter generates a first intermediate reconstructed sample at a first intermediate node (3511). The first intermediate reconstructed sample is an input of both the CDEF and the nonlinear mapping-based filter. Based on the first intermediate reconstructed sample, a CDEF is applied to generate a second intermediate reconstructed sample at a second intermediate node (3512). Based on the first intermediate reconstructed sample, the nonlinear mapping-based filter also generates a sample offset (SO). The second intermediate reconstructed sample is combined with a sample offset to generate a third intermediate reconstructed sample at a third intermediate node (3513), which is clipped to generate CLIP0, which is provided to an LR filter for further filtering.
[0219] FIG. 36 shows a flow chart outlining a process (3600) according to an embodiment of the present disclosure. The process (3600) can be used for video filtering. When the term block is used, the block can be interpreted as a prediction block, a coding unit, a luma block, a chroma block, etc. In various embodiments, the process (3600) is performed by a processing circuit of the terminal devices (310), (320), (330), and (340), a processing circuit performing the function of a video encoder (403), a processing circuit performing the function of a video decoder (410), a processing circuit performing the function of a video decoder (510), a processing circuit performing the function of a video encoder (603), etc. In some embodiments, the process (3600) is implemented in software instructions, such that the processing circuit performs the process (3600) when the processing circuit executes the software instructions. The process starts at (S3601) and proceeds to (S3610).
[0220] At (S3610), a first boundary pixel value of a first reconstructed sample at a first node along a loop filter chain is buffered, the first node being associated with a nonlinear mapping based filter applied in the loop filter chain before the loop restoration filter.
[0221] In some examples, the nonlinear mapping based filter is a cross-component sample offset (CCSO) filter. In some examples, the nonlinear mapping based filter is a local sample offset (LSO) filter.
[0222] At (S3620), a loop reconstruction filter is applied to the reconstructed sample to be filtered based on the buffered first boundary pixel values.
[0223] In some examples, the first reconstructed sample at the first node is an input of a nonlinear mapping-based filter.
[0224] In the examples of Figures 30, 31 and 33, the first node may be the first intermediate node (3011) / (3111) / (3311) in the descriptions of Figures 30, 31 and 33, respectively, and the first reconstructed sample may be the first intermediate reconstructed sample in the descriptions of Figures 30, 31 and 33, respectively.
[0225] Then, in the examples of Figures 30 and 33, a second boundary pixel value of a second reconstructed sample at a second node along the loop filter chain can be buffered, and the second reconstructed sample at the second node is generated after application of the sample offset generated by the nonlinear mapping-based filter. The loop restoration filter applies the loop restoration filter to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.
[0226] In the example of Figure 30, the sample offset generated by the nonlinear mapping-based filter is combined with the output of a constrained directional enhancement filter (e.g., the second intermediate reconstructed sample in the illustration of Figure 30) to generate a second reconstructed sample (e.g., the third intermediate reconstructed sample in the illustration of Figure 30).
[0227] In the example of Figure 33, the sample offset generated by the nonlinear mapping-based filter is combined with a first reconstructed sample (e.g., the first intermediate reconstructed sample in the description of Figure 33) to generate an intermediate reconstructed sample (e.g., the second intermediate reconstructed sample in the description of Figure 33), and then a constrained directional enhancement filter is applied to the intermediate reconstructed sample to generate a second reconstructed sample (e.g., the third intermediate reconstructed sample in the description of Figure 33).
[0228] In the example of Figure 31, a second boundary pixel value of a second reconstructed sample at a second node along the loop filter (e.g., a second intermediate reconstructed sample at a second intermediate node (3122) in the illustration of Figure 33) is buffered. The second reconstructed sample at the second node is combined with a sample offset generated by the nonlinear mapping-based filter to generate a reconstructed sample to be filtered. The loop restoration filter is applied to the reconstructed sample to be filtered based on the buffered first boundary pixel value and the buffered second boundary pixel value.
[0229] In the example of FIG. 32, the first reconstructed sample may be a third intermediate reconstructed sample at the third intermediate node (3213) in the illustration of FIG. 32, which is a result of applying a nonlinear mapping-based filter. Then, the second boundary pixel value of the second reconstructed sample (e.g., the second intermediate reconstructed sample in the illustration of FIG. 32) generated by the deblocking filter is buffered. The constrained directional enhancement filter is applied to the second reconstructed sample to generate an intermediate reconstructed sample (e.g., the second intermediate reconstructed sample in the illustration of FIG. 32). The intermediate reconstructed sample is combined with the sample offset generated by the nonlinear mapping-based filter to generate the first reconstructed sample (e.g., the third intermediate reconstructed sample in the illustration of FIG. 32). Then, the loop restoration filter can be applied based on the buffered first boundary pixel value and the buffered second boundary pixel value.
[0230] In some examples, the reconstructed samples to be filtered are clipped to an appropriate range, such as [0, 255] for 8 bits, before application of the loop reconstruction filter, as in the examples of Figures 34 and 35. In some examples (e.g., Figure 34), an intermediate reconstructed sample (e.g., the second intermediate reconstructed sample in the illustration of Figure 34) is clipped to an appropriate range, such as [0, 255] for 8 bits, before combination with the sample offset generated by the nonlinear mapping-based filter.
[0231] The process (3600) proceeds to (S3699) and ends.
[0232] In some examples, the nonlinear mapping based filter is a cross-component sample offset (CCSO) filter. In some examples, the nonlinear mapping based filter is a local sample offset (LSO) filter.
[0233] The process 3600 may be adapted as appropriate. Steps in the process 3600 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0234] The embodiments of the present disclosure may be used separately or in combination in any order. Furthermore, each method (or embodiment), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.
[0235] The techniques described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 37 illustrates a computer system (3700) suitable for implementing certain embodiments of the disclosed subject matter.
[0236] Computer software can be coded using any suitable machine code or computer language and can be subject to assembly, compilation, linking, or similar mechanisms to produce code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly, or through interpretation, microcode execution, etc.
[0237] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0238] 37 for computer system (3700) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Neither the arrangement of components should be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (3700).
[0239] The computer system (3700) may include certain human interface input devices that may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0240] The input human interface devices may include one or more (only one of each is shown) of a keyboard (3701), a mouse (3702), a trackpad (3703), a touch screen (3710), a data glove (not shown), a joystick (3705), a microphone (3706), a scanner (3707), and a camera (3708).
[0241] The computer system (3700) may also include some type of human interface output device. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (3710), data gloves (not shown), or joystick (3705); although there may be haptic feedback devices that do not act as input devices), audio output devices (e.g., speakers (3709), headphones (not shown)), visual output devices (e.g., screens (3710), including CRT screens, LCD screens, plasma screens, OLED screens; each may or may not have touch screen input capability, each may or may not have haptic feedback capability, some of which may output two-dimensional visual output or higher than three-dimensional output through such means as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0242] The computer system (3700) may also include human accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (3720) along with CD / DVD or similar media (3721), thumb drives (3722), removable hard drives or solid state drives (3723), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD based devices (not shown) such as security dongles, etc.
[0243] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0244] The computer system (3700) may also include an interface (3754) to one or more communication networks (3755). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, in-vehicle, and industrial, real-time, delay tolerant, and the like. Examples of networks include Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like, TV wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, in-vehicle and industrial including CANBus, and the like. Some networks require an external network interface adapter that is typically attached to some general-purpose data port or peripheral bus (3749) (e.g., a USB port of the computer system (3700)). Others are typically integrated into the core of the computer system (3700) by attachment to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3700) can communicate with other entities. Such communications may be unidirectional, receive only (e.g., broadcast television), unidirectional transmit only (e.g., CANbus to certain CANbus devices), or bidirectional, for example, to other computer systems using local or wide area digital networks. Certain protocols and protocol stacks may be used with each of these networks and network interfaces as described above.
[0245] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core (3740) of the computer system (3700).
[0246] The core (3740) may include one or more central processing units (CPUs) (3741), graphics processing units (GPUs) (3742), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (3743), hardware accelerators for certain tasks (3744), graphics adapters (3750), etc. These devices may be connected through a system bus (3748), along with read only memory (ROM) (3745), random access memory (3746), internal mass storage devices (3747), such as internal non-user accessible hard drives, solid state drives (SSDs), etc. In some computer systems, the system bus (3748) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (3748) or through a peripheral bus (3749). In one example, a display (3710) may be connected to the graphics adapter (3750). Architectures for peripheral buses include PCI, USB, and the like.
[0247] The CPU (3741), GPU (3742), FPGA (3743), and accelerator (3744) may execute certain instructions that may combine to constitute the above-mentioned computer code. The computer code may be stored in a ROM (3745) or a RAM (3746). Temporary data may also be stored in the RAM (3746), while persistent data may be stored, for example, in an internal mass storage device (3747). Rapid storage and retrieval from any of the memory devices may be enabled through the use of a cache memory that may be closely associated with one or more of the CPU (3741), GPU (3742), mass storage device (3747), ROM (3745), RAM (3746), etc.
[0248] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.
[0249] By way of example and not limitation, a computer system having the architecture (3700), and in particular the core (3740), can provide functionality as a result of the processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as introduced above as well as media associated with some type of storage of the core (3740) of a non-transitory nature, such as a mass storage device (3747) internal to the core or a ROM (3745). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (3740). The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core (3740) and in particular the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain specific portions thereof described herein, including defining data structures stored in RAM (3746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (3744)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Reference to software includes logic, and vice versa, as appropriate. Reference to a computer-readable medium may include circuitry (e.g., an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0250] Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding MPM: most probable mode WAIP: Wide-Angle Intra Prediction SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SDR: standard dynamic range SNR: Signal Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub-Partition SPS: Sequence Parameter Setting HDR: high dynamic range SDR: standard dynamic range JVET: Joint Video Exploration Team MPM: most probable mode WAIP: Wide-Angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Transform Unit CTU: Coding Tree Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub-Partition SPS: Sequence Parameter Setting PPS: Picture Parameter Set APS: Adaptation Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross-Component Sample Offset CCSAO: Cross-Component Sample Adaptive Offset LSO: Local Sample Offset LR: Loop Restoration Filter FSR: Frame Super-Resolution AV1: AOMedia Video 1 AV2: AOMedia Video 2
[0251] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents, which fall within the scope of this disclosure. Thus, it will be appreciated that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for filtering in video encoding or decoding, comprising the steps of: buffering, by a processor, first boundary pixel values of a first subset of reconstructed samples at a first node in a loop filter chain, where a first filter and a second filter are applied to the first boundary pixel values, the first boundary pixel values being values of pixels at a frame boundary; buffering, by the processor, second boundary pixel values of a subset of a second reconstructed sample at a second node in the loop filter chain, the second reconstructed sample at the second node being generated after applying at least one of the first filter or the second filter to the first boundary pixel values, the second boundary pixel values being values of pixels at the frame boundary; applying, by the processor, a loop reconstruction filter to reconstructed samples to be filtered based on the buffered first boundary pixel values of the subset of the first reconstructed samples and the buffered second boundary pixel values of the subset of the second reconstructed samples; The method includes:
2. The method described in claim 1, wherein the first filter is at least one of a cross-component sample offset (CCSO) filter and a local sample offset (LSO) filter.
3. A method as described in claim 1 or 2, wherein the first reconstructed sample at the first node is an input to the first filter.
4. The method described in claim 3, wherein the second reconstructed sample at the second node is generated after application of a sample offset generated by the first filter.
5. The method of claim 4, further comprising a step of the processor combining the sample offset generated by the first filter with the output of the second filter to generate the second reconstructed sample.
6. The method of claim 1, further comprising: combining, by the processor, the sample offset produced by the first filter and the first reconstructed sample to generate an intermediate reconstructed sample; applying, by the processor, the second filter to the intermediate reconstructed samples to generate the second reconstructed samples; The method of claim 4 further comprising:
7. The method of claim 3, further comprising a step of the processor combining the second reconstructed sample at the second node with a sample offset generated by the first filter to generate the reconstructed sample to be filtered.
8. The method according to claim 7, wherein the first reconstructed samples are generated by a deblocking filter; applying, by the processor, the second filter to the first reconstructed samples to generate intermediate reconstructed samples; combining, by the processor, the intermediate reconstructed samples with sample offsets produced by the first filter to produce the second reconstructed samples; The method of claim 1 further comprising:
9. A method according to any one of claims 1 to 8, further comprising the step of clipping, by the processor, the reconstructed samples to be filtered prior to the application of the loop reconstruction filter.
10. The method of claim 9, further comprising the step of clipping, by the processor, intermediate reconstruction samples prior to combination with the sample offset generated by the first filter.
11. An apparatus for video encoding or decoding including a processing circuit, the apparatus comprising: Apparatus, wherein the processing circuitry is configured to perform a method according to any one of claims 1 to 10.
12. A program that causes a processor to execute a method according to any one of claims 1 to 10.