In-loop filter having a plurality of regions
Region-based post-filtering techniques in video encoding and decoding improve compression efficiency by reducing coding artifacts and latency, leveraging existing HEVC and JEM codecs for adaptive filtering across defined regions.
Patent Information
- Application Number
- JP2020570112
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-08-14
- Filing Date
- 2019-07-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-07-10
AI Technical Summary
Existing video encoding and decoding technologies face challenges in effectively reducing coding artifacts through in-loop filtering, particularly in handling block artifacts and spatial-temporal redundancy, which affects compression efficiency.
Implement region-based post-filtering techniques, such as sample adaptive offset (SAO) and adaptive loop filter (ALF), by defining filter regions within a slice or picture and sharing filter parameters across these regions, allowing for adaptive filtering and reduced overhead in parameter signaling.
Enhances video compression efficiency by reducing coding artifacts and improving latency through region-based filtering, leveraging existing HEVC and JEM codec designs while minimizing implementation costs.
Smart Images

Figure 0007702786000002 
Figure 0007702786000003 
Figure 0007702786000004
Abstract
Description
Technical Field
[0001] At least one of the embodiments generally relates to a method or apparatus for video encoding or decoding.
Background Art
[0002] To achieve high compression efficiency, video and image coding schemes typically employ prediction including motion vector prediction and transform to exploit spatial and temporal redundancy in video content. Generally, intra or inter prediction is used to exploit the correlation of intra frames or inter frames, and then the difference between the original image and the predicted image, often called the prediction error or prediction residue, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy coding, quantization, transform, and prediction.
[0003] The in-loop filter enables post-filtering of the reconstructed picture to reduce coding artifacts. For example, by sample adaptive offset (SAO) filtering, an offset can be added to some categories (or classes) of the reconstructed samples to reduce coding artifacts. Another example is an adaptive loop filter (ALF) that implements Wiener linear post-filtering of the reconstructed samples. Another example is a deblocking filter (DBF) that uses smoothing of block boundaries to reduce block artifacts.
Summary of the Invention
[0004] The disadvantages and drawbacks of the prior art are addressed by the general aspects described herein, which are directed to the direction of block shape adaptive intra prediction in encoding and decoding.
[0005] According to a first aspect, a method is provided. The method includes determining a region of a picture that uses a common set of filter parameters for filtering at least one reconstructed block of the picture; obtaining a plurality of sets of filter parameters; filtering the region of the picture that includes at least one reconstructed block using the common set of filter parameters for the blocks within the region; and encoding information in a bitstream that includes a syntax indicating the set of filter parameters used for filtering the region and an encoded version of the region.
[0006] According to another aspect, a second method is provided. The method includes decoding syntax from a bitstream that indicates a plurality of sets of filter parameters used for filtering a region of a picture; determining the region of the picture from the bitstream using a common set of filter parameters for filtering at least one reconstructed block of the picture; filtering the at least one reconstructed block using the set of filter parameters associated with the region that includes the at least one reconstructed block; and decoding the filtered reconstructed block of the picture.
[0007] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to encode blocks of video or decode a bitstream by performing any of the aforementioned methods.
[0008] According to another general aspect of at least one embodiment, a device is provided that includes an apparatus according to any of the embodiments for decoding, and at least one of: (i) an antenna configured to receive a signal that includes a video block; (ii) a band limiter configured to limit the received signal to a frequency band that includes a video block; or (iii) a display configured to display an output representing a video block.
[0009] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that includes data content generated according to any of the described encoding embodiments or variations.
[0010] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated according to any of the described encoding embodiments or variations.
[0011] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.
[0012] According to another general aspect of at least one embodiment, a computer program product is provided that includes instructions that, when executed by a computer, cause the computer to perform any of the described decoding embodiments or variations.
[0013] These and other aspects, features, and advantages of the general aspects will become apparent from the following detailed description of the exemplary embodiments, which should be read in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Embodiments for Carrying Out the Invention
[0015] The general aspects described herein are in the field of video compression. The general aspects relate to in-loop filtering such as using sample adaptive offset (SAO) using the "advanced merge" (also known as the SAO palette) technique as described in the following EP applications by the same applicant, the teachings of which are specifically incorporated herein by reference, that is, (1) EP application No. 17305626.8 (agent docket number PF170034) entitled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING", (2) EP application No. 18305736.3 (agent docket number PF180072) entitled "ADVANCED MERGE PARALLELIZABLE SAO", (3) EP application No. 17305033.7 (agent docket number PF160213) entitled "A METHOD AND A DEVICE FOR IMAGE ENCODING AND DECODING", and (4) EP application No. 17305627.6 (agent docket number PF170089) entitled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING".
[0016] With an in-loop filter, the reconstructed picture can be post-filtered to reduce coding artifacts (see blocks 265 and 465 in the conventional decoder and encoder methods of FIG. 1). For example, with SAO, offsets can be added to some categories (or classes) of the reconstructed samples to reduce coding artifacts. As another example, there is an adaptive loop filter (ALF) that implements Wiener linear post-filtering of the reconstructed samples. As another example, there is a deblocking filter (DBF) that uses smoothing at block boundaries to reduce block artifacts.
[0017] Generally, the in-loop filter (k) process (on the decoder side) is composed of the following steps. 1) Syntax analysis of the C(k) filter parameter set P(c); c = 0 to C(k) 2) Classification of the reconstructed samples into C(k) classes 3) Filter the reconstructed samples belonging to class c using the parameter P(c)
[0018] Generally, the in-loop filter (k) process (on the encoder side) is composed of the following steps. 1) Classification of the reconstructed samples into C(k) classes 2) Derive the C(k) filter parameter set P(c); c = 0 to C(k) 3) Filter the reconstructed samples belonging to class c using the parameter P(c)
[0019] The object of the embodiments described herein is to improve the performance of the in-loop filter by using region-based post-filtering.
[0020] When HEVC (High Efficiency Video Coding) is effective, a Coding Tree Unit (CTU) can be coded using three SAO modes (SaoTypeIdx), namely, OFF (invalid), Edge Offset (EO), or Band Offset (BO). In the case of EO or BO, one set of parameters (Y, U, V) is coded for each channel and, in some cases, shared with adjacent CTUs (see SAO MERGE flag). The SAO mode is the same for the Cb component and the Cr component.
[0021] In the case of EO, as shown in Figure 2, each reconstructed sample is classified into a category (sao_eo_class) with NC = 5 according to the local gradient. (NC - 1) offset values are coded one by one for each category (one category has a zero offset).
[0022] In the case of BO, the range of pixel values (e.g., 0 to 255, 8 bits) is evenly divided into 32 bands, and the sample values belonging to (NC - 1) = 4 consecutive bands are corrected by adding an offset off(n). Figure 3 shows an example of four consecutive bands. (NC - 1) offset values are coded one by one for each (NC - 1) band (the remaining bands have a zero offset).
[0023] In the case of EO or BO, the offset is sometimes not coded and is copied from the adjacent upper or left CTU (merge mode). Figure 4 shows how SAO is processed across the picture (left) and the SAO filtering process itself for each CTU (right).
[0024] EP application No. 17305627.6 titled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING" proposes collecting all samples of the reconstructed picture using or sharing the same SAO parameters to derive an optimal set of SAO parameters.
[0025] ALF filter In JEM software, each 2x2 block is classified into one of 25 classes based on its directionality and its function using local gradients. Next, the ALF filter coefficients are derived for each class over the whole picture.
[0026] For the luma samples of each CTU (filter block), the encoder determines whether ALF is applied and an appropriate signalling flag is included in the slice header. For chroma samples, the decision to apply the filter is made at the picture level rather than at the CTU level.
[0027] The ALF filter parameters can be signalled within the first CTU or within the slice header. Up to 25 sets of luma filter coefficients can be signalled. To reduce overhead bits, filter coefficients of different classifications can be merged. Also, the ALF coefficients of reference pictures are saved and made available for reuse as the ALF coefficients of the current picture (ALF temporal prediction).
[0028] To assist ALF temporal prediction, a candidate list of ALF filter sets is maintained. When starting the decoding of a new sequence, the candidate list is empty. After decoding one picture, the corresponding set of filters can be added to the candidate list. Temporal prediction of ALF coefficients improves the coding efficiency of coded inter-frames. To improve the coding efficiency in case temporal prediction is not available (intra-frame), a set of 16 fixed filters is also assigned to each class.
[0029] Improved merge SAP and other features Feature 1: In EP application No. 17305626.8 titled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING", all SAO parameters (list of SAO parameter candidates) are first encoded (e.g., in the slice header or with the first CTU), and then, as shown in FIG. 6, SAO blocks including a (merge / candidate) index that refers to this list of already defined and encoded SAO parameters (new candidates) are encoded.
[0030] The number of SAO candidates (nb_sao_cand) and the list of SAO parameters are encoded in the same order as the order of use. In the encoder, the list of SAO parameter candidates is sorted after encoding each candidate index, and the last used parameter is placed at the top of the list. More precisely, the list of candidates is sorted such that the candidates used in the spatially closest vicinity are ordered first. This can be done by constructing a map of the last used candidates.
[0031] The OFF (where all offsets of all components are zero) candidate is implicitly placed in the list but is not explicitly coded at a position not too far from the top (e.g., position ≦ 2).
[0032] Feature 2: In EP application No. 18305736.3 titled "ADVANCED MERGE PARALLELIZABLE SAO", the principle of EP application No. 17305626.8 titled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING" is extended, and as a result, the area where the current SAO block can inherit from other SAO block parameters is - made to exist within a causal region parallel to the wavefront, - and / or such that the number of candidates in the sorted list of SAO candidates is less than a predefined value (see "list_reordered_size" in PF180072). - and / or constrained such that the number of candidates is limited by the maximum distance (dist_max) to the current SAO block.
[0033] Feature 3: In EP application No. 17305033.7 entitled "A METHOD AND A DEVICE FOR IMAGE ENCODING AND DECODING" and EP application No. 17305626.8 entitled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING", the size of the SAO block (i.e., the size to which the SAO parameters apply) is encoded in the slice header. In EP application No. 17305626.8 entitled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING", the width and / or height of the SAO block is a multiple N of the CTU size, where, for example, N = 1, 2, or 1 / 2 (see Figure 8).
[0034] The articles titled "Description of SDR, HDR and 360° video coding technology proposal by Qualcomm and Technicolor - low and high complexity versions", 10th JVET Meeting, San Diego, California, USA, April 2018, JVET-J0021, and "CE2:Tests on SAO design from JVET-J0021(CE2.3.2)", 11th JVET Meeting: Ljubljana, SI, 10 - 18 July 2018, JVET-K0324, by A.Gadde, D.Rusanovskyy, M.Karczewicz propose encoding two flags (merge_left_flag and merge_above_flag) to indicate whether the current SAO block inherits from the left or above, or another adjacent part (see Figure 9). At the syntax analysis stage, it is checked whether the left and upper SAO parameters are the same, and in that case, the upper flag is not syntax-analyzed (assumed to be an error). In other cases, the current SAO block is marked as "valid", and the SAO parameters are encoded / decoded without inheritance.
[0035] All information is decoded at the start of the slice or picture. After syntax-analyzing the merge flags of all SAO blocks, the SAO parameters of the SAO blocks marked as "valid" are decoded.
[0036] In the article titled "CE2-3.3 SAO_Palette results and discussion", by P. Bordes and F. Racape, presented at the 11th JVET Meeting: Ljubljana, SI, 10-18, July 2018, JVET-K0192, it is reported that the combination of the techniques just described provides BD-rate gains of 0.17% for AI, 0.38% for RA, and 0.52% for LDB for luminance, respectively, using the common test conditions (CTC) with the JVET reference software (VTM1.0).
[0037] In JVET-K0324, using the same conditions, BD-rate gains of 0.11% for AI, 0.30% for RA, and 0.43% for LDB are reported.
[0038] For at least some of the embodiments described herein, the purpose of the general aspects described is as follows. - Combine two approaches to leverage the coding gains of both techniques. - Avoid the problem of parsing all filter parameters at the start of a slice / picture, which results in a one-frame latency. - Optionally, share a group (region) of reconstructed samples used in some (two or more) intra-loop filters for post-filter parameter derivation.
[0039] In at least one embodiment, general aspects of in-loop post-filters such as SAO or ALF in EP application No. 17305626.8 entitled "A METHOD AND A DEVICE FOR PICTURE ENCODING AND DECODING" (features 1, 2, 3: first signaling SAO or ALF parameter sets and referring to them using indices to enable / disable the filter for the current block, etc., see reference) are separately extended to multiple regions within a slice or picture. These features can be combined with the ability to adapt the block size of the post-filter for each slice or picture. Several different post-filters can share the same region. As another concept of the encoder, for example, calculating post-filter parameters for each region.
[0040] In the first embodiment, a filter region associated with each filter block is defined. This filter block is the same size as the set of filter parameters. Then, each filter block belongs to one filter region. The filter region is a sub-division of the slice or picture (a subset of filter blocks). The filter region can be a rectangle (e.g., tile) as shown in FIG. 10, or one row / column of filter blocks. The filter region can also correspond to the entire slice or picture. The size or shape of the filter region (e.g., coded as a map) is typically coded within the slice, picture, or sequence header.
[0041] In the case of the SAO filter, the current filter block belonging to one region can inherit the SAO parameters from the candidate SAO parameters corresponding to the SAO blocks within the same region.
[0042] In the second embodiment, in the case of the SAO filter, the syntax is changed as follows. - The sao_palette_index is first parsed for each SAO block. - If the sao_palette_index is equal to the value idx_new (e.g., idx_new = 3), the SAO block is marked as NEW (new_flag = true), which means that this SAO parameter is being used within the current filter region for the first time.
[0043] In the parsing stage, when sao_palette_index = idx_new, the SAO parameter is parsed after the parsing of sao_palette_index and becomes available for use by other SAO blocks in the region during the filtering stage. In the filtering stage, the SAO parameter is added to the list of SAO parameters available for merging (inheritance) with other candidates in the current SAO region.
[0044] In a variation of this second embodiment, the value of idx_new can vary for each slice or region. The value can be coded within the slice header, or along with the first SAO block, or derived from other parameters such as being a function of the quantization parameter (QP).
[0045] In another variation of the second embodiment, for the ALF filter, the set of ALF filter parameters can be signaled within the first filter block of the region, or within the region header (e.g., if the region is a tile as defined within HEVC) or within the slice / picture header. The set of ALF filter parameters remains the same for all filter blocks in this region.
[0046] In another variation, if the filter supports temporal prediction (e.g., ALF temporal prediction), the list of sets of filter parameters is maintained for each region, and the filter blocks in the current region can use the filter parameters corresponding to the region at the same location in the reference picture.
[0047] In another variant of the second embodiment, the number of SAO blocks marked as NEW within a region is coded at the beginning of the region (e.g., together with the first SAO block of the region).
[0048] In the third embodiment, the value of sao_palette_index is coded with n1 + 1 + n2 bits, and new_flag is the n1-th bit (idx_new_bit) as shown in FIG. 11, where n1 and n2 are either redefined parameters or parameters that adaptively change according to the situation. In a variant of this embodiment, n1 ≤ 2 and the two n1 bits are merge_left_flags and merge_above_flag. Advantageously, when the upper and left parameters are the same, merge_above_flag is not coded as in JVET-J0021 and JVET-K0324.
[0049] In the fourth embodiment, for the SAO blocks in the first column of the SAO region (or the SAO blocks in the first row of the SAO region), the list of SAO parameters available for merging also includes the SAO block parameters on the left (or above) outside the SAO region.
[0050] Advantageously, when decoding the first SAO block of a region, the left column outside the current region and the row above the SAO block parameters are added to the list of the current SAO region.
[0051] In the fifth embodiment, the filter block size can vary from region to region. The filter block size is coded for each region based on a predefined table or derivation rule, or can be inferred from other coding parameters such as the quantization parameter (QP).
[0052] For example, the basic block size is defined in the SPS, PPS, or slice header (e.g., 128x128), and a QP table indicating the magnification factor to be applied to the width and height of the filter block is used. An example of such a table is given below.
Table 1
[0053] In the sixth embodiment, the filter block parameters can be parsed in classical raster scan order within a slice. Next, the filtering stage is executed for each region (see the left side of FIG. 13). The SAO filtering stage groups the SAO candidate lists, rearranges the association between the SAO parameters and each SAO block, classifies the samples, and applies SAO offsets to modify the reconstructed samples.
[0054] The ALF filtering stage groups the classification of samples and filters the reconstructed samples. Alternatively, the filter block parameters can be parsed for each region (usually using a raster scan of the filter blocks within the filter region), and then the filtering stage is executed for each region (see the right side of FIG. 13).
[0055] In the seventh embodiment, several different intra-loop filters k (k = 0 to N) can share the same filter region, and as a result, the parsing and filtering processes of several filters are performed on a region basis. As shown in FIG. 14, the order of parsing / classifying / filtering can be alternated between the filters within one region.
[0056] In the eighth embodiment, advantageously, several different post-filters can share the same classification process, and as a result, c(k1)=c(k2), where k1 is different from k2. This means that the sample set belonging to one class of filter k1 is the same as the sample set belonging to one class of filter k2. In that case, the classification of this class is performed once. Other variations of different filters sharing the same classification process can be used.
[0057] The proposed techniques can improve the overall video compression process. These techniques are lightweight in terms of memory access. These techniques improve post-filtering by grouping various filtering stages regionally and parallelizing post-filtering. This is achieved through the improvement of in-loop filtering.
[0058] The proposed modifications to the state-of-the-art SAO filter (existing standardized HEVC) or ALF reuse most of the conventional SAO or ALF block-level logic / operations. As a result, the existing designs of HEVC or JEM codecs using post-filters can be reused to the maximum extent, thereby reducing the implementation cost of the proposed techniques.
[0059] This document describes various aspects including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and are often described in a way that may seem restrictive, at least to show individual characteristics. However, this is for the purpose of clarifying the description and does not limit the application or scope of these aspects. In fact, all of the various aspects can be combined, exchanged, and further aspects can be provided. Furthermore, these aspects can also be combined with and exchanged with the aspects described in previous applications.
[0060] The aspects described and contemplated in this document can be implemented in many different forms. FIGS. 15, 16, and 17 provide some embodiments below, but other embodiments are envisioned, and the descriptions of FIGS. 15, 16, and 17 do not limit the frontage of those implementation forms. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the described methods.
[0061] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Although not necessarily, usually, the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.
[0062] Various methods are described herein, and each of those methods includes one or more steps or operations for achieving the described method. The order and / or use of specific steps and / or operations may be changed or combined, provided that a specific order of steps or operations is not required for the correct operation of the method.
[0063] Using the various methods and other aspects described in this document, as shown in FIGS. 15 and 16, modules such as the video encoder 100 and decoder 200 can be modified in terms of intra prediction, entropy coding, and / or decoding modules (160, 360, 145, 330). Further, this aspect is not limited to VVC or HEVC, and can be applied to other standards and recommendations, as well as extended versions of any such standards and recommendations (including VVC and HEVC), whether existing or to be developed in the future. Unless specifically indicated or technically excluded, the aspects described in this document can be used individually or in combination.
[0064] In this document, various numerical values such as {{1,0},{3,1},{1,1}} are used. The specific values are for illustrative purposes only, and the described aspects are not limited to these specific values.
[0065] FIG. 15 shows the encoder 100. Although variations of this encoder 100 are contemplated, the encoder 100 will be described below without explaining all the expected variations for clarity.
[0066] Before being encoded, the video sequence may go through pre-encoding processing (101), such as applying a color conversion (e.g., from RGB 4:4:4 to YCbCr 4:2:0) to the input color picture, or performing remapping of the input picture components to obtain a more resilient signal distribution for compression (e.g., using histogram equalization of one of the color components). Metadata can be attached to the bitstream in association with the pre-processing.
[0067] In encoder 100, as described below, a picture is encoded by encoder elements. The picture to be encoded is divided (102) and processed, for example, in units of CUs. Each unit is encoded, for example, using either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction is performed (160). In the inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) whether to use either the intra mode or the inter mode for encoding the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (110) the block predicted from the original image block.
[0068] Next, the prediction residual is transformed (125) and quantized (130). In addition to the quantized transform coefficients, motion vectors and other syntax elements are entropy-coded (145) to output a bitstream. The encoder may skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can also bypass both the transform and quantization, i.e., the residual is directly coded without applying the transform or quantization process.
[0069] The encoder decodes the encoded block and provides a reference for further prediction. The quantized transform coefficients are dequantized (140), inverse-transformed (150), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (155) to reconstruct the image block. The in-loop filter (165) is applied to the reconstructed picture, for example, to perform deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).
[0070] FIG. 16 shows a block diagram of the video decoder 200. In decoder 200, as described below, the bitstream is decoded by decoder elements. The video decoder 200 generally executes a decoding path that is opposite to the encoding path as described in FIG. 15. The encoder 100 also generally performs video decoding as part of the encoding of video data.
[0071] In particular, the input to the decoder includes a video bitstream that can be generated by the video encoder 100. First, the bitstream is entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoded information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (235). The transform coefficients are dequantized (240), inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (255) to reconstruct the image block. The predicted block can be obtained from intra prediction (260) or motion compensation prediction (i.e., inter prediction) (275) (270). The in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (280).
[0072] The decoded picture can further undergo post-decoding processing (285), such as inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (101). In post-decoding processing, metadata derived in the pre-encoding process and signaled in the bitstream can be used.
[0073] FIG. 17 shows a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device that includes various components described below and is configured to execute one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 1000 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.
[0074] System 1000 includes, for example, at least one processor 1010 configured to execute loaded instructions when implementing the various aspects described in this document. The processor 1010 may include an embedded memory, an input / output interface, and various other circuits, as is well known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0075] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module(s) that may be included in a device that performs an encoding function and / or a decoding function. As is well known, a device may include one or both of an encoding and a decoding module. Further, the encoder / decoder module 1030 may be implemented as a separate element of System 1000 or may be incorporated within the processor 1010 as a combination of hardware and software, as is well known to those skilled in the art.
[0076] The program code loaded into the processor 1010 or the encoder / decoder 1030 to execute the various aspects described in this document is stored in the storage device 1040 and can subsequently be loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of various items during the execution of the processes described in this document. Such stored items can include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstreams, matrices, variables, or equations, expressions, operations, and intermediate or final results from the processing of operational logic.
[0077] In some embodiments, the internal memory of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040 and can be, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations such as MPEG-2, HEVC, or VVC (Versatile Video Coding).
[0078] Inputs to the elements of system 1000 can be provided through various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted wirelessly, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI (registered trademark) input terminal.
[0079] In various embodiments, the input device of block 1130 has respective input processing elements as are well known in the art. For example, the RF section (i) selects a desired frequency (also called selecting a signal or bandlimiting a signal to a certain frequency band), (ii) downconverts the selected signal, (iii) bandlimits again to a narrower frequency band to select a signal frequency band, which may be called a channel in a particular embodiment, for example, (iv) demodulates the downconverted and bandlimited signal, (v) performs error correction, and (vi) may be associated with elements necessary to demultiplex and select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, signal selector, bandlimiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements perform frequency selection by receiving, filtering, downconverting, and filtering again to a desired frequency band an RF signal transmitted via a wired (e.g., cable) medium. In various embodiments, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements that perform similar or different functions are added. Adding elements may include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0080] Furthermore, the USB and / or HDMI (registered trademark) terminals may each include an interface processor for connecting the system 1000 to other electronic devices over a USB and / or HDMI (registered trademark) connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, as needed, within a separate input processing IC or within the processor 1010. Similarly, aspects of USB or HDMI (registered trademark) interface processing may be implemented, as needed, within a separate interface IC or within the processor 1010. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements including, for example, the processor 1010 and an encoder / decoder 1030 that operates in combination with memory and storage elements to process the data stream as needed for display on an output device.
[0081] The various elements of the system 1000 may be provided within an integrated housing, and within the integrated housing, the various elements may be interconnected and use an internal bus, such as an I2C bus, wiring, and a printed circuit board, well-known in the art, including a suitable connection configuration 1140, to transmit data between them.
[0082] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0083] In various embodiments, data is streamed to system 1000 using a wireless network such as IEEE 802.11. The wireless signals of these embodiments are received via, for example, communication channel 1060 and communication interface 1050 that are adapted for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router, which provides access to an external network including the Internet that enables streaming of applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that distributes data via an HDMI (registered trademark) connection of input block 1130. Still other embodiments provide streamed data to system 1000 using an RF connection of input block 1130.
[0084] System 1000 can provide output signals to various output devices, including display 1100, speaker 1110, and other peripheral devices 1120. In various embodiments, other peripheral devices 1120 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 1000. In various embodiments, signal notifications such as an AV link, CEC, or other communication protocols that enable device-to-device control are used to transmit control signals between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120, regardless of the presence or absence of user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 can be integrated within a single unit together with other components of system 1000, such as within an electronic device, e.g., a television. In various embodiments, display interface 1070 includes a display driver, e.g., a timing controller (T Con) chip.
[0085] Alternatively, display 1100 and speaker 1110 can be separate from one or more of the other components, e.g., if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, output signals can be provided via a dedicated output connection that includes, e.g., an HDMI (registered trademark) port, a USB port, or a COMP output section.
[0086] Embodiments may be implemented by computer software executed by a processor 1010, or by hardware, or by a combination of hardware and software. By way of non-limiting example, some embodiments may be implemented by one or more integrated circuits. Memory 1020 may be of any type suitable for the technical environment, and may be implemented using any suitable data storage technology, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. Processor 1010 may be of any type suitable for the technical environment, and may include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0087] Various embodiments involve decoding. As used in this application, "decoding" includes, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process typically includes one or more of the processes performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such a process also includes, or alternatively, a process performed by a decoder in various implementation forms described in this application, such as extracting weight indexes used for various intra prediction reference arrays.
[0088] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the context of a particular description and is considered to be fully understood by those skilled in the art.
[0089] Various implementations involve encoding. Similar to the above considerations regarding "decoding", "encoding" as used in this application can include all or part of the processes performed on an input video sequence, for example, to generate an encoded bitstream. In various embodiments, such processes typically include one or more processes performed by an encoder, such as partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes can also, or alternatively, include processes performed by the encoders of the various implementations described in this application, such as weighting of the intra prediction reference array.
[0090] As a further example, in one embodiment, "encoding" refers to only entropy encoding, in another embodiment, "encoding" refers to only differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will become clear based on the context of the particular description and is considered to be fully understood by those skilled in the art.
[0091] Note that the syntactic elements used in this specification are for illustrative purposes. Thus, they do not preclude the use of other syntactic element names.
[0092] It should be understood that when a figure is presented as a flowchart, it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0093] Various embodiments refer to rate distortion calculation or rate distortion optimization. In particular, during the encoding process, often considering the computational complexity constraints, a balance or trade-off between rate and distortion is usually considered. Rate distortion optimization is typically formulated to minimize a rate distortion function that is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, an approach is based on an extensive test of all encoding options including all modes or coded parameter values to be considered, and can fully evaluate the encoding cost and the associated distortion of the reconstructed signal after encoding and decoding. In particular, a faster approach can also be used to reduce the decoding complexity by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. These two approaches can also be used in combination, such as by using the approximate distortion only for some of the possible encoding options and the complete distortion for other encoding options. In other approaches, only a subset of the possible encoding options is evaluated. More generally, many approaches employ any of various techniques for performing the optimization, but the optimization is not necessarily a full evaluation of both the encoding cost and the associated distortion.
[0094] The implementation forms and modes described in this specification can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if considered only in the context of a single implementation form (e.g., considered only as a method), the implementation of the considered features may also be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. Those methods can be implemented, for example, within a processor, which refers to generally processing devices including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as a computer, a mobile phone, a portable / personal digital assistant (``PDA''), and other devices that facilitate the transmission of information between end users.
[0095] References to ``one embodiment'' or ``an embodiment'', or ``one implementation form'' or ``an implementation form'', and other variations thereof, mean that the specific features, structures, characteristics, etc. described in relation to the embodiment are included in at least one embodiment. Thus, the phrases ``in one embodiment'' or ``in an embodiment'' or ``in one implementation form'' or ``in an implementation form'', and the appearance of any other variations, which are seen at various places throughout this document, do not necessarily all refer to the same embodiment.
[0096] Furthermore, this document may refer to ``determining'' various parts of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0097] Furthermore, this document may refer to ``accessing'' various parts of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, judging information, predicting information, or evaluating information.
[0098] Furthermore, this document may refer to "receiving" various parts of the information. Receiving is intended to be a broad term, similar to "accessing". Receiving the information may include, for example, one or more of accessing the information or retrieving the information (e.g., from memory). Furthermore, "receiving" typically includes, in some way, operations such as storing the information, processing the information, transmitting the information, moving the information, copying the information, deleting the information, calculating the information, judging the information, predicting the information, or estimating the information.
[0099] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of any of the following, namely " / ", "and / or", and "at least one of", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A, B, and C). As will be apparent to those skilled in the art, this can be extended for as many items as are listed.
[0100] Also, as used herein, the word "signal" (or "signaling") refers to, among other things, instructing something to the corresponding decoder. For example, in certain embodiments, the encoder signals a particular one of a plurality of weights used in the intra prediction reference array. In this way, in an embodiment, the same parameter is used on both the encoder side and the decoder side. Thus, for example, the encoder can send (explicitly signal) a particular parameter to the decoder, and as a result, the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without performing a transmission (implicit signaling) to simply enable the decoder to recognize and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the word "signal", but the word "signaling" can also be used as a noun herein.
[0101] As will be apparent to those skilled in the art, implementations can generate a variety of signals that are formatted, for example, to carry information that can be stored or transmitted. The information can include, for example, instructions for executing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted via a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.
[0102] Implementations can include, alone or in combination, one or more of the following features or entities across various different claim categories and types. · Modifying the in-loop filter process applied to a decoder and / or encoder. · Enabling some sets of filtering parameters and regions within a decoder and / or encoder. · Inserting signaling syntax elements that enable a decoder to identify the regions to which a set of filter parameters is applied for in-loop filtering. · Selecting a set of filter parameters to apply in a decoder based on these syntax elements. · Applying in-loop filtering such as adaptive loop filtering and sample adaptive offset filtering in a decoder. · Executing in-loop filtering in an encoder according to any of the described embodiments. ·A bitstream or signal includes one or more of the described syntax elements, or variations thereof. ·Create and / or transmit and / or receive and / or decode a bitstream or signal that includes one or more of the described syntax elements, or variations thereof. ·A TV, set-top box, mobile phone, tablet, or other electronic device performs in-loop filtering according to any of the described embodiments. ·A TV, set-top box, mobile phone, tablet, or other electronic device performs in-loop filtering according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). ·A TV, set-top box, mobile phone, tablet, or other electronic device tunes a channel (e.g., using a tuner) to receive a signal including an encoded image and performs in-loop filtering according to any of the described embodiments. ·A TV, set-top box, mobile phone, tablet, or other electronic device receives a signal including an encoded image wirelessly (e.g., using an antenna) and performs in-loop filtering according to any of the described embodiments.
[0103] Various other generalized as well as specialized inventions and claims are also supported and contemplated throughout this disclosure.
[0104] An embodiment of a method 1800 for encoding a block of video data using the general aspects described herein is shown in FIG. 18. The method begins at start block 1801 and control proceeds to functional block 1810 to determine a region of a picture that uses a common set of filter parameters for filtering at least one reconstructed block of the picture. Control then proceeds from block 1810 to block 1920 to obtain a plurality of sets of filter parameters. Control proceeds from block 1820 to block 1830 to filter a region of a picture that includes at least one reconstructed block with a common set of filter parameters for the blocks within the region. Control then proceeds from block 1830 to block 1840 to encode information in a bitstream that includes a syntax indicating the set of filter parameters used to filter the region and an encoded version of the region.
[0105] An embodiment of a method 1900 for decoding a block of video data using the general aspects described herein is shown in FIG. 19. The method begins at start block 1901 and control proceeds to functional block 1910 to decode syntax from a bitstream that indicates a plurality of sets of filter parameters used to filter a region of a picture. Control then proceeds from block 1910 to block 1920 to determine a region of a picture from the bitstream using a common set of filter parameters for filtering at least one reconstructed block of the picture. Control proceeds from block 1920 to block 1930 to filter at least one reconstructed block using the set of filter parameters associated with the region that includes the at least one reconstructed block. Control then proceeds from block 1930 to block 1940 to decode the filtered reconstructed blocks of the picture.
[0106] FIG. 20 shows an embodiment of an apparatus 2000 for encoding or decoding a block of video data. The apparatus includes a processor 2010 and can be interconnected to a memory 2020 via at least one port. Both the processor 2010 and the memory 2020 can also have one or more additional interconnections to external connections.
[0107] The processor 2010 is configured to encode or decode video data by forming a plurality of reference arrays from the reconstructed samples of a block of video data, applying one weight set selected from a plurality of weight sets to one or more of the plurality of reference arrays to predict each target pixel of the block of video data respectively, calculating a final prediction of the target pixel of the block of video from one or more of the reference arrays as a function of the prediction respectively, and using the final prediction to encode or decode the block of video.
Claims
Claim 1 A method comprising: determining a region of a picture that uses a common set of filter parameters for filtering at least one reconstructed block of the picture; obtaining a plurality of sets of filter parameters; filtering the region of the picture that includes the at least one reconstructed block, using the common set of filter parameters for blocks within the region; encoding into a bitstream information including syntax indicating the set of filter parameters used for filtering the region, the information further including adaptative loop filter parameters encoded in a picture header, the shape of the region encoded as a map, and the encoded version of the region; wherein the filter block size is encoded per region based on a predefined table or derivation rule, or inferred from a quantization parameter, each filter block belongs to one filter region, the filter region is a sub-division of a slice, and the adaptative loop filter parameters are signaled at picture level or above and within the first filter block or in a slice header. Claim 2 An apparatus comprising: a processor configured to determine a region of a picture that uses a common set of filter parameters for filtering at least one reconstructed block of the picture; obtain a plurality of sets of filter parameters; filter the region of the picture that includes the at least one reconstructed block, using the common set of filter parameters for blocks within the region; encode into a bitstream information including syntax indicating the set of filter parameters used for filtering the region, the information further including adaptative loop filter parameters encoded in a picture header, the shape of the region encoded as a map, and the encoded version of the region; An apparatus comprising a processor configured such that a filter block size is coded for each region based on a predefined table or derivation rule, or is inferred from quantization parameters, each filter block belongs to one filter region, the filter region is a sub - division of a slice, and the adaptive loop filter parameters are signaled at picture level or above and within the first filter block or in a slice header. **Claim 3** A method comprising: decoding syntax from a bitstream that indicates a plurality of sets of filter parameters used to filter a region of a picture that includes adaptive loop filter parameters in a picture header; determining the region of the picture from the bitstream using a common set of filter parameters for filtering at least one reconstructed block of the picture, wherein the shape of the region is decoded as a map; filtering the at least one reconstructed block using the set of filter parameters associated with the region that includes the at least one reconstructed block; decoding the filtered reconstructed block of the picture, wherein a filter block size is coded for each region based on a predefined table or derivation rule, or is inferred from quantization parameters, each filter block belongs to one filter region, the filter region is a sub - division of a slice, and the adaptive loop filter parameters are signaled at picture level or above and within the first filter block or in a slice header. **Claim 4** An apparatus comprising: a processor configured to: decode syntax from a bitstream that indicates a plurality of sets of filter parameters used to filter a region of a picture that includes adaptive loop filter parameters in a picture header; determine the region of the picture from the bitstream using a common set of filter parameters for filtering at least one reconstructed block of the picture, wherein the shape of the region is decoded as a map; Filtering the at least one reconstructed block using the set of filter parameters associated with the region containing the at least one reconstructed block, Decoding the filtered reconstructed block of the picture, wherein the filter block size is coded for each region based on a predefined table or derivation rule, or is inferred from the quantization parameter, each filter block belongs to one filter region, the filter region is a sub-division of a slice, and the adaptable loop filter parameters are signaled at picture level or above and within the first filter block or in the slice header, a processor configured as such, an apparatus comprising. **Claim 5** The method according to claim 1 or 3, wherein the syntax is an index indicating a set of filter parameters. **Claim 6** The method according to claim 1 or 3, wherein filtering is performed on a sample adaptive offset filter and the index is applied to an SAO block. **Claim 7** The method according to claim 1 or 3, wherein the filter parameters for the SAO filter are available for filtering other blocks within the SAO region. **Claim 8** The method according to claim 1 or 3, wherein a set of filter parameters for the adaptable loop filter is signaled in the first filter block of the region. **Claim 9** The method according to claim 1 or 3, wherein for adaptable loop filter time prediction, the filter blocks of the current region use the filter parameters corresponding to the region at the same location in the reference picture. **Claim 10** The method according to claim 1 or 3, wherein for the SAO blocks in the first column or the first row of the SAO region, the sample adaptive offset filter parameters available for merge include the filter parameters of the SAO blocks to the left or above outside the SAO region. **Claim 11** The method according to claim 1 or 3, wherein the filter block size varies for each region. **Claim 12** A device, The apparatus according to claim 4, and A device comprising at least one of: (i) an antenna configured to receive a signal, wherein the signal includes a video block; (ii) a band limiter configured to limit the received signal to a frequency band including the video block; and (iii) a display configured to display an output representing the video block. **Claim 13** A computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to perform the method according to any one of claims 1, 3, and 5 to 11.
Citation Information
Patent Citations
Image decoding device, image encoding device, image filter device, and data structure of encoded data
JP2013141094A
Method and apparatus for improved loop-type filtering process
JP2014506061A
Peak Sample Adaptive Offset
JP2019534631A
Image encoder, image decoder, image encoding method and image decoding method
WO2010137322A1
Peak sample adaptive offset
WO2018067722A1