Codec mode dependent adaptive filtering

By performing local adaptive filtering in video encoding and decoding, the problems of blockiness and artifacts in the video compression process of existing technologies are solved, resulting in more efficient compression and better image quality.

CN121890069APending Publication Date: 2026-04-17INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480059172.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-18
Filing Date
2024-09-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively reduce blockiness and artifacts during efficient compression, leading to a decline in video quality.

Method used

Local adaptive filtering is performed using an adaptive filter (ALF). The filter parameters are determined by offline pre-trained filters and encoder-side optimization rate-distortion criteria and transmitted in the bitstream. The decoder side applies the filter to perform video image filtering.

Benefits of technology

It improves the compression efficiency and image quality of video encoding and decoding, reduces block artifacts and artifacts, and enhances the quality of the reconstructed signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121890069A_ABST
    Figure CN121890069A_ABST
Patent Text Reader

Abstract

Methods and apparatus for classification-based adaptive filtering are provided. At least one embodiment includes selecting a particular filter and set of filters depending on various classification criteria including a codec mode or a prediction mode. In a video encoding or decoding process, filtering is performed on a reconstructed video. In at least one of the embodiments, information is signaled from an encoder to a decoder, the information indicating filters to be used within a classification.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of European serial number 23306534.1, filed on 18 September 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] At least one of these embodiments generally relates to a method or apparatus for video encoding or decoding, compression or decompression. Background Technology

[0003] To achieve high compression efficiency, image and video codec schemes typically employ prediction (including motion vector prediction) and transform to utilize spatial and temporal redundancy in the video content. Intra-frame or inter-frame prediction is usually used to leverage intra- or inter-frame correlations, followed by transforming, quantizing, and entropy encoding / decoding of the differences between the original and predicted images (often referred to as prediction error or prediction residual). To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy encoding / decoding, quantization, transforming, and prediction. Summary of the Invention

[0004] At least one embodiment of the present invention generally relates to a method or apparatus for video encoding or decoding, and more specifically, to a method or apparatus for adaptive filtering when encoding or decoding video data.

[0005] According to a first aspect, a method is provided. The method includes the steps of: filtering a portion of a digital video image based on a classification of indicator filter parameters; and encoding the digital video image using the filtered portion of the digital video image.

[0006] According to the second aspect, another method is provided. This method includes the steps of: decoding a portion of a digital video image; and filtering the portion of the digital video image based on information indicating filter parameters.

[0007] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to operate on digital video data according to the aforementioned method.

[0008] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to encode blocks of video or decode video data by performing any of the methods described above.

[0009] According to another general aspect of at least one embodiment, an apparatus is provided comprising: means according to any embodiment of the decoding embodiment; and at least one of the following: (i) an antenna configured to receive a signal including a video block, (ii) a band limiter configured to restrict the received signal to a frequency band including the video block, or (iii) a display configured to display an output representing the video block.

[0010] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided, which contains data content generated according to any embodiment of the described encoding embodiments or variations.

[0011] According to another general aspect of at least one embodiment, a signal is provided, the signal comprising video data generated according to any embodiment or variant of the described encoding embodiment.

[0012] According to another general aspect of at least one embodiment, video data or bitstream is formatted to include data content generated according to any embodiment or variant of the described encoding embodiment.

[0013] According to another general aspect of at least one embodiment, a computer program product including instructions is provided that, when executed by a computer, cause the computer to implement any of the described decoding embodiments or variations.

[0014] These and other aspects, features, and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments in conjunction with the accompanying drawings. Attached Figure Description

[0015] Figure 1 The figure shows an example of a loop filter in a VVC.

[0016] Figure 2 The figure shows examples of (a) a luminance ALF 7x7 diamond filter and (b) a chrominance ALF 5x5 diamond filter.

[0017] Figure 3 The diagram illustrates an example of an ALF filter set that is signal-informed and contains more than one filter in a VVC.

[0018] Figure 4 The figure shows an example of a pre-trained ALF filter set in VVC.

[0019] Figure 5 The figure shows an example of a general pre-trained filter set.

[0020] Figure 6The figure illustrates an example of adaptive filtering that relies on encoding / decoding modes and uses an independent pre-trained filter bank.

[0021] Figure 7 The figure illustrates an example of mode-dependent adaptive filtering that utilizes shared pre-trained filters.

[0022] Figure 8 The figure illustrates an embodiment of the first method according to the described aspects.

[0023] Figure 9 The figure illustrates an embodiment of the second method according to the described aspects.

[0024] Figure 10 The figure illustrates one embodiment of the apparatus according to the described aspects.

[0025] Figure 11 The diagram illustrates a standard, universal video compression scheme.

[0026] Figure 12 The diagram illustrates a standard, universal video decompression scheme.

[0027] Figure 13 The figure illustrates a processor-based system for encoding / decoding, as described in the general description. Detailed Implementation

[0028] The embodiments described herein are in the field of video compression and generally relate to video compression as well as video encoding and decoding, and more specifically to adaptive filtering processes such as post-filters and in-loop filters, for example, adaptive loop filters (ALF) or any other Wiener-based or linear convolution-based post-filters.

[0029] To achieve high compression efficiency, image and video codec schemes typically employ block-based prediction (including motion vector prediction) and transform to utilize spatial and temporal redundancy in the video content. Generally, intra-frame or inter-frame prediction is used to leverage intra- or inter-frame correlations, followed by transforming, quantizing, and entropy encoding / decoding of the differences between the original and predicted blocks (often referred to as prediction error or prediction residual). To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy encoding / decoding, quantization, transforming, and prediction.

[0030] In the High Efficiency Video Coding (HEVC) video compression standard, motion-compensated time prediction is used to take advantage of the redundancy between consecutive frames in a video.

[0031] To achieve this, motion vectors are associated with each prediction unit (PU). Each CTU (code-decode tree unit) is represented by a code-decode tree in the compression domain. This is a quadtree partition of the CTU, where each leaf is called a code-decode unit (CU).

[0032] Each CU is then given some intra-frame or inter-frame prediction parameters (prediction information). For this purpose, the CU is spatially divided into one or more prediction units (PUs), and each PU is assigned some prediction information. Intra-frame or inter-frame coding / decoding modes are assigned at the CU level.

[0033] Block-based intra / inter-frame prediction and transform coding, along with quantization, can introduce various artifacts at low to medium bit rates. To reduce these artifacts, video codec standards (such as VVC) and video codec solutions (such as ECM) implement in-loop filters such as deblocking filter (DBF), bilateral filter (BIF), sample adaptive offset (SAO), and adaptive loop filter (ALF).

[0034] The parameters of adaptive filters (such as SAO and ALF) are determined on the encoder side by optimizing the rate-distortion criterion. These filter parameters are transmitted with the bitstream and then decoded and used on the decoder side. A significant portion of the gain in brightness provided by ALF is due to further local adaptation implemented in the classification: samples sharing the same local features are filtered with the same specific filter. Other types of classification have been proposed, for example, based on color or component frequency bands, based on residual energy, and based on the variance of the reconstructed samples.

[0035] The described embodiments address the problem of classifying samples in the most general way, making filter optimization as efficient as possible.

[0036] The relevant technologies are summarized below.

[0037] Adaptive Loop Filter (ALF) Loop filters in Universal Video Codec (VVC) The VVC standard implements three in-loop filters: Deblocking Filter (DBF), Sample Adaptive Shift (SAO), and Adaptive Loop Filter (ALF). The Deblocking Filter aims to reduce block discontinuities. The Sample Adaptive Shift primarily aims to reduce artifacts caused by the quantization of transform coefficients. The Adaptive Loop Filter and Cross Component Adaptive Loop Filter are adaptive filters that can enhance the reconstructed signal using coding methods such as Wiener filters.

[0038] The working process of the filter in the VVC loop is as follows: Figure 1As shown, if the local deblocking condition is met, the luminance and chrominance reconstructed samples located along the block boundaries are first filtered using a deblocking filter (DBF). Then, sample adaptive offset (SAO) is used to locally add an offset based on the classification using a band-based classifier or an edge classifier. Finally, an adaptive loop filter (ALF) and a cross-component adaptive loop filter (CCALF) are run before storing the resulting sample values ​​in the reference image buffer.

[0039] ALF is an adaptive filter that is applied to reduce the mean square error (MSE) between the original sample and the reconstructed sample using coding methods such as Wiener filters. The ALF filter is transmitted in the bitstream and decoded on the decoder side, derived from parameters (previously) signaled or from predefined parameters, and then applied to the reconstructed sample.

[0040] The Cross Component Adaptive Loop Filter (CCALF) uses luminance samples to refine chrominance sample values ​​during the ALF process.

[0041] ALF in VVC In VVC, the ALF filter is point-symmetric, DC-neutral, and has integer coefficients. The ALF in VVC uses a 7x7 diamond filter for luminance and a 5x5 diamond filter for chrominance, as shown below. Figure 2 As shown.

[0042] Filtering operation set up To reconstruct the sample, and This represents the value after ALF filtering. In the linear implementation of ALF, The calculation is as follows: (1) in Represents the filter coefficients. and: (2) in Indicates the relationship with the first coefficients The coordinate offset of the associated reconstructed sample.

[0043] In the nonlinear implementation of ALF, equation (1) becomes: (3) in: (4) in It is related to the coefficient The associated clipping parameters are determined by the clipping index. Confirm. Trimming parameters. The derivation is as follows: (5) in Indicates the sample bit depth, and It can be 0, 1, 2 or 3.

[0044] Classification and Filter Sets In terms of brightness, based on its directionality and Laplace activity, Sub-block level execution classification.

[0045] Let D and A represent directionality and Laplace activity, respectively, both of which are derived by calculating the local horizontal, vertical, and diagonal gradients. The directionality D ranges from 0 to 4 (inclusive), as shown in Table 1. Direction value Texture 0 Weak horizontal / vertical 1 Strong horizontal / vertical 2 weak diagonal 3 strong diagonal 4 Table 1: Directionality of ALF classification in VVC.

[0046] Activity A is obtained by locally accumulating the absolute values ​​of the local Laplacian operator, and its value range can be from 0 to 4 (inclusive).

[0047] This resulted in a classification of 25 categories. The derivation of the category index is as follows: (6) Each category has a specified specific filter.

[0048] A filter set contains a maximum of 25 filters; if there is more than one filter, a list of 25 indexed items is provided. This list maps (lookup tables) to which filter should be used for each category. Figure 3 An example of a set of filters that is signaled and contains more than one filter is depicted.

[0049] Online filter optimization and offline pre-training The ALF filter coefficients and clipping indices are determined on the encoder side, for example, by minimizing the MSE between the reconstructed sample and its original value by solving the Wiener-Hopf equation. If the rate-distortion condition is met, these coefficients and the corresponding clipping indices (if applicable) are encoded in an adaptive parameter set (APS). In VVC, the ALF APS contains a luma filter set and at most eight chroma filters.

[0050] VVC also allows the use of pre-trained (predefined) luminance filters and filter sets, which are hard-coded at both the encoder and decoder sides. In VVC, there are a total of 16 pre-trained filter sets and 64 pre-trained filters. Figure 4A schematic example of such a pre-trained filter set is provided.

[0051] Further Development of ECM The ALF model in Enhanced Compression Model (ECM) software has undergone further development, and its main classification-related features are summarized below.

[0052] Fixed-brightness filtering and related classifications In ECM, the "fixed" indicator filters are not derived online from image samples; instead, the filters are pre-trained offline. Therefore, fixed filters are not signaled in the bitstream; they are hard-coded in both the encoder and decoder.

[0053] To filter the brightness samples, three different classifiers were used. , and And three different sets of filters (F0, F1, and F2). Sets F0 and F1 contain fixed filters whose coefficients are tailored to the classifier. and Pre-training was performed. The coefficients of the filters in F2 are indicated by the signal. For a given sample, which filter from set Fi is used is determined by the classifier. The category assigned to this sample Decide.

[0054] Two classifiers corresponding to fixed filtering and Both are based on the Laplacian operator and applied to 2×2 sub-blocks. In both classifiers, activity and orientation values ​​are derived based on the vertical, horizontal, and diagonal gradients. The class index is then determined based on the activity and orientation values.

[0055] In ECM-2 through ECM-9, the number of classes in the two classifiers is 896. Eight QP adaptive filter banks are pre-trained for each classifier. Each QP adaptive bank contains 512 filters. Seven QP ranges are defined, and for each QP range, each classifier has two filter set candidates.

[0056] An alternative 2x2 ALF classifier for luminance For signal-notified filters, luminance classification is extended using an additional surrogate classifier (called a band-based classifier). The signal-notified luminance filter set contains a flag indicating whether a surrogate classifier has been applied. Geometric transformations are not applied to the surrogate band-based classifier. When applying a band-based classifier, the sum of sample values ​​for a 2×2 luminance block is first calculated, and then the class index is calculated as follows: (7).

[0057] Classification of ALF filters based on residual data and signal notification In addition to Laplacian classifiers and band-based classifiers, one proposal is to apply a new classifier to residual samples. This involves calculating the sum of the absolute values ​​of residual samples in adjacent windows, and the class index is derived as follows: (8) Similar to the previous two classifiers, the value of classIdx ranges from 0 to 24.

[0058] Classification of Extended ALF Fixed Filters with Variance In the JVET update, the classifier with a fixed filter is expanded. First, for each 2×2 block, the mean of the surrounding window is calculated. Then, for each sample in that window, the difference between the sample value and the mean is calculated. The scaling factor is determined based on the activity values ​​derived from the Laplacian classifier. The variance, i.e., the square root of the sum of squared differences, is further quantized by the scaling factor. . The value is an integer between 0 (inclusive) and 7 (inclusive). For ,set up This indicates that it comes from the first part of ECM-8.0. A classifier with a fixed number of filters. Then, the proposed class index. Deduced as (9).

[0059] Variance-based classification for in-loop filters Variance-based classifications for in-loop filtering have been proposed, namely bilateral filtering (BIF) and adaptive loop filtering (ALF).

[0060] For BIF, the variance of TU is used to identify texture intensity. More filter intensity is introduced into BF. The filter intensity is determined by the variance.

[0061] For ALF, each classification unit is jointly classified into two levels of texture intensity based on variance and boundary location. The texture intensity levels are then further combined with the existing classifier to output the final classification result.

[0062] The described embodiments include the use of adaptive filters and filter sets that depend on the encoding / decoding mode.

[0063] The encoding / decoding mode under consideration can refer to slices / images or codec units (CUs).

[0064] Depending on the encoding / decoding mode, different adaptive filters and filter sets are used.

[0065] The described embodiments include the selection of specific filters and filter sets depending on the encoding / decoding or prediction mode.

[0066] This section considers classification-based adaptive filtering. Let K represent the number of categories in the classification. Consistent with existing techniques, the adaptive filter set maps at most K filters to K categories; for example, when there is more than one filter, it has a list of K indexes. This list specifies the filter that should be used for each category.

[0067] Offline pre-trained filters and filter sets In the embodiments, pre-trained filters (i.e. filters not signaled in the bitstream) and pre-trained filter sets (i.e. filter sets not signaled in the bitstream) are considered.

[0068] The pre-trained filter set includes a list of K indices that specify the filter index for each class, as well as potential additional information that may, for example, specify the classifier to be used.

[0069] Let N be the number of filters that are pre-trained offline and hard-coded in both the encoder and decoder. Figure 5 A schematic example of a general pre-trained filter set is provided.

[0070] Let P be the number of filter sets that are pre-determined offline and are also hard-coded in both the encoder and decoder.

[0071] During encoding, the encoder selects which filter set should be used according to the rate-distortion criterion (e.g., at the CTU level) corresponding to the encoding / decoding mode, and signals the corresponding filter set index in the bitstream.

[0072] On the decoder side, the decoder sets the filter index, and the decoder selects the appropriate adaptive filter to use based on the encoding / decoding mode.

[0073] Determining the encoding / decoding mode Slice / Image Encoding / Decoding Mode In one variant, the encoding / decoding mode under consideration refers to the current slice or image. Depending on whether the current slice is inter-frame or intra-frame encoded, the filter set index to be used is selected from the P filter sets dedicated to inter-frame or intra-frame samples, respectively.

[0074] Codec Unit (CU) Prediction Mode In another variation, the encoding / decoding mode under consideration refers to the prediction mode of the encoding / decoding unit (CU) to which the current sample belongs.

[0075] Consider two prediction modes: intra-frame prediction and inter-frame prediction. Intra-frame block copy (IBC) and inter-frame / intra-frame combined prediction (CIIP) can be regarded as intra-frame prediction modes or inter-frame prediction modes.

[0076] In this variant, both inter-frame and intra-frame filter set indices are signaled, for example, at the CTU level.

[0077] Pre-trained filter set depending on encoding / decoding mode In ECM-2 through ECM-9, without implementing such codec mode dependency, there are two fixed filter classifiers, each with 8 QP adaptive pre-trained filter banks. For both classifiers, K is equal to 896. Each QP adaptive bank contains N=512 filters. Seven QP ranges are defined, and for each QP range, P is equal to 2 (each classifier has two filter set candidates for each QP range).

[0078] Independent inter-frame / intra-frame filter banks In this embodiment, the pre-trained filter sets for inter-frame samples and intra-frame samples are pre-trained independently, and therefore comprise two separate filter sets, such as... Figure 6 As shown. That is, there are two LUTs, one for intra-frame samples and the other for inter-frame samples, and there are two pre-trained filter sets, one for intra-frame samples and the other for inter-frame samples.

[0079] In this embodiment, the definitions of the pre-trained filters and filter sets on which the encoding / decoding mode depends can be illustrated by the following pseudocode: Here, " / / …" and " / * … * / " represent informational comments, classToFilterIdx represents the pre-trained filter set, filterCoef represents the pre-trained filter coefficient array, NB_CODING_MODE represents the number of codec modes considered, NB_QP represents the number of QPs considered during offline filter training, NB_CLASSIFIER represents the number of classifiers, NB_CLASS represents the number of classes (i.e., using the previous notation: K), NB_FILTER represents the number of filters pre-trained per QP and per classifier (i.e., using the previous notation: N), and NB_COEF represents the number of coefficients defined for each filter.

[0080] Shared pre-trained filter bank In another embodiment, both the inter-frame and intra-frame pre-trained filter sets are mapped to the same set of pre-trained filters via a lookup table (LUT), such as... Figure 7As shown, define one LUT for the intra-frame codec mode (samples used for intra-frame codec), and define another LUT for the inter-frame codec mode (samples used for inter-frame codec).

[0081] In this embodiment, the definitions of the pre-trained filters and filter sets on which the encoding / decoding mode depends can be illustrated by the following pseudocode: In this context, as in the previous embodiments, “ / / …” and “ / * … * / ” represent informational comments, classToFilterIdx represents the pre-trained filter set, filterCoef represents the pre-trained filter coefficient array, NB_CODING_MODE represents the number of codec modes considered, NB_QP represents the number of QPs considered during offline filter training, NB_CLASSIFIER represents the number of classifiers, NB_CLASS represents the number of classes (i.e., using the previous notation: K), NB_FILTER represents the number of filters pre-trained per QP and per classifier (i.e., using the previous notation: N), and NB_COEF represents the number of coefficients defined for each filter.

[0082] Filters that use signals for notification In another embodiment, consider a filter that is exported online from video samples and signaled in the bitstream.

[0083] The corresponding filter set must also be signaled in the bitstream and include at most K filters, or, if there is more than one, a list of K indices that specifies which filter should be used for each category, along with potential additional information, such as specifying the classifier to be used.

[0084] The adaptive filter set on which the encoding / decoding mode depends contains additional information, such as indices, which specify the encoding / decoding mode they involve.

[0085] During encoding, the encoder calculates one or more MSE-optimal filter sets for each encoding / decoding mode, for example, based on slices or images, and then determines whether to signal these filter sets according to the rate-distortion criterion.

[0086] On the decoder side, the filter set notified by the signal is decoded, and can then be used accordingly in subsequent slices / images based on its specific encoding / decoding mode.

[0087] Figure 8An embodiment of method 800 under the general aspects described herein is shown. The method begins at a start block 801 and control proceeds to block 810 for filtering a portion of a digital video image based on a classification indicating filter parameters. Control then proceeds from block 810 to block 820 for encoding the digital video image using the filtered portion.

[0088] Figure 9 An embodiment of method 900 under the general aspects described herein is shown. The method begins at a starting block 901 and control proceeds to block 910 for decoding a portion of a digital video image. Control then proceeds from block 910 to block 920 for filtering the portion of the digital video image based on information indicating filter parameters.

[0089] Figure 10 An embodiment of an apparatus 1300 for encoding, decoding, compressing, decompressing, or filtering video data using the aforementioned methods is shown. The apparatus includes a processor 1310 and can be interconnected with a memory 1320 via at least one port. Both the processor 1310 and the memory 1320 may also have one or more additional interconnects to external connections.

[0090] The processor 1310 is also configured to insert or receive information in the bitstream and to compress, encode, or decode using any of the aspects described.

[0091] The embodiments described herein encompass a wide variety of aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described, and often in a manner that may sound limiting, at least for the purpose of illustrating individual characteristics. However, this is for the purpose of clarity and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in earlier applications.

[0092] The aspects described and envisioned in this application can be realized in many different forms. Figure 11 , Figure 12 and Figure 13 Some embodiments have been provided, but other embodiments are contemplated, and... Figure 11 , Figure 12 and Figure 13The discussion does not limit the breadth of implementation methods. At least one aspect generally relates to video encoding and decoding, while at least another aspect generally relates to the transmission of the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having a bitstream generated according to any of the described methods stored thereon.

[0093] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "pixel" and "sample," and the terms "image," "picture," and "frame." Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0094] This document describes various methods, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions may be modified or combined.

[0095] The various methods and other aspects described in this application can be used to modify, for example... Figure 11 and Figure 12 The modules of the video encoder 100 and decoder 200 shown include, for example, intra-frame prediction, entropy encoding / decoding, and / or decoding modules (160, 260, 145, 230). Furthermore, this aspect is not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and any extensions to such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically impractical, the aspects described in this application may be used individually or in combination.

[0096] Various numerical values ​​are used in this application. Specific values ​​are for illustrative purposes only, and the aspects described are not limited to these specific values.

[0097] Figure 11 The encoder 100 is illustrated. Although variations of the encoder 100 are envisioned, for clarity, only the encoder 100 will be described below, without describing all anticipated variations.

[0098] Before being encoded, the video sequence may undergo pre-encoding processing (101), such as applying a color transformation to the input color image (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input image components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with the preprocessing and attached to the bitstream.

[0099] In encoder 100, the image is encoded by encoder elements as described below. The image to be encoded is divided (102) and processed in units, for example, CUs. Each unit is encoded using, for example, an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, intra-frame prediction (160) is performed. In inter-frame mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.

[0100] Then, the predicted residual is transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode and decode the residual without applying the transform or quantization process.

[0101] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined with the prediction block (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) filtering, thereby reducing coding artifacts. The filtered image is stored in a reference image buffer (180).

[0102] Figure 12 The figure shows a block diagram of a video decoder 200. In decoder 200, the bitstream is decoded by decoder elements, as described below. Video decoder 200 generally performs operations similar to... Figure 11 The encoding pass shown is the inverse of the decoding pass. Encoder 100 also typically performs video decoding as part of the encoded video data.

[0103] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 100. First, entropy decoding (230) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoding / decoding information. Image segmentation information indicates how the image is segmented. Therefore, the decoder can segment the image based on the decoded image segmentation information (235). The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. The prediction blocks can be obtained from intra-frame prediction (260) or motion-compensated prediction (i.e., inter-frame prediction) (275) (270). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (280).

[0104] The decoded image can undergo further post-decoding processing (285), such as inverse color transformation (e.g., from YcbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which is the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding process can use metadata derived in the pre-encoding process and signaled in the bitstream.

[0105] Figure 13 The figure illustrates a block diagram of an example system in which various aspects and embodiments are implemented. System 1000 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.

[0106] System 1000 includes at least one processor 1010 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this document. Processor 1010 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0107] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Furthermore, the encoder / decoder module 1030 may be implemented as a separate element of system 1000, or it may be incorporated into processor 1010 as a combination of hardware and software known to those skilled in the art.

[0108] To perform the various aspects described in this document, program code to be loaded onto processor 1010 or encoder / decoder 1030 may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, during the execution of the processes described in this document, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more items. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results generated by processing equations, formulas, operations, and operational logic.

[0109] In some embodiments, the memory within the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, such as dynamically volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, while 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Codec, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Codec, a new standard being developed by the Joint Video Experts Group JVET).

[0110] As shown in block 1130, inputs can be provided to the components of system 1000 through various input devices. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Other examples ( Figure 13 (Not shown in the image) Includes composite video.

[0111] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or limiting a signal band to a certain frequency band); (ii) down-converting the selected signal; (iii) further band-limiting to a narrower band to select, for example, a signal frequency band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0112] Furthermore, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented as needed, for example, within a separate input processing IC or within the processor 1010. Similarly, various aspects of USB or HDMI interface processing can be implemented as needed within a separate interface IC or within the processor 1010. Demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010, and encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.

[0113] Various components of System 1000 can be provided within an integrated housing. Within the integrated housing, the various components can be interconnected using a suitable connection arrangement and data can be transferred between them, for example, internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards.

[0114] System 1000 includes a communication interface 1050, which is capable of communicating with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or network interface card (NIC), while the communication channel 1060 may be implemented, for example, in a wired and / or wireless medium.

[0115] In various embodiments, data is streamed to system 1000 using a wireless network (such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) or otherwise provided. In these embodiments, the Wi-Fi signal is received via a communication channel 1060 and a communication interface 1050 adapted for Wi-Fi communication. The communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, thereby allowing streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 1000, with the set-top box delivering data via an HDMI connection to input block 1130. Still other embodiments use an RF connection to input block 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. Furthermore, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0116] System 1000 can provide output signals to various output devices, including display 1100, speaker 1110, and other peripheral devices 1120. Display 1100 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 1100 can be used in a television, tablet computer, laptop computer, mobile phone, or another device. Display 1100 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, both terms applicable), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functionality based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.

[0117] In various embodiments, signaling using communication protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention transmits control signals between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120. Output devices can be communicatively coupled to system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. Display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000 in electronic devices such as, for example, televisions. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0118] For example, if the RF portion of input 1130 is part of a separate set-top box, then display 1100 and speaker 1110 may alternatively be separate from one or more other components. In various embodiments where display 1100 and speaker 1110 are external components, output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0119] These embodiments can be implemented by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, these embodiments can be implemented by one or more integrated circuits. Memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. As a non-limiting example, processor 1010 can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0120] Various implementations involve decoding. As used in this application, "decoding" can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by the decoder of the various implementations described in this application.

[0121] As another example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to a broader decoding process will be clear from the specific context of the description and is considered well understood by those skilled in the art.

[0122] Various implementations involve encoding. Similar to the discussion of "decoding" above, "encoding," as used herein, can encompass all or part of the processes performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such processes also include, or alternatively include, processes performed by the encoders of the various implementations described herein.

[0123] As another example, in one embodiment, “encoding” refers only to entropy encoding; in another embodiment, “encoding” refers only to differential encoding; and in yet another embodiment, “encoding” refers to a combination of entropy encoding and differential encoding. Whether the phrase “encoding process” is intended to specifically refer to a subset of operations or to refer to a broader encoding process will be clear from the context of the specific description and is considered well understood by those skilled in the art.

[0124] Note that the grammatical elements used in this article are descriptive terms. Therefore, they do not preclude the use of other grammatical element names.

[0125] When a diagram is presented in flowchart form, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a diagram is presented in block diagram form, it should be understood that it also provides a flowchart of the corresponding method / process.

[0126] Various implementations may involve parametric models or rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, which usually imposes constraints on computational complexity. This can be measured by rate-distortion optimization (RDO), or by least mean square (LMS), mean absolute error (MAE), or other such measures. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving rate-distortion optimization problems. For example, these approaches might be based on extensive testing of all encoding options, including all considered modes or encoding / decoding parameter values, and a comprehensive evaluation of their encoding / decoding costs and the associated distortion of the encoded / decoded and reconstructed signals. To save encoding complexity, faster methods can also be used, particularly based on calculating approximate distortion based on the predicted signal or predicted residual signal (rather than the reconstructed signal). A hybrid of these two approaches can also be used, such as by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of the techniques from a wide variety of technologies to perform optimization, but optimization is not necessarily a comprehensive assessment of encoding / decoding costs and associated distortion.

[0127] The implementations and aspects described herein can be implemented, for example, in methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementations of the features in question can also be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. These methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0128] References to "an embodiment," "an embodiment," "an implementation," "an implementation," and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment," "in an embodiment," "in an implementation," "in an implementation," and any other variations that appear throughout this application do not necessarily all refer to the same embodiment.

[0129] Furthermore, this application may involve "determining" various pieces of information. Determining information may include, for example, one or more of estimated information, calculated information, predicted information, or information retrieved from memory.

[0130] Furthermore, this application may involve "accessing" various information fragments. Accessing information may include one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0131] Furthermore, this application may involve "receiving" various pieces of information. "Receiving," like "accessing," is intended to be a broad term. Receiving information may include, for example, accessing information or retrieving information (e.g., from memory) or one or more of them. Moreover, "receiving" is typically involved in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0132] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of…”, for example, in the cases of “A / B”, “A and / or B”, and “at least one of A and B”, is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B) simultaneously. As another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). As will be apparent to those skilled in the art and related fields, this can be extended to many of the listed items.

[0133] Furthermore, as used herein, the term "signal" refers, among other things, to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a specific one of a plurality of transforms, encoding / decoding modes, or flags. In this way, in embodiments, the same transform, parameter, or mode is used at both the encoder and decoder sides. Thus, for example, the encoder may transmit (explicit signaling) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select that specific parameter. In various embodiments, bit savings are achieved by avoiding the transmission of any actual function. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the foregoing refers to the verb form of the term "signal," the word "signal" may also be used as a noun herein.

[0134] As will be apparent to those skilled in the art, implementations can generate a wide variety of signals that are formatted to carry, for example, information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted over a wide variety of wired or wireless links. The signal may be stored on a processor-readable medium.

[0135] The preceding sections describe multiple embodiments across various claim classes and types. Features of these embodiments may be provided individually or in any combination. Furthermore, across various claim classes and types, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination: At least one embodiment includes adaptively implementing the filtering operation by using one of a plurality of filters.

[0136] At least one embodiment includes encoding or decoding video blocks by using filtering in a heavy video path, according to the described embodiment.

[0137] At least one embodiment includes the embodiments described above to implement an adaptive loop filter.

[0138] At least one embodiment further includes a classification to be used when determining the set of filters to be used for in-loop filtering.

[0139] At least one embodiment further includes the above-described embodiment, wherein an index is used to select a filter from within a filter set.

[0140] At least one embodiment further includes the above-described embodiments, wherein the classification is based on encoding / decoding modes or prediction modes.

[0141] At least one embodiment includes the embodiments described above, wherein the classification is based on the directionality of the video image.

[0142] At least one embodiment includes any encoding or decoding operation based on the above operations.

[0143] At least one embodiment includes using the above method to encode or decode a sub-block.

[0144] At least one embodiment includes a bitstream or signal that includes one or more of the described syntax elements or variations thereof.

[0145] At least one embodiment includes a bitstream or signal that includes a syntax for conveying information generated according to any of the described embodiments.

[0146] At least one embodiment includes creation and / or transmission and / or reception and / or decoding according to any of the described embodiments.

[0147] At least one embodiment includes a method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments.

[0148] At least one embodiment includes inserting syntax elements into the signaling that enable the decoder to determine the decoded information in a manner corresponding to that used by the encoder.

[0149] At least one embodiment includes creating, transmitting, receiving, and / or decoding a bitstream or signal, which includes one or more of the described syntax elements or variations thereof.

[0150] At least one embodiment includes a TV, set-top box, mobile phone, tablet computer or other electronic device that performs one or more transformation methods according to any of the described embodiments.

[0151] At least one embodiment includes a TV, set-top box, mobile phone, tablet computer or other electronic device that performs one or more transformation methods according to any of the described embodiments and displays (e.g., using a monitor, screen or other type of display) the resulting image.

[0152] At least one embodiment includes a TV, set-top box, mobile phone, tablet computer or other electronic device that selects, band-limits or tunes (e.g., using a tuner) a channel to receive a signal including an encoded image and performs one or more transformation methods according to any of the described embodiments.

[0153] At least one embodiment includes a TV, set-top box, mobile phone, tablet computer or other electronic device that receives signals including encoded images over the air (e.g., using an antenna) and performs one or more transformation methods.

Claims

1. A method comprising: A portion of a digital video image is filtered based on the classification of the indicator filter parameters; as well as The digital video image is encoded using the filtering portion of the digital video image.

2. An apparatus comprising: Memory, and The processor is configured as follows: A portion of a digital video image is filtered based on the classification of the indicator filter parameters; as well as The digital video image is encoded using the filtering portion of the digital video image.

3. A method comprising: Decoding a portion of a digital video image; as well as The portion of the digital video image is filtered based on information indicating the filter parameters.

4. An apparatus comprising: Memory, and The processor is configured as follows: Decoding a portion of a digital video image; as well as The portion of the digital video image is filtered based on information indicating the filter parameters.

5. The method of claim 1 or 3, or the apparatus of claim 2 or 4, wherein the filtering is performed on an in-loop component during the encoding or decoding process.

6. The method according to any one of claims 1 or 5, or the apparatus according to any one of claims 2 or 5, wherein the set of classification indicator filters is described.

7. The method according to any one of claims 1, 5, or 6, or the apparatus according to any one of claims 2, 5, or 6, wherein the classification is based on a prediction pattern.

8. The method according to any one of claims 1, 3, or 5-6, or the apparatus according to any one of claims 2, 4, or 5-6, wherein an index is used to indicate the filter parameters within the set.

9. The method according to any one of claims 1, 3, or 5-8, or the apparatus according to any one of claims 2, 4, or 5-8, wherein separate filter banks are used for inter-frame codec samples and intra-frame codec samples, and the filter banks are pre-trained independently.

10. The method according to any one of claims 1, 3 or 5-9, or the apparatus according to any one of claims 2, 4 or 5-9, wherein the filter parameters are determined at the codec unit level.

11. The method according to any one of claims 1, 3 or 5-10, or the apparatus according to any one of claims 2, 4 or 5-10, wherein information indicating the filter is signaled in the bitstream.

12. An apparatus comprising: The apparatus according to claim 2; as well as At least one of the following: (i) an antenna configured to receive a signal comprising a video block; (ii) a band limiter configured to restrict the received signal to a frequency band comprising the video block; or (iii) a display configured to display an output representing the video block.

13. A non-transitory computer-readable medium comprising data content generated by the method of any one of claims 1 or 3-11 or the apparatus of any one of claims 2 or 3-11, for playback using a processor.

14. A signal comprising video data generated by the method of any one of claims 1 or 3-11 or by the apparatus of any one of claims 2 or 3-11, for playback using a processor.

15. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 or 3-11.