Systems, methods, apparatuses, and computer program products for harmonizing neural network-based filtering with in-loop operations of next generation video codecs

WO2026201499A1PCT designated stage Publication Date: 2026-10-01NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/055831
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-03
Publication Date
2026-10-01

Smart Images

  • Figure EP2026055831_01102026_PF_FP_ABST
    Figure EP2026055831_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A method for video encoding, a method for video decoding, encoding and decoding apparatuses, and non-transitory computer-readable storage media thereof are provided. In the method for video encoding, the encoder may obtain pixel information associated with raw pixels in a picture for a first pixel that is to be filtered in the picture. The encoder may select and apply a plurality of filters for filtering the pixel information associated with the picture to obtain a filtered pixel. This approach may use a multi-stage adaptive loop filter (ALF). The encoder can, following filtering of the pixel information, generate adaptive filter parameters associated with the filtered pixel. Information associated with the adaptive filter parameters can be provided to the decoder in a bitstream. In the method for video decoding, the decoder receives the adaptive filter parameters and uses the same during decoding and regeneration of the filtered pixel.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS, METHODS, APPARATUSES, AND COMPUTER PROGRAM PRODUCTS FOR HARMONIZING NEURAL NETWORK- BASED FILTERING WITH IN-LOOP OPERATIONS OF NEXT GENERATION VIDEO CODECSTechnological Field

[0001] An example embodiment relates generally to video coding and decoding and, more particularly, but not exclusively, to harmonization of neural network-based filtering with in-loop operations of next generation video codecs.Background

[0002] This section is intended to provide a background or context to the invention that is recited in the claims. The description herein may include concepts that could be pursued, but are not necessarily ones that have been previously conceived or pursued. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description and claims in this application and is not admitted to be prior art by inclusion in this section.

[0003] A video coding system may include an encoder that transforms (e.g., encodes) a video sequence into a compressed representation suited for storage and / or transmission. For example, the encoder may discard certain information in the video sequence in order to represent the video sequence in a more compact form (e.g., a lower bitrate, etc.) for storage and / or transmission of the video sequence. Additionally, a video coding system may include a decoder that uncompress (e.g., decodes) the compressed representation of the video sequence to enable consumption of the video sequence in a viewable form via a display.Brief Summary

[0004] According to some aspects, there is provided the subject matter of the independent claims. Some further aspects are defined in the dependent claims. The embodiments that do not fall under the scope of the claims are to be interpreted as examples useful for understanding the disclosure.

[0005] A method for video encoding, a method for video decoding, encoding and decoding apparatuses, and non-transitory computer-readable storage media thereof are provided. In the method for video encoding, the encoder may obtain pixel information associated with raw pixels in a picture for a first pixel that is to be filtered in the picture. The encoder may select and apply aplurality of filters for filtering the pixel information associated with the picture to obtain a filtered pixel. This approach may use a multi-stage enhanced compression model-adaptive loop filter (ECM-ALF). The encoder can, following filtering of the pixel information, generate adaptive filter parameters associated with the filtered pixel. Information associated with the adaptive filter parameters can be provided to the decoder in a bitstream. In the method for video decoding, the decoder receives the adaptive filter parameters and uses the same during decoding and regeneration of the filtered pixel.

[0006] According to an aspect of the present disclosure, a method can be provided or carried out that comprises, inputting information associated with a set of pixels into a first stage of a multistage filtering process, the multi-stage filtering process comprising one or more first filters and one or more neural network-based (NN-based) filters; generating, using at least one of the one or more first filters or the one or more NN-based filters, one or more first filter outputs; inputting, into a second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first filter outputs, the second stage comprising one or more adaptive filters; generating, conditional on a presence of the one or more first filter outputs from at least one of the one or more first filters or the one or more NN-based filters, using the one or more adaptive filters of the second stage of the multi-stage filtering process, information about one or more parameters used during the multi-stage filtering process, and transmitting, in a bitstream, the information about the one or more parameters used during the multi-stage filtering process.

[0007] In some embodiments, the adaptive filter comprises an adaptive loop filter (ALF). In some embodiments, the information about the multi-stage filtering process comprises one or more coefficients associated with one or more filters used in the multi-stage filtering process. In some embodiments, the method further comprises: inputting, into the one or more second filters in the second stage of the multi-stage filtering process, supplemental information associated with the set of pixels, the supplemental information comprising one or more of: raw pixel data, information regarding how the set of pixels were produced, a quantization parameter (QP), or a coding mode.

[0008] In some embodiments, the one or more first filters comprise one or more of: a deblocking (DB) filter, a sample adaptive offset (SAO) loop filter, a bilateral filter (BIF), a residual filter (RF), a Gaussian filter (GF), transform-domain filtering, a fixed filter, a plurality of cascading fixed filters, a spatial filter, a texture-based filter, a band-based filter, a residual-based filter, a fixed-filter-output filter, a residuals-before-DB filter, a pre-smoothing filter, or a Laplacian filter.

[0009] In some embodiments, the method further comprises: generating a weighted average of the one or more first filter outputs and the NN-base filter output; and inputting the weighted average of the one or more first filter outputs and the NN-based filter output into the one or more adaptive filters in the second stage of the multi-stage filtering process.

[0010] In some embodiments, the method further comprises: inputting, into at least one of the one or more NN-based filters, the information about the set of pixels and supplemental information associated with the set of pixels, the supplemental information comprising one or more of: one or more values indicating one or more coding modes used in the first phase of the multi-phase filtering process, coding modes input channel (IPB) values, parameters of a loop filtering (BS) input channel, or quantization parameters (QP). In some embodiments, the supplemental information comprises one or more of: an intra coded mode identifier, an inter coded mode identifier, a prediction type identifier, a uni-hypothesis prediction identifier, a bi-hypothesis prediction identifier, a multi-hypothesis prediction identifier, a geometric prediction mode (GPM) identifier, an intra block copy (IBC) mode identifier, an intra template matching (IntraTMP) identifier, a palette mode (PLT) identifier, a bilateral filtering identifier, a deblocking filtering identifier, a transform domain filtering identifier, a Hadamard filter filtering identifier, a motion compensation filtering identifier, an interpolation filtering identifier, an overlapped block motion compensation (OBMC) identifier, an affine mode (AFFINE) identifier, one or more identifiers or parameters of in-loop filtering utilized during the multi-phase filtering process, one or more identifiers or parameters of a deblocking filter (DB) utilized during the multi-phase filtering process, one or more identifiers or parameters of a boundary strength associated with the set of pixels, one or more identifiers or parameters of a block strength associated with the set of pixels, an interpolation filter rescaling ratio, and / or one or more identifiers or parameters of models for intra-template matching (IntraTMP) selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a cross-component reconstruction mode, a convolutional cross-component model, a multi-cross component linear model.

[0011] In some embodiments, the one or more NN-based filters have a filter support area of NxM, where N and M are integer numbers. In some embodiments, N and M are predefined values ranging from 0 to 17. In some embodiments, at least one of the one or more first filters has a shapeselected from among: a rectangular, a diamond, or a cross. In some embodiments, a tap shape of the one or more NN-based filters has a height and width of NxN.

[0012] In some embodiments, the method further comprises: determining whether NN-based filtering is enabled for the set of pixels; and, in an instance in which NN-based filtering is enabled for the set of pixels, inputting one or more NN-based filter outputs with the information about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more first filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process. In some embodiments, the method further comprises: determining whether NN-based filtering is enabled for the set of pixels; and, in an instance in which NN-based filtering is not enabled for the set of pixels, inputting the one or more first filter outputs with the information about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more NN-based filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

[0013] In some embodiments, the one or more first filters comprise a DB filter, a sample adaptive offset (SAO) filter, and a bilateral filter (BIF). In some embodiments, the method further comprises, in an instance in which NN-based filtering is not enabled for the set of pixels, filtering the information associated with the set of pixels sequentially using the DB filter, the SAO filter, and the BIF filter of the one or more first filters in the first stage of the multi-stage filtering process; and inputting, into the one or more adaptive filters in the second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first outputs from the one or more first filters.

[0014] In some embodiments, the information transmitted in the bitstream about the one or more adaptive filter parameters used during the multi-stage filtering process comprises one or more of: a filter coefficient, a pixel value offset, or clipping range information.

[0015] According to another aspect of the present disclosure, a method can be provided or carried out that comprises: providing or receiving a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes; providing or receiving, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value from among the plurality of codeword values; and identifying, using the at least one codeword value in the informationassociated with the set of pixels, based on the mapping of the plurality of coding modes to the plurality of codeword values, at least one coding mode of the plurality of coding modes used during a multi-stage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter.

[0016] In some embodiments, the method can further comprise: extending the mapping to include one or more additional codeword values and one or more additional coding modes corresponding respectively to the one or more additional codeword values, the one or more additional coding modes being usable by, or associated with, the one or more NN-based filters in the first stage of the multi-stage filtering process. In some embodiments, the method can further comprise: modifying the mapping to associate one or more of the plurality of codeword values with one or more replacement coding modes, the one or more replacement coding modes being usable by, or associated with, the one or more NN-based filters in the first stage of the multi-stage filtering process. In some embodiments, the one or more additional coding modes or the one or more replacement coding modes comprise one or more of: a palette (PLT) mode, an intra template matching (intraTMP) mode, a geometric prediction mode (GPM), a bilateral filtering mode, a unihypothesis prediction mode, a bi-hypothesis prediction mode, a multi-hypothesis prediction mode, an intra block copy (IBC) mode, an overlapped block motion compensation (OBMC) mode, or an affine mode. In some embodiments, the one or more additional coding modes are indicated using one or more of: cgiTools, cu. preMode, MODE IBC, MODE PLT, MODE INTRA, Cu.tmpFlag, gpmTools, interGPM, sgpmFlag, ibcGPMFlag, MODE OBMC, gpmOBMC, cu.isobmcMC, bifValue, bifEnableFlag, bilateral filter strength, bsValueOut, bsValueln, bifid, bitOffsets, MODE AFFINE, affinelD, or cu. affine.

[0017] According to another aspect of the present disclosure, a method can be provided or carried out that comprises: providing or receiving a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes; providing or receiving, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value from among the plurality of codeword values; identifying, using the at least one codeword value in the information associated with the set of pixels, based on the mapping of the plurality of coding modes to theplurality of codeword values, at least one coding mode of the plurality of coding modes used during a multi-stage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter; and determining, based on the at least one coding mode associated with the at least one codeword value in the information associated with the set of pixels, whether the at least one coding mode comprises an aggregated coding mode, the aggregated coding mode indicating use of either a first coding mode or a second coding mode during the multi-stage filtering process.

[0018] In some embodiments, the method can further comprise, in an instance in which the at least one coding mode indicated by the at least one codeword value in the information associated with the set of pixels comprises an aggregated coding mode indicating use of either a first coding mode or a second coding mode during the multi-stage filtering process, executing a logical operation to determine whether the at least one codeword value indicates the first coding mode or the second coding mode. In some embodiments, one or more of the first coding mode or the second coding mode comprises one or more of: a palette (PLT) mode, an intra template matching (intraTMP) mode, a geometric prediction mode (GPM), a bilateral filtering mode, a uni-hypothesis prediction mode, a bi-hypothesis prediction mode, a multi-hypothesis prediction mode, an intra block copy (IBC) mode, an overlapped block motion compensation (OBMC) mode, or an affine mode. In some embodiments, the one or more additional coding modes are indicated using one or more of: cgiTools, cu.preMode, MODE IBC, MODE PLT, MODE INTRA, Cu.tmpFlag, gpmTools, interGPM, sgpmFlag, ibcGPMFlag, MODE OBMC, gpmOBMC, cu.isobmcMC, bifValue, bifEnableFlag, bilateral filter strength, bsValueOut, bsValueln, bifid, bitOffsets, MODE AFFINE, affinelD, or cu. affine.

[0019] According to another aspect of the present disclosure, a method can be provided or carried out that comprises: inputting, into at least one of one or more first filters or one or more neural network-based (NN-based) filters in a first stage of a multi-stage filtering process, information associated with a set of pixels to generate one or more first filter outputs; inputting, into an adaptive filter in a second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first filter outputs; generating or receiving filter parameter information associated with the multi-stage filtering process; and transmitting orreceiving, in a bitstream, the filter parameter information associated with the multi-stage filtering process. In some embodiments, the filter parameter information comprises one or more of: coding modes input channel (IPB), parameters of the loop filtering (BS) input channel, or quantization parameters (QP).

[0020] According to another aspect of the present disclosure, a method can be provided or carried out that comprises: providing or receiving, in a bitstream, information about a set of pixels and information about one or more filter parameters for a multi-stage filtering process associated with the set of pixels, wherein a first stage of the multi-stage filtering process comprises one or more first filters and / or one or more neural network-based (NN-based) filters, wherein a second stage of the multi-stage filtering process comprises one or more second filters, the one or more second filters comprising an adaptive filter; and, based on the information about the one or more filter parameters, determining whether the first stage of the multi-stage filtering process comprises the one or more first filters, the one or more NN-based filters, or both the one or more first filters and the one or more NN-based filters.

[0021] In some embodiments, the method can further comprise: reconstructing, based at least on the information about the one or more pixels and the information about the one or more filter parameters, the one or more pixels. In some embodiments, the information associated with the set of pixels comprises one or more of: raw pixel data, supplementary information about how the set of pixels were produced, a quantization parameter (QP), or one or more codeword values indicating one or more coding modes corresponding respectively to the one or more codeword values. In some embodiments, the one or more first filters comprise one or more of: a deblocking (DB) filter, a sample adaptive offset (SAO) loop filter, a bilateral filter (BIF), a residual filter (RF), a Gaussian filter (GF), transform-domain filtering, a fixed filter, a plurality of cascading fixed filters, a spatial filter, a texture-based filter, a band-based filter, a residual-based filter, a fixed-filter-output filter, a residuals-before-DB filter, a pre-smoothing filter, or a Laplacian filter. In some embodiments, an input to the adaptive filter in the multi-stage filtering process comprises a weighted average of one or more first filter outputs and an NN-base filter output. In some embodiments, the filter parameter information comprises one or more of: one or more values indicating one or more coding modes used in the first phase of the multi-phase filtering process, coding modes input channel (IPB) values, parameters of a loop filtering (BS) input channel, or quantization parameters (QP).

[0022] In some embodiments, the filter parameter information comprises one or more of: an intra coded mode identifier, an inter coded mode identifier, a prediction type identifier, a unihypothesis prediction identifier, a bi-hypothesis prediction identifier, a multi-hypothesis prediction identifier, a geometric prediction mode (GPM) identifier, an intra block copy (IBC) mode identifier, an intra template matching (IntraTMP) identifier, a palette mode (PLT) identifier, a bilateral filtering identifier, a deblocking filtering identifier, a transform domain filtering identifier, a Hadamard filter filtering identifier, a motion compensation filtering identifier, an interpolation filtering identifier, an overlapped block motion compensation (OBMC) identifier, an affine mode (AFFINE) identifier, an adaptive loop filtering identifier, one or more identifiers or parameters of in-loop filtering utilized during the multi-phase filtering process, one or more identifiers or parameters of a deblocking filter (DB) utilized during the multi-phase filtering process, one or more identifiers or parameters of a boundary strength associated with the set of pixels, one or more identifiers or parameters of a block strength associated with the set of pixels, an interpolation filter rescaling ratio, and / or one or more identifiers or parameters of models for intra-template matching (IntraTMP) selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a cross-component reconstruction mode, a convolutional cross-component model, a multi-cross component linear model.

[0023] In some embodiments, the one or more NN-based filters have a filter support area of NxM, where N and M are integer numbers. In some embodiments, N and M are predefined values ranging from 0 to 17. In some embodiments, at least one of the one or more first filters has a shape selected from among: a rectangular, a diamond, or a cross. In some embodiments, a tap shape of the one or more NN-based filters has a height and width of NxN. In some embodiments, the method can further comprise: determining whether NN-based filtering is enabled for the set of pixels; and, in an instance in which NN-based filtering is enabled for the set of pixels, inputting one or more NN-based filter outputs with the information about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more first filter outputs into the one or more adaptive filters in the second stage of the multistage filtering process. In some embodiments, the method can further comprise: determining whether NN-based filtering is enabled for the set of pixels; and, in an instance in which NN-based filtering is not enabled for the set of pixels, inputting the one or more first filter outputs with theinformation about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more NN-based filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

[0024] In some embodiments, the one or more first filters comprise a DB filter, a sample adaptive offset (SAO) filter, and a bilateral filter (BIF). In some embodiments, the method can further comprise, in an instance in which NN-based filtering is not enabled for the set of pixels, filtering the information associated with the set of pixels sequentially using the DB filter, the SAO filter, and the BIF filter of the one or more first filters in the first stage of the multi-stage filtering process; and inputting, into the one or more adaptive filters in the second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first outputs from the one or more first filters. In some embodiments, the information transmitted in the bitstream about the one or more adaptive filter parameters used during the multi-stage filtering process comprises one or more of: a filter coefficient, a pixel value offset, or clipping range information. In some embodiments, the multi-stage filtering process comprises one or more models for intra-template matching (IntraTMP) selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a crosscomponent reconstruction mode, a convolutional cross-component model, or a multi-cross component linear model. In some embodiments, the one or more coding modes are indicated, in a mapping of coding modes to codeword values, using one or more of: cgiTools, cu. preMode, MODE IBC, MODE PLT, MODE INTRA, Cu.tmpFlag, gpmTools, interGPM, sgpmFlag, ibcGPMFlag, MODE OBMC, gpmOBMC, cu.isobmcMC, bifValue, bifEnableFlag, bilateral filter strength, bsValueOut, bsValueln, bifid, bitOffsets, MODE AFFINE, affinelD, or cu. affine.

[0025] In some embodiments, the multi-stage filtering process is, comprises, or is comprised in, a multi-phase adaptive loop filter (ALF) process. In some embodiments, the multi-stage filtering process is, comprises, or is comprised in, an enhanced compression model (ECM).

[0026] According to other aspects of the present disclosure, an apparatus can be provided that is configured to perform at least a part of at least one of the methods described herein. For example, an apparatus can comprise means for performing one or more elements or operations of a method, such as those described herein. In some embodiments, the means can comprise at least one processor and at least one memory. In some embodiments, the at least one memory can compriseinstructions, program codes, computer program codes, programs, computer-executable codes, or the like stored therein. In some embodiments, the instructions stored in the at least one memory, when executed by the at least one processor, can cause the apparatus to perform at least one element or operation of a method, such as those described herein. In some embodiments, the apparatus is, comprises, or is comprised in, an encoder device. In some embodiments, the apparatus is, comprises, or is comprised in, a decoder device.

[0027] According to other aspects of the present disclosure, a computer program product can be provided that is or comprises a storage medium. For example, a computer program product can be provided that comprises at least one non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium can comprise instructions, program codes, computer program codes, programs, computer-executable codes, or the like stored therein. In some embodiments, the instructions stored in the non-transitory computer-readable storage medium, when executed by at least one processor of an apparatus, can cause the apparatus to perform at least one element or operation of a method, such as those described herein. In some embodiments, the apparatus is, comprises, or is comprised in, an encoder device. In some embodiments, the apparatus is, comprises, or is comprised in, a decoder device.

[0028] This summary is intended to provide a brief overview of some of the aspects and features according to the subject disclosure. Accordingly, it will be appreciated that the abovedescribed features are merely examples and should not be construed to narrow the scope of the subject disclosure in any way. Other features, aspects, and advantages of the subject disclosure will become apparent from the following detailed description, drawings and claims.Brief Description of the Drawings

[0029] A better understanding of the subject disclosure may be obtained when the following detailed description of various embodiments is considered in conjunction with the following drawings, in which:

[0030] FIG. 1 shows a block flow diagram illustrating a process for video encoding and decoding, in accordance with one or more embodiments of the present disclosure;

[0031] FIG. 2 shows schematically an example of a system for video encoding and decoding, in accordance with one or more embodiments of the present disclosure;

[0032] FIG. 3 shows schematically an example of a decoder-side device configured for carrying out video decoding, in accordance with one or more embodiments of the present disclosure;

[0033] FIG. 4 shows schematically an example of an encoder-side device configured for carrying out video encoding, in accordance with one or more embodiments of the present disclosure;

[0034] FIG. 5 is a block flow diagram illustrating a multi-stage filtering process, in accordance with one or more embodiments of the present disclosure;

[0035] FIG. 6 illustrates various tap shapes of various fixed filters, in accordance with one or more embodiments of the present disclosure;

[0036] FIG. 7 illustrates various tap shapes of various fixed filters and a process for aggregating or weight averaging across the various fixed filters, in accordance with one or more embodiments of the present disclosure;

[0037] FIG. 8 illustrates various tap shapes of various fixed filters, in accordance with one or more embodiments of the present disclosure;

[0038] FIG. 9 is a block flow diagram illustrating a multi-stage filtering process, in accordance with one or more embodiments of the present disclosure;

[0039] FIG. 10 is a block flow diagram illustrating a multi-stage filtering process, in accordance with one or more embodiments of the present disclosure;

[0040] FIG. 11 is a block flow diagram illustrating a multi-stage filtering process, in accordance with one or more embodiments of the present disclosure;

[0041] FIG. 12 is a block flow diagram illustrating a multi-stage filtering process, in accordance with one or more embodiments of the present disclosure;

[0042] FIG. 13 is a flowchart of operations according to certain example embodiments implemented, for example, by an apparatus as described herein;

[0043] FIG. 14 is a flowchart of operations according to certain example embodiments implemented, for example, by an apparatus as described herein;

[0044] FIG. 15 is a flowchart of operations according to certain example embodiments implemented, for example, by an apparatus as described herein;

[0045] FIG. 16 is a flowchart of operations according to certain example embodiments implemented, for example, by an apparatus as described herein; and

[0046] FIG. 17 is a flowchart of operations according to certain example embodiments implemented, for example, by an apparatus as described herein.Detailed Description

[0047] The following abbreviations that may be found in the specification and / or the drawing figures are defined as follows:

[0048] 4G fourth generation

[0049] 5G fifth generation

[0050] 5GC 5G core network

[0051] ANN artificial neural network

[0052] ALF adaptive loop filter

[0053] BIF bilateral filter

[0054] CCALF cross-component ALF

[0055] CCSAO cross-component SAO

[0056] CDMA code division multiple access

[0057] CPU central processing unit CTU coding tree unit

[0058] CU coding unit

[0059] DSP digital signal processor

[0060] ECM enhanced compression model

[0061] eNB / eNodeB evolved Node B (e.g., an LTE base station)

[0062] E-UTRA evolved universal terrestrial radio access, i.e., the LTE radio access technology

[0063] GDR gradual decoding refresh gNB (or gNodeB) base station for 5G / NR, i.e., a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC

[0064] FDMA frequency division multiple access

[0065] GPU graphical processing unit

[0066] GSM global systems for mobile communications

[0067] HMD head-mounted display

[0068] IEEE Institute of Electrical and Electronics Engineers

[0069] IMD integrated messaging device

[0070] IMS instant messaging service

[0071] loT Internet of Things

[0072] JVET joint video experts team

[0073] LTE long term evolution

[0074] MMS multimedia messaging service

[0075] MPEG-I Moving Picture Experts Group immersive codec family

[0076] MV motion vector

[0077] NG new / next generation

[0078] NG-eNB new / next generation eNB

[0079] NR new radio

[0080] PC personal computer

[0081] PDA personal digital assistant

[0082] QP quarter pixel

[0083] SAO sample adaptive offset

[0084] SMS short messaging service

[0085] SPS sequence parameter set

[0086] TCP-IP transmission control protocol-internet protocol

[0087] TDMA time division multiple access

[0088] UICC universal integrated circuit card

[0089] UMTS universal mobile telecommunications system

[0090] VVC versatile video coding

[0091] The examples and embodiments set forth below represent information to enable those skilled in the art to practice the subject disclosure. Upon reading the following description in light of the accompanying drawing figures, those skilled in the art will understand the concepts of the description and will recognize applications of these concepts not particularly addressed herein. It should be understood that these concepts and applications fall within the scope of the description.

[0092] In the following description, numerous specific details are set forth. However, it is understood that embodiments may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques have not been shown in detail in order not to obscure the understanding of the description. Those of ordinary skill in the art, with the included description, will be able to implement appropriate functionality without undue experimentation.

[0093] The following embodiments are examples. Although the specification may refer to “an”, “one”, or “some” embodiment(s) in several locations of the text, this does not necessarily mean that each reference is made to the same embodiment(s), or that a particular feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments. Further, when a particular feature, structure, or characteristic is described in connection of an embodiment, it is within the knowledge of one skilled in the art to apply such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described. It shall be understood that although the terms “first,” “second” and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0094] For the purposes of the present disclosure, the phrases “at least one of A or B”, “at least one of A and B”, and “A and / or B” means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

[0095] As used herein, "plurality" means two or more. As used herein, a "set" of items may include one or more of such items. As used herein, whether in the subject disclosure or the claims, the terms "comprising", "including", "carrying", "having", "containing", "involving", and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of' and "consisting essentially of', respectively, are closed or semiclosed transitional phrases with respect to claims. Use of ordinal terms such as "first", "second", "third", etc., in the claims or the subject disclosure to modify an element does not by itself connote any priority, precedence, or order of one element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one element having a certain name from another element having a same name (but for use of the ordinal term) to distinguish the elements. As used herein, "and / or" and "at least one of' means that the listed items are alternatives, but the alternatives also include any combination of the listed items.

[0096] Like reference numerals refer to like elements throughout. As used herein, the terms “data,” “content,” “information,” and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with an embodiment of the present disclosure. Thus, use of any such terms should not be taken to limit the spirit and scope of one or more embodiments of the present disclosure.

[0097] As used herein, the term ‘circuitry’ refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and / or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even if the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and / or portion(s) thereof and accompanying software and / or firmware. As defined herein, a “computer-readable storage medium,” which refers to a physical storage medium (e.g., volatile or non-volatile memory device), may be differentiated from a “computer-readable transmission medium,” which refers to an electromagnetic signal.

[0098] Systems, methods, apparatuses, and computer program products for providing media content based on context are provided in accordance with one or more embodiments of the subject disclosure. The systems, methods, apparatuses, and computer program products may be utilized in conjunction with a variety of video formats including High Efficiency Video Coding standard (HEVC or H.265 / HEVC), Advanced Video Coding standard (AVC or H.264 / AVC), Versatile Video Coding standard (WC or H.266 / VVC), and / or with a variety of video and multimedia file formats including International Standards Organization (ISO) base media file format (ISO / IEC 14496-12, which may be abbreviated as ISOBMFF), Moving Picture Experts Group (MPEG)-4 file format (ISO / IEC 14496-14, also known as the MP4 format), and file formats for NAL (Network Abstraction Layer) unit structured video (ISO / IEC 14496-15) and 3rdGeneration Partnership Project (3GPP file format) (3GPP Technical Specification 26.244, also known as the 3GP format). ISOBMFF is the base for derivation of all the above-mentioned file formats.

[0099] Some key definitions, bitstream and coding structures, and concepts of some video coding standards and specifications are described in this section for providing context relative to video encoders, video decoders, video encoding methods, video decoding methods, and bitstream structures, wherein the embodiments may be implemented. It is to be understood that embodiments are not limited to the referenced video coding standards or specifications.

[0100] An elementary unit for the input to an encoder and the output of a decoder, respectively, in many cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoder may be referred to as a decoded picture or a reconstructed picture.

[0101] The source and decoded pictures are each comprised of one or more sample arrays. The sample arrays of a picture may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr; regardless of the actual color representation method in use. The actual color representation method in use can be indicated e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax of HEVC or alike. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.

[0102] Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays. Chroma formats comprise monochrome format and non-monochrome formats, and these may be summarized as follows:

[0103] In monochrome sampling there is only one sample array, which may be nominally considered the luma array. In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array. In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.

[0104] Samples of a sample array have a certain bit depth, such as 8 bits per sample or 10 bits per sample. A bit depth implicitly specifies a value range, which may be referred to as the full range. For example, the full range is from 0 to 255, inclusive, for 8 bits per sample, or from 0 to 1023, inclusive, for 10 bits per sample. The source video may use allocate a narrower sample value range than the full range. A specific value range, sometimes referred to as the studio range, has been specified in the ITU-T H.273 standard specifying coding-independent code points for video. A source value range may interchangeably be referred to as a source sample value range, and may be defined as the sample value range of the video that is given as input to a video encoder to be encoded.

[0105] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternatesample rows of a frame and may be used as encoder input, when the source signal is interlaced. A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream. A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.

[0106] Syntax structures may be specified, for example, using arithmetic, logical, relational, bit-wise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise 'AND' operation. Furthermore, syntax structures may be specified with reference to mathematical functions.

[0107] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper-case letter and without any underscore characters. Variables starting with an upper-case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.

[0108] Video coding specifications may define an elementary unit that for the output an of an encoder and / or for the input to a decoder. For example, such an elementary unit may be an open bitstream unit (OBU), as specified e.g. in AVI, or a Network Abstraction Layer (NAL) unit, as specified e.g., in HEVC or VVC.

[0109] In some video codecs, an elementary unit for the output of an encoder and the input of a decoder, respectively, may be a Network Abstraction Layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bitstream format has been specified in some video coding standards for transmission or storage environments that do not provide framing structures. The bitstream format separates NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders run a byte-oriented start code emulation prevention algorithm, which adds an emulation prevention byte to the NAL unit payload if a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet- and stream-oriented systems, start code emulation prevention may always beperformed regardless of whether the bitstream format is in use or not. A NAL unit may be defined as a syntax structure containing an indication of the type of data to follow and bytes containing that data in the form of an RBSP interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure containing an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits containing syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.

[0110] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.

[0111] In some coding formats or standards, a bitstream may be in the form of a NAL unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.

[0112] Certain video coding system related to video coding and / or video decoding are explained with reference to FIGs. 1-4.

[0113] FIG. 1 shows a process 100 in which video compression is applied. The process 100 can comprise generating, recording, rendering, receiving, retrieving, or otherwise providing original video data 101. The original video data 101 can be encoded by a video encoder 102, using, e.g., one or more algorithms. Algorithms, such as Discrete Cosine Transform-based video compression algorithms, e.g., MPEG-2, MPEG-4, H.263, and H.264, can be used by the video encoder 102 to encode the original video data 101. The output from the video encoder 102 is compressed video data 103. Compressed video data 103 is sent to a network 104 that provides the compressed video data 105 to a video decoder 106. The video decoder 106 decodes the compressed video data 105 to generate decoded video data 107, which is approximately equivalent to the original video data 101.

[0114] The video encoder 102 compresses the original video data 101 in such a way that the compressed video data 103 does not exceed an available bandwidth of the network 104 in order for the video decoder 106 to be able to receive and decode the compressed video data 105.However, communication bandwidth may vary depending on the type of the network 104. For example, the available communication bandwidth of an Ethernet is different from that of a wirelesslocal area network (WLAN). The network 104, which may be e.g., a cellular communication network, may have a very narrow bandwidth. Thus, it can be important to generate compressed video data 103 at various bit-rates from the same original video data 101, such as by using scalable video coding. Scalable video coding is a video compression technique that allows video data to provide scalability. Scalability is the ability to generate video sequences at different resolutions, frame rates, and qualities from the same compressed bitstream. In some embodiments, the video encoder 102 can achieve temporal scalability can be provided using, e.g., Motion Compensation Temporal filtering (MCTF), Unconstrained MCTF, Successive Temporal Approximation and Referencing, and / or the like. In some embodiments, the video encoder 102 can achieve Signal-to-Noise Ratio (SNR) scalability or Signal-to-Noise-plus-Interference Ratio (SNIR) scalability using, e.g., Embedded ZeroTrees Wavelet (EZW), Set Partitioning in Hierarchical Trees (SPIHT), Embedded ZeroBlock Coding (EZBC), Embedded Block Coding with Optimized Truncation (EBCOT), etc. In some embodiments, the video encoder 102 can transmit only a portion of a scene, image, or picture need be transmitted as compressed video data 105 to the video decoder 106, which may improve bit-rate efficiency of video compression / coding.

[0115] In some embodiments, the video encoder 102 can achieve spatial scalability by using, e.g., a wavelet transform algorithm or multi-layer coding. For example, in some embodiments the video encoder 102 can use a multi-layer bitstream to transmit different portions of a scene, image, picture, inlay, overlay, background imagery, and / or the like, as separate layers, overlays, textures, alphas, or the like, to the video decoder 106 via the network 104.

[0116] When the video encoder 102 uses a multi-layer bitstream to transmit different portions of a scene, image, picture, inlay, overlay, background imagery, and / or the like, as separate layers, overlays, textures, alphas, or the like, to the video decoder 106 via the network 104, the video encoder 102 must typically provide metadata before, with, or after transmitting a portion or component of the multi-layer bitstream. For example, metadata may include bitstream information, such as attributes of a frame, overlay, layer, texture, alpha, or the like. In some embodiments, the metadata can be provided in a network abstraction layer (NAL) unit, e.g., in accordance with the H.264 / AVC and HEVC video coding standards, the entire disclosures of which are hereby incorporated herein by reference in their entireties for all purposes.

[0117] Among other elements of the metadata, supplemental enhancement information (SEI) can be provided as additional data inserted into the multi-layer bitstream to convey informationthat may be helpful for the network 104 to properly transmit the compressed video data 105 to the video decoder 106. Additionally or alternatively, SEI can be provided as additional data inserted into the multi-layer bitstream to convey information that may be helpful for the video decoder 106 to properly synchronize related audio and video content from the multi-layer bitstream, properly orient the video content from among the layers, determine a proper texture overlay order, determine characteristics about each layer such as whether a texture overlay / layer includes displayable content in every portion / region of the texture overlay / layer, etc.

[0118] In some embodiments, the video encoder 102 may insert at least a content item and one or more context identifiers into a file associated with the original video data 101 during encoding and / or transmission of one or more portions of the original video data 101. The file may be configured with a predefined digital container format for media content. For example, the file may be a HEIF file, an ISOBMFF file, or another type of container file data structure for one or more portions of the original video data 101. The content item may include one or more portions of the original video data 101. For example, the content item may include media content such as video content, audio content, one or more images, a sequence of images, and / or other media content. In some embodiments, a context identifier comprises a capture session identifier for a capture session associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises an event identifier for an event associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises a presentation identifier for a media presentation associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises a global session identifier associated with camera extrinsic matrix information for capture of one or more portions of the original video data 101. In some embodiments, the video encoder 102 may insert other information into the file, such as

[0119] In some embodiments, the video decoder 106 may receive a file that includes at least a content item and one or more context identifiers during decoding and / or transmission of one or more portions of the compressed video data 105. For example, the video decoder 106 may identify one or more context identifiers in a file during decoding and / or transmission of one or more portions of the compressed video data 105. The file may be configured with a predefined digital container format for media content. For example, the file may be a HEIF file, an ISOBMFF file, or another type of container file data structure for one or more portions of the original videodata 101. The content item may include one or more portions of the original video data 101. For example, the content item may include media content such as video content, audio content, one or more images, a sequence of images, and / or other media content. In some embodiments, a context identifier comprises a capture session identifier for a capture session associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises an event identifier for an event associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises a presentation identifier for a media presentation associated with capture of one or more portions of the original video data 101. In some embodiments, a context identifier comprises a global session identifier associated with camera extrinsic matrix information for capture of one or more portions of the original video data 101. In some embodiments, the video decoder 106 may identify other information, such as one or more identifiers, classifiers, parameters, or the like in the file that provides information helpful to the video decoder 106 during decoding and reconstruction of, approximately, the original video data 101 from the compressed video data 105.

[0120] A multi-layer bitstream (such as for scalable coded video) may contain different layers comprising different image sequences, overlay sequences, texture sequences, and / or the like. For example, a scalable coded video bitstream may comprise different layers each containing different representations of an original image sequence. In one specific example, a first layer in the multilayer bitstream can contain a relatively lower quality version of an original image sequence, and a second layer in the multi-layer bitstream can contain a relatively higher quality version of the original image sequence. In a second specific example, a first layer in the multi-layer bitstream can contain a first image sequence of background images, and a second layer in the multi-layer bitstream can contain a second image sequence of foreground images to be overlayed over respective background images from the first image sequence of background images. Other examples will be readily apparent to those skilled in the art, such as more sophisticated examples that include a plurality of different image layers within the multi-layer bitstream, one or more alpha representations, initialization values for a target picture to be rendered (e.g., at the video decoder 106), a combination of these, and / or the like. In some embodiments, SEI can be provided in a SEI Raw Byte Sequency Payload as a detached NAL unit.

[0121] Various SEI messages can be used to indicate / signal various information between the video encoder 102 and the video decoder 106. For example, information about camera-capturedcontent, such as a shutter interval used, can be conveyed to the video decoder 106 using a Shutter Interval Information SEI message. Information such as parameters of annotated regions using bounding boxes can be communicated to the video decoder 106 using an Annotated Regions SEI message. Other information or parameters associated with video content being transmitted by the video decoder 106, such as content light level information, equirectangular projection information, fisheye information, color content volume information, color remapping information, motion-constrained tile set (MCTS) information, cubemap projection information, sphere rotation information, region-wise packing information, omnidirectional viewport information, SEI manifest information, SEI prefix information, and / or the like, can be communicated by one or more other SEI messages.

[0122] Some metadata and / or other enhancement information can be provided via in-band and / or SEI messages. For example, information about a recovery point in a group of pictures (GOP) or sequence of NAL units, can be indicated suing a presentation time stamp or decoding time stamp identifier, NAL unit identifier, frame number, etc. A Recovery Point SEI message can be used to provide such information. An I-frame / slice (intra-coded picture) can be provided before / betweenP-frames / slices (predicted pictures) and / or B-frames / slices (bidirectional predicted pictures) to indicate a latest recovery point as the latest decoded I-frame / slice. Under the H.264 video coding standard, in addition to I-frames / slices, P-frames / slices and / or B-frames / slices, switching I-frames / slices, switching P-frames / slices, and multi-frame motion estimation frames / slices can be provided. In other instances, a recovery point can be indicated in, e.g., the Recovery Point SEI message, that indicates a recovery point objective, system restore point, orphaned recovery point, recovery point chain, etc.

[0123] FIG. 2 illustrates a system 200, according to an embodiment, within which embodiments of the present invention can be utilized is shown. The system 200 comprises multiple communication devices which can communicate through one or more networks. The system 200 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.1

[0124] The system 200 may include both wired and wireless communication devices and / or electronic devices suitable for implementing select, various, or all of the embodiments described herein. For example, the system 200, as shown in FIG. 2, is illustrated as comprising a mobile network 210 and a representation of the internet 220. The mobile network 210 can be or comprise, e.g., a fourth generation (4G) network, a Long Term Evolution (LTE), a fifth generation (5G) network, a sixth generation (6G) network, and / or the like. Connectivity to the internet 220 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0125] The example communication devices shown in the system 200 may include, but are not limited to, an electronic device or apparatus, such as a mobile device 211, a user equipment 212, etc. User equipment 212 can be connected to the internet 220 by way of at least an access point 213, which may be or comprise a gNodeB (gNB), an eNodeB (eNB), base station, access network node, radio access network (RAN) node, and / or the like. The mobile device 211 can be connected to the internet 220 by way of at least a cell tower 214, e.g., via radio signaling 215 with the cell tower 214, short messaging service (SMS) with the cell tower 214, and / or the like.

[0126] Additionally or alternatively, the mobile device 211 and / or user equipment 212 can be connected to the internet 220 by way of a WiFi access point 216 or the like. Access point 213, cell tower 214, and / or WiFi access point 216 can be configured to communicate directly with the internet 220 or with the internet 220 by way of a network server 219, which can comprise, be comprised in, hosted on, or otherwise functionalized via any suitable network-side device. Such network-side devices can include, but are not limited to, a server, a computing device, a centralized processing unit (CPU), a graphics processing unit (GPU), a processor, processing circuitry, a controller, a network element, a virtualized network function, a mobility management entity (MME), a serving gateway (SGW), a packet data network (PDN) gateway (PGW), a home subscriber server (HSS), a public data network (PDN), an access and mobility management function (AMF), a user plane function (UPF), a data network (DN), an authentication server function (AUSF), a session management function (SMF), a network slice selection function (NSSF), a network exposure function (NEF), a network function repository function (NRF), a policy control function (PCF), a unified data management (UDM) function, anapplication function (AF), or any other suitable network-side device, element, function, hardware, device, etc.

[0127] In the system 200, electronic devices such as, e.g., 211, 212, 217, etc., may be stationary or mobile when carried by an individual who is moving. For example, the user equipment 212 can be or comprise a head-mounted display, a body- worn display, an immersive gaming system, a smartphone, a laptop (e.g., 217), or the like. In other embodiments, electronic devices in the system 200, such as 211, 212, 217, can be mobile by virtue of being located in, mounted on, coupled to, or otherwise supported by a device configured for transportation, including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transport. In other embodiments, the computing device 218 can also be stationary or can be configured to be used while stationary and while mobile, or to be operationally or configurationally switched from a mobile use mode to a stationary use mode. For example, in embodiments in which the user equipment 212 is or comprises a head mounted display, the user equipment 212 may be effectively used by a user wearing the user equipment 212 while the user is stationary or while the user is moving.

[0128] Additionally or alternatively, electronic devices in the system 200 can be stationary. For example, the system 200 can comprise a computing device 218, which can be or comprise a gaming console, a desktop computer, a three-dimensional gaming system, a virtual reality display system, an augmented reality display system, an interactive-display system, an image projection system, and / or the like. In other embodiments, the computing device 218 can also be mobile or can be configured to be used while stationary and while mobile, or to be operationally or configurationally switched from a stationary use mode to a mobile use mode.

[0129] In some embodiments, one or more of the electronic devices in the system 200, e.g., one of 211, 212, 217, 218, may be or comprise a set-top box, a digital TV receiver, a device configured to transmit / receive streaming content or audio / video via a bitstream, etc., but which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware or software or combination of the encoder / decoder implementations, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0130] In some embodiments, certain electronic devices (e.g., 211, 212, 217, 218) in the system 200 may be configured to send and receive calls and messages and communicate withservice providers through a wireless connection, such as 215, to the cell tower 214 or the access point 213. The cell tower 214 and / or the access point 213 may be connected to the network server 219 that allows communication between the mobile network 210 and the internet 220. The system 200 may include additional communication devices and communication devices of various types, such as electronic devices 221, 222, and 223, which may be outside of the mobile network 210 but nevertheless connected to the internet 220, e.g., by way of a wired or wireless connection 224.

[0131] The various communication devices (e.g., 211, 212, 217, 218, 221, 222, 223) illustrated in the system 200 of FIG. 2 may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. A communications device involved in implementing various embodiments of the present invention may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0132] Among other transmissions between two or more of the communication devices (e.g., 211, 212, 217, 218, 221, 222, 223) illustrated in the system 200 of FIG. 2, video and / or audio transmissions can be carried out. In order to improve the efficiency and / or effectiveness of resource use, reduce bit-rate, reduce and / or improve signaling, and reduce transmission-side (TX-side) and / or receiver-side (RX-side) computational complexity, data to be transmitted can be compressed (i.e., encoded) at the TX-side and decoded at the RX-side. To carry out such video / audio data compression, a coder / decoder device (i.e., codec), or one or more codecs, can be used. In some embodiments, the video encoder 102 and / or video decoder 106 described above with reference to FIG. 1 can comprise at least one codec.

[0133] Real-time Transport Protocol (RTP) is widely used for real-time transport of timed media such as audio and video. RTP may operate on top of the User Datagram Protocol (UDP), which in turn may operate on top of the Internet Protocol (IP). RTP is specified in Internet Engineering Task Force (IETF) Request for Comments (RFC) 3550, available fromwww.ietf.org / rfc / rfc3550.txt. In RTP transport, media data is encapsulated into RTP packets. Each media type or media coding format may have a dedicated RTP payload format.

[0134] An RTP session is an association among a group of participants communicating with RTP. It is a group communications channel which can potentially carry a number of RTP streams. An RTP stream is a stream of RTP packets comprising media data.

[0135] Communication systems may include any number of media-aware network elements (MANEs). For example, many multipoint audio-visual conferences operate utilizing a centralized unit called Multipoint Control Unit (MCU). An MCU may implement the functionality of an RTP translator or an RTP mixer. An RTP translator may be a media translator that may modify the media inside the RTP stream. A media translator may for example decode and reencode the media content (i.e. transcode the media content). An RTP mixer is a middlebox that aggregates multiple RTP streams that are part of a session by generating one or more new RTP streams. An RTP mixer may manipulate the media data. One common application for a mixer is to allow a participant to receive a session with a reduced amount of resources compared to receiving individual RTP streams from all endpoints. A mixer can be viewed as a device terminating the RTP streams received from other endpoints in the same RTP session. Using the media data carried in the received RTP streams, a mixer generates derived RTP streams that are sent to the receiving endpoints. In another example, a MANE is a selective forward unit (SFU) that selectively forwards incoming RTP packets from one or more senders to one or more receivers.

[0136] According to some embodiments, a video codec can consist of an encoder that transforms the input video into a compressed representation suited for storage / transmission and / or a decoder that can uncompress the compressed video representation back into a viewable form. A video encoder (e.g., 102) and / or a video decoder (e.g., 106) may be combined within a singular device or can be separate from each other, i.e., need not form a codec within a singular device. Typically, a video encoder (e.g., 102) discards some information in the original video data 101 in order to represent the original video data 101 in a more compact form (that is, at a lower bit-rate).

[0137] Hybrid video encoders, for example many encoder implementations of ITU-T H.263 and H.264, often may encode the original video data 101 in two or more phases. According to some embodiments, pixel values in a certain picture area (or "block") are initially predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (usingthe pixel values around the block to be coded in a specified manner), and thereafter a prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This can be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT), a DCT algorithm, or a variant of the same), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the video encoder (e.g., 102) can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate), as reflected for example in the compressed video data 103 and / or the compressed video data 105.

[0138] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0139] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.

[0140] Referring now to FIG. 3, a video coding system or device is illustrated, according to an example embodiment, as a schematic block diagram of an exemplary apparatus, referred to herein as decoding device 300, which is configured to carry out at least a portion of the video compression / encoding / decoding processes and tasks described herein, e.g., 100.

[0141] The decoding device 300 may for example be configured to function as the video decoder 106. In other embodiments, the decoding device 300 can be, comprise, or be comprised within, e.g., mobile device 211, user equipment 212, computing device 217, or computing device 218 in the wireless network 210. In other embodiments, the decoding device 300 can be, comprise, or be comprised within a heads-up display, a head-mounted display, a gaming console,a user's computer, and / or the like. However, it will be appreciated that embodiments of the invention may be implemented within any electronic device or apparatus which may require encoding and decoding, or encoding or decoding, of video images, multi-layer bitstreams, scalable coded video, bitstreams comprising audio and video, bidirectional bitstreams, conversational content bitstreams, non-conversational content bitstreams, and / or the like.

[0142] The decoding device 300 may comprise a controller 301 in operable communication with a memory 302 and a radio interface 303. The decoding device 300 can be further may comprise a display 308, e.g., in the form of a liquid crystal display (LCD), light emitting diode (LED) display, organic LED (OLED) display, plasma display, Active-Matrix OLED (AMOLED) display, Quantum dot LED (QLED) display, micro-LED display, augmented reality display, virtual reality display, projected image display, any combination thereof, and / or the like. In other embodiments, the display 308 may be any other display technology suitable to display an image and / or video. The decoding device 300 may, optionally, further comprise a keypad 309. In other embodiments, any suitable data or user interface mechanism may be employed. For example, a user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0143] The decoding device 300 may comprise a microphone (not shown) or any suitable audio input which may be a digital or analogue signal input. The decoding device 300 may further comprise an audio output device which, in some embodiments, may be any one of: an earpiece, a speaker, or an analogue audio or digital audio output connection. The decoding device 300 may also comprise a battery (not shown). In other embodiments, the decoding device 300 may be powered by any suitable mobile energy device such as a solar cell, a fuel cell, a clockwork generator, etc. The decoding device 300 may, optionally, further comprise a camera 310 capable of recording or capturing images and / or video. The decoding device 300 may further comprise an infrared port (not shown) for short range line of sight communication to other devices. In other embodiments the decoding device 300 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0144] According to some embodiments, the controller 301 can comprise, e.g., a processor or the like configured for controlling at least some aspects, functionalities, equipment, subcomponents, or subsystems of the decoding device 300. The controller 301 may be connectedeither directly or indirectly to the memory 302 which, in some embodiments, may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 301. The controller 301 may further be connected to codec circuitry 305 and the codec circuitry 305 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for carrying out coding and / or decoding of audio and / or video data or assisting in coding and decoding carried out by the controller 301.

[0145] The decoding device 300 may further comprise a card reader (not shown) and / or a smart card (not shown), for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network (e.g., 210).

[0146] The decoding device 300 may comprise radio interface circuitry 303 connected to, or otherwise in operable communication with, the controller 301. The radio interface circuitry 303 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The decoding device 300 may further comprise an antenna array 304 connected to the radio interface circuitry 303 for transmitting radio frequency signals generated at the radio interface circuitry 303 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).

[0147] The antenna array 304 can be configured to receive radio signals comprising or representing the compressed / coded audio / video content from, e.g., an encoder-side device. The antenna array 304 can relay the radio signals to the radio interface 303, which can be configured to convert the radio signals to decodable information, which it then passes along to the codec circuitry 305. The codec circuitry 305 then decodes, e.g., with the aid of, and / or upon receiving instructions from, the controller 301. The codec circuitry 305 can then, once the decodable information is decoded, provide decoded information to the controller 301. The controller 301 can interpret the decoded information to synchronize the audio / video content, and otherwise determine howto reconstitute, build, reconstruct, render, overlay, display, emit, broadcast, and / or present the decoded audio, images, and / or video frames to one or more users, either directly on the decoding device 300 (e.g., via the display 308) or by transmitting / providing the decoded audio, images, and / or video frames to another device (e.g., 211, 212, 217, 218) for display thereon.

[0148] The decoding device 300 may, optionally, further comprise a camera 310 capable of recording or detecting individual frames which are then passed to the codec 305 or the controller 301 for processing. The decoding device 300 may receive the video image data for processing from another device prior to transmission and / or storage. The decoding device 300 may also receive either wirelessly or by a wired connection the image for coding / decoding. The incorporation of the camera 310 into / with the decoding device 300 can be helpful in certain circumstances when, e.g., the image and / or video content is or comprises virtual reality content or augmented reality content in which real world objects and imagery located about the decoding device 300 or other device (e.g., 211, 212, 217, 218) displaying thereon the content may need to record and return images and / or video to assist with rendering subsequent images and / or video frames of the content for the user(s).

[0149] Referring now to FIG. 4, a video coding system or device is illustrated, according to an example embodiment, as a schematic block diagram of an exemplary apparatus, referred to herein as an encoding device 400, which is configured to carry out at least a portion of the video compression / encoding / decoding processes and tasks described herein, e.g., 100.

[0150] The encoding device 400 may for example be configured to function as the video encoder 102. In other embodiments, the encoding device 400 can be, comprise, or be comprised within a network-side or encoder-side device, e.g., 219. However, it will be appreciated that the encoding device 400 can be, comprise, or be comprised within another device within the mobile network (e.g., 210) in which the decoding device 300 is located. In other embodiments, the encoding device 400 can be, comprise, or be comprised within a device operationally and / or physically located outside the system or network (e.g., 210) in which the decoding device 300 is located. For example, the encoding device 400 can be, comprise, or be comprised within, e.g., 221, 222, 223, or the like. However, it will be appreciated that embodiments of the invention may be implemented within any electronic device or apparatus which may require encoding and decoding, or encoding or decoding, of video images, multi-layer bitstreams, scalable coded video, bitstreams comprising audio and video, bidirectional bitstreams, conversational content bitstreams, non-conversational content bitstreams, and / or the like.

[0151] The encoding device 400 may comprise a controller 401 in operable communication with a memory 402 and a radio interface 403. The encoding device 400 can be further may comprise codec circuitry 405 in operable communication with one or both of the controller 401and / or the radio interface 403. The encoding device 400 can further comprise an antenna array 404 in operable communication with the radio interface 403.

[0152] In some embodiments, the encoding device 400 can be configured to capture audio and / or video content to be encoded and transmitted to a decoder-side device (e.g., 300). In such embodiments, the encoding device 400 can, optionally, comprise a camera 407, a microphone (not shown), and / or the like.

[0153] In other embodiments, the encoding device 400 can be configured to generate audio and / or video content to be encoded and transmitted to a decoder-side device (e.g., 300). In such embodiments, the encoding device 400 can, optionally, comprise a graphics generator 409.

[0154] In other embodiments, encoding device 400 can be configured to request, retrieve, or otherwise receive audio and / or video from one or more external devices, subcomponents, systems, etc. For example, the encoding device 400 can be configured to receive audio from an external microphone (not shown) or an external audio generation device (not shown). In some embodiments, the encoding device 400 can be configured to receive video from an external camera 408. Whether video content is captured by the camera 407 or by the external camera 408, these cameras are capable of recording or capturing images and / or video.

[0155] In some embodiments, the encoding device 400 can be configured to receive generated graphics or other rendered content from an external rendering device or graphics generating device, such as the graphics generator 409.

[0156] According to some embodiments, the controller 401 can be or comprise, e.g., a processor, processing circuitry, or the like, that is configured for controlling at least some aspects, functionalities, equipment, subcomponents, or subsystems of the encoding device 400. The controller 401 may be connected either directly or indirectly to the memory 402 which, in some embodiments, may store both data in the form of image and / or audio data and / or may also store instructions for implementation of the same using the controller 401. The controller 401 may further be connected to the codec circuitry 405 and the codec circuitry 405 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for carrying out coding and / or decoding of audio and / or video data or assisting in coding and decoding carried out by the controller 401.

[0157] The radio interface circuitry 403 of the encoding device 400 can be configured to be connected to, or otherwise in operable communication with, the controller 401. The radio interfacecircuitry 403 can be arranged and dimensioned for, operably programmed for, programmatically capable of, or otherwise suitably configured for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The encoding device 400 may further comprise the antenna 404 connected to the radio interface circuitry 403 for transmitting radio frequency signals generated at the radio interface circuitry 403 to other apparatus(es) and for receiving radio frequency signals from other apparatus(es).

[0158] The codec circuitry 405 of the encoding device 400 may be configured to receive images and / or video (e.g., as bitstream data) from the controller 401. The codec circuitry 405 can be further configured to encode / compress this image and / or video data, and optionally metadata for decoder-side use in decoding / interpreting the encoded / compressed image and / or video data. The codec circuitry 405 can then provide the encoded / compressed image and / or video data to the radio interface 403, which can convert the encoded / compressed image and / or video data to a form that can be provided / transmitted to a decoder-side device (e.g., 300) via radio signaling using the antenna array 404.

[0159] In order to ensure that the encoded / compressed image and / or video data being provided from, e.g., the encoder device 400 to the decoder device 300, is properly and synchronously encoded and decoded, standard means can be defined for how the encoded / compressed image and / or video data is encoded. Likewise, metadata that is associated with the encoded / compressed image and / or video data can likewise be provided between the encoder-side and the decoder-side (e.g., from the encoder device 400 to the decoder device 300).Standard means can also be defined for what the metadata associated with the encoded / compressed image and / or video data comprises, the form in which the metadata associated with the encoded / compressed image and / or video data is provided, how syntax used in the metadata associated with the encoded / compressed image and / or video data is selected and used, and / or how the metadata associated with the encoded / compressed image and / or video data is encoded / decoded.

[0160] In some embodiments, one or more context identifiers for a content item can be provided from the encoder device 400 as additional data inserted into a file and / or a bitstream (e.g., a multi-layer bitstream) to convey contextual information regarding capture of the content item that may be helpful for the decoder device 300 to properly decode and interpret the compressedvideo data 105 (e.g., encoded / compressed image and / or video data) received therewith from the encoder device 400. The one or more context identifiers can likewise be helpful for the decoder device 300, once the encoded / compressed image and / or video data is decoded, if the decoder device 300 generates displayable or renderable image(s) or video frame(s) from the encoded / compressed image and / or video data. Additionally and / or alternatively, the one or more context identifiers can be helpful for the decoder device 300 if the decoder device 300 generates data or information about the encoded / compressed image and / or video data that, when transmitted to a displaying device (e.g., 112, 113, 117, 118, etc.), enables the displaying device to generate and display displayable image(s) or video frame(s). Alternatively, the one or more context identifiers can be helpful for the decoder device 300 if the decoder device 300 generates data or information about the encoded / compressed image and / or video data that, when transmitted to a rendering device (e.g., 112, 113, 117, 118, etc.), enables the rendering device to render and display / present rendered imagery, rendered graphics, and / or rendered video frame(s).

[0161] Some or all of the elements, steps, or components of the approaches described herein can be carried out by a computing device or an apparatus comprising a processor and memory. Examples of such computing devices and apparatuses are described in more detail below. Referring now to both FIG. 3 and FIG. 4, various aspects related to the functionality of the decoder device 300 and / or the encoder device 400 and components / configurations thereof are described. Embodiments of the present invention can be implemented as an apparatus or device, such as described above with regard to one or more embodiments of the decoder device 300 and / or one or more embodiments of the encoder device 400. In other embodiments, the present invention can be implemented as a computer program product that is executable on a computing device - such as by execution of program codes or computer-readable instructions stored on at least one memory device.

[0162] Some embodiments of the present invention may be implemented in various other ways, such as an article of manufacture. One example of an article of manufacture in the context of the invention disclosed herein is a computer program product that includes one or more software components including, for example, software objects, methods, data structures, program codes, computer-readable instructions, application-specific software, and / or the like. A software component may be coded in any of a variety of programming languages. An illustrative programming language may be a lower-level programming language such as an assembly languageassociated with a particular hardware architecture and / or operating system platform. A software component comprising assembly language instructions may require conversion into executable machine code by an assembler prior to execution by the hardware architecture and / or platform. Another example programming language may be a higher-level programming language that may be portable across multiple architectures. A software component comprising higher-level programming language instructions may require conversion to an intermediate representation by an interpreter or a compiler prior to execution.

[0163] Other examples of programming languages include, but are not limited to, a macro language, a shell or command language, a job control language, a script language, a database query or search language, and / or a report writing language. In one or more example embodiments, a software component comprising instructions in one of the foregoing examples of programming languages may be executed directly by an operating system or other software component without having to be first transformed into another form. A software component may be stored as a file or other data storage construct. Software components of a similar type or functionally related may be stored together such as, for example, in a particular directory, folder, or library. Software components may be static (e.g., pre-established or fixed) or dynamic (e.g., created or modified at the time of execution).

[0164] A computer program product may include a non-transitory computer-readable storage medium storing applications, programs, program modules, scripts, source code, program code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like (also referred to herein as executable instructions, instructions for execution, computer program products, program code, and / or similar terms used herein interchangeably). Such non-transitory computer-readable storage media include all computer-readable media (including volatile and non-volatile media).

[0165] In one embodiment, a non-volatile computer-readable storage medium may include a floppy disk, flexible disk, hard disk, solid-state storage (SSS) (e.g., a solid-state drive (SSD), solid state card (SSC), solid state module (SSM), enterprise flash drive, magnetic tape, or any other non-transitory magnetic medium, and / or the like. A non-volatile computer-readable storage medium may also include a punch card, paper tape, optical mark sheet (or any other physical medium with patterns of holes or other optically recognizable indicia), compact disc read only memory (CD-ROM), compact disc-rewritable (CD-RW), digital versatile disc (DVD), Blu-ray disc (BD), anyother non-transitory optical medium, and / or the like. Such a non-volatile computer-readable storage medium may also include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory (e.g., Serial, NAND, NOR, and / or the like), multimedia memory cards (MMC), secure digital (SD) memory cards, SmartMedia cards, CompactFlash (CF) cards, Memory Sticks, and / or the like. Further, a non-volatile computer-readable storage medium may also include conductive-bridging random access memory (CBRAM), phase-change random access memory (PRAM), ferroelectric random-access memory (FeRAM), non-volatile random-access memory (NVRAM), magnetoresistive randomaccess memory (MRAM), resistive random-access memory (RRAM), Silicon-Oxide-Nitride-Oxide-Silicon memory (SONOS), floating junction gate random access memory (FJG RAM), Millipede memory, racetrack memory, and / or the like.

[0166] In one embodiment, a volatile computer-readable storage medium may include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), fast page mode dynamic random access memory (FPM DRAM), extended data-out dynamic random access memory (EDO DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), double data rate type two synchronous dynamic random access memory (DDR2 SDRAM), double data rate type three synchronous dynamic random access memory (DDR3 SDRAM), Rambus dynamic random access memory (RDRAM), Twin Transistor RAM (TTRAM), Thyristor RAM (T-RAM), Zero-capacitor (Z-RAM), Rambus in-line memory module (RIMM), dual in-line memory module (DIMM), single in-line memory module (SIMM), video random access memory (VRAM), cache memory (including various levels), flash memory, register memory, and / or the like. It will be appreciated that where embodiments are described to use a computer-readable storage medium, other types of computer-readable storage media may be substituted for or used in addition to the computer-readable storage media described above.

[0167] As should be appreciated, various embodiments of the present invention may also be implemented as methods, apparatus, systems, computing devices, computing entities, and / or the like. As such, embodiments of the present invention may take the form of an apparatus, system, computing device, computing entity, and / or the like executing instructions stored on a computer-readable storage medium to perform certain steps or operations. Thus, embodiments of the presentinvention may also take the form of an entirely hardware embodiment, an entirely computer program product embodiment, and / or an embodiment that comprises combination of computer program products and hardware performing certain steps or operations.

[0168] In some embodiments, the decoding device 300 and / or the encoding device 400 according to one embodiment of the present invention. In general, the terms computing device, computing entity, computer, entity, device, system, and / or similar words used herein interchangeably may refer to, for example, one or more computers, computing entities, desktops, mobile phones, tablets, phablets, notebooks, laptops, distributed systems, kiosks, input terminals, servers or server networks, blades, gateways, switches, processing devices, processing entities, set-top boxes, relays, routers, network access points, base stations, the like, and / or any combination of devices or entities adapted to perform the functions, operations, and / or processes described herein. Such functions, operations, and / or processes may include, for example, transmitting, receiving, operating on, processing, displaying, storing, determining, creating / generating, monitoring, evaluating, comparing, and / or similar terms used herein interchangeably. In one embodiment, these functions, operations, and / or processes can be performed on data, content, information, and / or similar terms used herein interchangeably.

[0169] In some embodiments, the decoding device 300 and / or the encoding device 400 may include or be in communication with one or more processing elements (also referred to as processors, processing circuitry, and / or similar terms used herein interchangeably) that communicate with other elements within the decoding device 300 and / or the encoding device 400 via a bus, for example. As will be understood, the processing element of the decoding device 300 and / or the encoding device 400 may be embodied as one or more complex programmable logic devices (CPLDs), microprocessors, multi-core processors, coprocessing entities, applicationspecific instruction-set processors (ASIPs), microcontrollers, and / or controllers. Further, the processing element may be embodied as one or more other processing devices or circuitry. The term circuitry may refer to an entirely hardware embodiment or a combination of hardware and computer program products. Thus, the processing element may be embodied as integrated circuits, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), hardware accelerators, other circuitry, and / or the like. As will therefore be understood, the processing element may be configured for a particular use or configured to execute instructions stored in volatile or non-volatile media or otherwise accessibleto the processing element. As such, whether configured by hardware or computer program products, or by a combination thereof, the processing element may be capable of performing steps or operations according to embodiments of the present invention when configured accordingly.

[0170] In one embodiment, the decoding device 300 and / or the encoding device 400 may further include or be in communication with non-volatile media (also referred to as non-volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). In one embodiment, the non-volatile storage or memory may include one or more non-volatile storage or memory media, including but not limited to hard disks, ROM, PROM, EPROM, EEPROM, flash memory, MMCs, SD memory cards, Memory Sticks, CBRAM, PRAM, FeRAM, NVRAM, MRAM, RRAM, SONOS, FJGRAM, Millipede memory, racetrack memory, and / or the like. As will be recognized, the non-volatile storage or memory media may store databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like. The term database, database instance, database management system, and / or similar terms used herein interchangeably may refer to a collection of records or data that is stored in a computer-readable storage medium using one or more database models, such as a hierarchical database model, network model, relational model, entity-relationship model, object model, document model, semantic model, graph model, and / or the like.

[0171] In one embodiment, the decoding device 300 and / or the encoding device 400 may further include or be in communication with volatile media (also referred to as volatile storage, memory, memory storage, memory circuitry and / or similar terms used herein interchangeably). In one embodiment, the volatile storage or memory may also include one or more volatile storage or memory media 404, including but not limited to RAM, DRAM, SRAM, FPM DRAM, EDO DRAM, SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, RDRAM, TTRAM, T-RAM, Z-RAM, RIMM, DIMM, SIMM, VRAM, cache memory, register memory, and / or the like. As will be recognized, the volatile storage or memory media may be used to store at least portions of the databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like being executed by, for example, the processing element. Thus, the databases, database instances, database management systems, data,applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like may be used to control certain aspects of the operation of the decoding device 300 and / or the encoding device 400 with the assistance of the processing element and operating system.

[0172] In some embodiments, the decoding device 300 and / or the encoding device 400 may also include one or more network interfaces, such as a transceiver for communicating with various computing entities, such as by communicating data, content, information, and / or similar terms used herein interchangeably that can be transmitted, received, operated on, processed, displayed, stored, and / or the like. Such communication may be executed using a wired data transmission protocol, such as fiber distributed data interface (FDDI), digital subscriber line (DSL), Ethernet, asynchronous transfer mode (ATM), frame relay, data over cable service interface specification (DOCSIS), or any other wired transmission protocol. Similarly, the decoding device 300 and / or the encoding device 400 may be configured to communicate via wireless external communication networks using any of a variety of protocols, such as general packet radio service (GPRS), Universal Mobile Telecommunications System (UMTS), Code Division Multiple Access 2000 (CDMA2000), CDMA2000 IX (IxRTT), Wideband Code Division Multiple Access (WCDMA), Global System for Mobile Communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), Time Division-Synchronous Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Evolved Universal Terrestrial Radio Access Network (E-UTRAN), Evolution-Data Optimized (EVDO), High Speed Packet Access (HSPA), High-Speed Downlink Packet Access (HSDPA), IEEE 802.11 (Wi-Fi), Wi-Fi Direct, 802.16 (WiMAX), ultra-wideband (UWB), infrared (IR) protocols, near field communication (NFC) protocols, Wibree, Bluetooth protocols, wireless universal serial bus (USB) protocols, and / or any other wireless protocol.

[0173] Although not shown, the decoding device 300 and / or the encoding device 400 may include or be in communication with one or more input elements, such as a keyboard input, a mouse input, a touch screen / display input, motion input, movement input, audio input, pointing device input, joystick input, keypad input, and / or the like. The decoding device 300 and / or the encoding device 400 may also include or be in communication with one or more output elements (not shown), such as audio output, video output, screen / display output, motion output, movement output, and / or the like.

[0174] The signals provided to and received from the decoding device 300 and / or the encoding device 400 may include signaling information / data in accordance with air interface standards of applicable wireless systems. In this regard, the decoding device 300 and / or the encoding device 400 may be capable of operating with one or more air interface standards, communication protocols, modulation types, and access types. More particularly, the decoding device 300 and / or the encoding device 400 may operate in accordance with any of a number of wireless communication standards and protocols, such as those described above. In a particular embodiment, the decoding device 300 and / or the encoding device 400 may operate in accordance with multiple wireless communication standards and protocols, such as UMTS, CDMA2000, IxRTT, WCDMA, GSM, EDGE, TD-SCDMA, LTE, E-UTRAN, EVDO, HSPA, HSDPA, Wi-Fi, Wi-Fi Direct, WiMAX, UWB, IR, NFC, Bluetooth, USB, and / or the like. Similarly, the decoding device 300 and / or the encoding device 400 may operate in accordance with multiple wired communication standards and protocols, such as those described above, via a network interface.

[0175] Via these communication standards and protocols, the decoding device 300 and / or the encoding device 400 can communicate with various other entities using concepts such as Unstructured Supplementary Service Data (USSD), Short Message Service (SMS), Multimedia Messaging Service (MMS), Dual-Tone Multi-Frequency Signaling (DTMF), and / or Subscriber Identity Module Dialer (SIM dialer). The decoding device 300 and / or the encoding device 400 can also download changes, add-ons, and updates, for instance, to its firmware, software (e.g., including executable instructions, applications, program modules), and operating system.

[0176] According to one embodiment, the decoding device 300 and / or the encoding device 400 may include location determining aspects, devices, modules, functionalities, and / or similar words used herein interchangeably. For example, the decoding device 300 and / or the encoding device 400 may include outdoor positioning aspects, such as a location module adapted to acquire, for example, latitude, longitude, altitude, geocode, course, direction, heading, speed, universal time (UTC), date, and / or various other information / data. In one embodiment, the location module can acquire data, sometimes known as ephemeris data, by identifying the number of satellites in view and the relative positions of those satellites (e.g., using global positioning systems (GPS)). The satellites may be a variety of different satellites, including Low Earth Orbit (LEO) satellite systems, Department of Defense (DOD) satellite systems, the European Union Galileo positioning systems, the Chinese Compass navigation systems, Indian RegionalNavigational satellite systems, and / or the like. This data can be collected using a variety of coordinate systems, such as the Decimal Degrees (DD); Degrees, Minutes, Seconds (DMS); Universal Transverse Mercator (UTM); Universal Polar Stereographic (UPS) coordinate systems; and / or the like.

[0177] Alternatively, the location information / data can be determined by triangulating a position of the decoding device 300 and / or the encoding device 400 in connection with a variety of other systems, including cellular towers, Wi-Fi access points, and / or the like. Similarly, the decoding device 300 and / or the encoding device 400 may include indoor positioning aspects, such as a location module adapted to acquire, for example, latitude, longitude, altitude, geocode, course, direction, heading, speed, time, date, and / or various other information / data. Some of the indoor systems may use various position or location technologies including RFID tags, indoor beacons or transmitters, Wi-Fi access points, cellular towers, nearby computing devices (e.g., smartphones, laptops) and / or the like. For instance, such technologies may include the iBeacons, Gimbal proximity beacons, Bluetooth Low Energy (BLE) transmitters, NFC transmitters, and / or the like. These indoor positioning aspects can be used in a variety of settings to determine the location of someone or something to within inches or centimeters.

[0178] The decoding device 300 and / or the encoding device 400 may also comprise a user interface (that can include a display coupled to the processing element / controller) and / or a user input interface (coupled to the processing element / controller). For example, the user interface may be a user application, browser, user interface, and / or similar words used herein interchangeably executing on and / or accessible via the decoding device 300 and / or the encoding device 400 to interact with and / or cause display of information / data from the decoding device 300 and / or the encoding device 400, as described herein. The user input interface can comprise any of a number of devices or interfaces allowing the decoding device 300 and / or the encoding device 400 to receive data, such as a keypad (hard or soft), a touch display, voice / speech or motion interfaces, or other input device. In embodiments in which the decoding device 300 and / or the encoding device 400 comprises a keypad, the keypad can include (or cause display of) the conventional numeric (0-9) and related keys (#, *), and other keys used for operating the decoding device 300 and / or the encoding device 400 and may include a full set of alphabetic keys or set of keys that may be activated to provide a full set of alphanumeric keys. In addition to providing input, the userinput interface can be used, for example, to activate or deactivate certain functions, such as screen savers and / or sleep modes.

[0179] The decoding device 300 and / or the encoding device 400 can also include volatile storage or memory and / or non-volatile storage or memory, which can be embedded and / or may be removable. For example, the non-volatile memory may be ROM, PROM, EPROM, EEPROM, flash memory, MMCs, SD memory cards, Memory Sticks, CBRAM, PRAM, FeRAM, NVRAM, MRAM, RRAM, SONOS, FJGRAM, Millipede memory, racetrack memory, and / or the like. The volatile memory may be RAM, DRAM, SRAM, FPM DRAM, EDO DRAM, SDRAM, DDR SDRAM, DDR2 SDRAM, DDR3 SDRAM, RDRAM, TTRAM, T-RAM, Z-RAM, RIMM, DIMM, SIMM, VRAM, cache memory, register memory, and / or the like. The volatile and nonvolatile storage or memory can store databases, database instances, database management systems, data, applications, programs, program modules, scripts, source code, object code, byte code, compiled code, interpreted code, machine code, executable instructions, and / or the like to implement the functions of the decoding device 300 and / or the encoding device 400. As indicated, this may include a user application that is resident on the entity or accessible through a browser or other user interface for the decoding device 300 to communicate with the encoding device 400 and / or for the encoding device 400 to communication with the decoder device 300.

[0180] In another embodiment, the decoding device 300 and / or the encoding device 400 may include other components or functionalities. As will be recognized, these architectures and descriptions are provided for exemplary purposes only and are not limiting to the various embodiments.

[0181] In some embodiments, a file can be modified (e.g., by the encoder device 400) to provide additional information (e.g., one or more context identifiers) associated with capture of related content items that enables more efficient decoding and target picture display by the decoder device 300 or a display device on the decoder-side.

[0182] In some embodiments, the encoder device 400 can comprise means, such as at least one processor and at least one memory storing instructions stored thereon that, when executed by the at least one processor, cause the encoder device 400 to generate a file that comprises at least a portion of the original video data 101. The file may include information such as filtering protocols, in-line filters used, on-line filters used, etc.

[0183] In some embodiments, the encoder device 400 can comprise means, such as at least one processor and at least one memory storing instructions stored thereon that, when executed by the at least one processor, cause the encoder device 400 to

[0184] In some embodiments, the decoder device 300 can comprise means, such as at least one processor and at least one memory storing instructions stored thereon that, when executed by the at least one processor, cause the decoder device 300 to receive the file that comprises media content, such as the compressed video data 105. In some embodiments, the file can further comprise information such as filtering protocols used and the like.

[0185] In some embodiments, the decoder device 300 can comprise means, such as at least one processor and at least one memory storing instructions stored thereon that, when executed by the at least one processor, cause the decoder device 300 to

[0186] As such, some aspects of this disclosure relate to video codecs that consist of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, encoder discards some information in the original video sequence to be able to represent the video in a more compact form (that is, at lower bitrate).

[0187] Referring now to FIG. 5, an example encoding process is illustrated. Typical hybrid video codecs, such as H.264 / AVC, H.265 / HEVC and H.266 / VVC, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded pictures that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the resulting transform coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).

[0188] In some video codecs, such as H.265 / HEVC and H.266 / VVC, the video pictures are divided into coding units (CU) covering the area of the picture. A CU consists of one or more prediction units (PU) defining the prediction process for the samples within the CU and one ormore transform units (TU) defining the prediction error coding process for the samples in the said CU. Typically, a CU consists of a rectangular block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size is typically named as LCU (largest coding unit) or CTU (coding tree unit) and the video picture is divided into nonoverlapping CTUs. A CTU can be further split into a combination of smaller CUs, e.g. by recursively splitting the CTU and resultant CUs. Each resulting CU typically has at least one PU and at least one TU associated with it. Each PU and TU can be further split into smaller PUs and TUs to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs). Similarly, each TU is associated with information describing the prediction error decoding process for the samples within the TU (including e.g. DCT coefficient information). It is typically signaled at CU level whether prediction error coding is applied or not for each CU. In the case there is no prediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and TUs is typically signaled in the bitstream allowing the decoder to reproduce the intended structure of these units.

[0189] Referring now to FIG. 6, an example decoding process is illustrated. The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.

[0190] Instead, or in addition to approaches utilizing sample value prediction and transform coding for indicating the coded sample values, a color palette based coding can be used. Palette based coding refers to a family of approaches for which a palette, i.e., a set of colors and associated indexes, is defined and the value for each sample within a coding unit is expressed by indicatingits index in the palette. Palette based coding can typically achieve good coding efficiency in coding units with a relatively small number of colors (such as image areas which are representing computer screen content, like text or simple graphics). In order to improve the coding efficiency of palette coding, different kinds of palette index prediction approaches can be utilized, or the palette indexes can be run-length coded to be able to represent larger homogenous image areas efficiently. Also, in the case the CU contains sample values that are not recurring within the CU, escape coding can be utilized. Escape coded samples are transmitted without referring to any of the palette indexes. Instead, their values are indicated individually for each escape coded sample.

[0191] In typical video codecs, the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. To represent motion vectors efficiently, those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs, the predicted motion vectors are created in a predefined way, for example by calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging or merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0192] Typically, video codecs support motion compensated prediction from at least one source image (uni-prediction) and two sources (bi-prediction). In the case of uni-prediction, asingle motion vector is applied whereas in the case of bi-prediction two motion vectors are determined and the motion compensated predictions from two sources are combined to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions can be adjusted, or a signaled offset can be added to the prediction signal.

[0193] In addition to applying motion compensation for inter picture prediction, a similar approach can be applied to intra picture prediction. In this case, the displacement vector indicates where from the same picture a block of samples can be copied to form a prediction of the block to be coded or decoded. This kind of intra block copying (IBC) methods can improve the coding efficiency substantially in the presence of repeating structures within the frame - such as text or other graphics.

[0194] In typical video codecs, the prediction residual after motion compensation or intra prediction is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can, in many cases, help reduce this correlation and provide more efficient coding.

[0195] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A. to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + Rwhere C is the Lagrangian cost to be minimized, D is the image distortion (e.g., Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0196] Despite the ongoing study and development of video coding technologies under ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5), there remains a need for video coding technologies with a compression capability that significantly exceeds that of the current WC standard. Accordingly, described herein in certain embodiments are systems, methods, approaches, devices, and computer program products for implementing advantageous algorithms into current video coding technologies.

[0197] Among other things adopted into ECM is advanced in-loop filtering, e.g., as shown in FIG. 5. In the current ECM16, the in-loop filtering of ECM comprise: deblocking filtering (DB), SAG and Bilateral Filter (BIF), followed by multi-stage Enhanced Compression Model- Adaptive Loop Filtering (ECM-ALF). The 1ststage of the ECM-ALF consists of:- set of cascaded fixed filters {FFO and FF1}, each of which consists of 2 separate classifiers, and allows selection around 7000 filters of different filter support, - Residual Filter (RF), and- Gaussian Filter (GF).

[0198] Out put of the first stages filters, as well as un-filtered reconstructed pixels prior to deblocking and after SAO / BIF are input to the on-line training stage of the ALF. In this stage of ALF, various adaptive filters are applied to the pixels and adaptive filter parameters, such as weights and offsets, are signalled in bitstream to the decoder side.

[0199] Additional information about ECM-ALF is available in JVET-AK2019, the entire disclosure of which is hereby incorporated herein by reference in its entirety for all purposes.

[0200] ECM-ALF classification of fixed filters operates at granularity of 2x2. Filter size for both luma and chroma, for which ALF coefficients are signalled, is increased to 9x9.

[0201] To filter a luma sample, a first set of three different classifiers (Co, Ci and C2), a second set of one classifier C3, and three different sets of filters (Fo, Fi and F2) are used. Sets Fo and Fi contain fixed filters, with coefficients trained for classifiers Co and Ci. Coefficients of filters in F2 are signalled. Which filter from a set Fi is used for a given sample is decided by a class assigned to this sample using classifier Ci. The first set includes texture-based, band-based, and residualbased classifiers and the second set includes the coding information-based classifier.

[0202] For fixed filtering, at first, two 13x13 diamond shape fixed filters Fo andFi are applied to derive two intermediate samples / ?0(x,y) and R (x, y). After that, F2 is applied to / ?0(x, y), / ?! (x, y), and neighboring samples to derive a filtered sample as- 19 21^(x,y) = R(x,y) + q(^0-i=0 U=20where ft j is the clipped difference between a neighboring sample and current sample R (x, y) and gi is the clipped difference between 7?j_2oCL y) and current sample. The filter coefficients q, i = 0,... 21, are signalled.

[0203] Texture-based classification, which is similar to that in WC, can be employed. Based on directionality Dtand activity A a class is assigned to each 2x2 block:C[ — A[ * MD i+ D[where MD irepresents the total number of directionalities Dt. The derived is further mapped to C- by a pre-defined LUT to decrease the class number from 25 to 12.

[0204] As in WC, values of the horizontal, vertical, and two diagonal gradients are calculated for each sample using 1-D Laplacian. The sum of the sample gradients within a 4x4 window that covers the target 2x2 block is used for classifier Co and the sum of sample gradients within a 12x 12 window is used for classifiers Ci and C2. The sums of horizontal, vertical and two diagonal gradients are denoted, respectively, as ghl, gvlgdlland gdl2. The directionality Dtis determined by comparingt _ max(ghl,gvl'). _ max(gdll, gdl2)rh,v ~ / i >rdl,d2 ~ / i i xmin{ghl,gvl) min{gdll, gdl2)with a set of thresholds. The directionality D2is derived as in WC using thresholds 2 and 4.5. For Doand D, horizontal / vertical edge strength E^vand diagonal edge strength Ep are calculated first. Thresholds Th = [1.25, 1.5, 2, 3, 4.5, 8] are used. Edge strength EHLVis 0 if r^v< 77i[0]; otherwise, EpVis the maximum integer such that r^v> Th[EpV— 1], Edge strength EDlis 0 ifrdi,d2 — otherwise, Ep is the maximum integer such that rdl d2> Th[Ep — 1], Whenrh,v >rdi,d2’ ’e-’ horizontal / vertical edges are dominant, the Dtis derived by using Table 1; otherwise, diagonal edges are dominant, the Dtis derived by using Table 2.Table 1. Mapping of EDlto D,0 0 0 0 0 0 0 0 1 1 2 0 0 0 0 0 2 3 4 5 0 0 0 0 3 6 7 8 9 0 0 0 4 10 11 12 13 14 0 0 5 15 16 17 18 19 20 0 6 21 22 23 24 25 26 27Table 2. Mapping of EHlVto^HVED1\ 0 1 2 3 4 5 60 28 0 0 0 0 0 0 1 29 30 0 0 0 0 0 2 31 32 33 0 0 0 0 3 34 35 36 37 0 0 0 4 38 39 40 41 42 0 0 5 43 44 45 46 47 48 0 6 49 50 51 52 53 54 55

[0205] To obtain At, the sum of vertical and horizontal gradients Atis mapped to the range of 0 to n, where n is equal to 4 for A2and 15 for Aoand A

[0206] In an ALF APS, up to 4 luma filter sets are signalled, each set may have up to 24 filters.

[0207] Additionally, classification in ALF is extended with an additional alternative classifier. For a signalled luma filter set, a flag is signalled to indicate whether the alternative classifier is applied. Geometrical transformation is not applied to the alternative band classifier. When the band-based classifier is applied, the sum of sample values of a 2x2 luma block is calculated at first. Then the class index is calculated as below,class_index = (sum * 12) » (sample bit depth + 2).

[0208] A third classifier based on luma residual sample values. For each 2x2 luma block, the sum of absolute values of the residual samples in a neighbouring 8x8 window is calculated, and the class index is derived as:classldx = sum » (sample bit depth - 3).

[0209] The value of classldx is in the range of 0 to 11. The classifier usage is signalled for each luma filter set in APS.

[0210] For the online filters (signaled filters), each 2x2 unit is classified into 2 noise levels according to the partitioning information. For a 2x2 unit, it is classified into noise level-1 if it is located at a CU or TU boundary, and into noise level-0 if it is not. The class number of existing classifiers is reduced from 25 to 12. For the texture-based classifier, the 25 classes are mapped to 12 classes with a pre-defined LUT. For the band-based and residual based classifiers, the 25 classes are decreased to 12 classes by enlarging the band width. These 12 classes are further combined with the proposed 2 noise levels to generate final 12x2 = 24 classes in total. The number of classifiers for online filters is kept as 3 and no additional encoder selection is introduced.

[0211] For the offline filters (fixed filters), each 2x2 unit is classified into two (2) noise levels in the same way as the classification of online filters. Besides, each 2x2 unit is further classified into two (2) residual levels based on a predefined threshold also in the same way as the classifier of online filters. Furthermore, the generated offset of offline filters is adjusted based on the boundary level and residual level accordingly, where a stronger offset is applied on the positions at boundaries or with higher residuals.

[0212] The residual samples are used as additional inputs to the ALF. A filtered sample is derived as:19 25 27 R (x, y) = 7?(x, y) + q + Lt=O Lt=20 Lt=26 29 30 31 32ZCirFilterdiLt=28 Lt = 30 Lt=31 Lt = 32 where rtis the clipped neighboring residual sample value and rFilteredi is the clipped residual sample filtered by the fixed-filter. For residual samples, the fixed filter reuses the offline fixed filter trained for reconstruction after SAO.

[0213] Referring now to FIG. 7, in ALF, online-trained filters typically consist of four (4) kinds of filter taps: spatial taps, reconstruction-before-deblocking filter (DBF)-based taps, residual based taps and fixed-filter-output based taps. Additional fixed filter with a shape of diamond 7x7 is introduced, the filter parameters are stored at both encoder and decoder. There is no classification for the newly added fixed filter.

[0214] Referring now to FIG. 8, an example of an online filter method is shown. In some embodiments, spatial taps (i.e., tap #0 ~ #19), reconstruction-before-DBF-based taps (i.e., tap #26, #27, #36), residual-based taps (i.e., #37 ~ #38) and fixed-filter-output-based taps (i.e., tap #20 ~ #25, #34, #35) can be used. In some embodiments, several extended taps (i.e., tap #28 ~ #33, #39) can be added into luma online-trained filters. The reconstruction before DBF is fed into the additional fixed filter to produce the filter outputs, then these filter outputs are used as input for newly extended taps. In some embodiments, this filter may be always enabled without any filter shape switching.

[0215] Referring now to FIG. 9, a set of ALF filter shapes with Laplacian information-based taps are illustrated. In some embodiments, ALF coefficient filter taps for online training can include spatial taps (#0 - #9), fixed filter output-based taps (#10 - #29), reconstruction buffer before DBF (#30 - #32), residual-based taps (#33 - #34), Gaussian smoothed before DBF reconstruction buffer (#35). In some embodiments, a filtering method is provided that introduces a new diversity of information, Laplacian of ALF input luma buffer in deriving online ALF filters. Laplacian information buffer is obtained by applying Laplacian operator to ALF input luma buffer. At block boundaries, Gaussian pre-smoothing is applied before computing Laplacian. For deriving discrete Laplacian values, the 2D kernel [0,1,0; 1,-4,1; 0,1,0] is employed. In some embodiments, a new filter shape, a 5x5 cross shape (taps #36 - #40), can be applied to the Laplacian buffer. In some embodiments, the new filter shape is an online-trained filter shape.

[0216] In some embodiments, two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to C by a scaling factor. The value of C isan integer between 0 and 7, inclusively. With z=0, 1, let 6) denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index C is derived asC = C * 896 + Ct.

[0217] In some embodiments, the total number of the fixed filters is not changed.

[0218] Then a class index is determined based on the activity and directionality values. Two diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in the table below.Table 3. Filter support area for fixed filters of ECM- ALFECM-9.0 Improved fixed filteringALF input Samples before DBF ALF inputFixed filter f013x13 9x9 9x9Fixed filter 13x13 9x9 13x13

[0219] In some embodiments, fixed filter is applied to outputs of f0(instead of ALF input) and samples before DBF.

[0220] Finally, a signaled filter is applied to the ALF input samples, samples before the deblocking filter (DBF), outputs of the two fixed filters, output of a gaussian filter and the residual data.

[0221] In some embodiments, a classifier based on Laplacian values and variance is applied to a 2x2 chroma block. Compared to the luma classifier of a fixed filter, when calculating the activity value, the sum of the chroma vertical and horizonal Laplacian values is multiplied by 2 before scaling. Similarly, the chroma variance is multiplied by 2 before scaling. The derived class index is then used to select a fixed filter from a chroma filter set. A chroma fixed filter is applied to chroma ALF input samples in a 13x13 diamond shape and DBF input samples in a 7x7 diamond shape. The first luma classifier is applied to each 2x2 chroma block. The derived class index is then used to select a fixed filter from the luma fixed filter set related to this classifier. A fixed filter is applied to chroma ALF input sample in a 9x9 diamond shape and DBF input samples in a 9x9diamond shape. In a signalled chroma filter, 5x5 crossing extra taps are introduced, which are applied to the fixed filter output.

[0222] According to some embodiments, to derive the coefficients {c ”=oencoder firstly solves the Wiener-Hopf equations to minimize mean squared error (MSE) between the original picture and the filtered picture for arbitrary real-valued coefficient values. Secondly, the coefficients are quantized and clipped. Further, the quantized coefficients are adjusted by an optimization procedure based on coordinate search: encoder tries to modify a coefficient and accepts the modification only if MSE is decreased.

[0223] Finally, according to some embodiments, the coefficients are fine-tuned by Simulated Annealing. At each iteration of annealing run, (1) several coefficients are pseudo-randomly changed, (2) an objective function based on RD-cost is calculated, (3) the new coefficients are accepted with a probability depending on the objective function value and the iteration number, (4) the annealing run stops if the number of iterations without improving the objective function reached the predefined maximum. In total, 100 annealing runs are launched, where each run takes one of the previous outputs as its starting point.

[0224] In some embodiments, BIF filter is carried out in the sample adaptive offset (SAO) loop-filter stage, as shown in FIG. 5 for example. The bilateral filter (BIF), SAO and CC-SAO are using samples from deblocking as input. Each filter creates an offset per sample, and these are added to the input sample and then clipped, before proceeding to ALF.

[0225] For example, the output sample S0UTcan be obtained using the following:OUT — clip(S]N+ 8SAO+ 8CCSAO+ SBIF^,where SINis the input sample from deblocking, 8BIF is the offset from the bilateral filter, 8SA0is the offset from SAO and 8CCSA0is offset from CC-SAO.

[0226] In some embodiments, bilateral filter strength is signalled in the pps and can be 0 or 1 to specify, respectively, specifying half strength BIF filtering or full strength BIF filtering.

[0227] Referring now to FIG. 10, an NNVC in-loop filtering approach is illustrated. In some embodiments, an in-loop filter design can include one or more NN-based filters. In some embodiments, the NNVC in-loop filtering approach uses one or more CNN-based filters. The CCN-based filters can have a range of complexities, e.g. 455kMAC / pixel for HOP, 17kMAC / pixel for LOP and 5kMAC / pixel for VLOP. In some embodiments, the CNN-based filter can beimplemented in parallel to deblocking and its output combined with output of deblocking through weighted average.

[0228] Referring now to FIG. 11, an example of a NN architecture at 5kMAC / pixel for NNVC, referred to as VLOP1 is shown. In some embodiments, an NN based in-line filter (ILF) for Very Low Operation Point (5kMAC / pixel) can be provided.

[0229] An integration of NN-ILF into ECM targeting luma-only filtering have been proposed in JVET-AJ0210 and further studied in JVET-AK0183 and JVET-AK0184. In this contribution, implementation of JVET-AK0183 has been modified and NNVC ILF filters (VLOP, LOP and HOP) have been placed in parallel to deblocking filter and tested for joint YCbCr filtering.

[0230] The VLOP filter the following results are reported:- Al: {-0.7%, -1.8% and -1.8%} and 101% of encoder and 306% decoder run time. - RA: {-0.9%, -1.7% and -1.2%} and 101% of encoder and 284% decoder run time.

[0231] The LOP filter the following results are reported:- Al: {-1.3%, -4.9% and -5.2%} and 102% of encoder and 527% decoder run time. - RA: {-1.7%, -5.8% and -5.3%} and 102% of encoder and 468% decoder run time.

[0232] The HOP filter the following results are reported:- Al: {-3.4%, -10.3% and -11.1%} and 110% of encoder and 6,533% decoder run time.- RA: {-5.2%, -13% and -12.9%} and 108% of encoder and 5,646% decoder run time.

[0233] In some embodiments, a Unified Filter Architecture for In-Loop Filtering of NNVC as well as inference process and training procedure can be defined for High-performance Operation Point (HOP, 460kMAC / Pixel) and further extended to Low Operation Point (LOP, 17kMAC / pixel) and Very Low Operation Point (VLOP, 5kMAC / pixel).

[0234] In some embodiments, in NN ILF for LOP and VLOP, which may be considered as relevant for integration into ECM, takes reconstructed pixels (Luma and chroma), prior to the deblocking, prediction pixels, coding context information (e.g. QP, BS, coding modes), extract input features and jointly process these features by Luma and Chroma backbones (sequences of convolution layers), independently.

[0235] The output of NN-ILF is combined with output of the deblocking filter, with the weights and the offsets derived at the encoder side and signalled in the bitstream. Resulting pixelsare further processed by WC SAO and ALF. In some embodiments, a low complexity luma-only filtering can be derived from VLOP1 and applied for luma filtering in ECM-ALF.

[0236] In some embodiments, the NN-ILF process of JVET-AK0183 can be extended and NNVC ILF filters (VLOP, LOP and HOP) tested for joint YCbCr filtering and placed in parallel to the deblocking filtering.

[0237] Referring now to FIG. 12, an NN-ILF process is illustrated. In some embodiments, the NN-ILF process can be part of an ECM-ALF extension. In some embodiments, the NN-ILF process can integrated NN-based filtering with ECM-ALF framework was proposed in JVET-AK0183. In some embodiments, the reduced complexity NN-filter (4kMAC / pixel) is derivable from the architecture illustrated in FIG. 11. Reconstructed luma pixel before deblocking are filtered by NN filter and may be alternatively switched (ON / OFF) with output of the Gaussian filter. Encoder evaluates the performance of the NN filter and makes a decision whether NN filter or Gaussian filters should be applied.

[0238] The input of the ALF online-trained filters, associated with NN or Gaussian filter outputs, to be filtered by adaptive Wiener Filter taps: 1x1 in the case of non LSlices and 5x5 spatial support in the case of I-Slice.

[0239] However, current video coding and decoding techniques that include NN filtering, such as NNVC or JVET-AK0183, exhibit undesirably high complexity, e.g., around 5kMAC / pixel operation, and as reported in JVET-AK0183, between about 130% and about 150% of the ECM decoder run-time.

[0240] Similarly, the fixed filters of the ECM also considered relatively complex, and feature a cascaded processing with classifiers, multi-stage filtering, and more importantly need to switch between around 7,000 classes (each being a different, separately selectable filter). The supplemental information associated with these 7,000 or more different filters and the processes / parameters for switching therebetween currently represents over 1.2 MB of regularly transmitted data. This complexity is also reflected in the ECM run-time, resulting in around 15% Encoder and Decoder run-time increases.

[0241] The on-line training stage of the ALF is operating as following:Reconstructed pixels after BIF / SAO are input to the ALF online-trained filters and processed with 9x9 filter of a cross-like shape with 10 individually derived coefficients.Reconstructed pixels prior to deblocking are input to the ALF online-trained filters and processed with 3x3 filter of a cross-like shape with 3 individually derived coefficients.Output of the Fixed Filter 2 is input to the ALF online-trained filters and processed with a 13x13 filter of a cross-like shape, with 19 individually derived coefficients and output of the Fixed Filter 1 was processed with 1x1 filter.Output of the Gaussian filter, or alternatively, NN-ILF filter in JVET-AK0183 implementation, is processed with 5x5 cross-like filter for I slices and 1x1 filter for other slices.Output of the residual filter is processed with 1x2 filter.

[0242] In some embodiments, the Fixed Filters of ECM- ALF and NN-filter are fixed filters and may be combined or interchanged to further improve performance-to-complexity tradeoff.

[0243] In some embodiments, NN-filter defined in NNVC takes as its inputs Reconstructed and Predicted pixels as well as a wide range of supplementary data, including QP information, block boundary strength (deblocking classification) and coding modes information. The latter is produced as representation of the coding modes through a limited number of the codewords of the IPB input channel:

[0244] In particular, neural network-based video coding (NNVC) tested several representations the coding modes, in one example the following mapping was utilized:Table 4. Example of the coding modes mapping to supplementary information for NN-ILF Coding mode Codeword valueInter 0Intra 1IBC 2

[0245] In another example the following mapping was utilized:Table 5. Example of the coding modes mapping to supplementary information for NN-ILF Coding mode Codeword valueIntra 0IBC 2IBC & SKIP 3Inter & UniPred 4Inter & SKIP & UniPred 5Inter & BiPred 6Inter & SKIP & BiPred 7As such, described herein are systems, methods, apparatuses, and computer program products for video encoding and video decoding, encoding and decoding apparatuses, and non-transitory computer-readable storage media thereof. In a method for video encoding, the encoder may obtain pixel information associated with raw pixels in a picture for a first pixel that is to be filtered in the picture. The encoder may select and apply a plurality of filters for filtering the pixel information associated with the picture to obtain a filtered pixel. This approach may use a multistage adaptive loop filter (ALF). The encoder can, following filtering of the pixel information, generate adaptive filter parameters associated with the filtered pixel. Information associated with the adaptive filter parameters can be provided to the decoder in a bitstream. In a method for video decoding, the decoder receives the adaptive filter parameters and uses the same during decoding and regeneration of the filtered pixel.

[0246] As illustrated in FIG. 12, a multi-stage filtering process can be used that integrates NN-ILF as a part of an ECM-ALF extension. In some embodiments, the multi-stage filtering process and architecture may address problems of NN-based filters integration with convention hybrid-based video coding architecture. NN-based filtering for in-loop processing. In some embodiments, the output of the NN filter is used as input to on-line filter independently from the output of the Gaussian filter. In other embodiments, the output of the NN filter can be used as input to the on-line filter alongside the output of the Gaussian filter.

[0247] In some embodiments, the output of NN-filter to be processed by ALF online-trained filters with an adaptive filter adaptNN that has filter support area NxM, where N and M are integer numbers. In some embodiments, N and M may take predefined values ranging from 0 to 17 and / or filter support are of shape of rectangular, diamond or cross shape, as shown in FIG. 6. In some embodiments, adaptNN filter tap length can be fixed to NxN, e.g., 1x1, and not being a subject of Slice-type.

[0248] In some embodiments, the NN output pixels used as input to ALF online-trained filters in alternation with other inputs, e.g. output of a fixed filter(s), and / or residual filter, or reconstructed pixels before deblocking, or after SAO / BIF. In some embodiments, pixels of NNoutput may be used instead of the pixels output of Fixed Filter 0 or Fixed Filter 1. In another embodiment, use of any component of the ECM-ALF Fixed Filters, e.g. classifier 0 / 1, fixed filter 0 / 1 and / or ALF online-trained filter for Fixed Filter outputs, is disallowed for current block if NN-filter is being enabled for this block. In yet another embodiment, use Classifier 1 and Fixed Filter 1 may be disallowed for current block if NN-filter is being enabled for this block.

[0249] In some embodiments, ECM can be adapted to include or provide supplementary information for the NN filter. For example, in some embodiments, supplemental information can be extended, e.g., for example in the IPB and BS input channels as shown in FIG. 6, of the CNN based filter, to include additional coding modes present in ECM.

[0250] In some embodiments, IPB supplementary information associated with a set or group of pixels within current block (unit) can be extended to include use of palette mode. In some embodiments, a palette mode ID can be added as an additional, not-occupied yet code word value, e.g., 8 on top of Table 4, or aggregated with existing IBC codewords through logical operations OR, or through mathematical or bit- wise operations, e.g.,cgiTools = (cu.predMode == MODE IBC) || (cu.predMode == MODE PLT)

[0251] In another example the following mapping in Table 6 was utilized:Table 6. Example mapping of Coding mode to codework values.Coding mode Codeword valueIntra 0cgiTools 2cgiTools & SKIP 3Inter & UniPred 4Inter & SKIP & UniPred 5Inter & BiPred 6Inter & SKIP & BiPred 7

[0252] In yet another embodiment, palette identificatory may be aggregated with MODE INTRA codeword:Intra = (cu.predMode == MODE Intra) || (cu.predMode == MODE PLT).

[0253] In some embodiments, IPB supplementary information associated with a set or group of pixels within current block (unit) can be extended to include use of intra template matching mode. In some embodiments, IntraTMP mode ID can be added as an additional, not-occupied codeword value, e.g. 8 on top of Table 4, or aggregated with existing codewords through logical operations OR, or through mathematical or bit-wise operations, e.g. below is shown aggregation with cgiTools:cgiTools = cgiTools || cu.tmpFlag

[0254] In some embodiments, IPB supplementary information associated with a set or group of pixels within current block (unit) can be extended to include use of Geometric Prediction Modes (GPM) for intra, inter prediction and cgi modes. In some embodiments, inter GPM and intra (Spatial) GPM modes can be represented with individual, added code word, e.g., 8 and / or 9 on top of Table 4, or aggregated into a single codeword representing blocks with similar characteristics.gpmTools = interGPM || sgpmFlag || ibcGpmFlag

[0255] In another example the following mapping in Table 7 was utilized:Table 7. Example mapping of coding modes to codeword values.Coding mode Codeword valueIntra 0IBC 2IBC & SKIP 3Inter & UniPred 4Inter & SKIP & UniPred 5Inter & BiPred 6Inter & SKIP & BiPred 7gpmTools 8

[0256] In some embodiments, IPB supplementary information associated with a set or group of pixels within current block (unit) can be extended to include use of OBMC mode. In some embodiments, OBMC mode ID can be added as an additional, not-occupied code word value, e.g., 8 on top of Table 4, or aggregated with existing codewords through logical operations OR, or through mathematical or bit-wise operations, e.g. below is shown aggregation with code word for gpm modes:gpmOBMC = gpmTools || cu.isobmcMC.

[0257] In another example the following mapping in Table 8 was utilized:Table 8. Example mapping of coding modes to codeword values.Coding mode Codeword valueIntra 0IBC 2IBC & SKIP 3Inter & UniPred 4Inter & SKIP & UniPred 5Inter & BiPred 6Inter & SKIP & BiPred 7OBMC 8

[0258] In some embodiments, OBMC information can be aggregated with deblocking Block Strength information either as an additional, not-occupied, yet, codeword value, or aggregated with existing codewords and provided to the NN filter through BS input channel, see, e.g., FIG. 11.

[0259] In some embodiments, BS supplementary information associated with a set or group of pixels within current block (unit), see, e.g., FIG. 11, can be extended to include a flag, field, value, or other indication that bilateral filtering (BIF) has been or is being enabled, and / or to indicate a strength of the BIF. In some embodiments, a BIF mode enabled ID, can be aggregated with a bilateral filter strength value through logical, mathematical or bit-wise operations, for example as shown below:bifValue = bifEnableFlag * (1 + bilateral filter strength)where bifEnableFlag specify if BIF is used for current block, bilateral filter strength is parameters of the BIF strength, may be signalled through bitstream e.g., in PPS or SPS, or Slice Header, and bifValue is codeword representing BIF information. Following this aggregation, BIF information can be aggregated with existing codewords for block strength through logical operations OR, or through mathematical or bit- wise operations:bsValueOut = bsValueln | (bifid « bitOffsets)where bsValueln is a codeword associated with deblocking block strength, bifid is codeword associated with aggregated BIF enabled flag and strength value, bitOffset is number of bits, e.g., bitOffset = 4, used for left shift to ensure separation of the deblocking and BIF information in thecode word, and bsValueOut defines codeword value utilized as input to NN filter as part of BS input.

[0260] In an embodiment, at least some of the processes described herein may be carried out by an apparatus comprising means for carrying out at least some of the described processes. Means for performing method operations as disclosed herein may include software and / or hardware components (e.g., 110, 120, 220, 230, 320, 330) of the apparatus. For example, at least one processor and at least one memory storing thereon computer program codes can comprise means for carrying out the method or methods as disclosed herein, and any of the embodiments thereof. As used herein the term “means” is to be construed in singular form, i.e. referring to a single element, or in plural form, i.e. referring to a combination of single elements. Therefore, terminology “means for [performing A, B, C]”, is to be interpreted to cover an apparatus in which there is only one means for performing A, B and C, or where there are separate means for performing A, B and C, or partially or fully overlapping means for performing A, B, C. Further, terminology “means for performing A, means for performing B, means for performing C” is to be interpreted to cover an apparatus in which there is only one means for performing A, B and C, or where there are separate means for performing A, B and C, or partially or fully overlapping means for performing A, B, C.

[0261] FIGs. 13-17 illustrate flowcharts depicting various methods according to example embodiments of the present disclosure. It will be understood that each block of the flowcharts and combination of blocks in the flowcharts can be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other communication devices associated with execution of software including instructions, for example one or more computer program instructions. For example, one or more of the procedures described above can be embodied by computer program instructions. In some embodiments, the computer program instructions which embody the procedures described above can be stored, for example, by the memory 302 of the decoding device 300 illustrated in FIG. 3 employing an embodiment of the present disclosure and executed by a processor (e.g., the controller 301). In some embodiments, the computer program instructions which embody the procedures described above can be stored, for example, by the memory 402 of the encoding device 400 illustrated in FIG. 4 employing an embodiment of the present disclosure and executed by a processor (e.g., the controller 401). As will be appreciated, any such computer program instructions can be loaded onto a computer or other programmable apparatus (forexample, hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture the execution of which implements the function specified in the flowchart blocks. The computer program instructions can also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.

[0262] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, can be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.

[0263] Referring now to FIG. 13, the operations performed enable providing media content based on context, in accordance with one or more embodiments of the present disclosure. For example, some or all of the method 600 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 600. Additionally, a computer program product can be provided that comprises a non-transitory computer readable storage medium storing instructions thereon that, when executed by a processor, cause a machine or apparatus to perform some or all of the method 600. In some embodiments, the method 600 is associated with functionality of the video encoder 102 and / or the encoding device 400.

[0264] As shown in block 602 of FIG. 13, an apparatus includes means, such as the controller 401, the memory 402, and / or the like, configured to input information associated with a set of pixels into a first stage of a multi-stage filtering process, the multi-stage filtering process comprising one or more first filters and one or more neural network-based (NN-based) filters. Asshown in block 604 of FIG. 13, an apparatus additionally or alternatively includes means, such as the controller 401, the memory 402, or the like, configured to generate, using at least one of the one or more first filters or the one or more NN-based filters, one or more first filter outputs. As shown in block 606 of FIG. 13, an apparatus additionally or alternatively includes means, such as the controller 401, the memory 402, or the like, configured to input, into a second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first filter outputs, the second stage comprising one or more adaptive filters. As shown in block 608 of FIG. 13, an apparatus additionally or alternatively includes means, such as the controller 401, the memory 402, or the like, configured to generate, conditional on a presence of the one or more first filter outputs from at least one of the one or more first filters or the one or more NN-based filters, using the one or more adaptive filters of the second stage of the multi-stage filtering process, information about one or more parameters used during the multi-stage filtering process. As shown in block 610 of FIG. 13, an apparatus additionally or alternatively includes means, such as the controller 401, the memory 402, or the like, configured to transmit, in a bitstream, the information about the one or more parameters used during the multi-stage filtering process.

[0265] Referring now to FIG. 14, the operations performed enable providing media content based on context, in accordance with one or more embodiments of the present disclosure. For example, some or all of a method 700 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 700. Additionally, a computer program product can be provided that comprises a non-transitory computer readable storage medium storing instructions thereon that, when executed by a processor, cause a machine or apparatus to perform some or all of the method 700. In some embodiments, the method 700 is associated with functionality of the video decoder 106 and / or the decoding device 300.

[0266] As shown in block 702 of FIG. 14, an apparatus includes means, such as the controller 301, the memory 302, and / or the like, configured to provide or receive a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes. As shown in block 704 of FIG. 14, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configuredto provide or receive, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value from among the plurality of codeword values. As shown in block 706 of FIG. 14, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to identify, using the at least one codeword value in the information associated with the set of pixels, based on the mapping of the plurality of coding modes to the plurality of codeword values, at least one coding mode of the plurality of coding modes used during a multi-stage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter. As shown in block 708 of FIG. 14, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to determine, based on the at least one coding mode associated with the at least one codeword value in the information associated with the set of pixels, whether the at least one coding mode comprises an aggregated coding mode, the aggregated coding mode indicating use of either a first coding mode or a second coding mode during the multi-stage filtering process.

[0267] Referring now to FIG. 15, the operations performed enable providing media content based on context, in accordance with one or more embodiments of the present disclosure. For example, some or all of a method 800 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 800. Additionally, a computer program product can be provided that comprises a non-transitory computer readable storage medium storing instructions thereon that, when executed by a processor, cause a machine or apparatus to perform some or all of the method 800. In some embodiments, the method 800 is associated with functionality of the video decoder 106 and / or the decoding device 300.

[0268] As shown in block 802 of FIG. 15, an apparatus includes means, such as the controller 301, the memory 302, and / or the like, configured to provide or receive a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes. As shown in block 804 of FIG. 15, an apparatus additionally oralternatively includes means, such as the controller 301, the memory 302, or the like, configured to provide or receive, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value from among the plurality of codeword values. As shown in block 806 of FIG. 15, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to identify, using the at least one codeword value in the information associated with the set of pixels, based on the mapping of the plurality of coding modes to the plurality of codeword values, at least one coding mode of the plurality of coding modes used during a multi-stage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter.

[0269] Referring now to FIG. 16, the operations performed enable providing media content based on context, in accordance with one or more embodiments of the present disclosure. For example, some or all of a method 900 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 900. Additionally, a computer program product can be provided that comprises a non-transitory computer readable storage medium storing instructions thereon that, when executed by a processor, cause a machine or apparatus to perform some or all of the method 900. In some embodiments, the method 900 is associated with functionality of the video decoder 106 and / or the decoding device 300.

[0270] As shown in block 902 of FIG. 16, an apparatus includes means, such as the controller 301, the memory 302, and / or the like, configured to input, into at least one of one or more first filters or one or more neural network-based (NN-based) filters in a first stage of a multistage filtering process, information associated with a set of pixels to generate one or more first filter outputs. As shown in block 904 of FIG. 16, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to input, into an adaptive filter in a second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first filter outputs. As shown in block 906 of FIG. 16,an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to generate or receive filter parameter information associated with the multi-stage filtering process. As shown in block 908 of FIG. 16, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to transmit or receive, in a bitstream, the filter parameter information associated with the multistage filtering process.

[0271] Referring now to FIG. 10, the operations performed enable providing media content based on context, in accordance with one or more embodiments of the present disclosure. For example, some or all of a method 1000 can be carried out using a system (e.g., 100), by an apparatus, or by a computing device, such as 102, 104, 106, 219, 211, 212, 217, 218, 300, 400, etc. For example, an apparatus can comprise at least one processor and at least one memory storing instructions thereon that, when executed by the at least one processor, cause the apparatus to perform some or all of the method 1000. Additionally, a computer program product can be provided that comprises a non-transitory computer readable storage medium storing instructions thereon that, when executed by a processor, cause a machine or apparatus to perform some or all of the method 1000. In some embodiments, the method 1000 is associated with functionality of the video decoder 106 and / or the decoding device 300.

[0272] As shown in block 1002 of FIG. 17, an apparatus includes means, such as the controller 301, the memory 302, and / or the like, configured to provide or receive, in a bitstream, information about a set of pixels and information about one or more filter parameters for a multistage filtering process associated with the set of pixels, wherein a first stage of the multi-stage filtering process comprises one or more first filters and / or one or more neural network-based (NN-based) filters, wherein a second stage of the multi-stage filtering process comprises one or more second filters, the one or more second filters comprising an adaptive filter. As shown in block 1004 of FIG. 17, an apparatus additionally or alternatively includes means, such as the controller 301, the memory 302, or the like, configured to, based on the information about the one or more filter parameters, determine whether the first stage of the multi-stage filtering process comprises the one or more first filters, the one or more NN-based filters, or both the one or more first filters and the one or more NN-based filters.

[0273] It is also noted herein that although the above describes example embodiments, there are several variations and modifications which may be made to the disclosed solution without departing from the scope of the subject disclosure.

[0274] In general, the various example embodiments may be implemented in hardware or special purpose circuits, software, logic or any combination thereof. Some aspects of the subject disclosure may be implemented in hardware, whereas other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device, although the subject disclosure is not limited thereto. Although various aspects of the subject disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0275] Example embodiments of the subject disclosure may be implemented by computer software executable by a data processor of the mobile device, such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or program, also called program product, including software routines, applets and / or macros, may be stored in any apparatus-readable data storage medium and they comprise program instructions to perform particular tasks. A computer program product may comprise one or more computer- executable components which, when the program is run, are configured to carry out embodiments. The one or more computer-executable components may be at least one software code or portions of it.

[0276] Further in this regard it should be noted that any blocks of the logic flow as in the figures may represent program processes, or interconnected logic circuits, blocks and functions, or a combination of program processes and logic circuits, blocks and functions. The software may be stored on such physical media as memory chips, or memory blocks implemented within the processor, magnetic media such as hard disk or floppy disks, and optical media such as for example DVD and the data variants thereof, CD. The physical media is a non-transitory media.

[0277] The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The data processors may be of any type suitable to the localtechnical environment, and may comprise one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), FPGA, gate level circuits and processors based on multi-core processor architecture, as non-limiting examples.

[0278] Example embodiments of the subject disclosure may be practiced in various components such as integrated circuit modules. The design of integrated circuits is by and large a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design ready to be etched and formed on a semiconductor substrate.

[0279] Moreover, in accordance with the foregoing description, the embodiments described herein reflect possible embodiments of the herein presented solution, which are further combinable between embodiments and / or sets of embodiments. These embodiments do not define the entire scope of the disclosure of this application; nor do these embodiments define or limit the scope of the disclosure.

[0280] The above-noted aspects and features may be implemented in systems, apparatuses, methods, articles and non-transitory computer-readable media depending on the desired configuration. The subject disclosure may be implemented in and used with a number of different types of devices, including but not limited to cellular phones, tablet computers, wearable computing devices, portable media players, and any of various other computing devices.

[0281] The foregoing description has provided by way of non-limiting examples a full and informative description of the example embodiment of the subject disclosure. However, various modifications and adaptations may become apparent to those skilled in the relevant arts in view of the foregoing description, when read in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure will still fall within the scope of the subject disclosure as defined in the appended claims. Indeed, there is a further embodiment comprising a combination of one or more embodiments with any of the other embodiments previously discussed.

[0282] Even though the disclosure has been described above with reference to an example according to the accompanying drawings, it is clear that the disclosure is not restricted thereto but can be modified in several ways within the scope of the appended claims. Therefore, all words and expressions should be interpreted broadly and they are intended to illustrate, not to restrict, theembodiment. It will be obvious to a person skilled in the art that, as technology advances, the inventive concept can be implemented in various ways. Further, it is clear to a person skilled in the art that the described embodiments may, but are not required to, be combined with other embodiments in various ways.

[0283] With regard to the illustrated signal flow diagrams and flowcharts depicting methods provided in the drawings and described above, it will be understood that each block or signal and combination of blocks and signals may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other communication devices associated with execution of software including one or more computer program instructions.

[0284] For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory of the apparatus employing an example embodiment and executed by processing circuitry, such as a processor. As will be appreciated, any such computer program instructions, such as instructions, may be loaded onto a computer or other programmable apparatus (for example, hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.

[0285] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, can be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.

[0286] Many modifications and other embodiments of the disclosures set forth herein will come to mind to one skilled in the art to which these disclosures pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the disclosures are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

CLAIMS1. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:inputting information associated with a set of pixels into a first stage of a multistage filtering process, the multi-stage filtering process comprising one or more first filters and one or more neural network-based (NN-based) filters;generating, using at least one of the one or more first filters or the one or more NN-based filters, one or more first filter outputs;inputting, into a second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first filter outputs, the second stage comprising one or more adaptive filters;generating, conditional on a presence of the one or more first filter outputs from at least one of the one or more first filters or the one or more NN-based filters, using the one or more adaptive filters of the second stage of the multi-stage filtering process, information about one or more parameters used during the multi-stage filtering process, andtransmitting, in a bitstream, the information about the one or more parameters used during the multi-stage filtering process.

2. The apparatus of claim 1, wherein the adaptive filter is a content-adaptive filter that operates in a loop of a video coding scheme with one or more filter parameters derived from the information associated with the set of pixels to minimize a target cost.

3. The apparatus of claim 2, wherein the target cost is minimized using Rate-Distortion Optimization.

4. The apparatus of any prior claim, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform: encapsulating the information about the one or more parameters used during the multi-stage filtering process into the bitstream.

5. The apparatus of claim 1, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:inputting, into the one or more second filters in the second stage of the multi-stage filtering process, supplemental information associated with the set of pixels, the supplemental information comprising one or more of: a set of pixels, information regarding how the set of pixels have been processed, a quantization parameter (QP), or a coding mode used to produce the set of pixels.

6. The apparatus of claim 1, wherein the one or more first filters comprise one or more of: a deblocking (DB) filter, a sample adaptive offset (SAO) loop filter, a bilateral filter (BIF), a residual filter (RF), a Gaussian filter (GF), transform-domain filtering, a fixed filter, a plurality of cascaded fixed filters, a spatial or temporal filter, a texture-based filter, a band-based filter, a residual-based filter, a pre-smoothing filter, or a Laplacian filter.

7. The apparatus of claim 1, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:deriving a weighted average of the one or more first filter outputs and the NN-base filter output; andinputting the weighted average of the one or more first filter outputs and the NN-based filter output into the one or more adaptive filters in the second stage of the multi-stage filtering process.

8. The apparatus of claim 1, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:inputting, into at least one of the one or more NN-based filters, the set of pixels and supplemental information associated with the set of pixels, the supplemental information comprising one or more of: one or more values indicating one or more coding operations utilized to generate the set of pixels, coding modes input channel (IPB) values, and / or parameters of a loop filtering operations applied to generate the set of pixels, (BS) input channel values, and / orinformation about quantization parameters (QP) utilized to generate the set of pixels.

9. The apparatus of claim 8, wherein the supplemental information comprises one or more parameter or identifier of:one or more parameters or identities associated with an intra coded mode,one or more parameters or identities associated with an inter coded mode,a prediction type identifier,one or more parameters or identities associated with a uni-hypothesis prediction, one or more parameters or identities associated with a bi-hypothesis prediction, one or more parameters or identities associated with a multi-hypothesis prediction, one or more parameters or identities associated with a geometric prediction mode (GPM), one or more parameters or identities associated with an intra block copy (IBC) mode, one or more parameters or identities associated with an intra templatematching (IntraTMP) mode,one or more parameters or identities associated with a palette mode (PLT),one or more parameters or identities associated with a bilateral filtering mode, one or more parameters or identities associated with a deblocking filtering mode, one or more parameters or identities associated with a transform domain filtering mode, one or more parameters or identities associated with a Hadamard filter filtering mode, one or more parameters or identities associated with a motion compensation filtering mode,one or more parameters or identities associated with an interpolation filtering mode, one or more parameters or identities associated with an overlapped block motion compensation (OBMC) mode,one or more parameters or identities associated with an affine mode (AFFINE) mode, one or more parameters or identities associated with in-loop filtering utilized during the multi-stage filtering process,one or more parameters or identities associated with a deblocking filter (DB) utilized during the multi-stage filtering process, one or more parameters or identities associated with a boundary strength for the set of pixels,an interpolation filter rescaling ratio, and / orone or more parameters or identities associated with models for intra-template matching (IntraTMP).

10. The apparatus of claim 9, wherein the models for IntraTMP are selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a cross-component reconstruction mode, a convolutional cross-component model, a multi-cross component linear model,11. The apparatus of claim 1, wherein the one or more NN-based filters have a filter support area of NxM, where N and M are integer numbers.

12. The apparatus of claim 11, wherein N and M are predefined independent values ranging from 0 to 17, or higher.

13. The apparatus of claim 12, wherein at least one of the one or more first filters has a shape selected from among: a rectangular, a diamond, or a cross.

14. The apparatus of claim 13, wherein a tap shape of the one or more NN-based filters has a height and width of NxN.

15. The apparatus of claim 1, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:determining whether NN-based filtering is enabled for the set of pixels; andin an instance in which NN-based filtering is enabled for the set of pixels, inputting one or more NN-based filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more first filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

16. The apparatus of claim 1, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:determining whether NN-based filtering is enabled for the set of pixels; andin an instance in which NN-based filtering is not enabled for the set of pixels, inputting the one or more first filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more NN-based filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

17. The apparatus of claim 1, wherein the one or more first filters comprise a DB filter, a sample adaptive offset (SAO) filter, and a bilateral filter (BIF).

18. The apparatus of claim 17, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which NN-based filtering is not enabled for the set of pixels, filtering the information associated with the set of pixels sequentially using the DB filter, the SAO filter, and the BIF filter of the one or more first filters in the first stage of the multi-stage filtering process; andinputting, into the one or more adaptive filters in the second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first outputs from the one or more first filters.

19. The apparatus of any prior claim, wherein the information transmitted in the bitstream about the one or more adaptive filter parameters derived during the multi-stage filtering process comprises one or more of: a filter coefficient, a pixel value offset, or clipping range information.

20. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:providing or receiving a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes;providing or receiving, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value fromamong the plurality of codeword values; andidentifying, using the at least one codeword value in the information associated with the set of pixels, based on the mapping of the plurality of coding modes to the plurality of codeword values, at least one coding mode of the plurality of coding modes used during a multistage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter.

21. The apparatus of claim 20, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:extending the mapping to include one or more additional codeword values and one or more additional coding modes corresponding respectively to the one or more additional codeword values, the one or more additional coding modes being usable by, or associated with, the one or more NN-based filters in the first stage of the multi-stage filtering process.

22. The apparatus of claim 20, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:modifying the mapping to associate one or more of the plurality of codeword values with one or more replacement coding modes, the one or more replacement coding modes being usable by, or associated with, the one or more NN-based filters in the first stage of the multi-stage filtering process.

23. The apparatus of claim 22, wherein the modifying the mapping comprises grouping or aggregating information about a plurality of coding modes into a single codeword value.

24. The apparatus of any one of claims 21 to 23, wherein the one or more additional coding modes or the one or more replacement coding modes comprise one or more of:an intra coding mode,an inter coding mode,a palette (PLT) mode,an intra template matching (intraTMP) mode,a geometric prediction mode (GPM),a bilateral filtering mode,a uni-hypothesis prediction mode,a bi-hypothesis prediction mode,a multi-hypothesis prediction mode,an intra block copy (IBC) mode,a LIC mode,an overlapped block motion compensation (OBMC) mode, oran affine mode.

25. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:providing or receiving a mapping of a plurality of coding modes to a plurality of codeword values that correspond respectively to the plurality of coding modes;providing or receiving, in a bitstream, information associated with a set of pixels, the information associated with the set of pixels comprising at least one codeword value from among the plurality of codeword values;identifying, using the at least one codeword value in the information associated with the set of pixels, based on the mapping of the plurality of coding modes to the plurality of codeword values, at least one coding mode of the plurality of coding modes used during a multistage filtering process for the set of pixels, the multi-stage filtering process comprising a first stage comprising one or more first filters and one or more neural network-based (NN-based) filters, the multi-stage filtering process further comprising a second stage comprising one or more second filters, the one or more second filters comprising at least one adaptive filter; and determining, based on the at least one coding mode associated with the at least one codeword value in the information associated with the set of pixels, whether the at least one coding mode comprises an aggregated coding mode, the aggregated coding mode indicating use of either a first coding mode or a second coding mode during the multi-stage filtering process.

26. The apparatus of claim 25, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which the at least one coding mode indicated by the at least one codeword value in the information associated with the set of pixels comprises an aggregated coding mode indicating use of either a first coding mode or a second coding mode during the multi-stage filtering process, executing a logical operation to determine whether the at least one codeword value indicates the first coding mode or the second coding mode.

27. The apparatus of claim 25, wherein one or more of the first coding mode or the second coding mode comprises one or more of:an intra coding mode,an inter coding mode,a palette (PLT) mode,an intra template matching (intraTMP) mode,a geometric prediction mode (GPM),a bilateral filtering mode,a uni-hypothesis prediction mode,a bi-hypothesis prediction mode,a multi-hypothesis prediction mode,an intra block copy (IBC) mode,a LIC mode,an overlapped block motion compensation (OBMC) mode, oran affine mode.

28. The apparatus of claim 25, wherein a single codework value from among the plurality of codeword values is operable to represent a group of coding modes from among the plurality of coding modes, the group of coding modes comprising at least one of:IPB and Palette,Intra, IPB, and Palette,IPB, IntraTMP, and Palette,a plurality of GPM modes,interGPM, spatial GPM, IBC, and GPM,a plurality of GPM modes and Intera plurality of GPM modes and bi-hypothesis prediction,a plurality of GPM modes and multi-hypothesis prediction, orOBMC mode and a plurality of GPM modes.

29. The apparatus of claim 28, wherein the single codeword value representing a group of coding modes is derivable via at least one of: one or more logical operations, one or more computational operations, one or more mathematical operations, or one or more bit-wise operations.

30. The apparatus of claim 28, wherein the single codeword value further represents a plurality of groups of coding modes from among the plurality of coding modes, the single codeword being derivable via at least one of: one or more logical operations, one or more computational operations, one or more mathematical operations, or one or more bit-wise operations.

31. The apparatus of claim 25, wherein one or more of parameters or identifiers of loop filtering operations preceding the one or more NN-based filters comprises:a deblocking filter,a bilateral filter,a transform-domain filter, ora SAG filter.

32. The apparatus of claim 31, wherein a single codeword value from among the plurality of codeword values represents a group of multiple parameters or identifiers of the loop filtering operations preceding the one or more NN-based filters, the group comprising one or more of:a boundary strength,a filter strength,a filter type, oran enable / disable flag.

33. The apparatus of claim 32, wherein the single codeword value representing the group of parameters or identifiers is derivable via at least one of: one or more logical operations, one or more computational operations, one or more mathematical operations, or one or more bit-wise operations.

34. An apparatus comprising:at least one processor; andat least one memory comprising instructions stored therein that, when executed by the at least one processor, cause the apparatus to perform at least:providing or receiving, in a bitstream, information about a set of pixels and information about one or more filter parameters for a multi-stage filtering process associated with the set of pixels,wherein a first stage of the multi-stage filtering process comprises one or more first filters and / or one or more neural network-based (NN-based) filters,wherein a second stage of the multi-stage filtering process comprises one or more second filters, the one or more second filters comprising an adaptive filter; andbased on the information about the one or more filter parameters, determining whether the first stage of the multi-stage filtering process comprises the one or more first filters, the one or more NN-based filters, or both the one or more first filters and the one or more NN-based filters.

35. The apparatus of claim 34, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:reconstructing, based at least on the information about the one or more pixels and the information about the one or more filter parameters, the one or more pixels.

36. The apparatus of claim 34, wherein the information associated with the set of pixels comprises one or more of: raw pixel data, supplementary information about how the set of pixels were produced, a quantization parameter (QP), or one or more codeword values indicating one or more coding modes corresponding respectively to the one or more codeword values.

37. The apparatus of claim 34, wherein the one or more first filters comprise one or more of: a deblocking (DB) filter, a sample adaptive offset (SAO) loop filter, a bilateral filter (BIF), a residual filter (RF), a Gaussian filter (GF), transform-domain filtering, a fixed filter, a plurality of cascading fixed filters, a spatial filter, a texture-based filter, a band-based filter, a residualbased filter, a fixed-filter-output filter, a residuals-before-DB filter, a pre-smoothing filter, or a Laplacian filter.

38. The apparatus of claim 34, wherein an input to the adaptive filter in the multi-stage filtering process comprises a weighted average of one or more first filter outputs and an NN-base filter output.

39. The apparatus of claim 34, wherein the filter parameter information comprises one or more of: one or more values indicating one or more coding modes used in the first phase of the multi-phase filtering process, coding modes input channel (IPB) values, parameters of a loop filtering (BS) input channel, or quantization parameters (QP).

40. The apparatus of claim 34, wherein the filter parameter information comprises one or more of:an intra coded mode identifier,an inter coded mode identifier,a prediction type identifier,a uni-hypothesis prediction identifier,a bi-hypothesis prediction identifier,a multi-hypothesis prediction identifier,a geometric prediction mode (GPM) identifier,an intra block copy (IBC) mode identifier,an intra template matching (IntraTMP) identifier,a palette mode (PLT) identifier,a bilateral filtering identifier,a deblocking filtering identifier,a transform domain filtering identifier,a Hadamard filter filtering identifier,a motion compensation filtering identifier,an interpolation filtering identifier,an overlapped block motion compensation (OBMC) identifier,an affine mode (AFFINE) identifier,an adaptive loop filtering identifier,one or more identifiers or parameters of in-loop filtering utilized during the multi-phase filtering process,one or more identifiers or parameters of a deblocking filter (DB) utilized during the multi-phase filtering process,one or more identifiers or parameters of a boundary strength associated with the set of pixels,one or more identifiers or parameters of a block strength associated with the set of pixels, an interpolation filter rescaling ratio, and / orone or more identifiers or parameters of models for intra-template matching (IntraTMP) selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a cross-component reconstruction mode, a convolutional crosscomponent model, a multi-cross component linear model.

41. The apparatus of claim 34, wherein the one or more NN-based filters have a filter support area of NxM, where N and M are integer numbers.

42. The apparatus of claim 41, wherein N and M are predefined values ranging from 0 to 17.

43. The apparatus of claim 42, wherein at least one of the one or more first filters has a shape selected from among: a rectangular, a diamond, or a cross.

44. The apparatus of claim 43, wherein a tap shape of the one or more NN-based filters has a height and width of NxN.

45. The apparatus of claim 34, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:determining whether NN-based filtering is enabled for the set of pixels; andin an instance in which NN-based filtering is enabled for the set of pixels, inputting one or more NN-based filter outputs with the information about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more first filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

46. The apparatus of claim 34, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:determining whether NN-based filtering is enabled for the set of pixels; andin an instance in which NN-based filtering is not enabled for the set of pixels, inputting the one or more first filter outputs with the information about the set of pixels into the one or more adaptive filters in the second stage of the multi-stage filtering process, and refraining from inputting the one or more NN-based filter outputs into the one or more adaptive filters in the second stage of the multi-stage filtering process.

47. The apparatus of claim 34, wherein the one or more first filters comprise a DB filter, a sample adaptive offset (SAO) filter, and a bilateral filter (BIF).

48. The apparatus of claim 47, wherein the instructions stored in the at least one memory, when executed by the at least one processor, further cause the apparatus to perform:in an instance in which NN-based filtering is not enabled for the set of pixels, filtering the information associated with the set of pixels sequentially using the DB filter, the SAO filter, and the BIF filter of the one or more first filters in the first stage of the multi-stage filtering process; andinputting, into the one or more adaptive filters in the second stage of the multi-stage filtering process, the information associated with the set of pixels and the one or more first outputs from the one or more first filters.

49. The apparatus of any one of claims 34 to 48, wherein the information transmitted in the bitstream about the one or more adaptive filter parameters used during the multi-stage filtering process comprises one or more of: a filter coefficient, a pixel value offset, or clipping range information.

50. The apparatus of claim 34, wherein the multi-stage filtering process comprises one or more models for intra-template matching (IntraTMP) selected from among: a cross-component prediction model, a cross-component linear model, a cross-component residual model, a crosscomponent reconstruction mode, a convolutional cross-component model, or a multi-cross component linear model.

51. The apparatus of any prior claim, wherein the apparatus is, comprises, or is comprised in, an encoder device.

52. The apparatus of any one of claims 1 to 50, wherein the apparatus is, comprises, or is comprised in, a decoder device.

53. The apparatus of any prior claim, wherein the multi-stage filtering process is, comprises, or is comprised in, a multi-phase adaptive loop filter (ALF) process.

54. The apparatus of any prior claim, wherein the multi-stage filtering process is, comprises, or is comprised in, an enhanced compression model (ECM).