Method and apparatus for performing an ai-based in-loop filtering

WO2026177503A1PCT designated stage Publication Date: 2026-08-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/002699
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-21
Filing Date
2026-02-13
Publication Date
2026-08-27

Smart Images

  • Figure KR2026002699_27082026_PF_FP_ABST
    Figure KR2026002699_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments herein disclose a method for performing an AI-based in-loop filtering in a decoder. The method may include obtaining a bitstream for each encoded pixel block of a source video frame, wherein the bitstream includes one or more quantization coefficients, at least one selection flag indicating availability of source error statistics as signaled in the bitstream. The method may include determining the source error statistics based on the at least one selection flag. The method may include generating a reconstructed video frame corresponding to a source video frame using the quantization coefficients. The method may include generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and the source error statistics to an AI model.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR PERFORMING AN AI-BASED IN-LOOP FILTERING

[0001] Embodiments disclosed herein relate to video codecs, and more particularly to generalizable Artificial Intelligence (AI)-based in-loop filters for video codecs, wherein the AI-based in-loop filters use quantization error estimation.

[0002] In-Loop Filter (ILF) in video codec is used to improve quality of reconstructed video frames by removing compression artifacts, such as, but not limited to, blockiness and ringing. There are several AI-based ILF proposals before the next generation video coding standardization bodies; for example, JVET (ISO / IEC & MPEG), AOMedia, and so on. However, most of them suffer from unequal compression gain (quality versus bitrate trade-off) across datasets containing diverse spatial resolutions (VGA, HD, UHD, 4K etc.), frame rates (24, 30, 60+fps), content / motion types (stationary, news-setting, sports-event, graphics, screen-content), capture-mode (pro-camera, UGC, CGI, Aerial Imagery), and so on.

[0003] The existing AI models tune "too well" to the training data & fail to perform well on a data of different characteristics. This overfitting can be resolved by introducing new training data for the intended use-case and re-training or fine-tuning an AI model. However, this is not only time / resource-intense, but many times, practically impossible due to unavailability of open training data that matches all of the above diverse data characteristics. This problem is exacerbated when the AI model is restricted by a specific allowed computational complexity which often limits the AI model's versatile learning ability.

[0004] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in a decoder. The method may include obtaining, by the decoder, a bitstream for each encoded pixel block of a source video frame, wherein the bitstream includes one or more quantization coefficients, at least one selection flag indicating availability of source error statistics as signaled in the bitstream. The method may include determining the source error statistics based on the at least one selection flag. The method may include generating a reconstructed video frame corresponding to a source video frame using the one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The method may include generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and the source error statistics to an artificial intelligence (AI) model.

[0005] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in an encoder is provided. The method may include generating a reconstructed video frame corresponding to a source video frame using one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The method may include obtaining first source error statistics by correlating the reconstructed pixel-blocks with corresponding encoded pixel-block of the source video frame. The method may include estimating second source error statistics by re-quantizing the one or more quantized coefficients in the bitstream (202) corresponding to each encoded pixel block. The method may include generating at least one selection flag indicating availability of source error statistics as signaled in a bitstream by selecting, for each encoded pixel block, one of the first source error statistics and the second source error statistics based on a difference between the first source error statistics and the second source error statistics, and the bits required to send the first source error statistics. The method may include generating the bitstream including the at least one selection flag, wherein the at least one selection flag indicates availability of source error statistics when the first source error statistics are selected, and wherein the at least one selection flag indicates non-availability of source error statistics when the second source error statistics are selected.

[0006] In an embodiment of the disclosure, the method for transmitting a bitstream is provided. The method may include performing the above image encoding method to generate the bitstream, and transmitting the bitstream.

[0007] Embodiments herein are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the following illustratory drawings. Embodiments herein are illustrated by way of examples in the accompanying drawings, and in which:

[0008] Figures 1a and 1b depict an overview of a standard compliant video decoder, according to related arts;

[0009] Figure 2 depicts a video decoder that enables generalizable AI-ILF, according to an embodiment of the disclosure;

[0010] Figures 3a and 3b depict a flow for performing an AI-based in-loop filtering in a video decoder, according to an embodiment of the disclosure;

[0011] Figure 4a and 4b depicts a flow for performing an AI-based in-loop filtering in a video encoder, according to an embodiment of the disclosure;

[0012] Figures 5a, 5b, and 5c conceptually illustrate how quantization error appears at the pixel-block level;

[0013] Figure 6 illustrates the usage of error statistics, according to an embodiment of the disclosure;

[0014] Figures 7a, 7b, 7c, and 7d illustrates re-quantization based error statistics estimation, according to an embodiment of the disclosure; and

[0015] Figures 8a, 8b, 8c, and 8d illustrates re-quantization based error statistics estimation: a study with JVET training dataset, according to an embodiment of the disclosure.

[0016] Figure 9 illustrates an embodiment of a system, according to an embodiment of the disclosure.

[0017] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.

[0018] The words / phrases "exemplary", "example", "illustration" "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," "i.e.," are merely used herein to mean "serving as an example, instance, or illustration. Any embodiment or implementation of the present subject matter described herein using the words / phrases "exemplary", "example", "illustration" "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.,", "i.e.," is not necessarily to be construed as preferred or advantageous over other embodiments.

[0019] Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.

[0020] It should be noted that elements in the drawings are illustrated for the purposes of this description and ease of understanding and may not have necessarily been drawn to scale. For example, the flowcharts / sequence diagrams illustrate the method in terms of the steps required for understanding of aspects of the embodiments as disclosed herein. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. Furthermore, in terms of the system, one or more components / modules which comprise the system may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

[0021] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any modifications, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings and the corresponding description. Usage of words such as first, second, third etc., to describe components / elements / steps is for the purposes of this description and should not be construed as sequential ordering / placement / occurrence unless specified otherwise.

[0022] In existing codec systems, the encoder generates, for each coding block, a predicted block using one or more prediction modes based on partially-coded current frames and / or fully-reconstructed reference frames available in the decoded picture buffer (DPB). The difference between the source block and the predicted block constitutes the residual signal, which is subjected to an orthogonal transformation (e.g., discrete cosine transform) to achieve energy compaction, followed by quantization of the transformed coefficients to reduce bit-depth and achieve compression, thereby introducing quantization error. The quantized transform coefficients along with control signals such as block partition flags, prediction modes, reference frame identifiers, and quantization parameters are entropy-coded to form the encoded video bitstream.

[0023] At the decoder, on receiving the encoded bitstream, a reconstructed residual is generated by inverse quantization and inverse transformation of the coded coefficients. The reconstructed block is obtained by adding the predicted block and the reconstructed residual, wherein the reconstructed block contains quantization error. This reconstructed block is subsequently processed using in-loop filters (ILF) such as deblocking, SAO, or ALF to reduce the perceptual impact of reconstruction error, and the filtered block is stored in the DPB for use in subsequent prediction. To avoid drift, the encoder includes an in-loop decoder that replicates the decoder-side reconstruction path so that prediction signals remain identical at both ends. Existing codecs lack access to source-level quantization error statistics at the decoder, and conventional in-loop filters are unable to adapt their behavior based on the true distortion characteristics of individual blocks.

[0024] Embodiments herein achieve generalizable Artificial Intelligence (AI)-based in-loop filters for video codecs, wherein the AI-based in-loop filters use quantization error estimation. Referring now to the drawings, and more particularly to Figures 2 through 8, where similar reference characters denote corresponding features consistently throughout the figures, there are shown embodiments.

[0025] Embodiments herein improve In Loop Filtering using an AI-based technique, targeting superior compression performance (% BD-BR) across variety of datasets, thus improving generalization of the trained neural network model.

[0026] Embodiments herein use quantization error statistics as input to a neural network (NN) model used for AI-based ILF, which allows the model to adapt and generalize on the characteristics of error that it aims to recover. The proposed estimation of error statistics contains a fundamental study, and a re-quantization trick.

[0027] Embodiments herein improve the compression performance in terms of Bjontegaard Data (BD) rate on a diverse set of content, as specified in the JVET common test conditions (CTC).

[0028] The principal object of embodiments herein is to disclose generalizable Artificial Intelligence (AI)-based in-loop filters for video codecs, wherein the AI-based in-loop filters use quantization error estimation.

[0029] Another object of embodiments herein is to disclose methods and systems for using quantization error statistics as input to a neural network (NN) model used for AI-based ILF, which allows the model to adapt and generalize on the characteristics of error that it aims to recover, wherein the proposed estimation of error statistics contains a fundamental study, and a re-quantization trick.

[0030] Another object of embodiments herein is to disclose methods and systems for deriving or estimating source error statistics for each encoded pixel block at the decoder, based on selection flags and coded quantized coefficients, thereby enabling block-wise characterization of quantization error during the decoding process.

[0031] Another object of embodiments herein is to provide encoder and decoder mechanisms for selectively transmitting or deriving source error statistics based on a blockwise decision, thereby minimizing signalling overhead while maintaining accuracy in AI-based in-loop filtering.

[0032] Another object of embodiments herein is to disclose AI-based in-loop filtering systems configured to consume reconstructed pixel-blocks together with source error statistics, thereby generating enhanced reconstructed frames with improved quality and generalization across diverse content types.

[0033] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the spirit thereof, and the embodiments herein include all such modifications.

[0034] Figures 1a and 1b depict an overview of a typical standard compliant video decoder and encoder respectively, according to related art. Video decoding process is specified such that all decoders that conform to a specified standard; for example, VVC will produce identical video frames, given an encoded bitstream. The bitstream typically contains several types of information, one of which is control signals that dictate decoder to invoke certain processes. For example, a split_cu flag indicates whether the decoder should split the basic processing unit known as coding tree-unit (CTU) into further coded-units (CUs), wherein each CU containscoding blocksas leaves. Another example is prediction_mode and ref_frame information that specifies how prediction signal for a coded block needs to be formed. Whereas, another type of bitstream information is actual frame / pixel data; for example, the residual signal often in the form of quantized transform coefficients.

[0035] For a coding block, the decoder first generates , the predicted block using the prediction mode and partially-decoded current frame and / or fully reconstructed reference-frame(s) available in the decoded picture buffer (DPB), based on the control signals relating to prediction mode. Since the encoder reconstructs each frame using an in-loop decoder, the prediction signals at encoder and decoder are identical. This is critical to avoid a phenomenon known as .

[0036] Then, the residual block is reconstructed from the coded pixel data, i.e. quantized transformed coefficients representing the difference between source and prediction, using the control signals such as quantization parameter (QP), transform size & type.

[0037] It is known that video encoder applies an orthogonal (for example, discrete cosine) transformation on the actual residual

[0038] , which is the difference between , the source block and :

[0039] ... Eq.(1)

[0040] Transformation is followed by the critical quantization operation in which each transformed-coefficient is represented in fewer bits to cause data-compaction. Due to the quantization process at the encoder, the reconstructed residual at the decoder is not identical to . Thus, the quantization error or block-error is the difference between residuals at decoder & encoder, which is given by:

[0041] ... Eq.(2)

[0042] Then, , the reconstructed block is generated by adding the predicted block to the reconstructed residual block. Thus, contains quantization error.

[0043] ... Eq.(3)

[0044] ... Eq.(4)

[0045] The reconstructed block is further enhanced using In-Loop filters (ILF) to reduce quantitative and visual impact of reconstruction errors. Let ILF output be such that a good ILF obeys , where is the expectation of the random variable used for modelling .

[0046] The resulting decoded block is stored in DPB which provides data for predicting blocks from subsequent video frames.

[0047] An example of a typical video encoder known as producing standard-compliant encoded bitstreams is shown in Figure 1B. In an embodiment herein, the coding block is compared with several predicted block-candidates in terms of their L1 / L2 distance to finally select the best-suited . Then, its corresponding is generated using Eq. (1). is the additional information that the current coding block contains. Hence it needs to be conveyed to the decoder. To efficiently do this, a combination of principles e.g. energy compaction, human visual response plays an important role. Typically, de-correlating transforms such as Karhunen-Loeve transform (KLT) or discrete cosine transform (DCT) are suitable for energy compaction. Higher spatial frequencies are less visible to human eyes and thus they can be sent with less data.

[0048] A transformed coding block undergoes quantization of the transform-coefficients to produce coded pixel data that is entropy coded using a Huffman or a context adaptive binary arithmetic coding (CABAC) method.

[0049] As mentioned earlier, the encoder also includes an in-loop decoder to avoid drift such that the losses incurred on the decoder-side are exactly mimicked at the encoder in this predictive coding control-system.

[0050] Once the decoder-mimicked residue is formed, it is added to the predicted coding block to obtain reconstructed coding block which forms a part of an entire reconstructed video frame. This partly reconstructed frame is also stored for generating INTRA prediction for a future coding block of a current video frame.

[0051] As the entire reconstructed video frame is formed, the next step of refining it using in-loop filtering begins. The "lossy"nature of video codec can be due to the quantization process where the bit-depth of each transformed residual coefficient can be limited using a pre-determined / content-adaptive strategy.

[0052] An encoder can take several decisions on prediction mode, transform type, size, quantization parameters, in-loop filtering applications, and so on, based on application requirement typically quantified using a mathematical equation known as rate-distortion optimizer (RDO). The RDO is inspired by classical information theory based on Shannon's entropy concepts where there's a trade-off between data size / rate and the received signal's distortion / quality, and is often implemented in the form of a Lagrangian function.

[0053] Hence, there is a need in the art for solutions which will overcome the above mentioned drawback(s), among others.

[0054] Figure 2 depicts a video decoder 200 that enables generalizable AI-ILF, according to an embodiment of the disclosure. Embodiments herein disclose an estimation & use of block-error statistics of each coded block, since it provides content & encoding-specific information to the neural network (AI model). This additional information about the current block improves the AI model's performance causing a better generalization across diverse contents & datasets. Embodiments herein perform re-quantization to estimate the quantization error, which can cause another quantization to the encoder-quantized residue to produce a "re-quantized residue". The re-quantized residue is compared to quantized residue to estimate the quantization error. For convenience of description, the video decoder 200 may be referred to as a decoder, and the encoded video bitstream 202 may be simply referred to as a bitstream.

[0055] As shown in Figure 2, the decoder 200 includes an entropy decoder 204 configured to parse an encoded video bitstream 202. In an example, the entropy decoder uses context-adaptive binary arithmetic coding (CABAC) in order to recover elements associated with each coding block. The encoded bitstream 202 comprises control signals 205 such as prediction_mode, ref_frame information, and transform / quantization parameters that guide the reconstruction process, as well as coded pixel data in the form of quantized transform coefficients. For each coding block, the decoder first generates the predicted block using INTRA / INTER Predictor 206 from fully reconstructed reference pictures stored in the decoded picture buffer (DPB) 207. The residual block is then reconstructed by performing inverse quantization and inverse transformation on the quantized coefficients based on the syntax elements transmitted in the bitstream. Adding this reconstructed residual to the predicted block yields the reconstructed block, which contains quantization error resulting from the encoder-side lossy quantization process. The reconstructed block is further processed by normative in-loop filters such as deblocking, SAO, and ALF to obtain to obtain the filtered reconstructed block, which is then stored in the DPB and forms the basis for generating subsequent predictions.

[0056] Meanwhile, in the disclosure, the encoding parameters may include other parameters required for decoding.

[0057] In accordance with the embodiments herein, the decoder 200 comprises an error-statistics estimator 208 configured to derive (obtain) or estimate the source error statistics for each reconstructed coding block. As illustrated in Figure 2, the quantized transform coefficients decoded from the encoded bitstream 202 are provided not only to the inverse-Q and inverse transform modules of the reconstruction path but also to a re-quantizer, together with the control signals 205 such as quantization parameter (QP), transform size / type, and prediction signalling. Let be the quantization parameter (QP) used for compressing the residual signal for the current coded block. The re-quantizer further quantizes the quantized-transformed coefficients to generate , using another QP: , where such that is a pre-determined function, mapping .

[0058] The re-quantized coefficients are scaled and transformed back to pixel domain to generate , a re-quantized residual pixel block. It is compared with to obtain error-estimate :

[0059] ... Eq.(5)

[0060] The re-quantizer simulates an alternative quantization behavior by producing re-quantized coefficients, which are then de-scaled and inverse-transformed to generate a re-quantized residual pixel block. This residual is compared with the reconstructed residual to obtain an estimate of the quantization error. The error-estimate is used to extract important error statistics such as min(), max(), mean(), standard-deviation(). The error statistics of each coded block is stored and fed to the AI-ILF to improve its generalizability. These statistics, together with the reconstructed pixel-blocks, are fed to a generalizable AI-based in-loop filter (AI-ILF) 210. By incorporating the block-level quantization error, the AI-ILF 210 produces an enhanced reconstructed block that reduces the visual and quantitative impact of compression artifacts and improves generalizability across diverse content types.

[0061] Figure 3a and 3b illustrate a flow for performing an AI-based in-loop filtering in a video decoder 200, according to an embodiment of the disclosure. In an embodiment of the disclosure, the decoder of apparatus may receive an encoded video bitstream 302 representing a source video frame and the bitstream may be for each encoded pixel block of a source video frame. The received bit-stream 302 includes one or more quantization coefficients, and at least one selection flag 304 indicating availability of source error statistics as signaled in the bitstream 302 corresponding to each encoded pixel block of the source video frame. the decoder may check if error statistics are available based on the at least one selection flag corresponding to each reconstructed pixel-block.

[0062] In an embodiment of the disclosure, the decoder may determine the source error statistics based on the at least one selection flag. If the error statistics are available, the electronic device may derive or obtain source error statistics 306 from the bitstream 302 when the at least one selection flag indicating availability of source error statistics. If the error statistics are not available, the decoder may estimate source error statistics by using the one or more quantized coefficients and one or more encoding parameters present in the bit-stream (e.g. a quantization parameter and transformation information) corresponding to each encoded pixel block. The decoder may estimate the source error statistics 308 by modifying the one or more encoding parameters or re-quantizing the values associated with the one or more quantized coefficients included in the bitstream 202 corresponding to each encoded pixel block when the at least one selection flag indicates non-availability of source error statistics.

[0063] In an embodiment of the disclosure, to re-quantize the one or more quantized coefficient, the decoder may determine a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter obtained from bitstream (202). And the decoder may divide transform coefficients by one or more quantization values based on the determined re-quantization parameter. And, the decoder may store the resulting re-quantized values in fixed-size data. The decoder may subtract the one or more quantized coefficients from the corresponding stored re-quantized values. The decoder may compute quantization error by performing an inverse transformation on the resulting subtracted values, wherein the inverse transformation comprises at least one of an inverse discrete cosine transform, an inverse discrete sine transform, an inverse wavelet transform, or an equivalent inverse signal domain transform. The decoder may estimate the source error statistics by using the computed quantization error.

[0064] For convenience of description, the source error statistics 306 may be referred to as first source error statistics, first error statistics, or computed error statistics, and the source error statistics 308 may be referred to as second source error statistics, second error statistics, or estimated error statistics. In an embodiment of the disclosure, in parallel with the error-statistics estimation, the decoder reconstructs each pixel block by applying inverse quantization and inverse transform to the decoded quantized coefficients and combining the resulting residual with the corresponding INTRA or INTER prediction. The set of all such reconstructed pixel-blocks 310 forms the reconstructed video frame corresponding to the source video frame.

[0065] In an embodiment of the disclosure, the decoder provides the reconstructed pixel-blocks 310 together with either the derived source error statistics 306 or the estimated source error statistics 308 as inputs to the artificial intelligence model. The AI model processes these inputs to generate an enhanced reconstructed video frame 312 with reduced compression artifacts and improved visual quality. The enhanced reconstructed video frame 312 may be stored in DPB.

[0066] In an embodiment of the disclosure, the AI model may be a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder 200 during in-loop filtering.

[0067] In an embodiment of the disclosure, within the AI model, a cross-attention operation may be performed. In general, cross-attention refers to an attention operation in which attention weights are computed based on one input (e.g., a query) and are applied to another input (e.g., a key / value), thereby enabling information from different inputs to be combined. The decoder may perform generating a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics; performing an element-wise product between the reconstructed pixel-blocks and the corresponding source error statistics (which may be referred to as, or may correspond to, a pointwise dot-product in this disclosure); and applying an attention mechanism by computing attention weights based on the reconstructed pixel-blocks or the corresponding source error statistics and by weighting the reconstructed pixel-blocks or the corresponding source error statistics using the attention weights as a cross attention operation.

[0068] Figure 3b illustrates the flow chart depicting the method for performing an AI-based in-loop filtering in the video decoder 200, according to an embodiment of the disclosure. At step S310, the decoder 200 may obtain, by the decoder (200), a bitstream (202) for each encoded pixel block of a source video frame, wherein the bitstream (202) includes one or more quantization coefficients and at least one selection flag indicating availability of source error statistics as signaled in the bitstream. The decoder 200 may receive an encoded video bit-stream representing a source video frame, the received bit-stream comprising, for each encoded pixel block of the source video frame, at least one selection flag indicating availability of source error statistics.

[0069] At step S320, the decoder 200 may determine the source error statistics based on the at least one selection flag. the decoder 200 may perform one of deriving the source error statistics from the encoded video bit-stream, when the at least one selection flag indicates the availability of the source error statistics, or estimating the source error statistics, when the at least one selection flag indicates non-availability of source error statistics, the estimation being performed by modifying one or more encoding parameters and re-quantizing one or more quantized coefficients present in the encoded video bit-stream corresponding to each encoded pixel block. The decoder may estimate the source error statistics by re-quantizing one or more quantized coefficients included in the bitstream (202) corresponding to each encoded pixel block when the at least one selection flag indicates non-availability of source error statistics. The decoder may obtaining the source error statistics from the bitstream (202), when the at least one selection flag indicates availability of source error statistics.

[0070] At step S330, the decoder 200 may generate a reconstructed video frame corresponding to a source video frame using the one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The decoder 200 generates, in parallel, a reconstructed video frame corresponding to the source video frame using the one or more quantization coefficients and the one or more encoding parameters, the reconstructed video frame including one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame.

[0071] At step S340, the decoder 200 may generate an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and the source error statistics to an artificial intelligence (AI) model (210). The decoder 200 generates an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and one of the derived source error statistics or the estimated source error statistics to an artificial intelligence (AI) model.

[0072] According to the embodiments herein, the decoder 200 further receives the encoded video bitstream 202, wherein, for each encoded pixel block of the source video frame, the encoded video bitstream includes at least one of a selection flag indicating availability or non-availability of source error statistics, and computed source error statistics corresponding to the encoded pixel block. The source error statistics comprise one or more of minimum, maximum, mean, variance, standard deviation, and parameters of a probabilistic distribution of the source error, wherein the probabilistic distribution can be one of Laplace, Gaussian, Generalized Gaussian, or similar distributions. Modifying the one or more encoding parameters comprises determining a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter extracted from the received encoded video bitstream. Re-quantizing the one or more quantized coefficients comprises dividing the transform coefficients by one or more quantization values and storing the resulting re-quantized values in fixed-size data, wherein the one or more quantization values are determined based on the determined re-quantization parameter.

[0073] Estimating the source error statistics comprises subtracting the one or more quantized coefficients from the corresponding stored re-quantized values, computing quantization error by performing an inverse transformation on the resulting subtracted values, wherein the transformation comprises at least one of a discrete cosine transform, sine transform, wavelet transform, or an equivalent signal-domain transform, and estimating the source error statistics by processing the computed quantization error, wherein such processing comprises computing one or more of minimum, maximum, mean, variance, standard deviation, and parameters of a probabilistic distribution of the computed quantization error. The artificial intelligence (AI) model is a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder during in-loop filtering. The pre-trained neural network model processes the reconstructed pixel-blocks and the corresponding source error statistics, wherein such processing comprises one or more of a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics or processed variants thereof, a pointwise dot-product of the reconstructed pixel-blocks and the corresponding source error statistics, and an attention mechanism in which one or more pixels from the reconstructed pixel-blocks are multiplied by one or more pixels from the reconstructed pixel-blocks or from the corresponding source error statistics.

[0074] Meanwhile, a description identical to that of Figure 3a is redundant and thus omitted.

[0075] Figure 4a and 4b illustrate a flow for performing an AI-based in-loop filtering in a video encoder, according to an embodiment of the disclosure.

[0076] In an embodiment of the disclosure, the encoder may receive or obtain a source video frame 402 and encoding it into an encoded video bit-stream using the conventional block-based prediction, transformation, quantization and entropy-coding pipeline. The encoder may generate a reconstructed video frame corresponding to the source video frame by applying inverse quantization, inverse transform, and prediction reconstruction on the encoded coefficients and coding parameters. The reconstructed video frame includes a plurality of reconstructed pixel-blocks 404 corresponding to the plurality of encoded pixel blocks of the source video frame.

[0077] In an embodiment of the disclosure, the encoder may compute, for each encoded pixel , source error statistics 406 by correlating the reconstructed pixel-block with the corresponding original pixel-block of the source video frame. These computed statistics characterize the true quantization error introduced by the encoder. In parallel, the encoder may estimate source error statistics 408 by using the quantized coefficients and / or encoding parameters present in the bit-stream. The estimation is performed by modifying one or more encoding parameters (e.g., quantization parameter or transform information), re-quantizing the quantized coefficients under the modified parameter conditions, and inverse-transforming the resulting re-quantized values to obtain an estimate of the quantization error for that block.

[0078] In an embodiment of the disclosure, the encoder may select, for each encoded pixel block, one of the computed source error statistics or the estimated source error statistics. The selection is based on evaluating which statistics provide the better trade-off between accuracy and signaling overhead. In an embodiment, a rate-distortion optimizer (RDO) compares (i) the difference between the computed statistics and the estimated statistics, and (ii) the bits required to signal the computed statistics. Based on this comparison, the encoder either selects the computed statistics for transmission or selects the estimated statistics and signals only a selection flag 412.

[0079] In an embodiment of the disclosure, the encoder may generate an enhanced reconstructed video frame 414 by feeding the reconstructed pixel-blocks and the selected source error statistics to an artificial intelligence model configured for AI-based in-loop filtering. The AI model processes the input to produce a refined reconstruction. The encoder may store the enhanced reconstructed video frame 414 in the DPB. And, the encoder may generate the encoded video bit-stream including, for each encoded pixel block, at least one of the selection flag indicating availability of source error statistics, and the computed source error statistics.

[0080] Referring to Figure 4b, at step S410, the encoder may generate a reconstructed video frame corresponding to a source video frame using one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame.

[0081] At step S420, the encoder may obtain first source error statistics by correlating the reconstructed pixel-blocks with corresponding encoded pixel-block of the source video frame.

[0082] At step S430, the encoder may estimate second source error statistics by re-quantizing one or more quantized coefficients corresponding to each encoded pixel block.

[0083] At step S440, the encoder may generate at least one selection flag indicating availability of source error statistics as signaled in a bitstream by selecting, for each encoded pixel block, one of the first source error statistics and the second source error statistics based on a difference between the first source error statistics and the second source error statistics, and the bits required to send the first source error statistics. The encoder may select one of the first source error statistics and the second source error statistics based on a Lagrangian-based objective function.

[0084] At step S440, the encoder may generate the bitstream including the at least one selection flag. The at least one selection flag may indicate availability of source error statistics when the first source error statistics are selected. The at least one selection flag may indicate non-availability of source error statistics when the second source error statistics are selected. The selection flag and the first source error statistics may be represented using one or more of a predictive coding or an entropy coding, wherein the entropy coding includes Huffman coding or an Exponential-Golomb coding.

[0085] In an embodiment of the disclosure, the encoder may generate an enhanced reconstructed video frame corresponding to the source video frame by feeding the one or more reconstructed pixel-blocks and the selected source error statistics to an AI model.

[0086] Meanwhile, a description identical to that of Figure 4a is redundant and thus omitted. In addition, since the operations performed by the decoder (including computing and estimating the source error statistics) of Figure 3a and 3b may likewise be performed by an encoder, a duplicate description thereof is also omitted. Figures 5a, 5b, and 5c conceptually illustrate how quantization error appears at the pixel-block level. As shown in Figure 5a, when a 128x128 block 511 from a test frame 510 is examined, the original signal 512 corresponding to the 128x128 block 511, the error representation 514 of the reconstructed signal 513 before filtering, and the error representation 516 of the enhanced reconstructed signal 515 after applying an AI-based ILF exhibit clear differences. These error maps highlight that even for the same block, the quantization error has a specific structure, and the residual distortion after filtering is not uniform across the block.

[0087] In an embodiment of the disclosure, a distribution 517 may represent, for each of a VTM Error (e.g., an error representation of the reconstructed pixel-block) and an AILF Error (e.g., an error representation of the enhanced reconstructed pixel-block), a distribution of a number of pixels corresponding to a (Rec-Orig) compression-error distribution. Further, a distribution 518 may represent a differential distribution of a number of pixels corresponding to a (VTM-AILF) compression-error distribution. In addition, a representation 519 may represent distribution 518 in the form of an error representation.

[0088] Figure 5b further shows that when the same block is coded with different QPs (for example, 22, 27, 32, 37, 42), the resulting reconstructed blocks and their corresponding error maps differ noticeably.

[0089] For example, the original block 522 may correspond the source pixel block 521 of the source frame 520. And, the reconstructed pixel-blocks 523, 525, 527, 529, and 531 may respectively represent pixel-blocks reconstructed when QPs are 22, 27, 32, 37, and 42. And the error representation may respectively represent the errors 524, 526, 528 ,530, and 532 corresponding to the reconstructed pixel-blocks 523, 525, 527, 529, and 531 when QPs are 22, 27, 32, 37, and 42.

[0090] Figure 5c shows the histogram or PDF curves plotted for these QPs also change shape, width, and peak height, indicating that the distribution of the quantization error is highly sensitive to QP.

[0091] From these observations, it becomes clear that important statistics such as minimum, maximum, mean, and standard deviation vary significantly across QPs and across content. The quantization error therefore captures meaningful information about how the coding process has affected each block. Since these variations are large and content-dependent, the quantization-error statistics become an important signal that should be extracted and used for improving in-loop filtering. A lightweight neural network operating inside an ILF cannot always learn or retain all relevant filter kernels needed to infer these statistics directly from the reconstructed pixels. Hence, providing these block-level error statistics explicitly helps the AI-based ILF operate more reliably and improves its ability to generalize across different content types and quantization conditions.

[0092] Figure 6 illustrates the usage of error statistics, according to an embodiment of the disclosure.

[0093] To summarize, as discussed earlier, embodiments herein propose a generalizable AI model which is the estimation and use of block-level error statistics for each coded block, since these statistics convey content-specific and encoding-specific information directly to the neural network. Providing this additional information about the current block enables the AI model to adapt better, leading to improved performance and stronger generalization across a wide range of contents and datasets. Figure 6 illustrates how the data input 602 comprising the proposed error statistics 603 and the existing inputs 604 are supplied as inputs to the neural network inside the AILF (AI In loop filter) 606 which then provides an enhanced output 608. The computed or estimated statistics act as additional cues that the neural network can make use of during enhancement. The AI model itself is pre-trained with this error information included, so its trainable weights, biases, and other parameters are fixed beforehand, making the model straightforward to deploy on hardware.

[0094] Meanwhile, the existing inputs 604 may be reconstructed pixel-blocks. So the AI model (e.g., AIIL) may output the enhanced reconstructed video frame by obtaining the reconstructed video frame and the source error statistics as inputs. But, the example of the existing inputs is not limited thereto.

[0095] Table 1 below further shows examples of data diversity that stand to benefit from this AILF generalization. In particular, screen-content and animation sequences differ significantly from natural images and often exhibit distinct error patterns. Without explicitly supplying error statistics to the AILF, a low-complexity neural network may not have the capacity to learn all such variations during training. By providing these statistics explicitly, the model receives a useful handcrafted signal that helps reduce model variance when encountering new or unseen types of content. This is especially important for a future-ready next-generation video coding standard, where new capture devices and applications may introduce content with characteristics not represented in today's training datasets.

[0096]

[0097] [Table 1]

[0098] Figures 7a-7d illustrate an example of how re-quantization can be used to estimate key quantization-error statistics, according to the embodiments herein.

[0099] In an embodiment of the disclosure, Figure 7a, illustrates an example of a pixel-block 712 from the input image 710 is shown.

[0100] In an embodiment of the disclosure, Figure 7b illustrates an example of pixel-blocks and difference representations, and transformed domain representation corresponding to the source block.

[0101] In an embodiment of the disclosure, the source pixel-block 720 may be an example of source pixel-block . The reconstructed pixel-block 721 may be an example of the reconstructed pixel-block corresponding to the source pixel-block 720, obtained after decoding. The predicted pixel-block 722 may be an example of the predicted pixel-block corresponding to the source pixel-block 720. The quantization error 723 may be computed as the difference between the reconstructed signal of reconstructed pixel-block 721 and the source signal of source pixel-block 720. And, the encoder may have a residual defined as , shown in residual pixel-block 724. This residual is then transformed and quantized to produce the quantized residual , depicted in decoded residual pixel-block 725. This residual may be same as .

[0102] As described earlier, the key step is to apply an additional quantization to the encoder-side quantized residue to generate a "re-quantized" residue. The DCT-transformed versions of these residues are shown in a transformed domain representation 726 and a rounded transformed domain representation 727 of the decoded residual pixel-block 725, where the rounded transformed domain representation 727 may correspond to the rounded (int16) representation. A transformed domain representation of re-quantized DCT coefficients 728 may show the re-quantized DCT coefficients on a logarithmic scale.

[0103] In an embodiment of the disclosure, Figure 7c illustrates an example of inverse transformed domain representation, and estimated quantization error corresponding to the source block.

[0104] In an embodiment of the disclosure, the inverse-transformed residues corresponding to these signals are shown in a rounded inverse transformed domain representation 730 and a inverse transformed domain representation 731 associated with the decoded residual pixel-block 725. Specifically, the rounded inverse transformed domain representation 730 may be a result of the inverse transformed associated with the rounded transformed domain representation 727. And, the inverse transformed domain representation 731 may be a result of the inverse transformed associated with the transformed domain representation 726. and the difference between the rounded inverse transformed domain representation 730 and the inverse transformed domain representation 731, representing the estimated quantization error, is highlighted in Estimated quantization error representation 732.

[0105] In an embodiment of the disclosure, Figure 7c illustrates an example of the quantization error and estimated quantization error. By comparing the re-quantized residual with the original quantized residual, we can estimate the quantization error of the encoded block. A result graph 740 plots the histogram (PDF) of the actual quantization error alongside the histogram of the estimated error. The two distributions closely resemble each other, indicating that the re-quantization-based approach provides a reliable estimate of the true quantization error. The statistics derived from these histograms, such as minimum, maximum, mean, and standard deviation, are therefore meaningful and can be used effectively within the proposed AI-based in-loop filtering framework.

[0106] An experiment was performed to study the accuracy of the proposed error-estimation method using the JVET-mandated DIV2K image dataset. For this evaluation, four key statistics, minimum, maximum, mean, and standard deviation of the quantization error, were estimated using the re-quantization approach described earlier. In this setup, the re-quantization parameter (re-QP) was chosen as the original QP with a negative offset. This is because quantization at the original QP level discards a considerable amount of information; therefore, using the same QP again would tend to overestimate the error. A slightly lower QP is thus preferred for generating a more realistic estimate of the underlying quantization error.

[0107] Figures 8a, 8b, 8c, and 8d show the estimated error statistics plotted against the ground-truth values for 2342 randomly sampled pixel-blocks from the dataset, according to the embodiments herein.

[0108] In an embodiment of the disclosure, Figure 8a is an example scatter plot of a correlation between minimum of the computed source error statistics and minimum of the estimated source error statistics. Figure 8b is an example scatter plot of a correlation between maximum of the computed source error statistics and maximum of the estimated source error statistics. Figure 8c is an example scatter plot of a correlation between mean of the computed source error statistics and mean of the estimated source error statistics. Figure 8d is an example scatter plot of a correlation between standard deviation of the computed source error statistics and standard deviation of the estimated source error statistics.

[0109] The 45-degree line in each plot represents the ideal case where the estimated value matches the ground truth for every block. As can be observed, the minimum, maximum, and standard-deviation estimates exhibit strong correlation with the corresponding ground-truth values (also reflected in the Pearson correlation coefficients shown in below table2. The mean error plot shows only a small spread around zero, which is expected because many blocks naturally have near-zero mean quantization errors.

[0110]

[0111] [Table 2]

[0112] These results show that the four statistics, minimum, maximum, mean, and standard deviation, can be estimated with high accuracy using the re-quantization process. These statistics therefore serve as reliable indicators of block-level quantization error and form meaningful auxiliary information for the AI-based in-loop filter. Below table shows the correlation coefficients.The embodiments herein further demonstrate that simulations performed using actual (true) error values at the decoder side lead to improved AILF performance and better generalization across diverse datasets. The proposed method has been implemented within the VTM software framework, and experimental results obtained using the CTC test dataset, as mandated under the JVET AhG11 neural network-based video coding (NNVC) activity, show that the method consistently outperforms state-of-the-art approaches across multiple classes of content. In addition, embodiments herein have been applied to AI models of varying complexity, and the results indicate that the improvement is model-agnostic, thereby confirming that the method is robust to architectural choices and can be integrated into a wide range of neural network-based in-loop filtering designs.

[0113] Embodiments herein can provide a better generalization of compression performance across variety of data, wherein an AI-ILF model generalization is provided using error estimation that helps generalizing compression (BD-BR) performance across varied datasets characterized by different content-capture techniques, spatial dimensions, fps, and so on.

[0114] Embodiments herein offer a better compression efficiency as compared to an existing video codec, wherein embodiments herein estimate quantization error(s), and uses it as input to the AI-ILF neural network model. This provides improvement over existing (SOTA) AI-ILF model in terms of % BD-BR compression efficiency.

[0115] Embodiments herein propose a re-quantization trick for error statistic-estimation, wherein embodiments herein provide a method to estimate key error statistics such as min, max, mean, standard deviation using a re-quantization trick that implicitly uses content properties as well as the encoder-specified quantization parameter.

[0116] Embodiments herein disclose a method to enhance generalization and performance of AI-model used in ILF.

[0117] The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the network elements. The elements include blocks which can be at least one of a hardware device, or a combination of hardware device and software module.

[0118] Figure 9 illustrates an embodiment of a system, according to an embodiment of the disclosure.

[0119] As shown in Figure 9, the system 1000 includes a processor 1002, a memory 1004, a storage component 1006, an input interface 1008, an output interface 1010, a communication interface 1012, an encoder 400, a decoder 200, and an artificial intelligence (AI) model 210.

[0120] The processor 1002, as used herein, refers to any type of computational circuit that may comprise hardware elements and software elements. The processor 1002 may be embodied as a multi-core processor, a single-core processor, or a combination thereof, a distributed processing system, or the like. The processor 1002 may be a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an application-specific integrated circuit (ASIC), or another type of processing component. The processor 1002 may be at least one.

[0121] The memory 1004 includes a non-transitory computer-readable medium. The memory 1004 may include random-access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 1002. The memory 1004 may store machine-readable instructions executable by the processor 1002, and such instructions, when executed, cause the processor 1002 to perform one or more method steps described herein.

[0122] The storage component 1006 stores information and / or software related to the operation and use of the system 1000. For example, the storage component 1006 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid-state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.

[0123] The input interface 1008 may be configured to receive information, such as user input and / or control signals. For example, the input interface 1008 may include, but is not limited to, a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone. The output interface 1010 may be configured to convey information from the system 1000 to a user or other systems, and may include, for example, display devices and / or audio output devices. The output interface 1010 may be configured to output notifications and / or visualizations related to encoding and / or decoding operations, without being limited thereto.

[0124] The communication interface 1012 may be configured to transmit data between the system 1000 and external devices or networks. For example, the communication interface 1012 may be configured to handle reception of a set of original video frames from a camera module and / or transmission of encoded video data to a streaming server or another external device.

[0125] The encoder 400 may be configured to process a set of original video frames to generate an encoded video bitstream 202. In an embodiment, the encoder 400 may generate, for each coded block, a selection flag indicating whether source error statistics are available in the bitstream, and may selectively signal computed source error statistics or omit the computed source error statistics and instead enable estimation of source error statistics at a decoder side. In this manner, the encoder 400 can control signaling overhead while enabling improved in-loop filtering performance.

[0126] The decoder 200 may be configured to reconstruct video frames from the encoded video bitstream 202. In an embodiment, the decoder 200 may parse the selection flag from the encoded video bitstream 202, and based on the selection flag, obtain computed source error statistics (e.g., statistics explicitly signaled in the bitstream) or estimate source error statistics (e.g., statistics derived from quantized coefficients and / or decoding parameters). The decoder 200 may provide reconstructed pixel-blocks and the obtained or estimated source error statistics to the AI model 210, and the AI model 210 may generate enhanced reconstructed video frames. Accordingly, the system 1000 can improve reconstruction quality and robustness across diverse content types while maintaining low complexity suitable for hardware deployment.

[0127] Meanwhile, although described herein as a system for convenience of description, the system 1000 may represent a apparatus.

[0128] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in a decoder (200) is provided. The method may include obtaining, by the decoder (200), a bitstream (202) for each encoded pixel block of a source video frame, wherein the bitstream (202) includes one or more quantization coefficients, at least one selection flag indicating availability of source error statistics as signaled in the bitstream. The method may include determining the source error statistics based on the at least one selection flag;. The method may include generating a reconstructed video frame corresponding to a source video frame using the one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The method may include generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and the source error statistics to an artificial intelligence (AI) model (210).

[0129] In an embodiment of the disclosure, the method may include estimating the source error statistics by re-quantizing the one or more quantized coefficients in the bitstream (202) corresponding to each encoded pixel block when the at least one selection flag indicates non-availability of source error statistics.

[0130] In an embodiment of the disclosure, the method may include obtaining the source error statistics from the bitstream (202), when the at least one selection flag indicates availability of source error statistics.

[0131] In an embodiment of the disclosure, the source error statistics may include one or more of minimum, maximum, mean, variance, standard deviation of source error.

[0132] In an embodiment of the disclosure, the source error statistics may represent probabilistic distributions of the source error, wherein probabilistic distributions can be one of Laplace, Gaussian, Generalized Gaussian.

[0133] In an embodiment of the disclosure, the method may include determining a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter obtained from bitstream (202).

[0134] In an embodiment of the disclosure, the method may include dividing the transform coefficients by one or more quantization values based on the determined re-quantization parameter. The method may include storing the resulting re-quantized values in fixed-size data.

[0135] In an embodiment of the disclosure, the method may include subtracting the one or more quantized coefficients from the corresponding stored re-quantized values. In an embodiment of the disclosure, the method may include computing quantization error by performing an inverse transformation on the resulting subtracted values, wherein the inverse transformation comprises at least one of an inverse discrete cosine transform, an inverse discrete sine transform, an inverse wavelet transform, or an equivalent inverse signal domain transform. In an embodiment of the disclosure, the method may include estimating the source error statistics by using the computed quantization error.

[0136] In an embodiment of the disclosure, the artificial intelligence (AI) model 210 may be a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder 200 during in-loop filtering.

[0137] In an embodiment of the disclosure, the pre-trained neural network model processes the reconstructed pixel-blocks and the corresponding source error statistics by performing one or more of generating a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics; performing an element-wise product between the reconstructed pixel-blocks and the corresponding source error statistics; or applying an attention mechanism by computing attention weights based on the reconstructed pixel-blocks or the corresponding source error statistics and by weighting the reconstructed pixel-blocks or the corresponding source error statistics using the attention weights.

[0138] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in an encoder (400). The method may include generating a reconstructed video frame corresponding to a source video frame using one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The method may include obtaining first source error statistics by correlating the reconstructed pixel-blocks with corresponding encoded pixel-block of the source video frame. The method may include estimating second source error statistics by re-quantizing the one or more quantized coefficients in the bitstream (202) corresponding to each encoded pixel block. The method may include generating at least one selection flag indicating availability of source error statistics as signaled in a bitstream by selecting, for each encoded pixel block, one of the first source error statistics and the second source error statistics based on a difference between the first source error statistics and the second source error statistics, and the bits required to send the first source error statistics. The method may include generating the bitstream including the at least one selection flag. The at least one selection flag indicates availability of source error statistics when the first source error statistics are selected. The at least one selection flag indicates non-availability of source error statistics when the second source error statistics are selected.

[0139] In an embodiment of the disclosure, the selection flag and the first source error statistics may be represented using one or more of a predictive coding or an entropy coding, wherein the entropy coding includes Huffman coding or an Exponential-Golomb coding.

[0140] In an embodiment of the disclosure, the method may include selecting one of the first source error statistics and the second source error statistics based on a Lagrangian-based objective function.

[0141] In an embodiment of the disclosure, the method may include generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the one or more reconstructed pixel-blocks and the selected source error statistics to an AI model.

[0142] In an embodiment of the disclosure, a method for transmitting a bitstream is provided. The method may include performing the image encoding method to generate the bitstream, and transmitting the bitstream.

[0143] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in a video decoder (200) is provided. The method may include receiving, by a decoder (200), an encoded video bit-stream (202) representing a source video frame, wherein the received bit-stream (202) comprises, for each encoded pixel block of the source video frame, at least one selection flag indicating availability of source error statistics. The method may include performing, by the decoder (200), at least one of deriving the source error statistics from the encoded video bit-stream (202), when the at least one selection flag indicates the availability of the source error statistics; and estimating the source error statistics, when the at least one selection flag indicates non-availability of source error statistics, wherein the estimation is performed by modifying one or more encoding parameters and re-quantizing one or more quantized coefficients present in the encoded video bit-stream corresponding to each encoded pixel block. The method may include generating, in parallel, a reconstructed video frame corresponding to the source video frame using the one or more quantization coefficients and the one or more encoding parameters, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame. The method may include generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and one of the derived source error statistics or the estimated source error statistics to an artificial intelligence (AI) model (210).

[0144] In an embodiment of the disclosure, the method may include receiving, by the decoder (200), the encoded video bitstream (202), wherein, for each encoded pixel block of the source video frame, the encoded video bitstream (202) includes at least one of: the selection flag indicating one of availability or non-availability of source error statistics; and computed source error statistics corresponding to the encoded pixel block.

[0145] In an embodiment of the disclosure, the source error statistics may comprise one or more of minimum, maximum, mean, variance, standard deviation and parameters of a probabilistic distribution of the source error, wherein probabilistic distribution can be one of Laplace, Gaussian, Generalized Gaussian and such.

[0146] In an embodiment of the disclosure, the method may include determining a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter extracted from the received encoded video bitstream (202).

[0147] In an embodiment of the disclosure, the method may include dividing the transform coefficients by one or more quantization values and storing the resulting re-quantized values in fixed-size data, wherein the one or more quantization values are determined based on the determined re-quantization parameter.

[0148] In an embodiment of the disclosure, the method may include subtracting the one or more quantized coefficients from the corresponding stored re-quantized values. The method may include computing quantization error by performing an inverse transformation on the resulting subtracted values, wherein transformation comprises at least one of a discrete cosine, sine, wavelet transform, or an equivalent signal domain transform. The method may include estimating the source error statistics by processing the computed quantization error, wherein processing refers to computing one or more of minimum, maximum, mean, variance, standard deviation and parameters of a probabilistic distribution of the computed quantization error.

[0149] In an embodiment of the disclosure, the artificial intelligence (AI) model (210) may be a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder (200) during in-loop filtering.

[0150] In an embodiment of the disclosure, the pre-trained neural network model may process the reconstructed pixel-blocks and the corresponding source error statistics, wherein processing comprises one or more of: a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics and processed variants thereof; a pointwise dot-product of the reconstructed pixel-blocks and the corresponding source error statistics; and an attention mechanism, wherein the one or more pixels from the reconstructed pixel-blocks are multiplied by the one or more pixels from the reconstructed pixel-blocks or the one or more source error statistics.

[0151] In an embodiment of the disclosure, the selection flag and the computed source error statistics are represented using one or more of a predictive or an entropy coder.

[0152] In an embodiment of the disclosure, a decoder (200) for performing an AI-based in-loop filtering, wherein the decoder is configured to receive an encoded video bit-stream (202) representing a source video frame, wherein the received bit-stream comprises, to each encoded pixel block of the source video frame, at least one selection flag indicating availability of source error statistics; perform at least one of: deriving the source error statistics from the encoded video bit-stream, when the at least one selection flag indicates the availability of the source error statistics; or estimating the source error statistics, when the at least one selection flag indicates non-availability of source error statistics, wherein the estimation is performed by modifying one or more encoding parameters and re-quantizing one or more quantized coefficients present in the encoded video bit-stream corresponding to each encoded pixel block; generate in parallel, a reconstructed video frame corresponding to the source video frame using the one or more quantization coefficients and the one or more encoding parameters, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame; and generate an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and one of the derived source error statistics and the estimated source error statistics to an artificial intelligence (AI) model (210).

[0153] In an embodiment of the disclosure, the decoder may be configured to: receive an encoded video bit-stream, wherein, for each encoded pixel block of the source video frame, the encoded video bit-stream includes at least one of: selection flag indicating one of availability or non-availability of source error statistics; and computed source error statistics corresponding to the encoded pixel block.

[0154] In an embodiment of the disclosure, the decoder may be configured to modify the one or more encoding parameters by determining a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter extracted from the received encoded video bitstream.

[0155] In an embodiment of the disclosure, the decoder (200) may be configured to re-quantize the one or more quantized coefficients by dividing the transform coefficients by one or more quantization values and storing the resulting quantization values in fixed-size data, wherein the one or more quantization values are determined based on the determined re-quantization parameter.

[0156] In an embodiment of the disclosure, the decoder (200) may be configured to estimate the source error statistics by: subtracting the one or more quantized coefficients from the corresponding stored re-quantized values; computing quantization error by performing an inverse transformation on the resulting subtracted values, wherein transformation comprises at least one of a discrete cosine, sine, wavelet transform, or an equivalent signal domain transform; and estimating the source error statistics by processing the computed quantization error, wherein processing refers to computing one or more of minimum, maximum, mean, variance, standard deviation and parameters of a probabilistic distribution of the computed quantization error.

[0157] In an embodiment of the disclosure, the artificial intelligence (AI) model (210) may be a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder during in-loop filtering.

[0158] In an embodiment of the disclosure, the pre-trained neural network model may process the reconstructed pixel-blocks and the corresponding source error statistics, wherein processing comprises one or more of: a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics and processed variants thereof; a pointwise dot-product of the reconstructed pixel-blocks and the corresponding source error statistics; and an attention mechanism, wherein the one or more pixels from the reconstructed pixel-blocks are multiplied by the one or more pixels from the reconstructed pixel-blocks or the one or more source error statistics.

[0159] In an embodiment of the disclosure, the selection flag and the computed source error statistics may be represented using one or more of a predictive or an entropy coder.

[0160] In an embodiment of the disclosure, a method for performing an AI-based in-loop filtering in an encoder (400) is provided. The method may include encoding a source video frame into an encoded video bitstream. The method may include determining for each encoded pixel block, whether to transmit source error statistics. The method may include embedding in the encoded video bitstream, for each encoded pixel block, at least one of: a selection flag indicating availability of source error statistics; and the source error statistics.

[0161] In an embodiment of the disclosure, the method may include selecting between derived or estimated error statistics is based on a Lagrangian-based objective function.

[0162] In an embodiment of the disclosure, an encoder (400) for performing an AI-based in-loop filtering is provided. The encoder (400) may be configured to: encode a source video frame into an encoded video bitstream; determine for each encoded pixel block, whether to transmit source error statistics; and embed in the encoded video bitstream, for each encoded pixel block, at least one of: a selection flag indicating availability of source error statistics; and the source error statistics.

[0163] In an embodiment of the disclosure, the encoder (400) may be configured to select the source error statistics based on a Lagrangian-based objective function, wherein the source error statistics comprises one or more of a minimum, a maximum, a mean, a variance, a standard deviation of the source error, and wherein the source error statistics represents probabilistic distribution of the source error.

[0164] The embodiments disclosed herein describe generalizable Artificial Intelligence (AI)-based in-loop filters for video codecs, wherein the AI-based in-loop filters use quantization error estimation. Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means contain program code means for implementation of one or more steps of the method, when the program runs on a server or mobile device or any suitable programmable device. The method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device. The hardware device can be any kind of portable device that can be programmed. The device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g. an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein. The method embodiments described herein could be implemented partly in hardware and partly in software. Alternatively, the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.

[0165] The foregoing descrption of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of embodiments, those skilled in the art will recognize that the embodiments herein can be practised with modification within the scope of the embodiments as described herein.

Claims

1.A method for performing an AI-based in-loop filtering in a decoder (200), the method comprising:obtaining, by the decoder (200), a bitstream (202) for each encoded pixel block of a source video frame, wherein the bitstream (202) includes one or more quantization coefficients and at least one selection flag indicating availability of source error statistics as signaled in the bitstream;determining the source error statistics based on the at least one selection flag;generating a reconstructed video frame corresponding to a source video frame using the one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame; andgenerating an enhanced reconstructed video frame corresponding to the source video frame by feeding the reconstructed pixel-blocks and the source error statistics to an artificial intelligence (AI) model (210).2.The method of claim 1, wherein the determining of the error statistics comprises:estimating the source error statistics by re-quantizing one or more quantized coefficients included in the bitstream (202) corresponding to each encoded pixel block when the at least one selection flag indicates non-availability of source error statistics.3.The method of claim 1, wherein the determining of the source error statistics comprises:obtaining the source error statistics from the bitstream (202), when the at least one selection flag indicates availability of source error statistics.4.The method of any one of claims 1 to 3, wherein the source error statistics comprise one or more of minimum, maximum, mean, variance, standard deviation of source error.5.The method of any one of claims 1 to 4, wherein the source error statistics represents probabilistic distributions of the source error, wherein probabilistic distributions can be one of Laplace, Gaussian, Generalized Gaussian.6.The method of claim 2, wherein the estimating of the source error statistics comprises:determining a re-quantization parameter for re-quantizing the one or more quantized transform coefficients based on at least one quantization parameter obtained from bitstream (202).7.The method of claim 2, wherein the estimating of the source error statistics comprises:dividing transform coefficients by one or more quantization values based on the determined re-quantization parameter; andstoring the resulting re-quantized values in fixed-size data.8.The method of claim 7, wherein estimating the source error statistics comprises:subtracting the one or more quantized coefficients from the corresponding stored re-quantized values;computing quantization error by performing an inverse transformation on the resulting subtracted values, wherein the inverse transformation comprises at least one of an inverse discrete cosine transform, an inverse discrete sine transform, an inverse wavelet transform, or an equivalent inverse signal domain transform; andestimating the source error statistics by using the computed quantization error.9.The method of any one of claims 1 to 8, wherein the artificial intelligence (AI) model (210) is a pre-trained neural network model having one or more trainable parameters including weights and biases, wherein the trainable parameters are predefined and fixed for use by the decoder (200) during in-loop filtering.10.The method of claim 9, wherein the pre-trained neural network model processes the reconstructed pixel-blocks and the corresponding source error statistics by performing one or more of:generating a weighted combination of the reconstructed pixel-blocks and one or more of the corresponding source error statistics;performing an element-wise product between the reconstructed pixel-blocks and the corresponding source error statistics; andapplying an attention mechanism by computing attention weights based on the reconstructed pixel-blocks or the corresponding source error statistics and by weighting the reconstructed pixel-blocks or the corresponding source error statistics using the attention weights.11.A method for performing an AI-based in-loop filtering in an encoder (400), the method comprising:generating a reconstructed video frame corresponding to a source video frame using one or more quantization coefficients, wherein the reconstructed video frame includes one or more reconstructed pixel-blocks corresponding to each encoded pixel block of the source video frame;obtaining first source error statistics by correlating the reconstructed pixel-blocks with corresponding encoded pixel-block of the source video frame;estimating second source error statistics by re-quantizing one or more quantized coefficients corresponding to each encoded pixel block;generating at least one selection flag indicating availability of source error statistics as signaled in a bitstream by selecting, for each encoded pixel block, one of the first source error statistics and the second source error statistics based on a difference between the first source error statistics and the second source error statistics, and the bits required to send the first source error statistics; andgenerating the bitstream including the at least one selection flag,wherein the at least one selection flag indicates availability of source error statistics when the first source error statistics are selected, andwherein the at least one selection flag indicates non-availability of source error statistics when the second source error statistics are selected.12.The method of claim 11, wherein the selection flag and the first source error statistics are represented using one or more of a predictive coding or an entropy coding, wherein the entropy coding includes Huffman coding or an Exponential-Golomb coding.13.The method of any one of claims 11 to 12, wherein the generating of the at least one selection flag indicating availability of source error statistics as signaled in the bitstream comprises:selecting one of the first source error statistics and the second source error statistics based on a Lagrangian-based objective function.14.The method of any one of claims 11 to 12, the method further comprises:generating an enhanced reconstructed video frame corresponding to the source video frame by feeding the one or more reconstructed pixel-blocks and the selected source error statistics to an AI model.15.A method for transmitting a bitstream, comprising:performing the image encoding method of claim 11 to generate the bitstream, andtransmitting the bitstream.