Method and apparatus for performing scalable de-quantization in a video codec
An AI-guided pixel classification system in video codecs addresses high complexity in ILF tools by categorizing pixels and applying class-specific offsets, enhancing video quality and reducing computational demands.
Patent Information
- Application Number
- PCT/KR2025/005416
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-22
- Filing Date
- 2025-04-22
- Publication Date
- 2025-10-30
Smart Images

Figure KR2025005416_30102025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR PERFORMING SCALABLE DE-QUANTIZATION IN A VIDEO CODEC
[0001] Embodiments disclosed herein relate to video processing systems, and more particularly to methods and systems for performing scalable de-quantization in video codec based on Artificial Intelligence (AI)-guided pixel classification.
[0002] In typical video codecs, data compression is achieved by quantizing transform-domain coefficients representing the residual data after prediction.
[0003] FIG. 1 depicts an example Versatile Video Coding (VVC) video decoder. Typically, an image / video compression method involves representing a signal in a smaller number of bits (digital storage units). At encoder, the prediction residue (i.e., additional information) from a source signal is first transformed into different domain such as a Fourier / spatial frequency space. The obtained transform coefficients become representation of the signal which is to be compressed. Based on human visual system response to signal of various frequencies, these coefficients are quantized using a higher data compression for lower visual response (typically, at higher frequencies). Due to quantization, effective number of code-words (values) that a signal takes can be reduced.
[0004] As the video decoder reconstructs the signal, the reconstructed signal different than the source signal causing distortions in objective sense (for example, a mean squared error), and also exhibiting as visual artifacts; for example, ringing, blockiness, texture-loss, and so on. To mitigate this, a set of tools, together known as an "In-Loop Filtering (ILF)" module is used to improve the quality of reconstruction; i.e., bring back some of the lost signal to make it closer to the source signal. For example, some video codecs use one or more of a Luma Mapping Chroma Scaling (LMCS) tool, a De-Blocking filter (DBK) tool, a Sample Adaptive Offset (SAO) tool, an Adaptive Loop Filter (ALF) tool, and a Chroma Component ALF (CC-ALF) tool, as a part of ILF.
[0005] There are some Artificial Intelligence (AI)-based ILF methods being proposed to replace / co-exist with the traditional ILF tools. However, the complexity of such tools as compared with the compression gain, measured in Bjontegaard Delta Bitrate (BD-BR) is prohibitively high. There is a need to reduce the complexity of AI tools used in ILF.
[0006] The principal object of embodiments herein is to disclose methods and systems for performing scalable de-quantization in video codec based on Artificial Intelligence (AI)-guided pixel classification.
[0007] Another object of embodiments herein is to disclose methods and systems that pose a quantization error-recovery problem as classification of pixels into several categories, such that for each category an offset value is to be added.
[0008] Another object of embodiments herein is to disclose methods and systems for performing a classification-based AI in-loop filtering in video codec for improving the overall reconstructed video quality.
[0009] Another object of embodiments herein is to disclose methods and systems for utilizing supplementary tools and techniques to make the framework classification error-resilient.
[0010] According to an embodiment of the disclosure, a decoding method may be disclosed. In an embodiment, the method may include receiving a bit-stream including an encoded video information. The method may include generating a reconstructed frame using the encoded video information. The method may include receiving, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame. The method may include determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The method may include identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The method may include generating an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0011] According to an embodiment of the disclosure, a decoding apparatus may be disclosed. In an embodiment, a video decoder may include a processor and a memory module, and the processor may be coupled with the memory module. In an embodiment, the processor may be configured to receive a bit-stream including an encoded video information. The processor may be configured to generate a reconstructed frame using the encoded video information. The processor may be configured to receive, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame. The processor may be configured to determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The processor may be configured to identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The processor may be configured to generate an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0012] According to an embodiment of the disclosure, an encoding method may be disclosed. In an embodiment, the method may include generating a reconstructed frame based on a source frame. The method may include determining a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame. The method may include determining one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error. The method may include determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The method may include identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The method may include determining an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame. The method may include generating an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0013] According to an embodiment of the disclosure, an encoding apparatus may be disclosed. In an embodiment, a video encoder may include a processor and a memory module, and the processor may be coupled with the memory module. In an embodiment, the processor may be configured to generate a reconstructed frame based on a source frame. The processor may be configured to determine a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame. The processor may be configured to determine one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error. The processor may be configured to determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The processor may be configured to identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The processor may be configured to determine an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame. The processor may be configured to generate an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0014] These and other aspects of the example embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating example embodiments and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the example embodiments herein without departing from the spirit thereof, and the example embodiments herein include all such modifications.
[0015] Embodiments herein are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the following illustratory drawings. Embodiments herein are illustrated by way of examples in the accompanying drawings, and in which:
[0016] FIG. 1 depicts an example Versatile Video Coding (VVC) video decoder, according to existing arts;
[0017] FIG. 2A depicts a process of quantization which is essentially a many-to-one mapping of signal code-words, according to existing arts;
[0018] FIG. 2B shows an example image (luma component) that is encoded using a VVC INTRA predictive codec, according to existing arts;
[0019] FIGS. 3A, 3B, and 3C depict an example image, a reconstructed image, and an error histogram respectively, according to existing arts;
[0020] FIG. 4 illustrates a block diagram of a system for performing an ILF processing of a video codec, according to embodiments as disclosed herein;
[0021] FIG. 5 depicts a method for performing in-loop filter processing of a video codec by a video encoder, according to embodiments as disclosed herein;
[0022] FIG. 6 depicts a method for processing a frame in a video codec by a video decoder, according to embodiments as disclosed herein;
[0023] FIG. 7 depicts an example scenario, where each pixel in an image is classified into one of the 3 classes, according to embodiments as disclosed herein;
[0024] FIG. 8 depicts an example process of performing class-based error recovery, according to embodiments as disclosed herein;
[0025] FIG. 9 depicts an example AI-Classifier flow, according to embodiments as disclosed herein;
[0026] FIG. 10 depicts an illustration of using 5 classes for classification, according to embodiments as disclosed herein; and
[0027] FIG. 11 depicts an example scenario of performing a sure pixel identification, according to embodiments as disclosed herein.
[0028] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0029] For the purposes of interpreting this specification, the definitions (as defined herein) will apply and whenever appropriate the terms used in singular will also include the plural and vice versa. It is to be understood that the terminology used herein is for the purposes of describing particular embodiments only and is not intended to be limiting. The terms “comprising”, “having” and “including” are to be construed as open-ended terms unless otherwise noted.
[0030] The words / phrases "exemplary", “example”, “illustration”, “in an instance”, “and the like”, “and so on”, “etc.”, “etcetera”, “e.g.,”, “i.e.,” are merely used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein using the words / phrases "exemplary", “example”, “illustration”, “in an instance”, “and the like”, “and so on”, “etc.”, “etcetera”, “e.g.,”, “i.e.,” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0031] Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
[0032] It should be noted that elements in the drawings are illustrated for the purposes of this description and ease of understanding and may not have necessarily been drawn to scale. For example, the flowcharts / sequence diagrams illustrate the method in terms of the steps required for understanding of aspects of the embodiments as disclosed herein. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. Furthermore, in terms of the system, one or more components / modules which comprise the system may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0033] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any modifications, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings and the corresponding description. Usage of words such as first, second, third etc., to describe components / elements / steps is for the purposes of this description and should not be construed as sequential ordering / placement / occurrence unless specified otherwise.
[0034] FIG. 2A depicts a process of quantization which is essentially a many-to-one mapping of signal code-words. In case of a predictive video coder, it is the transformed residues that undergo quantization. Once the mapping is done, recovering the lost code-word information may become a hard (difficult) inverse problem.
[0035] FIG. 2B shows an example image (luma component) that is encoded using VVC INTRA predictive codec. The plot on the right side shows the difference between original residue and quantized residue in pixel (i.e., inverse transformed)-domain shown on the X-axis, as a histogram plot with number of pixels on the Y-axis. For quantization parameter (QP) = 19, the curve is taller in center (near 0) indicating less spread of the error. Whereas, when QP increases to 32 and then 42, the curve flattens. This indicates there are more pixels having large positive and negative differences from the original residue.
[0036] The AI-model needs to recover the original residue from quantized residue and other encoder parameters, with a low computational complexity.
[0037] Recently, some AI-based ILF tools are being proposed to further improve the reconstructed video. Such tools augment or replace one or more ILF tools. However, the output video still contains significant pixel-differences causing major visual artifacts. The performance of recent proposals is limited possibly due to working / re-tuning of existing AI approaches.
[0038] The computational complexity of the existing AI-based ILF tools is quite high, making it practically infeasible to implement on current / near-future devices. For example, the two operating points that are being studied by the JVET standardization are: High: 477kMAC / pixel and Low: 17kMAC / pixel. This complexity is still quite high for any on-device implementation at this stage.
[0039] There exists a number of methods using lower-complexity Neural Networks (NNs) ~5kMAC / pixel, that employ a smaller and / or a shallower NN. However, the methods use regression as a core Machine Learning (ML) method which remains unchanged, and unchallenged.
[0040] As the network size reduces, smaller NN models may not solve the problem of same (high) difficulty. Currently, there is no proposal to reduce this difficulty using scalable, and classification-based approach.
[0041] An original image in FIG. 3A is compressed using a VVC INTRA coder. FIG. 3B shows a reconstructed image after quantization (QP = 42) i.e., by adding the transformed, quantized and inverse-transformed residue back to a prediction signal. As compared with the original image, it has incurred some error (PSNR: 31.30 dB). The error histogram is plotted in FIG. 3C. As can be seen, the error is centered around 0, and almost symmetric around the origin, indicating that Direct Current (DC) (average) i.e., the lowest frequency component has been reconstructed almost lossless; whereas, higher order statistics are not lossless, since higher-frequency transform coefficients have been quantized to a greater extent.
[0042] Hence, there is a need in the art for solutions which will overcome the above-mentioned drawback(s), among others.
[0043] The embodiments herein achieve a method and system for performing an Artificial Intelligence (AI)-based In-Loop Filtering (ILF) in a video codec to improve video quality. Referring now to the drawings, and more particularly to FIGS. 4 through 11, where similar reference characters denote corresponding features consistently throughout the figures, there are shown embodiments.
[0044] FIG. 4 illustrates a block diagram of a system 400 for performing an ILF processing of a video codec. The system 400 comprises a video encoder 402, and a video decoder 404. The video encoder 402 further comprises a processor 406, a communication module 408, and a memory module 410. The processor 406 of the video encoder 402 further comprises an encoder ILF module 418.
[0045] In an embodiment herein, the encoder ILF module 418 may receive at least one source video frame, and generate a reconstructed video frame which is derived from the source video frame. In the present disclosure, the term "reconstructed video frame" may be referred to as "reconstructed frame".
[0046] The encoder ILF module 418 may compare one or more pixel blocks of the reconstructed video frame with corresponding one or more pixel blocks of the source video frame. The encoder ILF module 418 may compute (determine) a reconstruction error for each pixel in one or more pixel blocks of the reconstructed video frame, based on the compared result of the pixel blocks of the reconstructed video frame and the source video frame. In the present disclosure, the term "reconstruction error" may be referred to as "residual". In an embodiment, the reconstruction error for each pixel may be analyzed using a probability distribution curve. A probability distribution curve is a graphical representation that shows the likelihood of different outcomes in a random variable.
[0047] In an embodiment herein, the encoder ILF module 418 may derive one or more pixel classes for each pixel block of the reconstructed video frame, based on the computed reconstruction error. The encoder ILF module 418 may determine, for each pixel in the pixel block of the reconstructed video frame, a pixel class and an associated confidence score for each pixel in the pixel block using the reconstructed video frame. In the present disclosure, the term "associated confidence score" may be referred to as "associated confidence parameter", "associated confidence variable".
[0048] In an embodiment, the pixel class, and the associated confidence score for each pixel may be determined by feeding the reconstructed video frame to an Artificial Intelligence (AI) model. The AI model is pre-trained to perform of pixels from the pixel blocks into one of N classes, where N is the number of classes that is dependent on a computational complexity of the AI model. The content characteristics of at least one video frame comprise at least one of an amount of texture, and a motion in the video frame.
[0049] In an embodiment herein, the encoder ILF module 418 may identify one or more pixels to be corrected (sure pixels) in each pixel block of the reconstructed video frame based on the associated confidence score of each pixel in the pixel block. In an embodiment herein, a binary flag may be assigned to each pixel, based on the confidence score of each pixel in the pixel block, to identify the pixels to be corrected such that only the identified sure pixels may undergo enhancement by employing an offset.
[0050] In an embodiment herein, the encoder ILF module 418 may determine an offset for each pixel class based on at least one of the computed reconstruction error, the identified pixels to be corrected, and one or more content characteristics of the source video frame. The offsets of the pixel classes may be determined using a rate-distortion framework by the video encoder 402. The offsets of the pixel classes are entropy-coded and transmitted to the video decoder 404 as bit-stream parameters at pixel-block, frame or sequence-level as determined by the video encoder 402. The offsets are categorized (classified) into at least one of a positive value offset, a negative value offset, and a zero value offset. The offsets are transmitted to the video decoder 404 as bit-stream parameters at one of a pixel block level, and a frame level as determined by the video encoder 400.
[0051] In an embodiment herein, the encoder ILF module 418 may generate an enhanced reconstructed video frame by applying the determined offset to each identified pixel to be corrected in the pixel block of the reconstructed video frame corresponding to the determined pixel class.
[0052] In an embodiment herein, the processor 406 of the video encoder 402 is configured to generate a bit-stream comprising of one or more encoded source video frames, and corresponding domain-transfer functions. The processor 406 of the video encoder 402 is configured to transmit the generated bit-stream to the video decoder 404.
[0053] In an embodiment herein, the video decoder 404 may further comprise a processor 412, a communication module 414, and a memory module 416. The processor 412 of the video decoder 404 further comprises a decoder ILF module 420.
[0054] In an embodiment herein, the decoder ILF module 420 may receive a bit-stream comprising of an encoded source video information. The decoder ILF module 420 can generate a reconstructed video frame corresponding to the encoded source video frame using the received bit-stream. The decoder ILF module 420 may receive a number of pixel classes, and one or more offsets corresponding to each pixel class from the bit-stream for each pixel block of the reconstructed video frame. For example, the decoder ILF module 420 may receive the pixel classes, and corresponding offsets from the encoder ILF module 418 of the video encoder 402.
[0055] The decoder ILF module 420 may determine, for each pixel in the pixel block of the reconstructed video frame, a pixel class and an associated confidence score for each pixel in the pixel block using the reconstructed video frame. The decoder ILF module 420 may identify one or more pixels to be corrected in each pixel block of the reconstructed video frame based on the associated confidence score for each pixel in the pixel block. Further, the decoder ILF module 420 may generate an enhanced reconstructed video frame by applying at least one offset to each identified pixel to be corrected in the pixel block of the reconstructed video frame corresponding to the determined pixel class.
[0056] In an embodiment herein, the processor 406, and the processor 412 may process and execute data of a plurality of modules of the video encoder 402, and the video decoder 404 respectively. The processor 406, and the processor 412 may be configured to execute instructions stored in the memory module 410, and the memory module 416 respectively. The processor 406, and the processor 412 may comprise one or more of microprocessors, circuits, and other hardware configured for processing. The processor 406, and the processor 412 may be at least one of a single processer, a plurality of processors, multiple homogeneous or heterogeneous cores, multiple Central Processing Units (CPUs) of different kinds, microcontrollers, special media, and other accelerators. The processor 406, and the processor 412 may be an application processor (AP), a graphics-only processing unit (such as a graphics processing unit (GPU), a visual processing unit (VPU)), and / or an Artificial Intelligence (AI)-dedicated processor (such as a neural processing unit (NPU)).
[0057] In an embodiment herein, the plurality of modules of the processor 406, and the processor 412 of the video encoder 402, and the video decoder 404 may communicate via the communication module 408, and the communication module 414 respectively. The communication module 408, and the communication module 414 may be in the form of either a wired network or a wireless communication network module. The wireless communication network may comprise, but not limited to, Global Positioning System (GPS), Global System for Mobile Communications (GSM), Wi-Fi, Bluetooth low energy, Near-field communication (NFC), and so on. The wireless communication may further comprise one or more of Bluetooth, ZigBee, a short-range wireless communication (such as Ultra-Wideband (UWB)), and a medium-range wireless communication (such as Wi-Fi) or a long-range wireless communication (such as 3G / 4G / 5G / 6G and non-3GPP technologies or WiMAX), according to the usage environment.
[0058] In an embodiment herein, the memory module 410, and the memory module 416 may comprise one or more volatile and non-volatile memory components which are capable of storing data and instructions of the modules of the video encoder 402, and the video decoder 404 to be executed. Examples of the memory module 410, and the memory module 416 can be, but not limited to, NAND, embedded Multi Media Card (eMMC), Secure Digital (SD) cards, Universal Serial Bus (USB), Serial Advanced Technology Attachment (SATA), solid-state drive (SSD), and so on. The memory module 410, and the memory module 416 may also include one or more computer-readable storage media. Examples of non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory module 410, and the memory module 416 may, in some examples, be considered a non-transitory storage medium. The term “non-transitory” may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term “non-transitory” should not be interpreted to mean that the memory module 410, and the memory module 416 is non-movable. In certain examples, a non-transitory storage medium may store data that can, over time, change (for example, in Random Access Memory (RAM) or cache).
[0059] FIG. 4 shows example modules of the video encoder 402, and the video decoder 404, but it is to be understood that other embodiments are not limited thereon. In other embodiments, the video encoder 402, and the video decoder 404 may include less or more number of modules. Further, the labels or names of the modules are used only for illustrative purpose and does not limit the scope of the invention. One or more modules may be combined together to perform same or substantially similar function in the video encoder 402, and the video decoder 404 respectively.
[0060] FIG. 5 depicts a method 500 for performing in-loop filter processing of a video codec by the video encoder 402. The encoder ILF module 418 of the video encoder 402 generates a reconstructed video frame which is derived from a source video frame by processing a pixel-block from the source video frame such that the residual pixel-block is added to the predicted pixel-block. The inputs for generating the reconstructed video frame comprise the source video frame, and video frames from other video coding modules 502. The method 500 comprises computing a reconstruction error for each pixel in one or more pixel blocks of the reconstructed video frame, as depicted in step 504. The encoder ILF module 418 compares the pixel blocks of the reconstructed video frame with corresponding pixel blocks of the source video frame for computing the reconstruction error.
[0061] The method 500 comprises deriving or determining one or more pixel classes or number of classes for each pixel block of the reconstructed video frame, as depicted in step 506, based on the computed reconstruction error. The encoder ILF module 418 uses allowed AI computations for determining the pixel classes and generate bit-stream parameter of number of classes. Thereafter, the method 500 comprises determining, for each pixel in the pixel block of the reconstructed video frame, a pixel class and an associated confidence score for each pixel in the pixel block using the reconstructed video frame, as depicted in step 508. The encoder ILF module 418 performs an AI-guided pixel classification for determining the pixel class.
[0062] The reconstructed pixel-block is fed to the AI model such as a neural network. The pixels of the reconstructed pixel-block are classified into a plurality of pixel classes along with their corresponding confidence scores using the AI model.
[0063] Later, the method 500 comprises identifying sure pixels i.e., one or more pixels to be corrected in each pixel block of the reconstructed video frame based on the associated confidence score of each pixel in the pixel block, as depicted in step 510. The method 500 comprises determining an offset for each pixel class i.e., class-correction offset, as depicted in step 512, based on at least one of the computed reconstruction error, the identified pixels to be corrected, and content characteristics of the source video frame such as standard deviation. The content characteristics may include, but not limited to predicted pixel-block, prediction type, boundary-strength map, and so on. The encoder ILF module 418 generates bit-stream parameter class offsets.
[0064] The method 500 comprises performing a per-pixel de-quantization, depicted in step 514, for generating an enhanced reconstructed video frame as an output by applying the determined offset to each identified pixel to be corrected in the pixel block of the reconstructed video frame corresponding to the determined pixel class. The offsets are coded into the bit-stream that is transmitted to the video decoder 404. The generated output is further stored at the other video coding modules 502.
[0065] The various actions in method 500 may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some actions listed in FIG. 5 may be omitted.
[0066] FIG. 6 depicts a method 600 for processing a frame in a video codec by a video decoder 404. The decoder ILF module 420 of the video decoder 404 receives a bit-stream comprising of an encoded source video information. The decoder ILF module 420 of the video decoder 404 generates a reconstructed video frame corresponding to the encoded source video frame using the received bit-stream. The reconstructed pixel-block of a video frame is generated by adding the residual pixel block to the predicted pixel block. The inputs for generating the reconstructed video frame comprise the source video frame, and video frames from other video coding modules 502. The decoder ILF module 420 receives a number of pixel classes, and one or more offsets corresponding to each pixel class from the bit-stream for each pixel block of the reconstructed video frame.
[0067] The method 600 comprises receiving number of pixel classes information from the bit-stream, and performing an AI-guided pixel classification for determining, for each pixel in the pixel block of the reconstructed video frame, a pixel class and an associated confidence score for each pixel in the pixel block using the reconstructed video frame, as depicted in step 602. The reconstructed pixel block is fed to the AI model such as a neural network.
[0068] The method 600 comprises identifying sure pixels i.e., one or more pixels to be corrected in each pixel block of the reconstructed video frame based on the associated confidence score of each pixel in the pixel block, as depicted in step 604. Later, the decoder ILF module 420 receives offsets for each pixel class from the bit-stream, and performs a per-pixel de-quantization, depicted in step 606, for generating an enhanced reconstructed video frame as an output by applying the determined at least one offset to each identified pixel to be corrected in the pixel block of the reconstructed video frame corresponding to the determined pixel class. The generated output is further stored at the other video coding modules 502.
[0069] The various actions in method 600 may be performed in the order presented, in a different order or simultaneously. Further, in some embodiments, some actions listed in FIG. 6 may be omitted.
[0070] In an embodiment herein, the proposed system 400 discloses a divide-n-conquer strategy to recover quantization error(s), which is computed by subtracting the source pixel value from the reconstructed pixel value. Considering the example image depicted in FIG. 3A, each pixel is categorized in N=3 classes, as depicted in FIG. 7. In this example, a “negative”, “positive” and “zero” classes indicated by λ(-1), λ(1), λ(0)respectively are shown. For each class {λ(-1), λ(0), λ(1)}, an additive offset {ρ(-1), ρ(0), ρ(1)} is assigned to cancel quantization error. For the λ(-1)pixels, a positive-values offset ρ(-1)is added to compensate the error, whereas for λ(1) pixels, a negative-valued offset ρ(1)is added. The λ(0)pixels are already perfectly reconstructed, thus ρ(0)= 0. This shows that the AI model needs to be trained for classification to output one of N (in this case, N = 3) labels for each pixel.
[0071] FIG. 8 depicts an example process of performing a class-based error recovery. An AI-classifier may obtain the reconstructed signal and other bit-stream parameters as inputs to output a 3-class pixel maps indicated by (d) DIFF-posmask, a positive mask with all λ(1)’s, (e) DIFF-negmask, a negative mask with all λ(-1)'s, and (f) DIFF-zeromask, a zeromask with all λ(0)labelled pixels shown. Based on the actual error mask shown by DIFF as the difference between reconstructed and original (source) signal and the 3 class-masks, a class to offset mapping is determined. This specifies for each class the value of additive offset: {ρ(-1), ρ(0), ρ(1)}. These values may be transmitted to the video decoder 404 as a bit-stream parameter. It is to be noted that DIFF is not available at the video decoder 404. The pixel-wise correction offsets for that image are then added to the reconstructed signal to obtain a Rec-s i.e., an enhanced image. The Peak Signal-to-Noise Ratio (PSNR) of the enhanced image is +1.2dB higher than the reconstructed image without an enhancement. The method may be applied at block of pixels which may be a frame or a smaller unit such as a Coding Tree Unit (CTU).
[0072] FIG. 9 depicts an example AI-Classifier flow. A pixel-block of a reconstructed frame is considered herein as a basic processing unit, for example 128x128 pixels or CTU. The AI-based pixel-classification model such as a Neural Network (NN) based architecture uses the following block information: reconstructed pixel block (Y, U, V channels), predicted pixel block (Y, U, V channels), quantization parameter, frame-Type (I,P,B), and so on. The model classifies each pixel into N categories / labels as shown in FIG. 9. For example, if the k’th pixel is being with one of 3 labels, the predicted class-label may be Hk∈ {λ(-1), λ(1), λ(0)}. The model also outputs a confidence score Ckfor each k’th pixel of the block, indicating the surety of the label Hk.
[0073] FIG. 10 depicts an illustration of using 5 classes. The AI model here needs to do a 5-class separation which is a somewhat difficult problem than 3-class separation of pixels. However, if the AI model can do a perfect classification, the reward i.e., the PSNR increase is multifold; 5.3dB in this case. As depicted, as N, the number of classes increases, the difficulty of AI-classifier increases so as the reconstruction quality. Thus, it is feasible to construct AI-models of varied complexity (in terms of computations per pixel) to achieve reconstruction quality in a scalable way. For each class {λ(-2), λ(-1), λ(0), λ(1), λ(2)}, an additive offset {ρ(-2), ρ(-1), ρ(0), ρ(1), ρ(2)} is assigned to cancel quantization error. Since an offset is added, it is theoretically equivalent to de-quantization; i.e., introducing new codewords that were lost due to quantization. Although the quantization process may have a loss of several codewords, the AI-guided de-quantization may not attempt to recover all the lost information. Instead, the AI-guided de-quantization can try to solve a simpler problem to introduce two new levels with a 3-class classifier, or four new levels with a 5-class classifier.
[0074] In an embodiment herein, to mitigate the possibility of degradation of output, the proposed system 400 proposes two methods. One method is conservative class-offsets, and other method is sure pixel identification.
[0075] In conservative class-offsets, the reconstruction error is centered at 0. A conservative set of ρ(i)’s for each "λ(i)" may be used to reduce the enhancement. For example, for ρ=0, there is no enhancement for any pixel. Thus, PSNR gain = 0. As the ρ-value increases, the PSNR value increases. At some ρ, the PSNR value reaches its maximum for a given misclassification rate. Table 1 shows that the simulations, of embodiments as disclosed herein, indicate for 10% misclassification, a ρ=σ / 2 gives the maximum PSNR gain of approx. (37.21 - 35.54 = ) 1.67dB, highest among all listed values of ρ. Here, σ is the standard deviation of the reconstruction error. This shows that in presence of non-zero AI-classification error, it is prudent to check and use a suitable offset.
[0076]
[0077] In sure pixel identification, as depicted in FIG. 11, embodiments herein take a binary (yes / no) decision whether it is modified by calling it a ‘sure’ pixel or leave it as-is, since it is an ‘unsure’ pixel. In an embodiment, a sure pixel may be referred as a pixel to be modified(corrected).
[0078] For each k’th pixel, embodiments herein assign it as ‘sure’ or ‘no-change’, based on the confidence score Ckas obtained by the AI-classification model, using a simple thresholding algorithm. The threshold may be either fixed or based on a pixel-percentile method, computed for the current block or for entire frame. Only all ‘sure’ pixels of a block are processed in the subsequent steps and ‘no-change’ pixels are not modified. Table 2 shows an example of % sure pixels and the PSNR impact for one classification algorithm, for a set of several test images.
[0079]
[0080] As can be seen, the best result can be achieved at some combination of ρ-value and % sure pixels. This may be determined at block / frame-level and sent as bit-stream parameters.
[0081] Table 3 and table 4 show the simulation of 3-class and 5-class de-quantization respectively using 27k images for different % classification errors for σ / 4, for 100% sure pixels.
[0082]
[0083]
[0084] Table 3 and Table 4 simulate the case where the AI-classifier can produce labels with a non-zero error indicated by the error %. From both tables, it can be seen that as the classification accuracy reduces, the PSNR gains are also reduced. The proposed framework can produce positive PSNR gains even with 30% classification errors. This shows that the method is error-resilient to a significant extent. As the error % reaches >40%, the gains become negative indicating it can degrade the quality of pixel-block if the error % increases further. At the same error %, the 3-class gains are less than 5-class gains since the AI is solving a harder problem. This shows the method (as disclosed herein) is highly scalable in terms of PSNR gains performance against the number of classes.
[0085] According to an embodiment of the disclosure, a decoding method may be disclosed. In an embodiment, the method may include receiving a bit-stream including an encoded video information. The method may include generating a reconstructed frame using the encoded video information. The method may include receiving, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame. The method may include determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The method may include identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The method may include generating an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0086] In an embodiment, the reconstruction error for each pixel may be determined using a probability distribution curve.
[0087] In an embodiment, the decoding method may include feeding the reconstructed frame to an Artificial Intelligence (AI) model.
[0088] In an embodiment, the AI model may be pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of pixel classes that is dependent on a computational complexity of the AI model.
[0089] In an embodiment, the one or more content characteristics of the source frame may include at least one of an amount of texture, and a motion in the source frame.
[0090] In an embodiment, a binary flag may be assigned to each pixel, based on the associated confidence parameter of each pixel in the pixel block, to identify the one or more pixels to be corrected.
[0091] In an embodiment, the offset for the each pixel class may be determined using a rate-distortion framework, and the offset for the each pixel class may be entropy-coded.
[0092] In an embodiment, the decoding method may include classifying the offset into at least one of a positive value offset, a negative value offset or a zero value offset.
[0093] In an embodiment, the decoding method may include receiving the offset for the each pixel class to the video decode from a bit-stream parameter at one of a pixel block level, or a frame level.
[0094] According to an embodiment of the disclosure, a decoding apparatus may be disclosed. In an embodiment, a video decoder may include a processor and a memory module, and the processor may be coupled with the memory module. In an embodiment, the processor may be configured to receive a bit-stream including an encoded video information. The processor may be configured to generate a reconstructed frame using the encoded video information. The processor may be configured to receive, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame. The processor may be configured to determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The processor may be configured to identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The processor may be configured to generate an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0095] In an embodiment, the reconstruction error for each pixel may be determined using a probability distribution curve.
[0096] In an embodiment, the processor may be configured to feed the reconstructed frame to an Artificial Intelligence (AI) model.
[0097] In an embodiment, the AI model may be pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of pixel classes that is dependent on a computational complexity of the AI model.
[0098] In an embodiment, the one or more content characteristics of the source frame may include at least one of an amount of texture, and a motion in the source frame.
[0099] In an embodiment, a binary flag may be assigned to each pixel, based on the associated confidence parameter of each pixel in the pixel block, to identify the one or more pixels to be corrected.
[0100] In an embodiment, the offset for the each pixel class may be determined using a rate-distortion framework, and the offset for the each pixel class may be entropy-coded.
[0101] In an embodiment, the processor may be configured to classify the offset into at least one of a positive value offset, a negative value offset or a zero value offset.
[0102] In an embodiment, the processor may be configured to receive the offset for the each pixel class to the video decode from a bit-stream parameter at one of a pixel block level, or a frame level.
[0103] According to an embodiment of the disclosure, an encoding method may be disclosed. In an embodiment, the method may include generating a reconstructed frame based on a source frame. The method may include determining a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame. The method may include determining one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error. The method may include determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The method may include identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The method may include determining an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame. The method may include generating an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0104] In an embodiment, the encoding method may include generating a bit-stream including encoded video information corresponding to the source frame, and corresponding number of pixel classes and one or more offsets, and transmitting the generated bit-stream to a video decoder.
[0105] In an embodiment, the reconstruction error for each pixel may be determined using a probability distribution curve.
[0106] In an embodiment, the encoding method may include feeding the reconstructed frame to an Artificial Intelligence (AI) model.
[0107] In an embodiment, the AI model may be pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of pixel classes that is dependent on a computational complexity of the AI model.
[0108] In an embodiment, the one or more content characteristics of the source frame may include at least one of an amount of texture, and a motion in the source frame.
[0109] In an embodiment, a binary flag may be assigned to each pixel, based on the associated confidence parameter of each pixel in the pixel block, to identify the one or more pixels to be corrected.
[0110] In an embodiment, the offset for the each pixel class may be determined using a rate-distortion framework, and the offset for the each pixel class may be entropy-coded.
[0111] In an embodiment, the encoding method may include classifying the offset into at least one of a positive value offset, a negative value offset or a zero value offset.
[0112] In an embodiment, the encoding method may include transmitting the offset for the each pixel class to the video decode as a bit-stream parameter at one of a pixel block level, or a frame level.
[0113] According to an embodiment of the disclosure, an encoding apparatus may be disclosed. In an embodiment, a video encoder may include a processor and a memory module, and the processor may be coupled with the memory module. In an embodiment, the processor may be configured to generate a reconstructed frame based on a source frame. The processor may be configured to determine a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame. The processor may be configured to determine one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error. The processor may be configured to determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame. The processor may be configured to identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter. The processor may be configured to determine an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame. The processor may be configured to generate an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.
[0114] In an embodiment, the encoding method may include generating a bit-stream including encoded video information corresponding to the source frame, and corresponding number of pixel classes and one or more offsets, and transmitting the generated bit-stream to a video decoder.
[0115] In an embodiment, the reconstruction error for each pixel may be determined using a probability distribution curve.
[0116] In an embodiment, the processor may be configured to feed the reconstructed frame to an Artificial Intelligence (AI) model.
[0117] In an embodiment, the AI model may be pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of pixel classes that is dependent on a computational complexity of the AI model.
[0118] In an embodiment, the one or more content characteristics of the source frame may include at least one of an amount of texture, and a motion in the source frame.
[0119] In an embodiment, a binary flag may be assigned to each pixel, based on the associated confidence parameter of each pixel in the pixel block, to identify the one or more pixels to be corrected.
[0120] In an embodiment, the offset for the each pixel class may be determined using a rate-distortion framework, and the offset for the each pixel class may be entropy-coded.
[0121] In an embodiment, the processor may be configured to classify the offset into at least one of a positive value offset, a negative value offset or a zero value offset.
[0122] In an embodiment, the processor may be configured to transmit the offset for the each pixel class to the video decode as a bit-stream parameter at one of a pixel block level, or a frame level.
[0123] The proposed system 400 and methods 500, 600 provide a quantization error-recovery problem as classification of pixels into several categories, such that an offset value is to be added for each category. This allows low-complex (shallow or lean) Neural Networks (NN) to solve a simpler problem than that of regression.
[0124] Therefore, the proposed system 400 and methods 500, 600 disclose a framework that poses a simpler problem to the AI in-loop filtering model with classification-based approach, thus improving the compression efficiency of video codec. The system 400 and methods 500, 600 disclose a scalable framework, which is scalable via the number of classes the model needs to categorize each pixel. The proposed system 400 and methods 500, 600 make it easier for the NN model of lower complexity to perform embodiments as disclosed herein.
[0125] The proposed framework may be useful for several on-device applications that use AI-models for image reconstruction and / or enhancement.
[0126] The embodiments disclosed herein may be implemented through at least one software program running on at least one hardware device. The elements include blocks which can be at least one of a hardware device, or a combination of hardware device and software module.
[0127] The embodiments disclosed herein describe a scalable ILF method to enhance reconstructed block of pixels using an AI-based pixel-classification and an offset addition in a video codec. Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means contain program code means for implementation of one or more steps of the method, when the program runs on a server or mobile device or any suitable programmable device. The method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device. The hardware device can be any kind of portable device that can be programmed. The device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g., an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein. The method embodiments described herein could be implemented partly in hardware and partly in software. Alternatively, the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.
[0128] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of embodiments and examples, those skilled in the art will recognize that the embodiments and examples disclosed herein can be practiced with modification within the scope of the embodiments as described herein.
Claims
1.A method (600) for processing a frame in a video codec by a video decoder (404), comprising:receiving a bit-stream including an encoded video information;generating a reconstructed frame using the encoded video information;receiving, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame;determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame;identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter; andgenerating an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.2.A video decoder (404), comprising:a processor (412), anda memory module (416);wherein the processor (412) is coupled with the memory module (416), and configured to:receive a bit-stream including an encoded video information;generate a reconstructed frame using the encoded video information;receive, from the bit-stream, a number of pixel classes, and one or more offsets corresponding to each pixel class for a pixel block of the reconstructed frame;determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame;identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter; andgenerate an enhanced reconstructed frame by applying an offset corresponding to the determined pixel class to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.3.A method (500) for performing in-loop filter processing of a video codec in an in-loop filtering module by a video encoder (402), comprising:generating a reconstructed frame based on a source frame;determining a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame;determining one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error;determining, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame;identifying one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter;determining an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame; andgenerating an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.4.The method (500) as claimed in claim 3, wherein the method (500) comprises:generating a bit-stream including encoded video information corresponding to the source frame, and corresponding number of pixel classes and one or more offsets; andtransmitting the generated bit-stream to a video decoder (404).5.The method (500) as claimed in any one of claims 3 to 4, wherein the reconstruction error for each pixel is determined using a probability distribution curve.6.The method (500) as claimed in any one of claims 3 to 5, wherein the method (500) of determining the pixel class, and the associated confidence parameter for each pixel, comprises:feeding the reconstructed frame to an Artificial Intelligence (AI) model, wherein the AI model is pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of pixel classes that is dependent on a computational complexity of the AI model.7.The method (500) as claimed in any one of claims 3 to 6, wherein the one or more content characteristics of the source frame comprises at least one of an amount of texture, and a motion in the source frame.8.The method (500) as claimed in any one of claims 3to 7, wherein a binary flag is assigned to each pixel, based on the associated confidence parameter of each pixel in the pixel block, to identify the one or more pixels to be corrected.9.The method (500) as claimed in any one of claims 3 to 8, wherein the offset for the each pixel class is determined using a rate-distortion framework, and wherein the offset for the each pixel class is entropy-coded.10.The method (500) as claimed in any one of claims 3 to 9, wherein the method (500) comprises:classifying the offset into at least one of a positive value offset, a negative value offset or a zero value offset.11.The method (500) as claimed in any one of claims 3 to 10, wherein the method (500) comprises:transmitting the offset for the each pixel class to the video decoder (404) as a bit-stream parameter at one of a pixel block level, or a frame level.12.A video encoder (402), comprising:a processor (406), anda memory module (410);wherein the processor (406) is coupled with the memory module (410), and configured to:generate a reconstructed frame based on a source frame;determine a reconstruction error for each pixel in a pixel block of the reconstructed frame, by comparing the pixel block of the reconstructed frame with corresponding pixel block of the source frame;determine one or more pixel classes for the pixel block of the reconstructed frame based on the determined reconstruction error;determine, for each pixel in the pixel block of the reconstructed frame, a pixel class and an associated confidence parameter using the reconstructed frame;identify one or more pixels to be corrected in the pixel block of the reconstructed frame based on the associated confidence parameter;determine an offset for each pixel class based on at least one of the determined reconstruction error, the identified one or more pixels to be corrected, and one or more content characteristics of the source frame; andgenerate an enhanced reconstructed frame by applying the determined offset to the identified one or more pixels to be corrected in the pixel block of the reconstructed frame.13.The video encoder (402) as claimed in claim 12, wherein the processor (406) is configured to:generate a bit-stream including encoded video information corresponding to the source frame, and corresponding number of pixel classes and one or more offsets; andtransmit the generated bit-stream to a video decoder (404).14.The video encoder (402) as claimed in any one of claims 12 to 13, wherein the reconstruction error for each pixel is determined using a probability distribution curve.15.The video encoder (402) as claimed in any one of claims 12 to 14, wherein the pixel class, and the associated confidence parameter for each pixel are determined by feeding the reconstructed frame to an Artificial Intelligence (AI) model, wherein the AI model is pre-trained to perform classification of pixels from the pixel block into one of N classes, where N is the number of classes that is dependent on a computational complexity of the AI model.
Citation Information
Patent Citations
Apparatus and method of adaptive offset restoration for video coding
KR101433501B1
Methods and apparatus for a classification-based loop filter
KR101810263B1
Rice cake manufacturing method with improved sensory properies and preservation
KR1020240007789A
Video encoding devic and driving method thereof
KR102276914B1
Image processing method and device for ai-based filtering
WO2023080464A1