Method and electronic device for performing in-loop filter processing of encoding and decoding
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-13
Smart Images

Figure KR2026002253_13082026_PF_FP_ABST
Abstract
Description
METHOD AND ELECTRONIC DEVICE FOR PERFORMING IN-LOOP FILTER PROCESSING OF ENCODING AND DECODING
[0001] Embodiments disclosed herein relate to content processing (e.g., video processing or the like), and more particularly to an Artificial intelligence (AI)-based in-loop processing in video codecs.
[0002] An image or video codec generally employs quantization to compress a visual signal, thereby reducing a number of bits needed for transmission. However, the quantization introduces visual distortions in a reconstructed frame at a receiver end. To alleviate such artefacts, modern video coding standards utilize a set of classical in-loop filters. Emerging standards now propose enhancing this process by integrating an artificial intelligence (AI)-based filters alongside traditional ones, aiming to more effectively reduce distortions and improve visual quality in compressed video content.
[0003] In conventional video codecs, compression is accomplished by quantizing transform-domain coefficients that represent the prediction residuals. This quantization causes the reconstructed frames to deviate from the original source frames, often resulting in visual artefacts such as ringing, blocking, and loss of texture. To mitigate these distortions, a set of tools, collectively referred to as the In-Loop Filtering (ILF) module, is employed to enhance the visual quality of the reconstructed frames and improve overall compression performance.
[0004] In an embodiment of the disclosure, a method for performing an in-loop filter processing by a video encoder included in an electronic device is provided. The method may include generating a reconstructed video frame corresponding to a source video frame. The method may include configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame. The method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model. The method may include generating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0005] In an embodiment of the disclosure, a method for performing an in-loop filter processing by a video decoder included in an electronic device is provided. The method may include obtaining a bitstream including information for selection of at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block to configure a scalable neural network. The method may include generating a reconstructed video frame using the bitstream, wherein the reconstructed video frame includes a plurality of reconstructed pixel blocks. The method may include configuring the scalable neural network by using a pre-defined scalable neural network model for each of the plurality of the reconstructed pixel blocks. The method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of the at least one pathway within the at least one core block. The method may include generating an enhanced reconstructed video frame by feeding each of the plurality of reconstructed pixel blocks in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0006] In an embodiment of the disclosure, a method for transmitting a bitstream generated by an encoding method is provided. The encoding method may include generating a reconstructed video frame corresponding to a source video frame. The encoding method may include generating a reconstructed video frame corresponding to a source video frame configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame. The encoding method may include generating a reconstructed video frame corresponding to a source video frame. The encoding method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model. The encoding method may include generating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0007] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating at least one embodiment and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
[0008] Embodiments herein are illustrated in the accompanying drawings, throughout which like reference letters indicate corresponding parts in the various figures. The embodiments herein will be better understood from the following description with reference to the following illustratory drawings. Embodiments herein are illustrated by way of examples in the accompanying drawings, and in which:
[0009] Figure 1 depicts an example of a VVC video decoder, according to related arts;
[0010] Figure 2 depicts the usage of AI-based model in parallel with the deblocking filter (DBF), according to related arts;
[0011] Figure 3 depicts the current configuration of an AI-based model, according to related arts;
[0012] Figure 4 shows various hardware components of an electronic device, according to the an embodiment;
[0013] Figure 5 depicts a method for performing in-loop filter processing of a video encoder, according to an embodiment;
[0014] Figure 6 depicts a flowchart for a method for performing in-loop filter processing of a video encoder, by an electronic device, according to an embodiment;
[0015] Figure 7 depicts a flowchart for a method for enhancing in-loop filter processing of a video codec in a neural network-based in-loop filtering controller by the electronic device, according to an embodiment;
[0016] Figure 8 depicts a method for performing in-loop filter processing of a video decoder, according to an embodiment;
[0017] Figures 9a -9b depict a flowchart for a method for performing in-loop filter processing of the video decoder by the electronic device according to an embodiment;
[0018] Figure 10 depicts a fundamental configurable building block, hereinafter referred to as an Aggregated Core Unit (ACU), according to an embodiment;
[0019] Figure 11 depicts a Tail Block (TB), according to an embodiment;
[0020] Figure 12 depicts a scalable model, wherein the ACU is used to design a scalable AI-ILF of desired complexity, according to an embodiment;
[0021] Figure 13 depicts the internal structure of the ACU, according to an embodiment;
[0022] Figure 14 describes the configurable inference-time framework of ACU block, according to an embodiment;
[0023] Figure 15 depicts an example scenario of a width-wise scalable framework of the idea with ACU as the building block, according to an embodiment;
[0024] Figure 16 depicts an example scenario of a depth-wise scalable framework of the idea with ACU as the building block, according to an embodiment;
[0025] Figure 17 depicts an example scenario of a scalable framework of the idea with ACU as the building block, according to an embodiment;
[0026] Figure 18 depicts an example scenario of the overall configurable and scalable module, according to an embodiment;
[0027] Figure 19 depicts an example scenario where the configurable and scalable model is trained with P = 3 ACU blocks and N = 5 Residual Unit (RU) blocks, according to an embodiment; and
[0028] Figure 20 depicts an example scenario where the overall trained model is embedded with P = 3 ACU blocks and N = 5 RU blocks, according to an embodiment.
[0029] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0030] In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
[0031] While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.
[0032] The terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a device or system or apparatus proceeded by "comprises... a" does not, without more constraints, preclude the existence of other elements or additional elements in the device or system or apparatus.
[0033] In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
[0034] For the purposes of interpreting this specification, the definitions (as defined herein) will apply and whenever appropriate the terms used in singular will also include the plural and vice versa. It is to be understood that the terminology used herein is for the purposes of describing particular embodiments only and is not intended to be limiting. The terms "comprising", "having" and "including" are to be construed as open-ended terms unless otherwise noted.
[0035] The words / phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," , "i.e.," are merely used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein using the words / phrases "exemplary", "example", "illustration", "in an instance", "and the like", "and so on", "etc.", "etcetera", "e.g.," , "i.e.," is not necessarily to be construed as preferred or advantageous over other embodiments.
[0036] Embodiments herein may be described and illustrated in terms of blocks which carry out a described function or functions. These blocks, which may be referred to herein as managers, units, modules, hardware components or the like, are physically implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by a firmware. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and discrete blocks without departing from the scope of the disclosure. Likewise, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the disclosure.
[0037] It should be noted that elements in the drawings are illustrated for the purposes of this description and ease of understanding and may not have necessarily been drawn to scale. For example, the flowcharts / sequence diagrams illustrate the method in terms of the steps required for understanding of aspects of the an embodiment. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. Furthermore, in terms of the system, one or more components / modules which comprise the system may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the present embodiments so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0038] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any modifications, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings and the corresponding description. Usage of words such as first, second, third etc., to describe components / elements / steps is for the purposes of this description and should not be construed as sequential ordering / placement / occurrence unless specified otherwise.
[0039] Accordingly, the embodiments herein provide a configurable and scalable AI-model architecture for in-loop filtering in video codecs consisting neural network blocks with multiple residual units. Moreover, the embodiments herein provide an encoder-decoder framework to scale and configure the network on-the-fly based on a desired compression performance (for example, quality vs. bitrate) and allowed complexity, based on content properties and encoder parameters.
[0040] Embodiments herein propose a method disclosed for performing in-loop filter processing in a video encoder using an electronic device, offering improved adaptability and efficiency. The method begins by generating a reconstructed video frame from a source frame, involves analyzing content-related properties (such as texture or motion characteristics) and encoder parameters (like quantization or prediction mode). Based on the analysis, a scalable neural network model is dynamically configured by selecting specific core blocks, residual units, and processing pathways tailored to the frame's content and encoding context. Unlike conventional fixed-complexity models, the proposed allows selective use of network components, reducing computational overhead, the method further updates pre-trained parameters of the neural network based on the selected pathways, enabling fine-grained control over filtering strength and complexity. Finally, the updated model is applied block-wise to the reconstructed frame, producing an enhanced output with reduced artefacts and improved visual quality. The method further represents a technical advancement by enabling content-aware, scalable neural in-loop filtering, offering both efficiency and performance adaptability across diverse use-cases and device capabilities.
[0041] The proposed method can be used for improving in-loop filter processing in the video codec using the neural network-based controller within the electronic device. The method involves deriving a scalable neural network model composed of core blocks, each including residual units representing multiple pixel blocks of a source video frame, with quantization levels applied across the blocks and pathways defined within each core block. For each reconstructed pixel block in the decoded video frame, a corresponding instance of the scalable model is configured, allowing dynamic adaptation based on frame content and complexity. The reconstructed frame is further enhanced using pre-trained neural network parameters, updating each pixel block accordingly. The proposed method introduces a flexible architecture allowing adaptive complexity control, supporting more efficient and hardware-friendly in-loop filtering while maintaining high visual quality across varying content and device capabilities.
[0042] Embodiments herein disclose a configurable and scalable neural network for in-loop filtering in video codecs. Embodiments herein disclose a hardware-friendly, scalable and configurable methods for low complexity (<=5kMAC / pixel) AI-based in-loop filters. Embodiments herein disclose a hardware-friendly Convolutional Neural Network (CNN) based core building block, which can used to design a scalable AI model of desired complexity, wherein the proposed CNN block offers re-configurability such that it can be configured on-the-fly during encoding / decoding.
[0043] Embodiments herein disclose a core CNN block with attention-based multi-branch learning paradigm in which the branches are re-configurable during the inference. Embodiments herein disclose a scalable AI model via parallelization / serialization of the arrangement of core blocks to suit the complexity budget & performance needs.
[0044] Embodiments herein disclose a flexible CNN framework which is adaptable and configurable based on applications, content, encoders and in general user-preferences. Embodiments herein disclose a hardware-friendly core NN block, wherein the core building block with all HW-friendly compute operations with a specific per-pixel complexity and can be used to construct a scalable low-complex AI-ILF model of desired complexity. The core building block contains switchable multi-pathways, and further stacking of core blocks serially / parallel allows for flexible in model structure thus providing a user-preferable configuration. The configurable CNN-based core building block that can be re-configured on-the-fly during encoding of a video by the rate-distortion optimizer.
[0045] Embodiments herein disclose a hardware-friendly CNN-based fundamental building block that can be used to construct AI-ILF of desired complexity, wherein such building blocks can be configured / adapted during inference stage by rate distortion optimizer. The core building framework can be useful to build low-complex AI model for several on-device applications, such as, but not limited to, image reconstruction and / or enhancement.
[0046] Embodiments herein disclose an attention framework that weighs the efficacy of each pathway, wherein such attention weights allow for auto-switching of the path based of the compression parameter.
[0047] Embodiments herein improve the BD-Rate performance, wherein preliminary experiments show significant improvement in compression performance as measured in BD-Rate as compared with AI-models of similar complexity. Preliminary simulation results show an impressive peak signal-to-noise ratio (PSNR) and enhance the visual quality of the reconstructed frames. Embodiments herein beat the related art among all JVET proposals in terms of the BD-BR performance and model complexity.
[0048] An embodiment of the disclosure provides a configurable and scalable neural network for in-loop filtering in video codecs.
[0049] An embodiment of the disclosure provides a hardware-friendly, scalable and configurable methods for low complexity (<=5kMAC / pixel) AI-based in-loop filters.
[0050] An embodiment of the disclosure provides a hardware friendly (HW-Friendly) Core neural network (NN) block with all HW-friendly compute operations with a fixed and deterministic per-pixel complexity.
[0051] An embodiment of the disclosure provides a hardware-friendly Convolutional Neural Network (CNN) based core building block, which can be used to design a scalable AI model of desired complexity, wherein the proposed CNN block offers re-configurability such that it can be configured on-the-fly during encoding / decoding.
[0052] An embodiment of the disclosure provides a core building block containing switchable multi-pathways and further stacking of core blocks serially / parallel that allows flexible in model structure, providing a user-preferable configuration.
[0053] An embodiment of the disclosure provides an attention framework that weighs the efficacy of each pathway, wherein such attention weights allow for auto-switching of the path based of the compression parameter.
[0054] An embodiment of the disclosure provides a core building framework that is useful to build low-complex AI-model for several on-device applications such as image reconstruction and / or enhancement
[0055] An embodiment of the disclosure provides a core CNN block with attention-based multi-branch learning paradigm in which the core blocks are re-configurable during the inference.
[0056] An embodiment of the disclosure provides an AI model that is scalable via parallelization / serialization of the arrangement of core blocks to suit the complexity budget & performance needs.
[0057] An embodiment of the disclosure provides a flexible CNN framework which is adaptable and configurable based on applications, content, encoders and in general user-preferences.
[0058] An embodiment of the disclosure provides a method for performing an in-loop filter processing in a video encoder using a scalable neural network model that adapts its architecture based on content-related properties and encoder parameters, thereby improving visual quality with controlled complexity.
[0059] An embodiment of the disclosure provides a neural network-based in-loop filter by selecting pre-defined core blocks, residual units, and pathways dynamically, so as to allow a flexible and efficient reconstruction of video frames
[0060] An embodiment of the disclosure provides a technique for enhancing reconstructed video frames using a configurable and content-aware neural network, wherein the model is updated using pre-trained parameters influenced by the number and type of selected processing pathways
[0061] An embodiment of the disclosure provides a method that enables dynamic adaptation of in-loop filtering complexity in a video encoder by using content characteristics and encoder settings to configure a low-complexity, high-performance neural network
[0062] An embodiment of the disclosure provides a system and method for selectively applying in-loop neural network filters on a per-pixel-block basis using a scalable neural network, thereby optimizing the trade-off between video quality and computational efficiency.
[0063] Figure 1 depicts an example of Versatile Video Coding (VVC) video decoder. The goal of image and video compression is to represent each image or video frame using fewer bits, enabling efficient transmission over communication channels. On an encoder side, the residual signal of the current frame to be encoded is first transformed into a spectral domain. Then, based on the response characteristics of the human visual system, the spectral coefficients are quantified by preserving lower-frequency components more than higher-frequency ones. This compression significantly reduces the number of bits required to represent the original signal.
[0064] Figure 1 illustrates the In-Loop Filtering (ILF) structure within a VVC video decoder. After initial decoding by preceding modules (based on standard coding methods), the reconstructed frame undergoes a sequence of filtering stages to reduce compression artefacts and enhance visual quality. These stages include: Luma Mapping with Chroma Scaling (LMCS) (102), Deblocking Filter (DBF) (104), Sample Adaptive Offset (SAO) (106), Adaptive Loop Filter (ALF) (108) and a Cross-Component ALF (CC-ALF) (110).
[0065] The filtered frame is then output or stored as a reference for predicting future frames, improving both coding efficiency and perceptual quality. Additionally, in-loop filters such as DBF (104), SAO (106), and ALF (108) are applied to the reconstructed frame to suppress artefacts. The filtered frame is stored in the reference buffer for use in future frame prediction.
[0066] On a decoder side, the received bitstream is decoded to reconstruct the video frame. However, due to compression at the encoder, particularly quantization, essential information is lost, resulting in visual artefacts (e.g., blocking, blurring) and pixel-level errors such as increased mean squared error. To mitigate these issues, the reconstructed frame passes through a series of in-loop filters, including DBF (104), SAO (106), and ALF (108). These filters aim to suppress artefacts and restore lost signal characteristics, thereby improving visual quality and fidelity. The filtered frame is used both for display and as a reference for decoding subsequent frames.
[0067] In recent developments, several AI-based in-loop filtering (ILF) tools have been proposed by members of the Moving Picture Experts Group / Joint Video Experts Team (MPEG / JVET) video standardization committee to further enhance the quality of reconstructed video frames. These tools are designed to either augment or replace one or more conventional classical ILF components. However, the computational complexity of such AI models―measured in kilo multiply-accumulate operations per pixel (kMAC / pixel)―is extremely high, making practical on-device implementation nearly impossible.
[0068] For instance, the computational complexity of current AI-based ILF tools is prohibitively high, rendering them infeasible for implementation on current or near-future devices. JVET AhG11 currently defines three operating points: High (477 kMAC / pixel), Low (17 kMAC / pixel), and Very Low (5.1 kMAC / pixel). Even the lowest setting remains too complex for efficient on-device deployment.
[0069] Very few implementations utilize low-complexity neural networks (e.g., <= 5kMAC / pixel). Not only is the performance of such models limited, but they also use fixed, rigid architectures. As a result, existing models require reconfiguration and retraining if the model's complexity needs to be adjusted during deployment. Furthermore, many low-complexity models still rely on core mathematical operations, such as MaxPooling or Transformer Attention, that are not yet hardware-friendly.
[0070] The current JVET proposals for next-generation video codecs suggest the use of various AI-based in-loop filters alongside the classical filters from earlier standards. However, not only are the proposed AI-ILF models computationally intensive, but they also operate within rigid frameworks and contain operations that are difficult to realize in hardware. Additionally, these models lack the flexibility to dynamically adjust computational complexity during inference.
[0071] Figure 2 depicts the usage of AI-based model in parallel with the deblocking filter (DBF) (104). The purpose of the AI-model (100) is to counter the deblocking artefacts that are present in the reconstructed signal (rec) (112) due to quantization. In an example, Figure 2 depicts an architecture where an AI model is integrated in parallel with the DBF (104) in a video decoding pipeline, typically in the VVC decoder. When a video is compressed at the encoder, the quantization of the transform-domain coefficients introduces artefacts in the reconstructed signal at the decoder. One common type of distortion is blocking artefacts, which appear as visible block boundaries in smooth regions due to independent block-based transformation and quantization. To reduce such artefacts, traditional video codecs apply the DBF (104) as part of the in-loop filtering (ILF) stage. However, while DBF improves visual quality, it may not always be sufficient, especially in complex or highly compressed scenarios.
[0072] However, the disclosed AI model is not limited to being arranged in parallel with the DBF (104). In an embodiment of the disclosure, the AI model may be arranged in parallel with one or more other in-loop filtering tools, such as the LMCS (102), the SAO (106), the ALF (108), and / or the CC-ALF (110). Additionally or alternatively, the AI model may be arranged in series with one or more of the foregoing in-loop filtering tools, for example, being positioned upstream or downstream of any of the LMCS (102), the SAO (106), the ALF (108), and the CC-ALF (110) in the video decoding pipeline.
[0073] Figure 3 depicts the current configuration of an AI model (100) proposed in JVET AhG11. In the depicted example, VLOP is considered as an example of the AI model, wherein the AI model (100) is a 5.16 KMAC / pixel pre-trained model that feeds reconstructed signal (rec) (112) along with prediction data (Pred) (116), boundary strength (BS) (118), quantization parameter (QP) (120), and type of frame (122), such as an intra (I) frame, a predicted (P) frame, or a bi-predicted (B) frame. The pre-trained neural network operates at a complexity of 5.16 KMAC / pixel and takes multiple inputs including the reconstructed signal (rec) (112), prediction data (Pred) (116), boundary strength (BS) (118), quantization parameter (QP) (120), and frame type. The inputs help the model identify and correct compression artefacts more effectively than traditional filters. The goal is to enhance visual quality in the reconstructed frame by intelligently restoring lost detail using contextual information.
[0074] In an embodiment of the disclosure, the AI model (100) may comprise a head network (1002) and a body network (1004). The head network (1002) may be an input-processing network configured to obtain one or more inputs associated with video encoding or video decoding (e.g., a reconstructed signal (112), prediction data (116), boundary strength (118) , a quantization parameter (120), and / or a type of frame (122)) and to generate one or more feature representations based on the obtained inputs.
[0075] In an embodiment of the disclosure, the body network (1004) may be a backbone network configured to process the feature representations output from the head network to perform in-loop filtering on a reconstructed video frame or a reconstructed pixel block. In an embodiment of the disclosure, the body network (1004) may comprise one or more neural network blocks, such as a transformer-based block, an attention-based block, a pooling-based block (e.g., a max-pooling block), and / or a convolution-based block, to enhance visual quality by reducing compression artefacts. And, the AI model (100) may output a output (114) from the body network (1004).
[0076] However, the AI models (100) like VLOP come with significant limitations, like fixed framework and the complexity is quite high with respect to entire VVC codec (1-1.5 KMAC / pixel). Moreover, the AI-models (100) cannot dynamically scale the complexity based on deployment scenarios, and any change in performance or efficiency requires complete re-design and re-training. Thus, there is a need to design a lower complex model (<5kMAC / pixel) with a scalable framework according to use-case. To reduce the complexity of models, it is required to be re-designed and re-trained. Hence, there is a need in the art for solutions which will overcome the above mentioned drawback(s), among others.
[0077] Referring now to the drawings, and more particularly to Figures 4 through 20, where similar reference characters denote corresponding features consistently throughout the figures, there are shown embodiments.
[0078] Figure 4 shows various hardware components of the electronic device (400), according to the an embodiment. The electronic device (400) can be, for example, but not limited to a laptop, a desktop computer, a notebook, a Device-to-Device (D2D) device, a vehicle to everything (V2X) device, a smartphone, a foldable phone, a smart TV, a tablet, an immersive device, and an internet of things (IoT) device.
[0079] The electronic device (400) includes at least one of a processor (410), a communicator (420), a memory (430), a sensor (440), a video encoder (450), a video decoder (460) and / or a neural network-based in-loop filtering controller (470). The processor (410) is coupled with the communicator (420), the memory (430), the sensor (440), the video encoder (450), the video decoder (460) and the neural network-based in-loop filtering controller (470). In an embodiment, the video encoder (450) or the video decoder (460) include the neural network-based in-loop filtering controller (470).
[0080] During encoder operations, the neural network-based in-loop filtering controller (470) generates a reconstructed video frame corresponding to a source video frame. Further, the neural network-based in-loop filtering controller (470) may configure a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame.
[0081] In an embodiment of the disclosure, a pre-defined scalable neural network model may be provided for in-loop filtering. The pre-defined scalable neural network model may include at least one pre-defined core block, and each pre-defined core block may comprise at least one residual unit. A scalable neural network model may be configured from the pre-defined scalable neural network model by selecting one or more of the pre-defined core blocks and, within each selected core block, selecting at least one residual unit. The selected core block(s) and the selected residual unit(s) may define a configured scalable neural network model to be applied for in-loop filtering. The electronic device (encoder or decoder) may configure the scalable neural network model by using the pre-defined scalable neural network model.
[0082] In an embodiment, the scalable neural network model for each reconstructed pixel block for the reconstructed video frame is configured by correlating at least one content property with at least one encoder parameter.
[0083] In an embodiment, the neural network-based in-loop filtering controller (470) selects the at least one pre-defined core block, and at least one residual unit within each selected core block. Further, the neural network-based in-loop filtering controller (470) selects at least one pathway within each selected core block. Further, the neural network-based in-loop filtering controller (470) configures the scalable neural network model for each reconstructed pixel block for the reconstructed video frame based on the selection.
[0084] The pre-defined scalable neural network model is defined by a number of core blocks, where each core block includes at least one residual unit. The pre-defined scalable neural network model is trained offline prior to a video encoding operation or a video decoding operation. The pre-defined scalable neural network model used by a video encoder and a video decoder may be the same.
[0085] In an embodiment, the residual unit handles a two dimensional (2D) convolutional filter operation and a rectified linear unit (ReLU) operation. The 2D convolutional filter operation applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, where the spatial feature includes at least one of: an edge, a texture, and a pattern, where the ReLU operation is performed after the 2D convolutional filter operation to introduce non-linearity in the source video frame.
[0086] In an embodiment, each core block includes at least one of: a series of 2D convolutional block and a series of 1D convolutional block with a normalization block and at least one of: a linear activation and a non-linear activation.
[0087] In an embodiment, a smaller number of the at least one core block is selected for a reconstructed pixel block with a lower quantization parameter (QP), wherein the at least one pre-trained neural network parameter is modified based on the at least one residual unit, wherein the pre-trained neural network parameter determines a weightage significance of each pathway within each selected core block, wherein the pre-trained neural network parameter is recalibrated when a configuration of the pre-defined scalable neural network model is changed post-training.
[0088] In an embodiment, the pre-defined scalable neural network model is configured using a rate-distortion-complexity optimizer, where the rate-distortion-complexity optimizer determine which configuration will provide a least cost of encoding, and where the rate-distortion-complexity optimizer determines a correct balance between distortion caused due to quantization in the video encoder (450) and a bitrate to encode the reconstructed pixel block.
[0089] Further, the neural network-based in-loop filtering controller (470) modifies at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within the at least one core block. Further, the neural network-based in-loop filtering controller (470) generates an optimized reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0090] Further, the neural network-based in-loop filtering controller (470) generates the reconstructed video frame includes at least one pre-defined core block, at least one residual unit within each selected core block and at least one pathway within the at least one selected core blocks. In an embodiment,
[0091] The neural network-based in-loop filtering controller (470) determines at least one content property associated with the source video frame and at least one encoder parameter associated with the source video frame. The at least one content property comprises at least one of: edge information and texture information. The at least one encoder parameter includes a quantization parameter. Further, the neural network-based in-loop filtering controller (470) generates the reconstructed video frame corresponding to the source video frame based on at least one of: the at least one content property and the at least one encoder parameter.
[0092] During the video decoder operation, the neural network-based in-loop filtering controller (470) generates a reconstructed video frame corresponding to a source video frame. The neural network-based in-loop filtering controller (470) receives an encoded video bit-stream comprising at least one neural network configuration, where the at least one neural network configuration comprises selection of at least one core block, at least one residual unit within the at least one selected core block and at least one pathway within the at least one selected core block.
[0093] The neural network-based in-loop filtering controller (470) generates a reconstructed video frame using the encoded video bitstream, where the reconstructed video frame is divided into a plurality of reconstructed pixel blocks. In an embodiment, the neural network-based in-loop filtering controller (470) determines at least one content property associated with the encoded video bitstream and at least one encoder parameter associated with the encoded video bitstream. The at least one content property includes at least one of: edge information and texture information. The at least one encoder parameter includes a quantization parameter. The neural network-based in-loop filtering controller (470) generates the reconstructed video frame corresponding to the encoded video bitstream based on at least one of: the at least one content property and the at least one encoder parameter.
[0094] The neural network-based in-loop filtering controller (470) configures a pre-defined scalable neural network model for each reconstructed pixel block using the at least one neural network configuration.
[0095] In an embodiment, the pre-defined scalable neural network model is defined by a number of core blocks, wherein each core block comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained in an offline prior to a video encoding operation.
[0096] the residual unit handles a two dimensional (2D) convolutional filter operation and a rectified linear unit (ReLU) operation, wherein the 2D convolutional filter operation applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein the ReLU operation is performed after the 2D convolutional filter operation to introduce non-linearity in the source video frame.
[0097] In an embodiment, each core block comprises at least one of: a series of 2D convolutional block and a series of 1D convolutional block with a normalization block and at least one of: a linear activation and a non-linear activation.
[0098] In an embodiment, the pre-defined scalable neural network model is configured using a rate-distortion-complexity optimizer, wherein the rate-distortion-complexity optimizer determine which configuration will provide a least cost of encoding, and wherein the rate-distortion-complexity optimizer determines a correct balance between distortion caused due to quantization in the video encoder (450) and the bitrate to encode the reconstructed pixel block.
[0099] In an embodiment, a smaller number of the at least one core block is selected for a reconstructed pixel block with a lower quantization parameter (QP), wherein the at least one pre-trained neural network parameter is modified based on the at least one residual unit, wherein the pre-trained neural network parameter determines a weightage significance of each pathway within each selected core block, wherein the pre-trained neural network parameter is recalibrated when a configuration of the pre-defined scalable neural network model is changed post-training.
[0100] Further, the neural network-based in-loop filtering controller (470) modifies at least one pre-trained neural network parameter associated with the configured scalable neural network model based on the at least one pathways within the at least one selected core block. The neural network-based in-loop filtering controller (470) generates an enhanced reconstructed video frame by feeding each pixel block in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter. The enhanced reconstructed video frame may also be referred to as an optimized reconstructed video frame.
[0101] The neural network-based in-loop filtering controller (470) is implemented by analog and / or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits and the like, and may optionally be driven by firmware.
[0102] The processor (410) may include one or a plurality of processors. The one or the plurality of processors may be a general-purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The processor (410) may include multiple cores and is configured to execute the instructions stored in the memory (430).
[0103] Further, the processor (410) is configured to execute instructions stored in the memory (430) and to perform various processes. The communicator (420) is configured for communicating internally between internal hardware components and with external devices via one or more networks. The memory (430) also stores instructions to be executed by the processor (410). The memory (430) may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory (430) may, in some examples, be considered a non-transitory storage medium. The term "non-transitory" may indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted that the memory (230) is non-movable. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in Random Access Memory (RAM) or cache).
[0104] Further, at least one of the plurality of modules / controller may be implemented through the AI model / the ML model using a data driven controller (not shown). The data driven controller can be an ML model based controller and an AI model based controller. A function associated with the AI model may be performed through the non-volatile memory, the volatile memory, and the processor.
[0105] The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or AI model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning.
[0106] Here, being provided through learning means that a predefined operating rule or AI model of a desired characteristic is made by applying a learning algorithm to a plurality of learning data. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / o may be implemented through a separate server / system.
[0107] The AI model may include of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
[0108] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0109] Although Figure 4 shows various hardware components of the electronic device (400) but it is to be understood that other embodiments are not limited thereon. In other embodiments, the electronic device (400) may include less or more number of components. Further, the labels or names of the components are used only for illustrative purposes and does not limit the scope of the invention. One or more components can be combined together to perform the same or substantially similar function in the electronic device (400).
[0110] Figure 5 depicts a method for performing in-loop filter processing of a video encoder, according to an embodiment. In an embodiment herein discloses a method for performing in-loop filter processing of the video encoder (450), as depicted in Figure 5.
[0111] In an embodiment of the disclosure, the flowchart (S500) depicts a method outlines a process for enhancing video quality, in steps S502-S512, by using the scalable neural network model in the video codec pipeline. The method uses content-aware configuration and adaptive neural network components to improve reconstructed video frames.
[0112] In step S502, the electronic device may obtain guiding encoder information and generating the reconstructed video frame corresponding to the source video frame, wherein generating includes determining one or more content properties (for example, edge and texture information) of the source video frame and encoder parameters (for example, quantization parameter) used to generate the reconstructed video frame. Meanwhile, Orig may represent the source video frame. In step S504, the electronic device may obtain the pre-trained neural network model trained. In step S506, the electronic device may select core blocks for the residual units and pathways.
[0113] In step S508, the electronic device may configure the pre-defined scalable neural network model for each reconstructed pixel block of the reconstructed video frame by correlating the extracted content properties and encoder parameters. The configuring operation includes selecting one or more pre-defined core blocks, one or more residual units within each selected core block, and pathways within the selected core blocks. In step S510, the electronic device may update one or more pre-trained neural network parameters of the configured scalable neural network model based on the number of pathways within the selected core blocks. In step S512, the electronic device may generate an enhanced reconstructed video frame by feeding each reconstructed pixel block of the reconstructed video frame to the configured scalable neural network model with the updated pre-trained neural network parameters.
[0114] In step S514, the electronic device may generate one or more encoded video bit-streams containing the selected one or more pre-defined core blocks, one or more residual units within each selected core block, and pathways within the selected core blocks with other video encoding modules.
[0115] Figure 6 depicts a flowchart (S600) for a method for performing in-loop filter processing of the video encoder (450) by the electronic device (400), according to an embodiment. Embodiments herein discloses a method for performing in-loop filter processing in a video encoder (450) using a scalable neural network model, wherein the model is dynamically configured based on content-related properties and encoder parameters to enhance the quality of reconstructed video frames.
[0116] In step S602, the electronic device may generate a reconstructed video frame corresponding to a source video frame. the electronic device may determining one or more content properties of the source video frame and at least one encoder parameter used to generate the reconstructed video frame. The at least one content property includes at least one of edge information and texture information. And the at least one encoder parameter includes a quantization parameter.
[0117] In step S604, the electronic device may configure a scalable neural network model by using pre-defined scalable neural network model. The scalable neural network may be configured based on the determined content-related properties and encoder parameters. The electronic device may select at least one of: one or more pre-defined core blocks, one or more residual units (RU) within each selected core block, and one or more pathways within the selected core blocks. The configured scalable neural network model for each reconstructed pixel block may be determined by correlating the at least one content property and the at least one encoder parameter.
[0118] In step S606, the electronic device may update or modify at least one pre-trained neural network parameters of the configured scalable neural network model based on the number of pathways within the selected core blocks in the configured scalable neural network model.
[0119] In step S608, the electronic device may generate an enhanced reconstructed video frame by feeding each pixel block of the reconstructed video frame with the at least one updated or modified pre-trained neural network parameter.
[0120] Figure 7 depicts a flowchart (S700) for a method for enhancing in-loop filter processing of the video codec by an electronic device (400), according to an embodiment. The method involves deriving and configuring a scalable neural network model to process reconstructed pixel blocks, thereby improving video quality while maintaining computational efficiency.
[0121] At step S702, the electronic device may generate the reconstructed video frame corresponding to the source video frame within at least one core block. At step S704, the electronic device may configure the pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame.
[0122] At step S706, the electronic device may modify the at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within the at least one core block. At step S708, the electronic device may generate the optimized reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0123] In an embodiment of the disclosure, the optimized reconstructed video frame may represent a reconstructed video frame to which in-loop filtering has been applied. Figure 8 depicts a method for performing in-loop filter processing of a video decoder, according to an embodiment. Embodiments herein disclose a method for performing in-loop filter processing of the video decoder (460) in the AI-based in-loop filtering module as depicted in Figure 8.
[0124] In an embodiment of the disclosure, the following flowchart (S800) depicts a method describing in steps of S802-S810, a decoder-side process for enhancing video quality using a scalable neural network model. The method utilizes neural network configuration data embedded in the encoded video bit-stream to dynamically reconstruct and refine each video frame with improved visual quality.
[0125] In step S802, the electronic device may obtain bitstream including information for neural network configurations. The information for the neural network configuration may include information for at least one of selection of core blocks, residual units within each selected core block and pathways within the selected core blocks to configure a scalable neural network.
[0126] In step S804, the electronic device may obtain a pre-trained scalable neural network model. The electronic device may generate a reconstructed video frame using the bitstream. The reconstructed video frame may include a plurality of reconstructed pixel blocks.
[0127] In step S806, the electronic device may configure the scalable neural network by using a pre-defined scalable neural network model for each of the plurality of reconstructed pixel blocks. the scalable neural network model may be determined by using the corresponding neural network configuration of the pre-defined scalable neural network model from the information included in the bitstream.
[0128] In step S808, the electronic device may update or modify at least one pre-trained neural network parameters of the configured scalable neural network model based on the number of pathways within the selected core blocks. And the electronic device may generate an enhanced reconstructed video frame by feeding each pixel block of the reconstructed video frame to the configured scalable neural network model with the at least one modified or updated pre-trained neural network parameters.
[0129] In step S810, the electronic device may perform operation of other video decoding module.
[0130] Figure 9a depicts a flowchart for a method for performing in-loop filter processing of the video decoder (460) by the electronic device (400) according to an embodiment. Embodiment herein discloses a method for enhancing video decoding by performing in-loop filter processing using a scalable neural network model (100). The method uses neural network configurations received in the encoded bitstream to dynamically reconstruct and improve the visual quality of each pixel block in the decoded video frame.
[0131] In step S902, the electronic device may obtain a bit-stream including information for selection of at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block to configure a scalable neural network. The information may comprise of one or more encoded source video frames of neural network configurations, wherein a neural network configuration consists of selection of at least one of: one or more core blocks, one or more residual units within each selected core block, and one or more pathways within the selected core blocks.
[0132] In step S904, the electronic device may generate a reconstructed video frame corresponding to a source frame using the bitstream, wherein the reconstructed video frame is divided into at least one reconstructed pixel blocks. The reconstructed video frame may include a plurality of the reconstructed pixel blocks.
[0133] In step S906, the electronic device may configure a scalable neural network model by using a pre-defined scalable neural network model for each of the plurality of the reconstructed pixel blocks l.
[0134] In step S908, the electronic device may update or modify at least one pre-trained neural network parameters of the configured scalable neural network model based on the number of pathways within the selected core blocks.
[0135] In step S910, the electronic device may generate an enhanced reconstructed video frame by applying each of the plurality of the reconstructed pixel blocks of the reconstructed video frame to the configured scalable neural network model with the at least one modified or updated pre-trained neural network parameter.
[0136] In an embodiment, the method includes generating the reconstructed video frame corresponding to the source video frame. Further, the method includes receiving the encoded video bit-stream comprising at least one neural network configuration. The at least one neural network configuration includes selection of at least one core block, at least one residual unit within the at least one selected core block and at least one pathway within the at least one selected core block. Further, the method includes generating the reconstructed video frame using the encoded video bitstream. The reconstructed video frame is divided into a plurality of reconstructed pixel blocks. Further, the method includes configuring the pre-defined scalable neural network model for each reconstructed pixel block using the at least one neural network configuration. Further, the method includes modifying the at least one pre-trained neural network parameter associated with the configured scalable neural network model based on the at least one pathways within the at least one selected core block. Further, the method includes generating the enhanced reconstructed video frame by feeding each pixel block in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0137] Figure 9b depicts a flowchart for a method for performing in-loop filter processing of the video decoder (460) by the electronic device (400) according to an embodiment. Embodiment herein discloses a method for enhancing video decoding by performing in-loop filter processing using a scalable neural network model (100). The method uses neural network configurations received in the encoded bitstream to dynamically reconstruct and improve the visual quality of each pixel block in the decoded video frame.
[0138] In step S912, the electronic device may generate a reconstructed video frame corresponding to a source video frame.
[0139] In step S914, the electronic device may obtain an encoded video bit-stream comprising at least one neural network configuration, wherein the at least one neural network configuration comprises selection of at least one core block, at least one residual unit within the at least one selected core block and at least one pathway within the at least one selected core block.
[0140] In step S916, the electronic device may generate a reconstructed video frame using the encoded video bitstream, wherein the reconstructed video frame is divided into a plurality of reconstructed pixel blocks.
[0141] In step S918, the electronic device may configure a pre-defined scalable neural network model for each reconstructed pixel block using the at least one neural network configuration.
[0142] In step S920, the electronic device may modify at least one pre-trained neural network parameter associated with the configured scalable neural network model based on the at least one pathways within the at least one selected core block.
[0143] In step S922, the electronic device may generate an enhanced reconstructed video frame by feeding each pixel block in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0144] Figure 10 depicts a fundamental configurable building block, hereinafter referred to as an Aggregated Core Unit (ACU) (1005), according to an embodiment. The ACU (1005) (specifically including the residual units (RUs) (1006)) involves only 2D convolutional filter and rectified linear unit (ReLU) operation both of which are hardware friendly.
[0145] Embodiments herein disclose a fundamental configurable building block, referred to as an Aggregated Core Unit (ACU) (1005), serves as a key component in the scalable neural network architecture. The ACU (1005) is designed to be modular and efficient, enabling flexible configuration based on video content characteristics and processing requirements. At the core, each ACU (1005) consists of one or more Residual Units (RUs) (1006), here depicted as RU (1006a) and RU (1006b), connected with N paths, in Figure 10 which perform only 2D convolutional filtering followed by a linear a Rectified Linear Unit (ReLU) activation function. The simple yet effective design ensures the computational operations within the ACU (1005) remain hardware-friendly, making it highly suitable for on-device implementation with limited resources, such as mobile or embedded platforms.
[0146] Figure 11 depicts the Tail Block (TB) (1008), which comprises of a series of 2D convolutional block (2D CNN) with linear activation. As depicted in Figure 11, the TB (1008) comprises of a linear series of K 2D CNN (1008a) and 2D CNN (1008b) and connected with linear activation As shown in Figure 10, the ACU (1005) has multiple switchable pathways (N paths) which are turned on / off by the rate-distortion optimizer based on a content properties of the source video frame and encoder parameters. The dynamic switching mechanism ensures the only the most relevant computational paths are used, optimizing both visual quality and computational efficiency. When combined with the TB (1008), the overall design enables a highly adaptable and resource-conscious neural network architecture for in-loop video enhancement.
[0147] When S is the set of all switched-on paths. Equation 1 depicts the output y if S is the set of all switched-on paths, as follow:
[0148] [Equation 1]
[0149]
[0150] Figure 12 depicts a scalable model, wherein the ACU (1005) is used to design a scalable AI-ILF (100) of desired complexity, wherein such a model can be implemented partly with hardware and software design where the trailing fundamental blocks can be turned on / off.
[0151] In an embodiment of the disclosure, a body network (1004) may include at least one ACU(1005) and at least one TB (1008).
[0152] Let, be the ACU (1005) operation and be the TB (1008) operation. Equation 2 depicts designing of scalable AI-ILF of desired complexity, as follow:
[0153] [Equation 2]
[0154] Consider the case where number of ACU (1005) are P. Equations 3-4 depict the loss function of the AI-ILF, as follow:
[0155] [Equation 3]
[0156]
[0157]
[0158] ...
[0159]
[0160] [Equation 4]
[0161]
[0162]
[0163] ...
[0164]
[0165] The loss function used to train this network minimizes the difference between the source signal with each of .
[0166] Figure 13 depicts the internal structure of the ACU (1005), according to an embodiment. The ACU block (1005) comprises of a multiway merge of residual signal with the original input signal. The residual unit (1006) comprises of hardware-friendly 2D convolutional layer with either e.g., filter kernel size which requires very low MAC / pixel. In addition, the residual paths are weighted with learnable scalar attention viz. branch attention for . , weighs the efficacy of each residual branch thus emphasizing which band-pass filter channels are more useful for reconstruction. comprises of learnable scaler weight and bias parameter and, the ACU (1005) output is given by Equation 5. Equation 5 depicts internal structure of the core blocks, as follow:
[0167] [Equation 5]
[0168]
[0169] Embodiment herein discloses the internal structure of the Aggregated Core Unit (ACU) (1005) is designed to efficiently enhance video frame reconstruction while maintaining low computational overhead. The ACU (1005) performs a multiway merging of the residual signal with the original input signal, enabling to selectively correct artefacts introduced during compression. Each Residual Unit (RU) (1006) within the ACU (1005) comprises a hardware-friendly 2D convolutional layer, typically with a small kernel size such as , which significantly reduces multiply-accumulate operations (MACs) per pixel, making it ideal for real-time and on-device deployment. Furthermore, each residual path within the ACU (1005) is modulated by a learnable scalar attention mechanism, defined by branch attention parameters and , where . These parameters dynamically adjust the influence of each branch based on content, effectively learning which band-pass filter channels contribute most to accurate reconstruction. The fine-grained weighting of residual branches allows the ACU (1005) to focus computational effort that matters most, enhancing both efficiency and reconstruction quality. The final output of the ACU (1005) is thus an adaptively refined signal that balances performance with resource constraints, making it well-suited for integration in scalable neural network architectures for video coding.
[0170] Figure 14 describes the configurable inference-time framework of ACU block (1005). Based on the ordered-statistics of the batch-attention scores / weights , the importance of the signal carried by each pathway is calculated. Let, be the order-statistic of the Branch Attention scores like Equation 6. Equation 6 depicts configurable inference-time framework of core blocks, as follow:
[0171] [Equation 6]
[0172]
[0173] Let, M<N number of pathways chosen for inference based on BA ordering, then inference-time scaling factors are calculated by Equations 7-8. Equations 7-8 depict inference-time scaling factors, as follow:
[0174] [Equation 7]
[0175]
[0176] [Equation 8]
[0177]
[0178] The pathway parameters needs to be normalized based on the ratio of the number of paths chosen, i.e., (N / M). Based upon the application such total number of pathways chosen `M' could be selected at a frame level or group of pictures or video sequence level.
[0179] Figure 15 depicts an example scenario of a width-wise scalable framework of the idea with ACU (1005) as the building block. The architecture information shown in the example is as follows:
[0180] Network width = 4
[0181] Network depth = 1
[0182] Total ACU blocks = 1
[0183] A wider network configuration can capture fine-grained details better than thinner network due to more parallel channels (RUs) (1006a, 1006b, 1006c, 1006d). However, if the depth of the network is made to shallow while width is increased, accuracy will reduce.
[0184] The figure illustrates an example scenario of a width-wise scalable neural network framework using the Aggregated Core Unit (ACU) (1005) as the fundamental building block. In the depicted configuration, the network has a width of 4, containing four parallel Residual Units (RUs) (1006a, 1006b, 1006c, 1006d) within the single ACU block, and a depth of 1, indicating a shallow network with only one ACU layer in the processing pipeline. The width-wise scalability allows the architecture to be adapted based on available computational resources or visual fidelity requirements. Width-wise scalable denotes the number of parallel residual block (RU links) are more than the serial residual block (ACU block). By increasing the width (i.e., the number of parallel RUs), the network can process and represent finer-grained features more effectively, as each parallel channel can focus on different frequency bands or spatial patterns. Such configuration proves particularly advantageous for complex video content where capturing subtle textures, edges, and visual artefacts is critical to high-quality reconstruction. Thus, a wider configuration, while computationally more intensive, offers improved performance in visual quality compared to thinner networks, making it ideal for applications where precision and detail preservation are prioritized.
[0185] Figure 16 depicts an example scenario of a depth-wise scalable framework of the idea with ACU (1005) as the building block. Depth-wise scalable denotes the number of serial residual blocks (ACU blocks) are more than the number of parallel residual blocks (RU links). The architecture information shown in the example is as follows:
[0186] Network width = 1
[0187] Network depth = 4
[0188] Total ACU blocks = 4
[0189] A deeper network configuration can better model noise / high frequency characteristics as noise generation is a Markov process so high depth will be better. However, the increased depth not always helps if the network is thinner, i.e., number of channels per depth level is less. Additionally, low channel width can cause loss in valuable information.
[0190] Figure 16 depicts an example of a scalable framework of the idea with ACU as the building block. The architecture information is as follows:
[0191] Network width = 4
[0192] Network depth = 3
[0193] Total ACU blocks = 4
[0194] This hybrid between depth-wise and width-wise scalable architecture offers a perfect blend of fine-grained information and noise characteristic generation. Application wise (for example, when the content is smooth), a shallow network could be chosen while for textured content a deeper network would be preferred.
[0195] Figure 17 depicts an example of the overall configurable and scalable module. Turning ON / OFF the switch either during inference in the encoder end the complexity of the model can be altered on-the-fly. The I / P signal is omni-present in ACU, for ease of visualization, the I / P signal has been removed from Figure 17. Configurable and scalable module denotes an operation where the RU links, more precisely the number of parallel residual blocks and the ACU blocks, or the number of serial residual blocks are increased or decreased compared to the initial configuration. The complexity could be made adaptable to a frame or group of pictures or video sequence itself. A mode signal needs to be added to the bitstream to signal the decoder the configuration of the model selected. Adaptable complexity model is required since, for smooth content a shallow network is enough to reconstructed the frame signal; for highly textured content a deeper network is required as noise is modelled by CNN as a Markov process, and if the compression ratio is high, the network needs to be wider as it can retrieve back more fine-grained details. The complexity of the network chosen at the encoder end can be pre-decided in a fixed fashion way. Alternatively, rate-distortion complexity optimizer could decide the complexity based upon the input video signal.
[0196] Figure 18 depicts depicts an example scenario of the overall configurable and scalable module.
[0197] In an embodiment of the disclosure, the in-loop filtering module comprises one or more configurable neural network units in which one or more switches are selectively turned ON or OFF to activate or deactivate corresponding processing paths and / or blocks. By turning ON / OFF the switch during inference at an encoder side or a decoder side, a computational complexity of the neural network can be altered on-the-fly without re-training, thereby enabling a configured scalable neural network model to be formed from a pre-defined scalable neural network model. The complexity could be made adaptable to a frame or group of pictures or video sequence itself. A mode signal needs to be added to the bitstream to signal the selected model configuration to the decoder.
[0198] In an embodiment of the disclosure, An adaptable-complexity model is desirable because different content types and compression conditions require different network capacities. For example, a shallow network may be sufficient to enhance smooth content and reconstruct the frame signal. In contrast, highly textured content may benefit from a deeper network to better suppress noise-like artefacts and recover fine structures (e.g., when such artefacts can be modeled as a Markov process by a CNN). Additionally, when the compression ratio is high, the network may be configured to be wider so as to retrieve finer-grained details from the reconstructed signal.
[0199] In an embodiment of the disclosure, network complexity may be selected in different ways. In some embodiments, the encoder may use a pre-decided, fixed network configuration (e.g., a fixed operating point) for selecting the network complexity. Alternatively, a rate-distortion(-complexity) optimizer may determine the network complexity based on the input video signal (and, optionally, encoder parameters), thereby selecting an appropriate configuration for a given content and compression condition.
[0200] For ease of visualization, an I / P signal that may be omni-present within an Aggregated Core Unit (ACU) is omitted from Figure 18.
[0201] Figure 19 depicts an example illustrating an initial training configuration where the number of ACU blocks, i.e., P = 3 and the number of RU links i.e; N = 5. During inference, this model could be configured based on the following objective:
[0202] Scalability w.r.t Quantization Parameter:
[0203] If quantization parameter (QP) , P =2 is chosen as the error in Rec will be less thus less scalable unit is required. While if QP , P = 3 is chosen to model the high error component in Rec
[0204] Configurability w.r.t Content:
[0205] If the frame is highly texture, a wider network (for example, N=5) is needed to model fine-grained details. While if the frame is smooth, N = 3 could to chosen to model the information in Rec.
[0206] Adaptability w.r.t Residual Signal:
[0207] If the difference between source and Rec signal is less, a lesser scalable unit (specifically P<3 to model) is needed to minimize the error signal. To minimize high difference in residual signal, all i.e., P=3 scalable blocks are needed.
[0208] In the example, the configured network is shown when the frame content is smooth as QP = 22 thus, requiring P =2 and N <=3 across ACUs'
[0209] Figure 20 depicts an example where the model overall trained model is embedded with P = 3 ACU blocks and N = 5 RU links. In continuation with the example depicted in Figure 15, where the encoder decided P = 2 and N <=3 re-configuration, mode information of the switch states is sent, where, i) a first mode denotes the switch-mode On, ii) a second mode denotes the switch-mode Off, and iii) a third mode represent the mode information needs not be sent in bit-stream as the final ACU is absent.
[0210] The mode information is sent as a part of the bitstream.
[0211] The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the network elements. The elements include blocks which can be at least one of a hardware device, or a combination of hardware device and software module.
[0212] The embodiments disclosed herein describe a configurable and scalable neural network for in-loop filtering in video codecs. The embodiment discloses a hardware-friendly, CNN-based building block tailored for AI-based in-loop filtering (AI-ILF) in video codecs. The core module is not only efficient but also configurable and scalable, enabling the construction of low-complexity neural network models, can be tailored to meet specific computational budgets. Furthermore, the rate-distortion optimizer during encoding can be reconfigured dynamically, based on content complexity and encoder parameters, offering fine-grained control over inference cost and visual fidelity. Unlike rigid models, the architecture supports adaptive deployment across a range of devices, including resource-constrained hardware. Preliminary results demonstrate significant improvements in Peak Signal-to-Noise Ratio (PSNR) and enhanced visual quality of reconstructed frames. Moreover, the method outperforms all existing JVET proposals in terms of BD-Bitrate (BD-BR) performance versus model complexity, establishing a new benchmark for efficient and high-performing AI-ILF solutions. Therefore, it is understood that the scope of the protection is extended to such a program and in addition to a computer readable means having a message therein, such computer readable storage means contain program code means for implementation of one or more steps of the method, when the program runs on a server or mobile deviceor any suitable programmable device. The method is implemented in at least one embodiment through or together with a software program written in e.g., Very high speed integrated circuit Hardware Description Language (VHDL) another programming language, or implemented by one or more VHDL or several software modules being executed on at least one hardware device. The hardware device can be any kind of portable device that can be programmed. The device may also include means which could be e.g., hardware means like e.g., an ASIC, or a combination of hardware and software means, e.g. an ASIC and an FPGA, or at least one microprocessor and at least one memory with software modules located therein. The method embodiments described herein could be implemented partly in hardware and partly in software. Alternatively, the invention may be implemented on different hardware devices, e.g., using a plurality of CPUs.
[0213] In an embodiment of the disclosure, a method for performing an in-loop filter processing by a video encoder (450) included in an electronic device (400) is provided. The method may include generating a reconstructed video frame corresponding to a source video frame. The method may include configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame. The method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model. The method may include generating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0214] In an embodiment of the disclosure, the method may include selecting the at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block.
[0215] In an embodiment of the disclosure, the method may include generating a bitstream including information for at least one of the at least one core block, at least one residual unit within each of the at least one core block or at least one pathway within the at least one core block.
[0216] In an embodiment of the disclosure, the method may include determining at least one content property associated with the source video frame and at least one encoder parameter used to generate the reconstructed video frame. The at least one content property may include at least one of: edge information and texture information. The at least one encoder parameter may include a quantization parameter.
[0217] In an embodiment of the disclosure, the configured scalable neural network model for each reconstructed pixel block may be determined by correlating the at least one content property and the at least one encoder parameter.
[0218] In an embodiment of the disclosure, the configured scalable neural network model is defined by a number of the least one core block. Each of the at least one core block may comprise at least one residual unit. The pre-defined scalable neural network model be trained offline prior to a video encoding operation.
[0219] In an embodiment of the disclosure, the residual unit may include a two dimensional (2D) convolutional layer and a rectified linear unit (ReLU) layer, wherein a 2D convolutional filter operation performed in the 2D convolution layer applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein a ReLU operation performed in the ReLU layer is applied after the 2D convolutional layer to introduce non-linearity into an output of the 2D convolution layer.
[0220] In an embodiment of the disclosure, each of the at least one core block may comprise at least one of: a series of 2D convolutional block and a normalization block.
[0221] In an embodiment of the disclosure, the configured scalable neural network model may be determined using a rate-distortion-complexity optimizer, wherein the rate-distortion-complexity optimizer determines at least one configuration of the pre-defined scalable neural network model that minimizes a cost function associated with encoding the reconstructed pixel block, and a balance between distortion caused by quantization in the video encoder (450) and a number of bits to encode the reconstructed pixel block.
[0222] In an embodiment of the disclosure, a smaller number of the at least one core block may be selected for a reconstructed pixel block with a lower quantization parameter (QP), or a lower difference with the reconstructed pixel block and a source pixel block corresponding to the reconstructed pixel block.
[0223] In an embodiment of the disclosure, the at least one pre-trained neural network parameter may be modified based on a number of the at least one residual unit within the configured scalable neural network model, wherein the pre-trained neural network parameter is normalized using ratio of a number of the at least one residual units within the pre-defined scalable neural network model to the number of the at least one residual unit within the configured scalable neural network.
[0224] In an embodiment of the disclosure, a method for performing an in-loop filter processing by a video decoder (460) included in an electronic device (400) is provided. The method may include obtaining a bitstream including information for selection of at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block to configure a scalable neural network. The method may include generating a reconstructed video frame using the bitstream, wherein the reconstructed video frame includes a plurality of reconstructed pixel blocks. The method may include configuring the scalable neural network by using a pre-defined scalable neural network model for each of the plurality of the reconstructed pixel blocks. The method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of the at least one pathway within the at least one core block. The method may include generating an enhanced reconstructed video frame by feeding each of the plurality of the reconstructed pixel blocks in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0225] In an embodiment of the disclosure, the configured scalable neural network model may be determined by a number of core blocks, wherein each of the core blocks comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained offline prior to a video decoding operation.
[0226] In an embodiment of the disclosure, the residual unit may include a two dimensional (2D) convolutional layer and a rectified linear unit (ReLU) layer, wherein a 2D convolutional filter operation performed in the 2D convolution layer applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein a ReLU operation performed in the ReLU layer is applied after the 2D convolutional layer to introduce non-linearity into an output of the 2D convolution layer.
[0227] In an embodiment of the disclosure, a method for transmitting a bitstream generated by an encoding method is provided. The encoding method may include generating a reconstructed video frame corresponding to a source video frame. The encoding method may include configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame. The encoding method may include modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model. The encoding method may include generating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0228] In an embodiment of the disclosure, a method for performing an in-loop filter processing of a video encoder (450) by an electronic device (400) is provided. The method may include generating, by the electronic device (400), a reconstructed video frame corresponding to a source video frame within at least one core block. The method may include configuring, by the electronic device (400), a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame. The method may include modifying, by the electronic device (400), at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within the at least one core block. The method may include generating, by the electronic device (400), an optimized reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0229] In an embodiment of the disclosure, the method may include generating, by the electronic device (400), the reconstructed video frame comprising at least one pre-defined core block, at least one residual unit within each selected core block and at least one pathway within the at least one selected core blocks.
[0230] In an embodiment of the disclosure, the method may include determining at least one content property associated with the source video frame and at least one encoder parameter associated with the source video frame, wherein the at least one content property comprises at least one of: edge information and texture information, wherein the at least one encoder parameter comprises a quantization parameter. In an embodiment of the disclosure, the method may include generating, by the electronic device (400), the reconstructed video frame corresponding to the source video frame based on at least one of: the at least one content property and the at least one encoder parameter.
[0231] In an embodiment of the disclosure, the pre-defined scalable neural network model for each reconstructed pixel block for the reconstructed video frame may be configured by correlating at least one content property with at least one encoder parameter.
[0232] In an embodiment of the disclosure, the method may include selecting the at least one pre-defined core block. The method may include selecting at least one residual unit within each selected core block. The method may include selecting at least one pathway within each selected core block. The method may include configuring, by the electronic device (400), a pre-defined scalable neural network model for each reconstructed pixel block for the reconstructed video frame based on the selection.
[0233] In an embodiment of the disclosure, the pre-defined scalable neural network model may be defined by a number of core blocks, wherein each core block comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained in an offline prior to a video encoding operation.
[0234] In an embodiment of the disclosure, the residual unit may handle a two dimensional (2D) convolutional filter operation and a rectified linear unit (ReLU) operation, wherein the 2D convolutional filter operation applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein the ReLU operation is performed after the 2D convolutional filter operation to introduce non-linearity in the source video frame.
[0235] In an embodiment of the disclosure, each core block may comprise at least one of: a series of 2D convolutional block and a series of 1D convolutional block with a normalization block and at least one of: a linear activation and a non-linear activation.
[0236] In an embodiment of the disclosure, the pre-defined scalable neural network model may be configured using a rate-distortion-complexity optimizer, wherein the rate-distortion-complexity optimizer determine which configuration will provide a least cost of encoding, and wherein the rate-distortion-complexity optimizer determines a correct balance between distortion caused due to quantization in the video encoder (450) and a bitrate to encode the reconstructed pixel block.
[0237] In an embodiment of the disclosure, a smaller number of the at least one core block may be selected for a reconstructed pixel block with a lower quantization parameter (QP), wherein the at least one pre-trained neural network parameter is modified based on the at least one residual unit, wherein the pre-trained neural network parameter determines a weightage significance of each pathway within each selected core block, wherein the pre-trained neural network parameter is recalibrated when a configuration of the pre-defined scalable neural network model is changed post-training.
[0238] In an embodiment of the disclosure, a method for performing an in-loop filter processing of a video decoder (460) by an electronic device (400) is provided. The method may include generating, by the electronic device (400), a reconstructed video frame corresponding to a source video frame. The method may include receiving, by the electronic device (400), an encoded video bit-stream comprising at least one neural network configuration, wherein the at least one neural network configuration comprises selection of at least one core block, at least one residual unit within the at least one selected core block and at least one pathway within the at least one selected core block. The method may include generating, by the electronic device (400), a reconstructed video frame using the encoded video bitstream, wherein the reconstructed video frame is divided into a plurality of reconstructed pixel blocks. The method may include configuring, by the electronic device (400), a pre-defined scalable neural network model for each reconstructed pixel block using the at least one neural network configuration. The method may include modifying, by the electronic device (400), at least one pre-trained neural network parameter associated with the configured scalable neural network model based on the at least one pathways within the at least one selected core block. The method may include generating, by the electronic device (400), an enhanced reconstructed video frame by feeding each pixel block in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0239] In an embodiment of the disclosure, the method may include determining at least one content property associated with the encoded video bitstream and at least one encoder parameter associated with the encoded video bitstream, wherein the at least one content property comprises at least one of: edge information and texture information, wherein the at least one encoder parameter comprises a quantization parameter. The method may include generating, by the electronic device (400), the reconstructed video frame corresponding to the encoded video bitstream based on at least one of: the at least one content property and the at least one encoder parameter.
[0240] In an embodiment of the disclosure, the pre-defined scalable neural network model may be defined by a number of core blocks, wherein each core block comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained in an offline prior to a video encoding operation.
[0241] In an embodiment of the disclosure, the residual unit may handle a two dimensional (2D) convolutional filter operation and a rectified linear unit (ReLU) operation, wherein the 2D convolutional filter operation applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein the ReLU operation is performed after the 2D convolutional filter operation to introduce non-linearity in the source video frame.
[0242] In an embodiment of the disclosure, each core block may comprise at least one of: a series of 2D convolutional block and a series of 1D convolutional block with a normalization block and at least one of: a linear activation and a non-linear activation.
[0243] In an embodiment of the disclosure, the pre-defined scalable neural network model may be configured using a rate-distortion-complexity optimizer, wherein the rate-distortion-complexity optimizer determine which configuration will provide a least cost of encoding, and wherein the rate-distortion-complexity optimizer determines a correct balance between distortion caused due to quantization in the video encoder (450) and a bitrate to encode the reconstructed pixel block.
[0244] In an embodiment of the disclosure, a smaller number of the at least one core block may be selected for a reconstructed pixel block with a lower quantization parameter (QP), wherein the at least one pre-trained neural network parameter is modified based on the at least one residual unit, wherein the pre-trained neural network parameter determines a weightage significance of each pathway within each selected core block, wherein the pre-trained neural network parameter is recalibrated when a configuration of the pre-defined scalable neural network model is changed post-training.
[0245] In an embodiment of the disclosure, the encoded video bit-stream is received by: generating the reconstructed video frame corresponding to a source video frame; configuring a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame; modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block; and predicting an optimized reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0246] In an embodiment of the disclosure, the method may include generating the encoded video bit-stream comprising at least one pre-defined core block, at least one residual unit within each selected core block and at least one pathway within the at least one selected core blocks.
[0247] In an embodiment of the disclosure, an electronic device (400), comprising: a processor (410); a memory (430); a video encoder (450); and a neural network-based in-loop filtering controller (470), coupled with the processor (410), the memory (430) and the video encoder (450), is provided. The processor (410) may be configured to: generate a reconstructed video frame corresponding to a source video frame within at least one core block; configure a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame; modify at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within the at least one core block; and generate an optimized reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0248] In an embodiment of the disclosure, an electronic device (400), comprising: a processor (410); a memory (430); a video decoder (460); and a neural network-based in-loop filtering controller (470), coupled with the processor (410), the memory (430) and the video decoder (460), is provided. The processor may be configured to: generate a reconstructed video frame corresponding to a source video frame; receive an encoded video bit-stream comprising at least one neural network configuration, wherein the at least one neural network configuration comprises selection of at least one core block, at least one residual unit within the at least one selected core block and at least one pathway within the at least one selected core block; generate a reconstructed video frame using the encoded video bitstream, wherein the reconstructed video frame is divided into a plurality of reconstructed pixel blocks; configure a pre-defined scalable neural network model for each reconstructed pixel block using the at least one neural network configuration; modify at least one pre-trained neural network parameter associated with the configured scalable neural network model based on the at least one pathways within the at least one selected core block; and generate an enhanced reconstructed video frame by feeding each pixel block in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.
[0249] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the embodiments as described herein.
Claims
1.A method for performing an in-loop filter processing by a video encoder (450) included in an electronic device (400), comprising:generating a reconstructed video frame corresponding to a source video frame;configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame;modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model; andgenerating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.2.The method of claim 1, wherein the configuring of the scalable neural network model comprising:selecting the at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block.3.The method of claim 2, further comprising:generating a bitstream including information for at least one of the at least one core block, at least one residual unit within each of the at least one core block or at least one pathway within the at least one core block.4.The method of any one of claims 1 to 3, wherein the generating the reconstructed video frame corresponding to the source video frame comprising:determining at least one content property associated with the source video frame and at least one encoder parameter used to generate the reconstructed video frame, wherein the at least one content property includes at least one of: edge information and texture information, wherein the at least one encoder parameter includes a quantization parameter.5.The method of claim 4, wherein the configured scalable neural network model for each reconstructed pixel block is determined by correlating the at least one content property and the at least one encoder parameter.6.The method of any one of claims 1 to 5, wherein the configured scalable neural network model is defined by a number of the least one core block, wherein each of the at least one core block comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained offline prior to a video encoding operation.7.The method of claim 6, wherein the residual unit includes a two dimensional (2D) convolutional layer and a rectified linear unit (ReLU) layer, wherein a 2D convolutional filter operation performed in the 2D convolution layer applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein a ReLU operation performed in the ReLU layer is applied after the 2D convolutional layer to introduce non-linearity into an output of the 2D convolution layer.8.The method of any one of claims 6 to 7, wherein each of the at least one core block comprises at least one of: a series of 2D convolutional block and a normalization block.9.The method of any one of claims 1 to 8, wherein the configured scalable neural network model is determined using a rate-distortion-complexity optimizer, wherein the rate-distortion-complexity optimizer determines at least one configuration of the pre-defined scalable neural network model that minimizes a cost function associated with encoding the reconstructed pixel block, and a balance between distortion caused by quantization in the video encoder (450) and a number of bits to encode the reconstructed pixel block.10.The method of any one of claims 1 to 9, wherein a smaller number of the at least one core block is selected for a reconstructed pixel block with a lower quantization parameter (QP), or a lower difference with the reconstructed pixel block and a source pixel block corresponding to the reconstructed pixel block.11.The method of any one of claims 2 to 10, wherein the at least one pre-trained neural network parameter is modified based on a number of the at least one residual unit within the configured scalable neural network model, wherein the pre-trained neural network parameter is normalized using ratio of a number of the at least one residual units within the pre-defined scalable neural network model to the number of the at least one residual unit within the configured scalable neural network.12.A method for performing an in-loop filter processing by a video decoder (460) included in an electronic device (400), comprising:obtaining a bitstream including information for selection of at least one core block, at least one residual unit within each of the at least one core block and at least one pathway within the at least one core block to configure a scalable neural network;generating a reconstructed video frame using the bitstream, wherein the reconstructed video frame includes a plurality of reconstructed pixel blocks;configuring the scalable neural network by using a pre-defined scalable neural network model for each of the plurality of the reconstructed pixel blocks;modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of the at least one pathway within the at least one core block; andgenerating an enhanced reconstructed video frame by feeding each of the plurality of the reconstructed pixel blocks in the reconstructed video frame to the configured scalable neural network model with the at least one modified pre-trained neural network parameter.13.The method of claim 12, wherein the configured scalable neural network model is determined by a number of core blocks, wherein each of the core blocks comprises at least one residual unit, wherein the pre-defined scalable neural network model is trained offline prior to a video decoding operation.14.The method of claim 13, wherein the residual unit includes a two dimensional (2D) convolutional layer and a rectified linear unit (ReLU) layer, wherein a 2D convolutional filter operation performed in the 2D convolution layer applies a set of trainable filters to a source video frame to extract a spatial feature of the source video frame, wherein the spatial feature comprises at least one of: an edge, a texture, and a pattern, wherein a ReLU operation performed in the ReLU layer is applied after the 2D convolutional layer to introduce non-linearity into an output of the 2D convolution layer.15.A method for transmitting a bitstream generated by an encoding method, the encoding method comprising:generating a reconstructed video frame corresponding to a source video frame;configuring a scalable neural network by using a pre-defined scalable neural network model for each reconstructed pixel block in the reconstructed video frame;modifying at least one pre-trained neural network parameter of the configured scalable neural network model based on a number of pathways within at least one core block in the configured scalable neural network model; andgenerating an enhanced reconstructed video frame by feeding each reconstructed pixel block in the reconstructed video frame into the configured scalable neural network model with the at least one modified pre-trained neural network parameter.