System and method for dynamically selecting artificial intelligence based in-loop filter in video codec

By dynamically selecting AI-based In-Loop filter models based on frame dependencies and wait-times, the method addresses high complexity and latency issues in video codecs, improving quality and efficiency in constrained hardware environments.

WO2026024134A1PCT designated stage Publication Date: 2026-01-29SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/011050
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-07-25
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

AI-based In-Loop filters in video codecs exhibit high computational complexity, leading to increased latency and suboptimal performance in current hardware environments, particularly in devices with constrained resources.

Method used

A method for dynamically selecting an AI-based In-Loop filter model based on frame dependencies and wait-times in the decoded picture buffer, applying higher complexity models to frames with longer wait-times to balance latency and quality while reducing bit requirements.

Benefits of technology

This approach enhances video quality at lower bitrates without additional latency, optimizing decoder performance and enabling efficient on-device applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011050_29012026_PF_FP_ABST
    Figure KR2025011050_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for dynamically selecting an AI ILF model by a video decoder is provided. The method may include obtaining a bit-stream including information for GOP of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure includes interdependency among the plurality of the video frames. The method may include determining a wait-time in a DPB for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The method may include generating a plurality of enhanced reconstructed video frames by enhancing the plurality of video frames based on the selected AI ILF model.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR DYNAMICALLY SELECTING ARTIFICIAL INTELLIGENCE BASED IN-LOOP FILTER IN VIDEO CODEC

[0001] The present disclosure relates to a field of video codec, and more particularly, relates to a method and a system for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model in a video codec. This application is based on and derives the benefit of Indian Provisional Application No. 202441056832 filed on 26th July 2024, the contents of which are incorporated herein by reference. Furthermore, the Complete Specification of the said Provisional Application No. 202441056832 have been successfully filed with the Indian Patent Office on 17th January 2025.

[0002] A video codec is a software or hardware tool that is used to compress and decompress digital video files. The term "codec" stands for compressor-decompressor, which refers to its primary function of reducing the file size for easier storage or transmission and then restoring it to a viewable format during playback. Video codecs play a crucial role in deciding how digital video content is delivered across different platforms, ranging from streaming services to video conferencing apps. Thus, the understanding of video codecs is essential for creators, consumers, and anyone working with video content.

[0003] As a next-generation codec, AI-based video codecs have emerged. However, due to the inherent complexity of AI technologies, the processing time of the codec may increase. For example, an AI-based in-loop filter has significantly higher complexity compared to other blocks within the encoder and decoder, and thus can substantially increase the decoder complexity. Accordingly, various methods have been proposed to reduce the decoder's latency.This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the invention. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended for determining the scope of the invention.

[0004] In an embodiment of the disclosure, a method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video decoder. The method may include obtaining a bit-stream including information for a group of pictures (GOP) structure of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure includes interdependency among the plurality of the video frames. The method may include determining a wait-time in a decoded picture buffer (DPB) for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The method may include generating a plurality of enhanced reconstructed video frames by enhancing the plurality of video frames based on the selected AI ILF model.

[0005] In an embodiment of the disclosure, a method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder may be provided. The method may include obtaining information for a group of pictures (GOP) structure of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure includes interdependency among the plurality of the video frames. The method may include determining a wait-time in a decoded picture buffer (DPB) for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time.

[0006] In an embodiment of the disclosure, a method for transmitting a bitstream generated by the method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder may be provided.

[0007] To further clarify the advantages and features of the present invention, a more particular description of the invention will be rendered by reference to specific embodiments thereof, which is illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the invention and are therefore not to be considered limiting its scope. The invention will be described and explained with additional specificity and detail with the accompanying drawings.

[0008] In an embodiment of the disclosure, a method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video decoder. The method may include obtaining a bit-stream including information for a group of pictures (GOP) structure of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure includes interdependency among the plurality of the video frames. The method may include determining a wait-time in a decoded picture buffer (DPB) for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The method may include generating a plurality of enhanced reconstructed video frames by enhancing the plurality of video frames based on the selected AI ILF model.

[0009] In an embodiment of the disclosure, a method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder may be provided. The method may include obtaining information for a group of pictures (GOP) structure of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure includes interdependency among the plurality of the video frames. The method may include determining a wait-time in a decoded picture buffer (DPB) for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time.

[0010] In an embodiment of the disclosure, a method for transmitting a bitstream generated by the method for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder may be provided.

[0011] The foregoing and other features of embodiments will become more apparent from the following detailed description of embodiments when read in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements.

[0012] Figure 1a illustrates an existing Versatile Video Coding (VVC) video decoder.

[0013] Figures 1b and 1c illustrate a decoder and an encoder.

[0014] Figure 1d illustrates an example of a GOP structure for encoding / decoding.

[0015] Figures 1e and 1f illustrate an example of CPU core and GPU core for decoding.

[0016] Figure 2 illustrates an environment for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model in a video codec, in accordance with an embodiment of the present disclosure.

[0017] Figures 3a and 3b illustrate a block diagram of the system for dynamically selecting the artificial intelligence (AI) based In-Loop filter (AI ILF) model in the video codec, in accordance with an embodiment of the present disclosure.

[0018] Figure 4 illustrates a block diagram of an operation performed by a video encoder ILF module, in accordance with an embodiment of the present disclosure.

[0019] Figure 5 illustrates a block diagram of an operation performed by a video decoder ILF module, in accordance with an embodiment of the present disclosure.

[0020] Figure 6 illustrates a flowchart depicting a method for dynamically selecting the AI ILF model in the video codec by the video decoder, in accordance with an embodiment of the present disclosure.

[0021] Figure 7 illustrates a flowchart depicting a method for dynamically selecting the AI ILF model in the video codec by the video encoder, in accordance with an embodiment of the present disclosure.

[0022] Figure 8 illustrates a flowchart depicting a method for dynamically selecting the AI ILF model in the video codec by at least one of the video decoder and the video encoder, in accordance with an embodiment of the present disclosure.

[0023] Figure 9a illustrates a use case of the system, in accordance with an embodiment of the present disclosure.

[0024] Figure 9b illustrates a use case of the system, in accordance with an embodiment of the present disclosure.

[0025] Figure 10 illustrates a use case of the system, in accordance with an embodiment of the present disclosure.

[0026] Figure 11 illustrates a use case of the system, in accordance with an embodiment of the present disclosure.

[0027] For the purpose of promoting an understanding of the principles of the present disclosure, reference will now be made to the various embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the present disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the present disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the present disclosure relates.

[0028] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the present disclosure and are not intended to be restrictive thereof.

[0029] Whether or not a certain feature or element was limited to being used only once, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element does not preclude there being none of that feature or element, unless otherwise specified by limiting language including, but not limited to, "there needs to be one or more..." or "one or more elements is required."

[0030] Reference is made herein to some "embodiments." It should be understood that an embodiment is an example of an implementation of any features and / or elements of the present disclosure. Some embodiments have been described for the purpose of explaining one or more of the potential ways in which the specific features and / or elements of the proposed disclosure fulfil the requirements of uniqueness, utility, and non-obviousness.

[0031] Use of the phrases and / or terms including, but not limited to, "a first embodiment," "a further embodiment," "an alternate embodiment," "one embodiment," "an embodiment," "multiple embodiments," "some embodiments," "other embodiments," "further embodiment", "furthermore embodiment", "additional embodiment" or other variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and / or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and / or elements may be described herein in the context of only a single embodiment, or in the context of more than one embodiment, or in the context of all embodiments, the features and / or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and / or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

[0032] Any particular and all details set forth herein are used in the context of some embodiments and therefore should not necessarily be taken as limiting factors to the proposed disclosure.

[0033] The terms "comprises," "comprising," or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by "comprises.. a" does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.

[0034] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings.

[0035] Figure 1a is an example diagram, illustrating an existing Versatile Video Coding (VVC) video decoder.

[0036] As shown in Figure 1a, a bitstream may be processed by various video decoder modules (e.g., entropy decoder, inverse quantization, inverse transform, prediction modules, etc.), collectively referred to as "Other video decoder modules." The output of such modules is a reconstructed picture (Rec), which undergoes a series of in-loop processing operations to improve the visual quality and enhance prediction efficiency for subsequent frames.

[0037] The in-loop processing module includes a sequence of filtering stages applied to the reconstructed picture before it is stored in a decoded picture buffer (DPB). The in-loop processing may include, in order, a Luma Mapping with Chroma Scaling (LMCS) module configured to apply non-linear mapping to luma components and chroma scaling, a Deblocking Filter (DBF or DBK) configured to reduce blocking artifacts along block boundaries, a Sample Adaptive Offset (SAO) module that adjusts sample values to compensate for edge distortions or quantization errors, an Adaptive Loop Filter (ALF) that applies a filter across the picture to improve subjective and objective quality, and a Cross-Component ALF (CC-ALF) that utilizes luma information to refine the filtering of chroma components.

[0038] In an embodiment of the disclosure, the output of the in-loop processing is provided to the Decoded Picture Buffer (DPB). The DPB stores decoded pictures that have completed in-loop processing and are used as reference pictures (Ref) for decoding subsequent inter-predicted frames.

[0039] Currently, in typical video codecs, data compression is performed by quantizing the transform-domain coefficients representing the residual data after prediction by the encoder. The quantizing results in a difference of reconstructed video frames from the source frames, thus, often leading to visual artifacts e.g. ringing, blockiness, texture-loss, etc. To mitigate this, a set of tools, together known as the "In-Loop Filtering (ILF)" model, is used to improve the quality of reconstruction.

[0040] Recently, many AI-based ILF models have been suggested to further improve the reconstructed video. Such models augment or replace one or more ILF models. However, the output video still contains significant pixel differences causing major visual artifacts. This limits the performance of the suggested AI-based ILF models due to the working / re-tuning of existing AI-based ILF tools models. Further, the computational complexity of the suggested AI-based ILF models is high and, thus, infeasible to implement on current / near-future devices. Further, a leading video standardization body specifies two operating points High: 477kMAC / pixel and Low: 17kMAC / pixel. The computational complexity is still high for any on-device implementation. Further, the current state-of-the-art AI ILF models are typically more complex than the existing decoder. Thus, the latency caused by such AI ILF models is not acceptable in current standards. Further, the operation performed by the encoder and decoder along with the disadvantages which are existing currently in the video codec are discussed in detail in the subsequent paragraphs.

[0041] Particularly, the design of typical video encoders / decoders allows different frames to have different wait-times in a decoder picture buffer (DPB) before being displayed. In the decoder, the bitstream is entropy decoded, further followed by inverse quantization and inverse transform to generate a reconstructed residual signal. Based on intra or inter mode, prediction is performed using frames from DPB. The DPB stores the previously reconstructed frames which are to be used by current and future frames for prediction. Thereafter, In Loop Filter (e.g., Traditional and AI based) is used to decode and generate frames known as reconstructed pictures. Particularly, DPB is used to store decoded pictures in both encoder and decoder. The decoded pictures are used as reference pictures in the encoder and reference pictures or display pictures in the decoder. Subsequent frames may refer to the decoded picture during motion estimation at the encoder and / or motion compensation at the decoder. Further, some frames stay in DBP for more time before the frames are used as reference or display than others.

[0042] Figures 1b and 1c illustrate a decoder and an encoder of the video codec according to an embodiment of the disclosure.

[0043] In an embodiment, the decoder may perform entropy decoding on a received compressed bitstream to reconstruct syntax elements and quantized transform coefficients. The decoder may apply inverse quantization to recover the transform-domain coefficients and may further apply inverse transform (e.g., IDCT or similar) to reconstruct the residual signal in the spatial domain.

[0044] To generate the predicted signal, the decoder may perform either intra prediction based on spatially adjacent pixels within the same frame, or inter prediction using reference pictures stored in the Decoded Picture Buffer (DPB). The residual signal may be added to the prediction signal to reconstruct the current frame. The reconstructed frame may then pass through a series of in-loop filters, including a deblocking filter (DBF) to reduce block boundary artifacts, a sample adaptive offset (SAO) filter to minimize quantization distortions, and an adaptive loop filter (ALF) to globally enhance perceptual quality.

[0045] In an embodiment of the disclosure, the decoder may further apply an AI-based in-loop filter (AILF) utilizing neural networks (e.g., CNNs, RNNs, Transformers, ViTs, GNNs) to refine the visual quality beyond traditional filtering stages. The final filtered frame may be stored in the DPB, serving as a reference pictures for decoding future frames and potentially for output display. Furthermore, either the AI-based in-loop filtering or In-Loop Filtering may be selectively performed. The decoder may determine whether AI-based in-loop filtering is to be applied based on information obtained from the bitstream.

[0046] In an embodiment, the encoder may begin by receiving an input video frame and generating a prediction using either intra prediction or inter prediction, based on reference frames previously stored in the DPB. The encoder may subtract the predicted signal from the input signal to produce a residual signal, and may apply a transform (e.g., DCT) to convert it to the frequency domain. This transformed residual may then be quantized, and the resulting coefficients may be entropy encoded to form a compressed bitstream.

[0047] To enable consistent reference picture reconstruction between the encoder and decoder, the encoder may also perform inverse quantization and inverse transform on the residual signal and apply the same in-loop filtering chain as used by the decoder. Specifically, the encoder may apply DBF, SAO, and ALF filters in sequence. Additionally, the encoder may apply the AI-based in-loop filter (AILF) to further improve the quality of the reconstructed frame. The encoder may generate a bitstream that includes information indicating whether the AI-based in-loop filtering is to be applied. This filtered frame may then be stored in the DPB, so it can be used as a reference for the prediction of subsequent frames during encoding.

[0048] Meanwhile, certain operations performed in the encoder or decoder may be omitted, not limited to the examples disclosed above. In addition, the encoder may be referred to as a video encoder, and the decoder may be referred to as a video decoder. And, the AILF may refereed to as an AI ILF(AI-ILF), and the filtered frame may be referred to as an enhanced frame or an enhanced reconstructed video frame.

[0049] Figure 1d illustrates an example of a group of pictures (GOP) structure for encoding / decoding.

[0050] Referring to Figure 1d, TIDs represent the Temporal IDs of each temporal layer. Frames are numbered from 0 to 8. As illustrated, at least one frame from a plurality of frames is encoded / decoded independent of other frames (e.g., Intra frames, 0) while other frames from the plurality of frames depend on one frame (e.g., P frame, 8) or more frames (e.g., B frames, 1 to 7). The encoded / decoded frames are kept in the DPB as some future frames may use them as the reference frame. Further, the display order refers to the sequence in which the frames have to be displayed to the user. The coding order refers to the sequence in which the frames are encoded / decoded. Thus, as illustrated in Figure 1d, the display order and encoding / decoding order may not be the same for a group of frames. Further, as illustrated, each frame spends a varying amount of time, known as the wait-time, in DPB depending upon the GOP structure. Furthermore, the traditional encoding / decoding steps are performed on a processor and a convolution neural network (CNN) based AI ILF models are applied using graphic processing units (GPUs) on a reconstructed frame to encode / decode the reconstructed frame.

[0051] Figures 1e and 1f illustrate an example of CPU core and GPU core for decoding.

[0052] In an embodiment of the disclosure, decoding of video frames may be performed by a CPU core, and subsequently, the reconstructed frames are transferred to a GPU core for AI ILF processing. To maximize parallelism of decoding and post-processing, the system may apply AI ILF in a pipelined manner across multiple GPU cores. However, due to inter-frame dependencies inherent in video coding structures (e.g., P / B frames referencing prior frames), decoding of a particular frame cannot begin until its reference frames are completely reconstructed and, in some cases, post-processed.

[0053] In such scenarios, using a high-complexity AI ILF model may introduce significant latency. For instance, if frame N requires filtering by a complex model and frame N+1 depends on frame N, the delay in completing AI ILF for frame N may postpone the decoding of frame N+1. This can hinder real-time decoding performance, particularly in constrained hardware environments. Using a low-complexity AI ILF model may mitigate such latency and enable continuous decoding without postponing the decoding of frame N+1. However, this may result in suboptimal quality enhancement compared to high-complexity models.

[0054] Therefore, the system may selectively apply an AI ILF model based on frame type, dependency structure, available hardware resources (e.g., number of GPU cores), or real-time performance constraints.

[0055] In an embodiment of the disclosure, due to the dependencies among the plurality of the video frames during encoding / decoding, preceding frames in encoding / decoding order have to be completely ready before the start of encoding / decoding of a frame. Thus, using a highly complex AI ILF model, on the frames before the encoding / decoding of the frames, causes the later frames to suffer latency as the previous frame may not be ready on time. Additionally, using a low complex AI ILF model does not cause latency, however, results in inferior quality improvement. In an embodiment of the disclosure below, information for a group of pictures (GOP) structure 502 of a plurality of video frames corresponding to a source video to be decoded may include interdependency of each of the plurality of the video frames and a sequence to be followed by each video frame, based on the interdependency, during decoding or encoding.

[0056] Thus, the method implementing AI ILF model using one fixed AI ILF tool (either the highly complex AI ILF model or the low complex AI ILF model) to perform filtering of the frames before decoding affects the enhancement of each frame. Particularly, if AI ILF model is highly complex, the quality of the reconstructed video is better, but the latency of generating that frame is also higher. Further, the low complex AI ILF model generates reconstructed frame with less latency, however by compromising on reconstructed picture quality. Thus, there is an inherent trade-off between model complexity and picture quality when the one fixed (either highly complex AI ILF or low complex AI ILF) AI ILF model is used to perform the filtering.

[0057] Hence, there is a need in the art for solutions that will overcome the above mentioned drawback(s), among others.

[0058] The system and method as disclosed in the present disclosure provide an artificial intelligence (AI) based In-Loop filter (AI ILF) model selection for a Hierarchical group of pictures (GOP) for parallel processing and enhanced output using frame wait-time in Decoder Picture Buffer (DPB). Further, the system and method as disclosed also ensure signaling of a frame wait-time to AI ILF model complexity & delta QP table and related signals from encoder to decoder. Particularly, the system and method exploit a wait-time of a frame in DPB to choose the complexity and latency of AI-ILF so that the average latency of a decoder / encoder is balanced. The system and method as disclosed apply higher complexity or latency AI-ILF on frames that have longer wait-time in DPB, before being used as reference pictures or displayed to a user. Further, the system and method as disclosed reduces the bits required to send certain frames which can afford to have more complex AI-ILF applied at the end. This increases the visual quality that can be achieved at a lower bitrate and without any additional latency / delay.

[0059] The system and method as disclosed are a unique approach within the in-loop filtering domain from the perspective of decoder / encoder implementation. In any video codec standardization, generally its proposed hardware implementation plays a major role. The system and method as disclosed ensures improved compression performance as measured in BD-Rate as compared with current approaches. Further, the system and the method as disclosed ensure that the framework is useful for several on-device applications that use AI models for AI ILF as the simplification suggested in this method enables the decoder optimization.

[0060] For the sake of clarity, the first digit of a reference numeral of each component of the present disclosure is indicative of the Figure number, in which the corresponding component is shown. For example, reference numerals starting with digit "1" are shown at least in Figure 1. Similarly, reference numerals starting with digit "2" are shown at least in Figure 2.

[0061] Figure 2 illustrates an environment 200 for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model in a video codec, in accordance with an embodiment of the present disclosure.

[0062] In an embodiment, the system 204 may be in communication with at least one user equipment 202 (referred to herein interchangeably as a UE 202), without departing from the scope of the present disclosure.

[0063] In an embodiment, the system 204 may be deployed in the UE 202, without departing from the scope of the present disclosure. the UE 202 may be a smartphone, a tablet, a camera, etc., without departing from the scope of the present disclosure. In an embodiment, the UE 202 may be any electronic device compatible with streaming the media / videos, without departing from the scope of the present disclosure. In other words, the UE 202 may comprise the system 204 according to the present disclosure.

[0064] Typically, various streaming platforms stream the media / videos for the consumption of the user. Further, the user generally consumes the media / video on the UE 202. Thus, to ensure an optimum streaming of the media / video on the UE 202, the system 204 is disclosed in the present disclosure. The system 204 dynamically selects at least one AI ILF model from a plurality of AI ILF models based on one or more inputs, ensuring an efficient reconstruction of at least one of a plurality of video frames associated with a source video, thus enabling the optimum streaming of the media / video on the UE 202 while reducing latency of the video.

[0065] Further, the constructional detail of the system 204 is explained in the subsequent paragraphs.

[0066] Figures 3a and 3b illustrate a block diagram 300 of the system 204 for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model in the video codec, in accordance with an embodiment of the present disclosure.

[0067] In an embodiment of the disclosure, the system 204 may include, but is not limited to, at least one of a video encoder 302, and a video decoder 304. The video encoder 302 may be defined as a software or hardware tool that compresses and converts a source video into a digital format that may be easily stored, transmitted, and decoded, without departing from the scope of the present disclosure. The video decoder 304 may be defined as a software / hardware tool that decodes compressed video data, without departing from the scope of the present disclosure. Further, in an embodiment of the disclosure, the system 204 is implemented as a standalone entity at a server / cloud architecture, the system 204 may be in communication with multiple devices to receive data from each of the multiple devices, and the details provided below with respect to the system 204 and the UE 202 are applicable for the system 204 and the multiple UEs as well. Furthermore, the video encoder 302 may include, but is not limited to, at least one processor (referred to here as a processor) 306, a memory 308, and at least one module 312, among other examples which are explained in detail in subsequent paragraphs. Similarly, the video decoder 304 may include, but is not limited to, at least one processor (referred to here as a processor) 316, a memory 318, and at least one module 320, among other examples which are explained in detail in subsequent paragraphs.

[0068] Additionally, the video encoder 302 may include an Input / Output (I / O) interface 338. Similarly, the video decoder 304 may include an Input / Output (I / O) interface 340.

[0069] In an embodiment of the disclosure, the processor 306 of the video encoder 302 may be operatively coupled to each of the I / O interface 338, the at least one module 312, and the memory 308. In an embodiment of the disclosure, the processor 306 may include a graphical processing unit (GPU) and / or an artificial intelligence engine (AIE). In an embodiment of the disclosure, the processor 306 may include at least one data processor for executing processes in a virtual storage area network. The processor 306 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. In an embodiment of the disclosure, the processor 306 may include a central processing unit (CPU), a graphics processing unit (GPU), or both. The processor 306 may be one or more general processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other now-known or later developed devices for analyzing and processing data. The processor 306 may execute a software program, such as code generated manually (i.e., programmed) to perform the desired operation.

[0070] In an embodiment of the disclosure, the processor 306 may be disposed in communication with one or more input / output (I / O) devices via the I / O interface 338. In some embodiments, the processor 306 may communicate with the UE 202 using the I / O interface 338. In an embodiment of the disclosure, the I / O interface 338 may be implemented within the UE 202. The I / O interface 338 may employ communication code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like. In an embodiment, the I / O interface 338 may enable input and output to and from the system 204 using suitable devices such as, but not limited to, display, keyboard, mouse, touch screen, microphone, speaker, and so forth.

[0071] Using the I / O interface 338, the video encoder 302 may communicate with one or more I / O devices, specifically, the UE 202, to which the video encoder 302 dynamically selects the AI ILF model in the video codec. For example, the input device may be an antenna, microphone, touch screen, touchpad, storage device, transceiver, video device / source, etc. The output devices may be a video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma Display Panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc.

[0072] In an embodiment of the disclosure, the processor 306 may be disposed in communication with a communication network via a network interface. In an embodiment of the disclosure, the network interface may be the I / O interface 338. The network interface may connect to the communication network to enable the connection of the video encoder 302 with the UE 202. The network interface may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc. The communication network may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. Using the network interface and the communication network, the video encoder 302 may communicate with other UEs. The network interface may employ connection protocols including, but not limited to, direct connect, Ethernet (e.g., twisted pair 10 / 100 / 1000 Base T), transmission control protocol / internet protocol (TCP / IP), token ring, IEEE 802.11a / b / g / n / x, etc.

[0073] In an embodiment of the disclosure, the database may be configured to store the information as required by the at least one module 312 and the processor 306 to perform one or more functions for dynamically selecting the AI ILF model in the video codec.

[0074] In an embodiment of the disclosure, the memory 308 may be communicatively coupled to the processor 306. The memory 308 may be configured to store data, and instructions executable by the processor to perform the one or more methods disclosed herein throughout the present disclosure. In one embodiment, the memory 308 may be provided within the UE 202. In an embodiment of the disclosure, the memory 308 may be provided within the video encoder 302 being remote from the UE 202. In an embodiment of the disclosure, the memory 308 may communicate with the processor 306 via a bus within the video encoder 302. In an embodiment of the disclosure, the memory 308 may be located remote from the processor 306 and may be in communication with the processor 306 via a network. The memory 308 may include, but is not limited to, a non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media including, but not limited to, random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media and the like.

[0075] In an embodiment of the disclosure, the memory 308 may include a cache or random-access memory for the processor 306. In alternative examples, the memory 308 is separate from the processor 306, such as a cache memory of a processor, the system memory, or other memory. The memory 308 may be an external storage device or database for storing data. The memory 308 may be operable to store instructions executable by the processor 306. The functions, acts, or tasks illustrated in the figures or described may be performed by the programmed processor for executing the instructions stored in the memory 308. The functions, acts, or tasks are independent of the particular type of instruction set, storage media, processor, or processing strategy and may be performed by software, hardware, integrated circuits, firmware, micro-code and the like, operating alone or in combination. Likewise, processing strategies may include multiprocessing, multitasking, parallel processing, and the like.

[0076] Meanwhile, the configurations of the video encoder and the video decoder are illustrated separately in the Figure 3b. However, the I / O interface 338 of the video encoder 302 may be the same as, or correspond to, the I / O interface 340 of the video decoder 304. Similarly, the processor(s) 306 of the video encoder 302 may be the same as, or correspond to, the processor(s) 316 of the video decoder 304, and the memory 308 of the video encoder may be the same as, or correspond to, the memory 318 of the video decoder.

[0077] In addition, the modules 312 of the video encoder and the modules 320 of the video decoder may be the same as, or correspond to, each other. The Video Encoder ILF Module 314 included in the encoder-side modules may be the same as, or correspond to, the Video Decoder ILF Module 322 included in the decoder-side modules. Furthermore, the Video Encoder ILF Module 314 may be included within the modules 312 of the video encoder 302, and the Video Decoder ILF Module 322 may be included within the modules 320 of the video decoder 304.

[0078] In an embodiment of the disclosure, the at least one module 312 may be included within the memory 308. The memory 308 may further include a database to store data. The at least one module 312 may include a set of instructions that may be executed to cause the video encoder 302, in particular, the processor 306 of the video encoder 302, to perform any one or more of the methods / processes disclosed herein. The at least one module 312 may be configured to perform the steps of the present disclosure using the data stored in the database. For instance, the at least one module 312 may be configured to perform the operation of video encoder according to an embodiment of the disclosure.

[0079] In an embodiment of the disclosure, the at least one module 312 may be a hardware unit which may be outside the memory 308. Further, the memory 308 may include an operating system for performing one or more tasks of the video encoder 302, as performed by a generic operating system.

[0080] In an embodiment of the disclosure, the at least one module 312 may include a video encoder ILF module 314. Further, the video encoder ILF module 314 may be in communication with the processor 306.

[0081] Further, the present disclosure contemplates a computer-readable medium that includes instructions or receives and executes instructions responsive to a propagated signal. Further, the instructions may be transmitted or received over the network via a communication port or interface or using a bus (not shown). The communication port or interface may be a part of the processor or may be a separate component. The communication port may be created in software or may be a physical connection in hardware.

[0082] The communication port may be configured to connect with a network, external media, the display, or any other components in the system, or combinations thereof. The connection with the network may be a physical connection, such as a wired Ethernet connection, or may be established wirelessly. Likewise, the additional connections with other components of the video encoder 302 may be physical or may be established wirelessly. The network may alternatively be directly connected to a bus. For the sake of brevity, the architecture and standard operations of the memory 308, the processor 306, and the I / O interface 338 are not discussed in detail.

[0083] Further, in an embodiment of the disclosure, the working of the video encoder 302 to dynamically select the AI ILF model in the video codec is explained in detail. The processor 306, in conjunction with the video encoder ILF module 314 may be configured to perform specific operations explained in subsequent paragraphs.

[0084] Further, the constructional details of the at least one processor 316, the memory 318, are same as the constructional details of the processor 306 and the memory 308 of the video encoder 302. Thus, the same has not been explained for the sake of brevity. Additionally, the at least one module 320 may include, but is not limited to, a video decoder ILF module 322. The processor 316, in conjunction with the video decoder ILF module 322 may be configured to perform specific operations explained in subsequent paragraphs.

[0085] In an embodiment of the disclosure, at least one of the video encoder ILF module 314 and the video decoder ILF module 322 may be configured to receive a group of pictures (GOP) structure of a plurality of video frames corresponding to the source video to be at least one of encoded and decoded. In an embodiment of the disclosure, the source video may be defined as an original, uncompressed video, without departing from the scope of the present disclosure. In an embodiment of the disclosure, each of the plurality of video frames may be defined as individual images / still pictures that form the video sequence, without departing from the scope of the present disclosure. The GOP structure indicates interdependency of each of the plurality of video frames and a sequence to be followed by each of the plurality of video frames, based on the interdependency, during at least one of the encoding and decoding.

[0086] In an embodiment of the disclosure, interdependency may be defined as a relationship between each of the plurality of video frames in a video sequence, without departing from the scope of the present disclosure. The at least one of the video encoder ILF module 314 and the video decoder ILF module 322 may be configured to determine a wait-time in a decoded picture buffer (DPB) for each of the plurality of video frames, based on the interdependency. In an embodiment of the disclosure, the wait-time may be defined as a time period in which each of the plurality of video frames (either encoded or decoded) may remain in the DPB, without departing from the scope of the present disclosure. The at least one of the video encoder ILF module 314 and the video decoder ILF module 322 may be configured to select, dynamically, at least one AI ILF model from a plurality of AI ILF models for at least one of the plurality of video frames, based on the determined wait-time. The at least one of the video encoder ILF module 314 and the video decoder ILF module 322 may be configured to select the at least one AI ILF such that the selected at least one AI ILF model enhances reconstruction of the at least one of the plurality of video frames. The at least one of the video encoder ILF module 314 and the video decoder ILF module 322 may be configured to generate a reconstructed at least one of the plurality of video frames based on the selected AI ILF model. In an embodiment of the disclosure, the reconstructed at least one of the plurality of video frames may be defined as a video frame that may be decoded and reconstructed from its compressed data during the video decoding process.

[0087] Further, the detailed operation performed by the at least one of the video encoder ILF module 314 and the video decoder ILF module 322 are explained in subsequent paragraphs in conjunction with Figure 4 and Figure 5.

[0088] Figure 4 illustrates a block diagram of an operation 400 performed by the video encoder ILF module 314, in accordance with an embodiment of the present disclosure.

[0089] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to receive or obtain the GOP structure 402 of the plurality of video frames corresponding to the source video to be encoded. The GOP structure 402 may indicate the interdependency of each of the plurality of video frames and the sequence to be followed by each of the plurality of video frames, based on the interdependency, during encoding. For example, some frames are encoded independently (Intra frames, 0) while others depend on 1 frame (P frame, 8) or more frames (B frames, 1 to 7). For example, referring to Figure 1d, the relationship among the plurality of video frames may be referred to as interdependency ― for instance, frame 3 may depend on frame 2 or frame 4.

[0090] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to generate information for the GOP structure includes interdependency among the plurality of the video frames.

[0091] In an embodiment of the disclosure, the wait-time associated with a frame from among the plurality of the videos may be calculated based on a difference between the order in which the frame is decoded and the order in which the frame is to be displayed, or based on the time difference between when the frame is decoded and when the frame is scheduled for display. The video encoder ILF module 314 may be configured to determine the wait-time for each of the plurality of the video frames based on the display sequence of the plurality of video frames. And, a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence. Meanwhile, the first video frame and the first video frame may be included in the plurality of the video frame.

[0092] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to determine the wait-time 404 in the DPB 406 for each of the plurality of video frames, based on the interdependency. In such an embodiment, the video encoder ILF module 314 may be configured to perform the encoding on each of the plurality of video frames, based on the interdependency. The video encoder ILF module 314 may be configured to determine a display sequence of the plurality of video frames, in response to performing the encoding of each of the plurality of video frames. In an embodiment, the display sequence may be a sequence in which each of the plurality of video frames has to be displayed to the user, without departing from the scope of the present disclosure. In such an embodiment, the display sequence may be different from the sequence to be followed by each video frame during encoding. In an example, the number of the plurality of video frames may be 9. Then, the display sequence may be frame 0, 1, 2, ..., 8. However, the sequence followed for encoding may be frame 0, 8, 2, 3, 6, 5, 7, etc.

[0093] The video encoder ILF module 314 may be configured to determine the wait-time 404 based on the display sequence of the plurality of video frames, where at least one of the plurality of video frames having an immediate display sequence has a shorter wait-time and another video frame from the plurality of video frames having later display sequence has a longer wait-time. In an example, the number of the plurality of video frames may be 9. Further, the display sequence may be frame 0, 1, 2, ..., 8. However, the sequence followed for encoding may be frame 0, 8, 2, 3, 6, 5, 7, etc, indicating interdependency. So, here, frame 8, as encoded earlier, has the longer wait-time, and frame 0 has the shorter wait time.

[0094] Meanwhile, in an embodiment of the disclosure, a wait-time of a frame may be referred to as a shorter wait-time if it is shorter than an average or median wait-time of a plurality of frames in a sequence. Conversely, a wait-time of a frame may be referred to as a longer wait-time if it is longer than the average or median wait-time of the plurality of frames in the sequence.

[0095] In an embodiment of the disclosure, a quantization parameter applied to the frame can be derived based on table information from at least one of sequence header, or frame header of plurality of video frames. The table may be predefined such that both the quantization parameter (or the differential quantization parameter) and the AI ILF model are determined according to the wait-time.

[0096] Meanwhile, the table may be preset and may be information shared between the video encoder and video decoder, or the information within the table may be transmitted via the bitstream.

[0097] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select a first quantization parameter for a first video frame from among the plurality of video frames. The video decoder ILF module 322 may be configured to select a second quantization parameter for a second video frame from among the plurality of video frames. when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter may be higher than the second quantization parameter.

[0098] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to select, a higher quantization parameter (QP1,...,QPn), for the at least one of the plurality of video frames. In an embodiment, the higher quantization parameter refers to values or settings that majorly reduce the precision of the plurality of video frames, during encoding. The video encoder ILF module 314 may be configured to select the higher quantization parameter (QP1,...,QPn), when the determined wait-time 404 of the at least one of the plurality of video frames indicates a longer wait-time. In an embodiment, the video encoder ILF module may be configured to select, a lower quantization parameter (QP1,...,QPn), for the at least one of the plurality of video frames. In an embodiment, the lower quantization parameter refers to values or settings that slightly reduce the precision of the plurality of video frames, during encoding. The video encoder ILF module may be configured to select the lower quantization parameter (QP1,...,QPn), when the determined wait-time 404 of the at least one of the plurality of video frames indicates a shorter wait-time. This configuration results in a reduced bitrate and thus increased compression ratio may be achieved.

[0099] Meanwhile, in an embodiment of the disclosure, the quantization parameter of the video frame may be referred to as a lower quantization parameter if it is lower than an average or median quantization parameter (or, differential quantization parameter or delta quantization parameter) of a plurality of frames in a sequence. Conversely, the quantization parameter of the video frame may be referred to as a higher quantization parameter if it is higher than the average or median quantization parameter of the plurality of frames in the sequence.

[0100] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to select an average quantization parameter for the at least one of the plurality of video frames, when the determined wait-time 404 of the at least one of the plurality of video frames indicates at least one of the longer wait-time and the shorter wait-time.

[0101] In an embodiment of the disclosure, the quantization parameter associated for the at least one of the plurality of video frames may be derived from a header of the at least one of the plurality of video frames. In an embodiment, the header may contain information and meta data about the plurality of frames, without departing from the scope of the present disclosure.

[0102] In an embodiment of the disclosure, the at least one of the plurality of video frames undergoes an inverse quantization operation and inverse transform operation 408 to form a reconstructed frame 410 corresponding to the at least one of the plurality of video frames. In an embodiment, the inverse quantization operation may be performed on the at least one of the plurality of video frames as an attempt to restore the reconstructed frame 410, as formed, in its original scale. In an embodiment, the inverse transform operation may be performed on the reconstructed frame 410 to assess the compression and optimization of the reconstructed frame 410. This operation ensures an accurate assessment of the effect of the quantization and transformation on predicted data associated with the at least one of the plurality of video frames. This operation results in efficient compression and bit rate control of each of the plurality of video frames while maintaining effective visual quality. This results in an optimized encoding, and error checking by the encoder.

[0103] Meanwhile, the video encoder can determine the quantization parameter or AI ILF corresponding to the wait-time based on the table, as the quantization parameter or AI ILF of the video frame.

[0104] In an embodiment of the disclosure, an in-loop filter 412 may be applied on the reconstructed frame 410 to improve the visual quality of the reconstructed frame 410 and also reduce compression artifacts, if arise. In an example, the in-loop filter includes a deblocking filter, a sample adaptive offset (SAO) filter, adaptive loop filter (ALF). Meanwhile, a frame to which ALF or AI ILF is applied can be referred to as an enhanced reconstructed video frame. And, the "enhance" or "enhancement" means performing ALF or ILF.

[0105] In an embodiment of the disclosure the video encoder ILF module 314 may be configured to select, dynamically, at least one AI ILF model 414 from the plurality of AI ILF models for at least one of the plurality of video frames. Particularly, the video encoder ILF module 314 may be configured to select, dynamically, the at least one AI ILF model 414 for the reconstructed frame 410, without departing from the scope of the present disclosure. The video encoder ILF module 314 may be configured to select, dynamically, the at least one AI ILF model 414 based on the determined wait-time 404. The video encoder ILF module 314 may be configured to select, dynamically, the at least one AI ILF model 414 such that the selected at least one AI ILF model enhances reconstruction of the at least one of the plurality of video frames / the reconstructed frame 410.

[0106] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to select, a first AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames. The video encoder ILF module 314 may be configured to select the first AI ILF model when the determined wait-time 404 for the at least one of the plurality of video frames indicates the longer wait-time. Further, the first AI ILF model may indicate a complex AI ILF model. The selection of the complex AI ILF model provides leverage to the video encoder ILF module 314 to compress the plurality of video frames having the longer wait-time with reduced bit-rate / quality, where the quality of the plurality of video frames may be further retained by the complex AI ILF, thus reducing latency. Further, the video encoder ILF module 314 may be configured to select, a second AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames. The video encoder ILF module 314 may be configured to select the second AI ILF model when the determined wait-time 404 for the at least one of the plurality of video frames indicates the shorter wait-time. The second AI ILF model may indicate a simple AI ILF model as compared to the first AI ILF model. The selection of the simple AI ILF model ensures the efficient visual quality of the plurality of frames having the shorter wait-time, thus, providing an optimum visual experience to the user. Further, in an embodiment, the video encoder ILF module 314 may be configured to dynamically select AI ILF model for each video frame which may be present on the same layer of the GOP structure 502. In an example, at Temporal ID (TID) 2, the second AI ILF model may be applied to frame 2, and similarly, the first AI ILF model may be applied to frame 6 for further processing.

[0107] In an embodiment of the disclosure, after selecting the at least one AI ILF 414 model (interchangeably referred to here as the AI ILF 414) for at least one of the plurality of video frames, the at least one of the plurality of video frames with the selected AI ILF 414 may again be saved in the DPB 406 for further processing, without departing from the scope of the present disclosure. Thus, the at least one of the plurality of video frames with the selected AI ILF model 414, as saved in the DPB 406, may be used as a reference frame in motion compensation for inter-frame prediction for compressing each of the plurality of video frames. In an example, the motion compensation enables predicting parts of at least one video frame based on the reference frame. The use of the selected AI ILF model 414 results in smoother transitions and reduced artifacts in the predicted frames. This configuration ensures accurate prediction, reduces residual, improves overall compression, and also results in better decoding by the video decoder 304. Additionally, the video encoder ILF module 302 may be configured to transmit the selected AI ILF 414 via encoded bit stream to the video decoder 304, without departing from the scope of the present disclosure.

[0108] In an embodiment of the disclosure, when the video encoder ILF module 302 may be configured to perform the operation as discussed above, in that case, thereafter, the video decoder 304 decodes the at least one of the plurality of video frames having the selected AI ILF model 414, accordingly such that the video may be displayed on the UE 202 in an enhanced manner while removing any latency.

[0109] Figure 5 illustrates a block diagram of an operation 500 performed by the video decoder ILF module 322, in accordance with an embodiment of the present disclosure.

[0110] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to receive or obtain a bit-stream. The bit-stream may include information for the GOP structure 502 of the plurality of video frames corresponding to the source video to be decoded. The information for the GOP structure 502 includes interdependency among the plurality of the video frames. For convenience of explanation, the information for the GOP structure 502 of the plurality of video frames may be referred to as the GOP structure. The GOP structure 502 may indicate the interdependency of each of the plurality of video frames and the sequence to be followed by each of the plurality of video frames, based on the interdependency, during decoding.

[0111] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to determine the wait-time 504 in the DPB 506 for each of the plurality of video frames, based on information for the GOP structure 502 (e.g., the interdependency). In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to perform the decoding on each of the plurality of video frames, based on the interdependency. The video decoder ILF module 322 may be configured to determine a display sequence of the plurality of video frames, in response to performing the decoding of each of the plurality of video frames.

[0112] In an embodiment of the disclosure, the wait-time associated with a frame from among the plurality of the videos may be calculated based on a difference between the order in which the frame is decoded and the order in which the frame is to be displayed, or based on the time difference between when the frame is decoded and when the frame is scheduled for display. The video decoder ILF module 322 may be configured to determine the wait-time for each of the plurality of the video frames based on the display sequence of the plurality of video frames. And, a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence. Meanwhile, the first video frame and the first video frame may be included in the plurality of the video frame.

[0113] In an embodiment of the disclosure, the display sequence may be different from the sequence to be followed by each video frame during decoding. In an example, the number of the plurality of video frames may be 9. Further, the display sequence may be frame 0, 1, 2,...,8. However, the sequence for decoding may be frame 0, 8, 2, 3, 6, 5, 7, etc. The video decoder ILF module 322 may be configured to determine the wait-time 504 based on the display sequence of the plurality of video frames, where at least one of the plurality of video frames having an immediate display sequence has a shorter wait-time and another video frame from the plurality of video frames having later display sequence has a longer wait-time. In an example, the number of the plurality of video frames may be 9. Further, the display sequence may be frame 0, 1, 2,...,8. However, the sequence for decoding may be frame 0, 8, 4, 2, 6, 1, 3, 5, 7, etc, indicating interdependency. So, here, frame 8, as decoded earlier, has the longer wait-time, and the frame 0 has the shorter wait time. For example, the below table shows an example of display order, decoding order, and wait time for a frame in DPB for a GOP size of 8. Frames spend varying amount of time in DPB depending upon the GOP structure. This wait time can be exploited for selecting AI models of varying latency and complexity.

[0114] Display orderDecoding orderWait-time010154240352430550642751820

[0115] Meanwhile, we may assume frames can be decoded in parallel if they don't have any dependency. Wait-time can be optimized by exploiting relationship between AI models and need to the frame. By applying higher complexity AI ILF model on the higher wait-time frames, better quality enhancement can be achieved. By applying low complexity AI ILF model on the lower wait-time frame, the decoder latency may reduce. For example, as some frames (e.g., higher wait-time frames) can afford to have a more complex AI ILF model applied, and to apply a higher QP resulting in more degradation. The degraded frame can be improved by more complex AI ILF. This will result in reduced bitrate and thus increased compression ratio can be achieved.Meanwhile, in an embodiment of the disclosure, a wait-time of a frame may be referred to as a shorter wait-time if it is shorter than an average or median wait-time of a plurality of frames in a sequence. Conversely, a wait-time of a frame may be referred to as a longer wait-time if it is longer than the average or median wait-time of the plurality of frames in the sequence.

[0116] In an embodiment of the disclosure, the at least one of the plurality of video frames undergoes an inverse quantization operation and inverse transform operation 508 to form each of the at least one of the plurality of reconstructed frames 510 corresponding to each of the at least one of the plurality of video frames, by restoring a lossy compression performed during the encoding, ensuring accurate reconstruction of the at least one of the plurality of video frames and thus, display as intended.

[0117] In an embodiment of the disclosure, an in-loop filter 512 may be applied on the reconstructed frame 510 to improve the visual quality of the reconstructed frame 510 and also reduce compression artifacts. In an example, the in-loop filter includes a deblocking filter, a sample adaptive offset (SAO) filter, adaptive loop filter (ALF).

[0118] In an embodiment of the disclosure, the bit-stream may include information on whether to dynamically select AI ILF for the plurality of video frames. Information on whether to dynamically select AI ILF for the plurality of video frames may be derived from at least one of a slice header, a sequence header, or a frame header of plurality of video frames. The video decoder ILF module 322 may be configured to obtain information on whether to dynamically select AI ILF for the plurality of video frames.

[0119] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select, dynamically, at least one AI ILF model 514 (interchangeably referred to here as the AI ILF 514) from among the plurality of AI ILF models for the at least one of the plurality of video frames, particularly, the reconstructed frame 510. The video decoder ILF module 322 may be configured to select, dynamically, the at least one AI ILF model 514 based on the determined wait-time. The video decoder ILF module 322 may be configured to select, dynamically, the at least one AI ILF model 514 such that the selected at least one AI ILF model 514 enhances reconstruction of the at least one of the plurality of video frames / reconstructed frame 510.

[0120] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The video decoder ILF module 322 may be configured to select a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames. The video decoder ILF module 322 may be configured to select a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of the video frames. When a determined wait-time for the first video frame is longer than a determined wait-time for the second video frames, the first AI ILF model may be more complex than the second AI ILF model.

[0121] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select, a first AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames. The video decoder ILF module 322 may be configured to select the first AI ILF model when the determined wait-time 504 for the at least one of the plurality of video frames indicates the longer wait-time. Further, the first AI ILF model may indicate a complex AI ILF model. The selection of a complex AI ILF model provides leverage to the video encoder 302 to compress the plurality of video frames having the longer wait-time with reduced bit-rate / quality, where the quality of the plurality of video frames may be retained by the complex AI ILF, thus reducing latency. Further, the video decoder ILF module 322 may be configured to select, a second AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames. The video decoder ILF module 322 may be configured to select the second AI ILF model when the determined wait-time 504 for the at least one of the plurality of video frames indicates a shorter wait-time. The second AI ILF model may indicate a simple AI ILF model as compared to the first AI ILF model. The selection of the simple AI ILF model ensures the efficient visual quality of the plurality of frames having the shorter-wait time, thus, providing an optimum visual experience to the user. Further, in an embodiment, the video decoder ILF module 322 may be configured to dynamically select AI ILF for each video frame which may be present on the same layer of the GOP structure 502.

[0122] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to generate a plurality of enhanced reconstructed video frames by performing AI ILF the plurality of the reconstructed frame 510 based on the selected AI ILF model 514, as intended. Thereafter, the plurality of generated enhanced reconstructed video frames may be stored in the DPB 506, where the DPB 506 ensures smooth, accurate, and high-quality video decoding and playback.

[0123] In an embodiment of the disclosure, a quantization parameter applied to the frame can be derived based on table information from at least one of sequence header, or frame header of plurality of video frames. The table may be predefined such that both the quantization parameter (or the differential quantization parameter) and the AI ILF model are determined according to the wait-time.

[0124] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select a first quantization parameter for a first video frame from among the plurality of video frames. The video decoder ILF module 322 may be configured to select a second quantization parameter for a second video frame from among the plurality of video frames. when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter may be higher than the second quantization parameter.

[0125] In this case, it is not necessary to obtain the QP or the QP difference separately, thereby allowing a reduction in the amount of data to be transmitted.

[0126] Figure 6 illustrates a flowchart depicting a method 600 for dynamically selecting the AI ILF model 514 in the video codec by the video decoder 304, in accordance with an embodiment of the present disclosure. The method 600 may include a series of operations shown at step 602 through step 608 of Figure 6. The method 600 may be performed by the video decoder 304 of the system 204 in conjunction with the video decoder ILF module 322, the details of which are explained in conjunction with Figure 5, and the same are not repeated here for the sake of brevity in the present disclosure. The method 600 begins at step 602.

[0127] At step 602, the method 600 may include receiving or obtaining the bit-stream comprises the group of pictures (GOP) structure 502 of the plurality of video frames corresponding to the source video to be decoded. The GOP structure 502 indicates interdependency of each of the plurality of video frames and the sequence to be followed by each video frame, based on the interdependency, during decoding. The bit-stream may include information for the GOP structure 502 of the plurality of video frames corresponding to the source video to be decoded. The information for the GOP structure 502 includes interdependency among the plurality of the video frames. For convenience of explanation, the information for the GOP structure 502 of the plurality of video frames may be referred to as the GOP structure.

[0128] At step 604, the method 600 may include determining the wait-time 504 in the decoded picture buffer (DPB) 506 for each of the plurality of video frames, based on the interdependency. The method 600 includes performing the decoding on each of the plurality of video frames, based on the interdependency. The method 600 includes determining the display sequence of the plurality of video frames, in response to performing the decoding of each of the plurality of video frames. The method 600 includes determining the wait-time 504 based on the display sequence of the plurality of video frames, where at least one of the plurality of video frame having immediate display sequence has the shorter wait-time and another video frame from the plurality of video frames having later display sequence has a longer wait-time.

[0129] In an embodiment of the disclosure, the wait-time associated with a frame from among the plurality of the videos may be calculated based on a difference between the order in which the frame is decoded and the order in which the frame is to be displayed, or based on the time difference between when the frame is decoded and when the frame is scheduled for display. The video decoder ILF module 322 may be configured to determine the wait-time for each of the plurality of the video frames based on the display sequence of the plurality of video frames. And, a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence. Meanwhile, the first video frame and the first video frame may be included in the plurality of the video frame.

[0130] At step 606, the method 600 may include selecting, dynamically, the at least one AI ILF model 514 from the plurality of AI ILF models for at least one of the plurality of video frames, based on the determined wait-time 504, such that the selected at least one AI ILF model 514 enhances reconstruction of the at least one of the plurality of video frames. The method 600 includes selecting, the first AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames, when the determined wait-time for the at least one of the plurality of video frames indicates the longer wait-time The first AI ILF model indicates the complex AI ILF model. The method 600 includes selecting, the second AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames, when the determined wait-time for the at least one of the plurality of video frames indicates the shorter wait-time. The second AI ILF model indicates the simple AI ILF model as compared to the first AI ILF model.

[0131] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The video decoder ILF module 322 may be configured to select a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames. The video decoder ILF module 322 may be configured to select a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of the video frames. When a determined wait-time for the first video frame is longer than a determined wait-time for the second video frames, the first AI ILF model may be more complex than the second AI ILF model.

[0132] At step 608, the method 600 includes generating the enhanced reconstructed at least one of the plurality of video frames based on the selected AI ILF model. In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to generate a plurality of the enhanced reconstructed video frames by performing AI ILF the plurality of the reconstructed frame 510 based on the selected AI ILF model 514.

[0133] Figure 7 illustrates a flowchart depicting a method 700 for dynamically selecting the AI ILF model 414 in the video codec by the video encoder 302, in accordance with an embodiment of the present disclosure. The method 700 may include a series of operations shown at step 702 through step 706 of Figure 7. The method 700 may be performed by the video encoder 302 of the system 204 in conjunction with the video encoder ILF module 314, the details of which are explained in conjunction with Figure 4, and the same are not repeated here for the sake of brevity in the present disclosure. The method 700 begins at step 702.

[0134] At step 702, the method 700 may include receiving the group of pictures (GOP) structure 402 of the plurality of video frames corresponding to the source video to be encoded. The GOP structure 402 indicates interdependency of each of the plurality of video frames and the sequence to be followed by each video frame, based on the interdependency, during encoding. In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to generate information for the GOP structure includes interdependency among the plurality of the video frames.

[0135] At step 704, the method 700 may include determining the wait-time 404 in the decoded picture buffer (DPB) 406 for each of the plurality of video frames, based on the interdependency. The method 700 may include performing the encoding on each of the plurality of video frames, based on the interdependency. The method 700 includes determining the display sequence of the plurality of video frames, in response to performing the encoding of each of the plurality of video frames. The method 700 includes determining the wait-time 404 based on the display sequence of the plurality of video frames, where at least one of the plurality of video frame having immediate display sequence has the shorter wait-time and another video frame from the plurality of video frames having later display sequence has the longer wait-time.

[0136] In an embodiment of the disclosure, the wait-time associated with a frame from among the plurality of the videos may be calculated based on a difference between the order in which the frame is decoded and the order in which the frame is to be displayed, or based on the time difference between when the frame is decoded and when the frame is scheduled for display. The video encoder ILF module 314 may be configured to determine the wait-time for each of the plurality of the video frames based on the display sequence of the plurality of video frames. And, a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence. Meanwhile, the first video frame and the first video frame may be included in the plurality of the video frame. d

[0137] The method 700 may include selecting, the higher quantization parameter (QP1,...,QPn) for the at least one of the plurality of video frames, when the determined wait-time 404 of the at least one of the plurality of video frames indicates the longer wait-time. The method 700 includes selecting the lower quantization parameter (QP1,...,QPn) for the at least one of the plurality of video frames, when the determined wait-time 404 of the at least one of the plurality of video frames indicates the shorter wait-time.

[0138] In an embodiment of the disclosure, a quantization parameter applied to the frame can be derived based on table information from at least one of sequence header, or frame header of plurality of video frames. The table may be predefined such that both the quantization parameter (or the differential quantization parameter or delta QP) and the AI ILF model are determined according to the wait-time. Meanwhile, the table may be preset and may be information shared between the video encoder and video decoder, or the information within the table may be transmitted via the bitstream. Also, The QP information and AI ILF model according to the wait-time are The longer the wait-time, the larger the delta qp, and the more complex the AI ILF model can be.

[0139] In an embodiment of the disclosure, the video decoder ILF module 322 may be configured to select a first quantization parameter for a first video frame from among the plurality of video frames. The video decoder ILF module 322 may be configured to select a second quantization parameter for a second video frame from among the plurality of video frames. when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter may be higher than the second quantization parameter.

[0140] The method 700 may includes selecting, the average quantization parameter for the at least one of the plurality of video frames, when the determined wait-time 404 of the at least one of the plurality of video frames indicates at least one of the longer wait-time and the shorter wait-time.

[0141] The method 700 may include the quantization parameter associated for the at least one of the plurality of video frames is derived from the header of the at least one of the plurality of video frames. Meanwhile, in an embodiment of the disclosure, the quantization parameter of the video frame may be referred to as a lower quantization parameter if it is lower than an average or median quantization parameter (or, differential quantization parameter or delta quantization parameter) of a plurality of frames in a sequence. Conversely, the quantization parameter of the video frame may be referred to as a higher quantization parameter if it is higher than the average or median quantization parameter of the plurality of frames in the sequence.

[0142] At step 706, the method 700 may include selecting, dynamically, at least one AI ILF model 414 from the plurality of AI ILF models for at least one of the plurality of video frames, based on the determined wait-time 404, such that the selected at least one AI ILF model enhances reconstruction of the at least one of the plurality of video frames. Thereafter, the method 700 includes transmitting the selected AI ILF model via the encoded bit stream to the video decoder 304.

[0143] The method 700 may include selecting, the first AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames, when the determined wait-time 404 for the at least one of the plurality of video frames indicates the longer wait-time. The first AI ILF model indicates the complex AI ILF model. The method 700 may include selecting, the second AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames, when the determined wait-time 404 for the at least one of the plurality of video frames indicates the shorter wait-time. The second AI ILF model indicates the simple AI ILF model as compared to the first AI ILF model.

[0144] In an embodiment of the disclosure, the video encoder ILF module 314 may be configured to select an AI ILF model from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time. The video encoder ILF module 314 may be configured to select a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames. The video encoder ILF module 314 may be configured to select a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of the video frames. When a determined wait-time for the first video frame is longer than a determined wait-time for the second video frames, the first AI ILF model may be more complex than the second AI ILF model.

[0145] Figure 8 illustrates a flowchart depicting a method 800 for dynamically selecting the AI ILF model in the video codec by the at least one of the video encoder 302 and the video decoder 304, in accordance with an embodiment of the present disclosure. The method 800 may include a series of operations shown at step 802 through step 808 of Figure 8. The method 800 may be performed by the at least one of the video encoder 302 and the video decoder 304 of the system 204 in conjunction with the at least one of the video decoder ILF module 314 and the video decoder ILF module 322, the details of which are explained and the same are not repeated here for the sake of brevity in the present disclosure. The method 800 begins at step 802.

[0146] At step 802, the method 800 may include receiving the group of pictures (GOP) structure of the plurality of video frames corresponding to the source video to be at least one of encoded and decoded. The GOP structure indicates interdependency of each of the plurality of video frames and the sequence to be followed by each video frame, based on the interdependency, during at least one of the encoding and decoding.

[0147] At step 804, the method 800 may include determining the wait-time in the decoded picture buffer (DPB) for each of the plurality of video frames, based on the interdependency.

[0148] At step 806, the method 800 may include selecting, dynamically, at least one AI ILF model from the plurality of AI ILF models for at least one of the plurality of video frames, based on the determined wait-time, such that the selected at least one AI ILF model enhances reconstruction of the at least one of the plurality of video frames.

[0149] At step 808, the method 800 may include generating the reconstructed at least one of the plurality of video frames based on the selected AI ILF model.

[0150] Figure 9a illustrates a use case of the system 204, in accordance with an embodiment of the present disclosure.

[0151] In an embodiment of the disclosure, different AI ILF models (with different complexities and latency) 902 in the video decoder 304 results in different decoding time per frame. In an embodiment of the disclosure, different AI ILF models (with different complexities and latency) 902 in the video encoder 302 results in different encoding time per frame. Including the time to copy (or transmit) between CPU / GPU, some AI ILF models can be applied don't increase the decoder complexity much, while on the other hand some AI ILF models increase the latency by 1 or more frame.

[0152] Meanwhile, AI ILF or In-Loop Filtering may be performed on the GPU, and other decoding processes may be performed using the CPU. For example, a base video decoder may operate in a standalone mode without GPU for low-complexity processing, or in combination with GPU modules implementing Very Low Operation Processing (VLOP'), Low Operation Processing (LOP'), or High Operation Processing (HOP').

[0153] Figure 9b illustrates a use case of the system 204, in accordance with an embodiment of the present disclosure.

[0154] In an embodiment of the disclosure, a scenario with three frames is disclosed as shown by 904, where different complexity AI ILF models may be applied based on the 'frame wait-time' in DPB. Thus, a similar enhancement quality may be achieved with reduced latency unlike existing art. Further, the overall delay for the frames where similar quality is required over a few frames may be reduced if the complexity of the AI ILF may be aligned with the wait-time of frames. The same process may be followed by the video decoder 304 and the video encoder 302, therefore, avoiding a requirement to send a bitstream signal for each frame. The table 904 shows selectively adjusting the complexity of enhancement models applied to each frame based on their respective wait times. Unlike the current method that applies a fixed model complexity across all frames, the table 904 shows selectively assigning higher-complexity models to frames with longer wait times. As a result, an embodiment of the present disclosure reduces complexity to an acceptable level (e.g., 1.3x) over existing approaches while reducing the overall delay to zero, thereby improving processing efficiency and responsiveness in real-time decoding scenarios.

[0155] Figure 10 illustrates a use case of the system 204, in accordance with an embodiment of the present disclosure.

[0156] Referring to Figure 10, the AI ILF model may be applied based on the dependency of at least one of the plurality of frames on another frame from the plurality of frames, as shown by 1002. Thus, when the at least one of the plurality of frames may not have an immediate dependency in the DPB, in that case, a higher complexity AI ILF model may be applied to the at least one of the plurality of frames and therefore a full capacity of GPU processing may be performed unlike the existing art where decoding performed by the video decoder through CPU cores while AI ILF model runs on GPU cores. Further, all the AI ILF models may be of same complexity, then the overall advantages of the GPU cores may not be utilized, thus resulting in latency.

[0157] Figure 11 illustrates a use case of the system 204, in accordance with an embodiment of the present disclosure.

[0158] Referring to Figure 11, the system 204 ensures the use of different AI ILF models for different frames based on 'wait-time' in DPB results in the additional enhancement, as shown by images 1102. The additional enhancement may be achieved by applying the higher complexity AI ILF model to the frames that have to wait idle in the DPB before being used. Further, this provides leverage to the video encoder 302 to assign higher QP to such frames knowing that the higher complexity AI ILF model may be able to recover the quality at the In-loop stage. Further, there are two ways in which the frame wait-time in the DPB may be used to (1) enhance the frames with higher 'wait-time' with a more complex model. Though this method might be good for individual frame enhancement, in a video it might cause temporal inconsistency, and (2) reduce the encoder QP for a frame with higher 'wait-time', as it has the freedom to apply a more complex model. This may provide temporal consistency and may save bits required to encode the frame.

[0159] For example, among the images 1102, the three images in the first row may correspond to the first complexity table 1104, and the three images in the second row may correspond to the second complexity table 1106. The three images in the first row are images processed by exist method, but the three images in the second row are images processed by an embodiment of the disclosure. Therefore, in an example of the disclosure, for the two frames in the second column of images 1102, the image quality of the second row can be improved or the bitrate can be reduced by using a more complex AI ILF or a larger QP for the corresponding frame of the first row.

[0160] The present disclosure ensures a technical advancement that the system 204 and method 600, 700, 800 as disclosed improve the overall latency of the video encoder 302 / video decoder 304 by applying high complex AI ILF model to the frames which are idle in the DPB 406, 506 for longer duration and simple AI ILF model to the frames which have smaller wait-time. This ensures the enhanced quality of the frames streamed on the UE 202. The system 204 and method 600, 700, 800 also provide better compression gain by exploiting the wait-time of the frames in the DPB 406, 506. Particularly, when the operation is performed by the video decoder 304, in that case, the plurality of frames with the longer wait-time having complex AI ILF models are applied in the video decoder 302. Further, on the basis of the same, the video encoder 302 chooses to apply higher quantization parameter for such frames, knowing that the more complex AI ILF model enhances such frames, thus reducing the latency. Additionally, even if the operation is performed by the video encoder 302, the video encoder 302 chooses to apply higher quantization parameter for the plurality of frames with longer wait-time and further, such frames have the complex AI ILF models. Further, the plurality of frames with the complex AI ILF and the higher quantization parameter is provided to the video decoder 304. Thereafter, the video decoder 304 decodes each of the plurality of frames accordingly, thus enhancing such frames while reducing latency. Further, the plurality of frames with the shorter wait-time having simple AI ILF models, in the operations as performed by the video encoder 302 and the video decoder 34, ensures the optimum visual experience for the user. Further, a plurality of parameters, for example, the wait-time, the quantization parameter, the AI ILF model may be standardized by a service provider such that the video decoder 304 receives the standardized parameter and decodes the plurality of video frames on the basis of the standardized parameter, thus improving the efficiency while decoding the plurality of video frames.

[0161] In an embodiment of the disclosure, a method 600 for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model 514 by a video decoder 304. The method may include obtaining 602 a bit-stream including information for a group of pictures (GOP) structure 502 of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure 502 includes interdependency among the plurality of the video frames. The method may include determining 604 a wait-time 504 in a decoded picture buffer (DPB) 506 for each of the plurality of video frames based on the information for the GOP structure 502. The method may include selecting 606 an AI ILF model 514 from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time 504. The method may include generating 608 a plurality of enhanced reconstructed video frames by enhancing the plurality of video frames based on the selected AI ILF model.

[0162] In an embodiment of the disclosure, the method may include determining the wait-time 504 for each of the plurality of the video frames based on the display sequence of the plurality of video frames. A first video frame having later display sequence may have a longer wait-time than a second video frame having immediate display sequence.

[0163] In an embodiment of the disclosure, the method may include selecting a first quantization parameter for a first video frame from among the plurality of video frames. The method may include selecting a second quantization parameter for a second video frame from among the plurality of video frames. when a determined wait-time for the first video frame may be longer than a determined wait-time for the second video frame, the first quantization parameter is higher than the second quantization parameter.

[0164] In an embodiment of the disclosure, a quantization parameter applied to the frame may be derived based on table information from at least one of sequence header, or frame header of plurality of video frames.

[0165] In an embodiment of the disclosure, the table information is predefined such that both the quantization parameter and the AI ILF model are determined according to the wait-time.

[0166] In an embodiment of the disclosure, information on whether to dynamically select AI ILF for the plurality of video frames is derived from at least one of a slice header, a sequence header, or a frame header of plurality of video frames.

[0167] In an embodiment of the disclosure, the method may include selecting a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames. The method may include selecting a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of the video frames. When a determined wait-time for the first video frame is longer than a determined wait-time for the second video frames, the first AI ILF model is more complex than the second AI ILF model.

[0168] In an embodiment of the disclosure, a method 700 for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model 414 by a video encoder 302 may be provided. The method may include obtaining 702 information for a group of pictures (GOP) structure 402 of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure 402 includes interdependency among the plurality of the video frames. The method may include determining 704 a wait-time 404 in a decoded picture buffer (DPB) 406 for each of the plurality of video frames based on the information for the GOP structure. The method may include selecting 706 an AI ILF model 414 from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time 404.

[0169] In an embodiment of the disclosure, the method 700 may include determining the wait-time 404 for each of the plurality of the video frames based on the display sequence of the plurality of video frames. The first video frame having later display sequence may have a longer wait-time than a second video frame having immediate display sequence.

[0170] In an embodiment of the disclosure, the method 700 may include selecting a first quantization parameter for a first video frame from among the plurality of video frames. The method may include selecting a second quantization parameter for a second video frame from among the plurality of video frames. When a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter may be higher than the second quantization parameter.

[0171] In an embodiment of the disclosure, a quantization parameter applied to the frame may be derived based on a table decoded by table information from at least one of sequence header, or frame header of plurality of video frames.

[0172] In an embodiment of the disclosure, the table information is predefined such that both the quantization parameter and the AI ILF model are determined according to the wait-time.

[0173] In an embodiment of the disclosure, information on whether to dynamically select AI ILF for the plurality of video frames may be derived from at least one of a slice header, a sequence header, or a frame header of plurality of video frames.

[0174] In an embodiment of the disclosure, the method may include selecting a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames. The method may include selecting, a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of video frames. When a determined wait-time for the first video frames is longer than a determined wait-time for the second video frame, the first AI ILF model may be more complex than the second AI ILF model.

[0175] In an embodiment of the disclosure, a method for transmitting a bitstream generated by the method 700 for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder 302 may be provided.

[0176] In this application, unless specifically stated otherwise, the use of the singular includes the plural and the use of "or" means "and / or." Furthermore, use of the terms "including" or "having" is not limiting. Any range described herein will be understood to include the endpoints and all values between the endpoints. Features of the disclosed embodiments may be combined, rearranged, omitted, etc., within the scope of the invention to produce additional embodiments. Furthermore, certain features may sometimes be used to advantage without a corresponding use of other features.

[0177] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be appreciated that a vast number of variations exist.

Claims

1.A method (600) for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model (514) by a video decoder (304), comprising:obtaining (602) a bit-stream including information for a group of pictures (GOP) structure (502) of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure (502) includes interdependency among the plurality of the video frames;determining (604) a wait-time (504) in a decoded picture buffer (DPB) (506) for each of the plurality of video frames based on the information for the GOP structure (502);selecting (606) an AI ILF model (514) from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time (504); andgenerating (608) a plurality of enhanced reconstructed video frames by enhancing the plurality of video frames based on the selected AI ILF model.2.The method (600) of claim 1, wherein the determining the wait-time (504) comprises:determining the wait-time (504) for each of the plurality of the video frames based on the display sequence of the plurality of video frames,wherein a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence.3.The method (600) of any one of claims 1 to 2, wherein the method (600) further comprises:selecting a first quantization parameter for a first video frame from among the plurality of video frames; andselecting a second quantization parameter for a second video frame from among the plurality of video frames,wherein, when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter is higher than the second quantization parameter.4.The method (600) of any one of claims 1 to 3, wherein a quantization parameter applied to the frame is derived based on table information from at least one of sequence header, or frame header of plurality of video frames.5.The method (600) of claims 4, wherein the table information is predefined such that both the quantization parameter and the AI ILF model are determined according to the wait-time.6.The method (600) of any one of claims 1 to 5, wherein information on whether to dynamically select AI ILF for the plurality of video frames is derived from at least one of a slice header, a sequence header, or a frame header of plurality of video frames.7.The method (600) of any one of claims 1 to 6, wherein the selecting the AI ILF model (514) comprises:selecting a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames; andselecting a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of the video frames,wherein, when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frames, the first AI ILF model is more complex than the second AI ILF model.8.A method (700) for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model (414) by a video encoder (302), comprising:obtaining (702) information for a group of pictures (GOP) structure (402) of a plurality of video frames corresponding to a source video, wherein the information for the GOP structure (402) includes interdependency among the plurality of the video frames;determining (704) a wait-time (404) in a decoded picture buffer (DPB) (406) for each of the plurality of video frames based on the information for the GOP structure; andselecting (706) an AI ILF model (414) from among a plurality of AI ILF models for each of the plurality of video frames based on the determined wait-time (404).9.The method (700) of claim 8, wherein the determining the wait-time (404) comprises:determining the wait-time (404) for each of the plurality of the video frames based on the display sequence of the plurality of video frames,wherein a first video frame having later display sequence has a longer wait-time than a second video frame having immediate display sequence.10.The method (700) of any one of claims 8 to 9, wherein the method (700) further comprises:selecting a first quantization parameter for a first video frame from among the plurality of video frames; andselecting a second quantization parameter for a second video frame from among the plurality of video frames,wherein, when a determined wait-time for the first video frame is longer than a determined wait-time for the second video frame, the first quantization parameter is higher than the second quantization parameter.11.The method (700) of any one of claims 8 to 10, wherein a quantization parameter applied to the frame is derived based on table information from at least one of sequence header, or frame header of plurality of video frames.12.The method (700) of claim 11, wherein the table information is predefined such that both the quantization parameter and the AI ILF model are determined according to the wait-time.13.The method (700) of any one of claims 8 to 12, wherein information on whether to dynamically select AI ILF for the plurality of video frames is derived from at least one of a slice header, a sequence header, or a frame header of plurality of video frames.14.The method (700) of any one of claims 8 to 13, wherein the selecting the AI ILF model (414) comprises:selecting a first AI ILF model from among the plurality of AI ILF models for a first video frame from among the plurality of the video frames; andselecting, a second AI ILF model from among the plurality of AI ILF models for a second video frame from among the plurality of video frames,wherein, when a determined wait-time for the first video frames is longer than a determined wait-time for the second video frame, the first AI ILF model is more complex than the second AI ILF model.15.A method for transmitting a bitstream generated by the method (700) for dynamically selecting an artificial intelligence (AI) based In-Loop filter (AI ILF) model by a video encoder (302), of any one of claims 8 to 14.

Citation Information

Patent Citations

  • Video frame encoding scheme selection

    US10171804B1

  • Methods, devices, and systems for decoding portions of video content according to a schedule based on user viewpoint

    US20200128279A1

  • Using neural network filtering in video coding

    US20240048775A1

  • Method, apparatus, and medium for video processing

    US20240244201A1

  • Apparatuses and methods for encoding and decoding a video using in-loop filtering

    WO2024013356A1