Method and System for Content-Based Scaling of an In-Loop Filter for an Artificial Intelligence Infrastructure

An AI-based in-loop filter with content-based scaling addresses the bias and variance issues in existing models by adapting to new data domains, enhancing video quality and reducing artifacts across diverse content types without frequent retraining.

JP2025524131APending Publication Date: 2025-07-25SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504412
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-27
Filing Date
2023-07-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing AI-based video compression models suffer from inherent bias and variance due to limited training datasets, leading to unreliable performance and the need for frequent retraining, especially when encountering new or unseen data types, and existing in-loop filters lack adaptability and intelligence in addressing various content types.

Method used

Implementing an AI-based in-loop filter that uses content-based scaling through pixel mapping and statistical regression to adapt to new data domains, reducing data domain gaps and enhancing generalization without requiring pre-training, while maintaining low complexity.

Benefits of technology

The solution provides improved video quality by adaptively reducing compression artifacts across diverse content types, ensuring reliable performance and reducing the need for frequent model retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524131000001_ABST
    Figure 2025524131000001_ABST
Patent Text Reader

Abstract

A method for encoding an artificial intelligence (AI) base of a media includes compressing an input image frame associated with an input video, generating a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI base, determining an offset value based on the input image frame and the reconstructed image frame, and encoding the reconstructed image frame based on the determined offset value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to image processing, and more specifically, to a method and system for content-based scaling for an in-loop filter based on artificial intelligence (AI).

Background Art

[0002] Video compression techniques are the basis of video content for consumption. The basic goal of video compression techniques is to reduce transmission time and bandwidth requirements. With the advent of artificial intelligence (AI) technology, video codec standard organizations expect to obtain potential benefits of AI technology in video compression techniques. However, there are other obstacles to implementing an AI-based model (e.g., deep neural networks, linear regression, etc.).

[0003] Video compression techniques include video codec pipelines for compressing a very diverse range of content, and new content is constantly being developed, such as content with higher resolutions and different screen contents. When an AI-based coding tool is introduced into the video codec pipeline, the AI-based model must first be trained and then deployed into the video codec pipeline. However, the AI-based model is trained only on a limited training dataset collected at a specific time point, which can introduce inherent bias and variance into the AI-based model and potentially lead to unexpected performance and unreliable responses.

Summary of the Invention

Means for Solving the Problems

[0004] ​This summary is provided to introduce a selection of concepts in a simplified format that will be further described in the detailed description of the invention. This summary is not intended to identify the core or essential inventive concepts of the invention and is not intended to determine the scope of the invention.

[0005] According to one embodiment of the present disclosure, a method for AI-based encoding of media is disclosed. In one embodiment of the present disclosure, the method may include compressing an input image frame associated with an input video. In one embodiment of the present disclosure, the method may include generating a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI base. In one embodiment of the present disclosure, the method may include determining an offset value based on the input image frame and the reconstructed image frame. In one embodiment of the present disclosure, the method may include encoding the reconstructed image frame based on the determined offset value.

[0006] According to an embodiment of the present disclosure, a method of using AI-based decoding of media is disclosed. In one embodiment of the present disclosure, the method may include receiving bitstream information including offset information from an encoder. In one embodiment of the present disclosure, the method may include generating a reconstructed image frame based on the bitstream information using an in-loop filter of the AI base. In one embodiment of the present disclosure, the method may include performing a scaling operation on the reconstructed image frame based on the offset information to generate a scaled image frame. In one embodiment of the present disclosure, the method may include generating an output video based on the scaled image frame.

[0007] According to one embodiment of the present disclosure, an electronic device for encoding an AI base of media is disclosed. The electronic device includes a system, and the system includes a processor to which a memory and a communication unit are coupled. In one embodiment of the present disclosure, the processor may be configured to compress an input image frame associated with an input video. In one embodiment of the present disclosure, the processor may be configured to generate a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI base. In one embodiment of the present disclosure, the processor may be configured to determine an offset value based on the input image frame and the reconstructed image frame. In one embodiment of the present disclosure, the processor may be configured to encode the reconstructed image frame based on the determined offset value.

[0008] According to one embodiment of the present disclosure, an electronic device for decoding an AI base of media is disclosed. The electronic device includes a system, and the system includes a processor to which a memory and a communication unit are coupled. In one embodiment of the present disclosure, the processor may be configured to receive bitstream information including offset information from an encoder. In one embodiment of the present disclosure, the processor may be configured to generate a reconstructed image frame based on the bitstream information using an in-loop filter of the AI base. In one embodiment of the present disclosure, the processor may be configured to perform a scaling operation on the reconstructed image frame based on the offset information to generate a scaled image frame. In one embodiment of the present disclosure, the processor may be configured to generate an output video based on the scaled image frame.

[0009] To further clarify the advantages and features of the present invention, the present invention will be described in more detail with reference to specific embodiments shown in the accompanying drawings. It will be understood that these drawings show only typical embodiments of the present invention and thus are not intended to limit the scope of the present invention. The present invention will be described and explained with additional specificity and detail together with the accompanying drawings.

Brief Description of the Drawings

[0010] These and other features, aspects, and advantages of the present invention will be better understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, throughout the drawings, like numerals represent like parts.

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Embodiments for Carrying Out the Invention

[0011] To facilitate understanding of the principles of the present invention, reference is made to the embodiments illustrated in the drawings and described with reference to specific terms. However, it is not intended thereby to limit the scope of the present invention, and such modifications and additional amendments in the illustrated devices and further applications of the principles of the present invention as illustrated in the present application are contemplated as being ordinarily within the scope of those skilled in the art of the present invention.

[0012] It will be understood by those skilled in the art that the foregoing general description and the following detailed description are for explaining the present invention and are not intended as limitations thereof.

[0013] References throughout this specification to "one embodiment", "another embodiment", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment", "in another embodiment", and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0014] The terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or subsystems or elements or structures or components proceeding with "comprising" does not, without more constraints, preclude the existence of other devices or other subsystems or other elements or other structures or other components or additional devices or additional subsystems or additional elements or additional structures or additional components.

[0015] Throughout the present disclosure, the expression "at least one of a, b, or c" indicates only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0016] Embodiments and their various features and advantageous details in the present disclosure are illustrated in the accompanying drawings and will be further described in more detail with reference to non-limiting embodiments detailed in the following description. Descriptions of widely known components and processing techniques are omitted so as not to unnecessarily obscure the embodiments of the present disclosure. Also, the various embodiments described in the present disclosure are not necessarily mutually exclusive, and some embodiments may be combined with one or more other embodiments to form new embodiments. The term "or" as used in the present disclosure means non-exclusive unless otherwise indicated. The examples used in the present disclosure are merely for facilitating the understanding of how the embodiments of the present disclosure may be implemented and for further enabling a person of ordinary skill in the art to implement the embodiments of the present disclosure. Therefore, the examples should not be construed as limiting the scope of the embodiments of the present disclosure.

[0017] As is customary in the art, embodiments may be described and illustrated in terms of blocks that perform the described functions or operations. These blocks, which may be referred to in the present disclosure as units or modules, etc., are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, and may be optionally driven by firmware and software. The circuits may be implemented, for example, within one or more semiconductor chips or on a substrate support such as a printed circuit board. The circuits constituting the blocks may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuits), or by a combination of dedicated hardware performing some functions of the block and a processor performing other functions of the block. Each block of the embodiments may be physically separated into two or more interacting and individual blocks without departing from the scope of the present invention. Similarly, the blocks of the embodiments may be physically combined into more complex blocks without departing from the scope of the present invention.

[0018] The accompanying drawings are used to facilitate the understanding of various technical features, and it should be understood that the embodiments presented in this disclosure are not limited by the accompanying drawings. Thus, this disclosure should be construed as extending to any modifications, equivalents, and substitutions in addition to those particularly presented in the accompanying drawings. Although terms such as first, second, etc. may be used in this disclosure to describe various elements, these elements should not be limited by these terms. These terms are generally used only for the purpose of distinguishing one element from another.

[0019] Throughout this disclosure, the terms "desired compression level" and "desired level of compression" are used interchangeably and have the same meaning. Throughout this disclosure, the terms "input video", "input video data", and "input media" are used interchangeably and have the same meaning. The terms "reconstructed video", "output video", and "optimized compressed media" are used interchangeably and have the same meaning.

[0020] Bias and variance are two core characteristics of a model (e.g., an AI-based model, a machine learning (ML) model, etc.). The bias of a model means how well the model can represent all potential outcomes. On the other hand, variance means how much the model's predictions are affected by small changes in the input data. The trade-off between bias and variance is a fundamental problem in models, and the trade-off usually requires experimenting with various model types to find the optimal balance in the training data. Figure 1A illustrates an exemplary scenario for visualizing the trade-off between bias and variance in an AI-based model according to the related art. A person expects to reach a point indicating a target value and throws one or more stones. However, the one or more stones land on the ground at various positions far from the target value. The bias is here depicted as the gap between the target value (i.e., the actual value) and the average of the points where the one or more stones landed. The dashed line reflects the average of the predictions provided by one or more AI-based models trained using different training data sets extracted from the same population. Here, the dashed line may reflect the predicted positions where the one or more stones land based on the calculations of the one or more AI-based models. However, since each AI-based model is trained on different training data, the one or more stones land at various positions around the target value. The dashed line can serve as an approximation of the result expected based on the one or more AI-based models considering the variation in the positions where the one or more stones land. The shorter the distance, the smaller the bias between the average (i.e., the dashed line) and the target value, leading to a better model. Variance can be expressed as the spread of the individual points around the average (i.e., the dashed line). Lower spread and variance indicate a better model among the one or more AI-based models. This is because lower spread means a higher level of accuracy and precision in the model's predictions. By minimizing variance, the model can predict the output and accurately provide more reliable results.

[0021] Figure 1B illustrates the limitations of an AI foundation model based on biases and variances learned from training data according to the related art. To obtain desired performance and reliable responses for the AI foundation model, the source data (i.e., training data) and the target data (i.e., test data) must belong to the same sample space. However, in a real-time environment, the source data and the target data can belong to different sample spaces. In that case, the AI foundation model may not be able to provide the expected performance and reliable responses. For example, consider a scenario where an AI foundation model is trained only on Indian faces (limited training data) for face classification and detection. In this situation, if the trained AI foundation model is executed on American faces (i.e., new, unseen data), the trained AI foundation model may not be able to perform as expected. This is due to certain differences (e.g., anchor point changes) between Indian faces and American faces. Thus, the trained AI foundation model has biases and variances that are learned with a bias towards Indian faces. Therefore, the trained AI foundation model cannot provide the expected performance and reliable responses. To solve this problem, the AI foundation model must be retrained each time a new sample space is discovered, which can be extensive and lead to continuous manual work.

[0022] Furthermore, the AI base model is undergoing further generation of new forms of data, which tends to become obsolete over time and must be retrained. However, it is inconvenient to retrain the AI base model regularly. Video compression techniques are, for example, deployed globally, and frequent adjustment of the weights of the AI base model is not desirable. Additionally, existing in-loop filters (ILFs) connected to the AI base model identify certain types of artifacts using image and signal processing principles. Due to limited options, existing ILFs employ various modes / filters that are optimal for a given content as illustrated in FIGS. 2A and 2B and as described hereinafter. Experimenting with various modes / filters for a single input increases the encoding time. Some existing ILFs do not use measurement results as a measure of quality and instead focus only on visible artifacts; due to such an inherent nature of existing ILFs, existing ILFs are rigidified.

[0023] FIG. 2A illustrates a block diagram of an existing general-purpose video coding (VVC) decoder according to the related art. The VVC decoder is an essential part of a video coding system that enables efficient compression and decompression of video data. The VVC decoder is configured to provide high-quality video playback while minimizing the amount of data required for transmission and storage. One of the fundamental modules of the VVC decoder is context-adaptive binary arithmetic coding (CABAC), which is a sophisticated entropy coding technique that adapts to the statistical characteristics of the input data. CABAC allows the VVC decoder to achieve a high compression ratio without sacrificing video quality. In addition to CABAC, the VVC decoder also includes an inverse quantization module, an inverse transform module, luma mapping chroma scaling, etc. scaling, LMCS) module, in-loop filter module, intra prediction module, inter prediction module, LMCS (forward luma mapping) module, combined inter intra prediction, CIIP) module, decoded picture buffer, DPB) module, and several other important modules including the reconstructed image. The inverse quantization module and the inverse transform module play the role of converting the quantized and transformed video data back to its original form. The LMCS module is used to adjust the chroma and luma components of the video data to improve the overall image quality.

[0024] Furthermore, the in-loop filter is configured to remove artifacts and noise generated during the video compression process. The in-loop filter generally includes inverse luma mapping chroma scaling (LMCS), deblocking filter, DBF), sample adaptive offset (SAO) filter, adaptive loop filter, ALF), and cross-component It includes adaptive loop filtering (CC-ALF). The inverse LMCS module is configured to adjust chroma scaling to match the luma mapping and ensure appropriate color representation in the video compression process. The DBF is used at block boundaries to remove blocking artifacts. The SAO filter is applied to the samples that have subsequently been block-removed. SAO is very effective in removing ringing artifacts and correcting local average intensity variations. The ALF and CC-ALF perform block-based linear filtering and adaptive filtering to minimize the mean squared error ( MSE) between the original data (e.g., the original image) and the regenerated data (e.g., the reconstructed image).

[0025] Finally, the DPB module and the reconstructed image module work together to store and display the decoded video data. The DPB module stores previously decoded frames to enable efficient inter-frame prediction, while the reconstructed image module generates a high-quality display image from the decoded video data.

[0026] The intra prediction module and the inter prediction module are configured to predict the values of video data based on spatial and temporal correlations. This helps reduce the amount of data that needs to be transmitted or stored by encoding only the difference between the predicted value and the actual value. The CIIP module combines both intra prediction techniques and inter prediction techniques to further improve the compression efficiency.

[0027] In existing VVC decoders, each in-loop filter targets specific artifacts or default artifacts. For example, as described above, DBF only removes blocking artifacts. As a result, existing filtering processes do not have the "intelligent" elements present in the existing filtering process, so the performance cannot be adjusted based on the type of input content. To obtain the best performance, existing filtering processes use heuristics to select from among a number of different filter versions (e.g., LMCS, DBF, SAO, etc.), which degrades the user experience.

[0028] Figure 2B illustrates an existing VVC decoder pipeline with an AI-based in-loop filter according to the related art. Compared with the block diagram of Figure 2A, the AI-based in-loop filter is an additional block in the existing VVC decoder pipeline. The AI-based in-loop filter removes the problems of the in-loop filter (i.e., each filter targets specific artifacts or default artifacts). However, as shown in Figure 1B, the AI-based in-loop filter can introduce substantial bias and distortion due to the limited amount of training data used by the AI-based model. As a result, the AI-based in-loop filter can produce unexpected results in scenarios where the AI-based model has not been trained, making the solution unreliable. Furthermore, the AI-based model is in a state where new forms of data generation are ongoing and tend to become obsolete over time and must be retrained. However, it is inconvenient to retrain the AI-based model regularly.

[0029] The method according to an embodiment of the present disclosure can provide a unique strategy for enhancing an AI base model with additional real-time modules that perform content-based scaling to reduce data domain gaps and further adapt to new and unseen data, as described with reference to FIGS. 3 to 16. The method according to an embodiment of the present disclosure uses pixel mapping to fill the data domain gap between new / unseen data and training data in the in-loop filter of the AI base. As a result, the method according to an embodiment of the present disclosure is superior to existing methods that use limited data targeted at specific artifacts and introduce biases and volatility inherent in the well-trained AI base training model. The method according to an embodiment of the present disclosure enables content-based scaling to identify pixel mapping for offset scaling, operates on-the-fly, requires no pre-training, and can provide more generalization for different types of data. Further, the method according to an embodiment of the present disclosure has a very low complexity when compared to existing methods, but appropriately uses a very complex deep neural network (DNN) model to achieve a higher gain for existing methods and is not adaptable by content. The method according to an embodiment of the present disclosure can solve this problem by introducing statistical regression of the video data base during the encoding process and using the same during the decoding process.

[0030] Here, referring to FIGS. 3 to 16, in which like reference numerals consistently denote corresponding features throughout the drawings, one or more embodiments of the present disclosure are illustrated.

[0031] FIG. 3 illustrates a block diagram of a general-purpose video coding (VVC) encoder pipeline 300 for AI base encoding of media (e.g., input video) according to an embodiment of the present disclosure.

[0032] The VVC encoder pipeline 300 includes a plurality of modules. The plurality of modules may include a compression module 301, an in-loop filter 302, a CABAC module 303, and an RD cost module 304. The compression module 301 is a traditional module that can receive (or identify) an input video (e.g., one or more input image frames). When receiving the input video, the compression module 301 can perform one or more predefined operations on the received input video, such as conversion, quantization, inverse quantization, and inverse conversion.

[0033] a. Conversion: The conversion operation involves converting data (e.g., the input video) from its original domain representation to a different domain, where the data can be more efficiently compressed and the data includes the converted coefficients. The conversion operation is typically achieved using mathematical techniques, so-called discrete cosine transform (DCT) or discrete wavelet transform (DWT).

[0034] b. Quantization: Once the data is converted, the quantization operation is performed. The quantization operation involves reducing the precision or dynamic range of the converted coefficients. The converted coefficients are essentially mapped from a continuous range of values to a finite set of discrete levels. By quantizing the coefficients, information is lost, which contributes to compression.

[0035] c. Inverse Quantization: The inverse quantization operation is the inverse of the quantization operation. Inverse quantization involves approximately restoring the quantized coefficients back to their original values, which is achieved by multiplying each quantized coefficient by an inverse quantization coefficient. The inverse quantization operation is performed during decompression to restore an approximation of the original converted coefficients.

[0036] d. Inverse Transformation: The inverse transformation operation involves converting the transformed and inverse-quantized coefficients back to the original domain. The inverse transformation operation performs the inverse operation of the initial transformation and reconstructs the compressed data to be nearly identical to the original input (e.g., the input video). The inverse transformation operation, so-called inverse DCT or inverse DWT, effectively reverses the energy decorrelation and concentration performed during the initial transformation stage.

[0037] The in-loop filter 302 can receive (identify) the compressed data (e.g., quantized data, reconstructed data) from the compression module 301. The in-loop filter 302 is configured to remove artifacts and noise that may occur during the compression operation. The in-loop filter 302 can include an LMCS 302a, a DBF 302b, an SAO filter 302c, an ALF, and a CC-ALF 302d. The functions of the various blocks (i.e., 302a, 302b, 302c, and 302d) are the same as the description in FIG. 2A, and these various blocks are traditional modules.

[0038] In one or more embodiments of the present disclosure, the in-loop filter 302 can include an AI-based in-loop filter 302e, a data distribution model 302f, a ground-truth data distribution model 302g, a rate-distortion optimization (RDO) control mapping range module 302h, a pixel mapping module 302i for offset scaling, and a final AI-based in-loop filter output module 302j.

[0039] In one or more embodiments of the present disclosure, the AI-based in-loop filter 302e can be configured to apply AI techniques to the loop filtering process of video coding, as described in connection with FIG. 5. The in-loop filter modules 302a, 302b, 302c, 302d can typically be designed using fixed techniques to reduce compression artifacts such as blocking and ringing artifacts. However, due to advancements in AI and machine learning, researchers have explored the use of AI-based approaches to improve the performance of the in-loop filter modules 302a, 302b, 302c, 302d. The AI-based in-loop filter 302e can utilize a machine learning model to learn the characteristics of video content (e.g., image input from VVC, input video, etc.) and apply adaptive filtering techniques. Instead of relying on predefined rules or heuristics, the AI-based in-loop filter 302e can analyze and process video content using a neural network or other AI techniques. The AI-based in-loop filter 302e can use an AI model to learn from a large amount of training data and make intelligent decisions regarding the filtering operation. By using the AI-based in-loop filter 302e, the effect of artifact reduction can be improved, especially in conductive scenarios where traditional in-loop filter modules 302a, 302b, 302c, 302d experience difficulties. The AI model can adapt to different content types, handle complex textures and motions, and preserve details in a better way while reducing artifacts. Generally, the AI-based in-loop filter 302e can aim to improve the quality of encoded video content by more adaptively and intelligently reducing compression artifacts, leveraging the capabilities of AI. The encoded video content can include the reconstructed video and / or one or more reconstructed image frames, as described in connection with FIG. 5.

[0040] In one or more embodiments of the present disclosure, the data distribution model 302f may generate a first representative data point set, as described in connection with FIG. 6. The data distribution model 302f may perform various operations shown below to generate the first representative data point set.

[0041] The data distribution model 302f can receive (or identify) one or more reconstructed image frames (e.g., one or more enhanced image frames, model output data) from the in-loop filter 302e of the AI infrastructure, where the one or more reconstructed image frames pertain to the reconstructed video. The data distribution model 302f can then receive (or identify) user input that indicates the desired compression level for each of the one or more reconstructed image frames. The data distribution model 302f can then determine, based on the received user input (or the desired compression level), the number of segments (e.g., 5) for each of the one or more reconstructed image frames. The data distribution model 302f can then fragment each of the one or more reconstructed image frames based on the determined number of segments. Thereafter, the data distribution model 302f uses one or more distribution mechanisms (e.g., normal distribution, etc.) to analyze the content changes within each of the one or more reconstructed image frames to calculate the optimal segments for each of the one or more reconstructed image frames (i.e., generate one or more optimally fragmented image frames). The data distribution model 302f then performs a pixel binning operation on the one or more optimally fragmented image frames using one or more statistical mechanisms (e.g., Gaussian mixture model (GMM)), where each of the one or more optimally fragmented image frames includes groups of pixels with unique characteristics (so-called regions with similar textures, regions with similar deviations from the mean, etc.). The data distribution model 302f performs a pixel binning operation on one or more segments (or regions) of each of the one or more optimally fragmented image frames to generate a first representative data point set for each of the one or more optimally fragmented image frames, where the first representative data point set can indicate scalar values associated with each of the one or more segments (or optimally fragmented image frames).

[0042] In one or more embodiments of the present disclosure, the ground truth data distribution model 302g can generate a second representative data point set, as described in connection with FIG. 7. The ground truth data distribution model 302g can perform various operations shown below to generate a second representative data point set.

[0043] The ground truth data distribution model 302g can receive (or identify) one or more input image frames associated with an input video (e.g., the original video, an uncompressed video). The ground truth data distribution model 302g can then receive (or identify) user input that indicates a desired compression level for each of the one or more input image frames. The ground truth data distribution model 302g can then determine the number of segments for each of the one or more input image frames based on the received user input (or the desired compression level). The ground truth data distribution model 302g can then fragment each of the one or more input image frames based on the determined number of segments. The ground truth data distribution model 302g can then analyze the content changes within each of the one or more input image frames using one or more distribution mechanisms and calculate the optimal segments for each of the one or more input image frames (i.e., generate one or more optimally fragmented input image frames). The ground truth data distribution model 302g can then perform a pixel binning operation on the one or more optimally fragmented input image frames using one or more statistical mechanisms, where each of the one or more optimally fragmented input image frames includes groups (or regions) of pixels with unique characteristics. The ground truth data distribution model 302g performs a pixel binning operation on one or more segments (or regions) of each of the one or more optimally fragmented input image frames to generate a second representative data point set for each of the one or more optimally fragmented input image frames, where the second representative data point set can indicate scalar values associated with each of the one or more segments (or optimally fragmented image frames).

[0044] In one or more embodiments of the present disclosure, the RDO control mapping range module 302h may determine the overall RD cost based on the output of the RD cost module 304 and user input regarding the number of fragments, as described in connection with FIG. 4. The RDO control mapping range module 302h can construct an objective function expressed as a mathematical formula. The objective function receives inputs such as updated rate and distortion, along with the previous RDO cost. Based on the updated RDO cost, the RD cost module 304 may make a decision to optimize the overall RD cost by minimizing either one or both of the rate or distortion.

[0045] In one or more embodiments of the present disclosure, the pixel mapping module 302i for offset scaling may determine a data distribution dissimilarity metric between a first representative data point set associated with each fragmented image frame associated with the reconstructed video and a second representative data point set associated with each fragmented input image frame associated with the input video, as described in connection with FIG. 8. The pixel mapping module 302i for offset scaling may further determine an offset value associated with each fragmented image frame by using the data distribution dissimilarity metric, as described in connection with FIG. 8. The offset values are calculated and used at the per-fragment granularity, where the number of fragments is received as input from the user and each fragment is calculated during a data distribution modeling operation (e.g., the data distribution dissimilarity metric).

[0046] In one or more embodiments of the present disclosure, the pixel mapping for the offset scaling module 302i can determine (or calculate) the (predicted) pixel mapping in the form of a per-slice offset using the model output data distribution, the ground truth data distribution, and the mapping range (e.g., the RDO control mapping range). The pixel mapping is also a pixel-level domain mapping based on the differences in the modeled distributions. The mapping range is also a determination of whether to apply offset scaling to a particular slice based on the RD cost, and the mapping range is determined using the user-defined number of slices and the output of the RD cost module 304 (e.g., the codec RD cost).

[0047] In one or more embodiments of the present disclosure, the final AI base in-loop filter output module 302j can perform a scaling operation on each pixel value associated with each fragmented image frame by using the determined offset values, as described in connection with FIG. 9. Examples of the scaling operation can include, but are not limited to, addition operations, multiplication operations, division operations, or exponentiation operations. The final AI base in-loop filter output module 302j can further improve the reconstructed video based on the performed scaling operation.

[0048] In one or more embodiments of the present disclosure, the encoder transmits bitstream information associated with the reconstructed video to a decoder (see FIG. 12), where the bitstream information can include the determined offset values (e.g., offset signaling).

[0049] Furthermore, the CABAC module 303 and the RD cost module 304 are two traditional modules used in the VVC encoder, which is the latest video coding standard. The CABAC module 303 is an entropy coding technique adopted from the VVC encoder to efficiently compress the transformed and quantized coefficients. In the VVC encoder, the CABAC module 303 is configured to compress the input video data and perform the following functions: a. Context modeling: The CABAC module 303 can analyze the input video data (e.g., input video, image frame, etc.) and generate a statistical model that captures the relationships and dependencies between symbols in the input video data. The CABAC module 303 can estimate the probabilities of different symbols that occur based on the context associated with the input video data.

[0050] b. Binary arithmetic coding: The CABAC module 303 can convert the probabilities into cumulative distribution functions (CDFs). The CDFs are used to assign binary codewords to the symbols of the input video data. The binary arithmetic coding technique efficiently encodes symbols based on probabilities.

[0051] c. Bitstream generation: The CABAC module 303 can generate a compressed bitstream containing the encoded symbols along with the necessary information for the decoder to reconstruct the input video data. The compressed bitstream is then transmitted or stored for further processing or transmission, as will be described in detail below.

[0052] The RD cost module 304 is configured to determine an optimal tradeoff between rate (i.e., bit rate) and distortion (i.e., visual quality), which helps the VVC encoder evaluate the rate and distortion of different encoding options to find the best coding parameters for each coding unit. The RD cost module 304 performs various operations for encoding the input video data, such operations being listed below.

[0053] The RD cost module 304 may estimate the number of bits required to represent the encoded data for a particular coding option, which includes bits used to code syntax elements, motion information, and residual data.

[0054] b. To evaluate the visual quality, the RD cost module 304 can calculate the distortion by comparing the reconstructed video with the original input. The mean squared error (MSE) or structural similarity index (SSE) can be used to evaluate the visual quality. A common metric such as the linear similarity index (SSIM) can be used to measure distortion.

[0055] c. The RD cost module 304 may perform rate-distortion optimization by evaluating various coding options for each coding unit, exploring different modes, motion vectors, transform options, and quantization parameters to find the combination that achieves the best tradeoff between rate and distortion.

[0056] Based on the rate and distortion calculations, the RD cost module 304 may select the coding options that minimize the overall RD cost. These optimal choices are used for the final encoding of the input video data.

[0057] By integrating the RD cost module 304, the VVC encoder can efficiently allocate bits based on content characteristics and perceptual importance, improving video compression performance while maintaining satisfactory visual quality.

[0058] FIG. 4 illustrates a rate-distortion diagram 400 for determining the RD cost value for each compressed image frame for AI-based encoding of media according to the related art.

[0059] The RD cost module 304 can obtain user preferences regarding the amount of media compression and the quality of the compressed media. The RD cost module 304 can determine the rate (R) value (i.e., bitrate) for each compressed image frame based on the amount of compression (or range) as shown in Equation 1. The RD cost module 304 can determine the distortion (D) value for each compressed image frame based on the quality of the compressed content as shown in Equation 2.

[0060]

Equation

[0061]

Equation

[0062] Based on the user's current requirements, the user can lower either one of the R value or the D value as illustrated in the rate distortion diagram 400. In other words, there is a trade-off between the R value and the D value. For example, as shown in element 403, if the user desires high quality for an output of a video codec pipeline with little distortion, the D value can be selected to be 40.5. If so, the R value can increase to reduce the distortion. The R value is associated with either one of the second curve 402 (e.g., R value = 1.25) or the first curve 401 (e.g., R value = 1.75). If the user desires to optimize the memory of the electronic device, the RD cost module 304 can select the second curve 402 with the lowest R value.

[0063] Figure 5 illustrates a typical architecture of a deep convolutional neural network (DCNN) 500 associated with an in-loop filter 302e of an AI base for improving the quality of each compressed image frame (e.g., reconstructed image frame) for encoding of a media according to the related art. Neural Network, DCNN) 500.

[0064] The in-loop filter 302e of the AI foundation receives one or more input image frames from at least one of the compression module 301, LMCS 302a, DBF 302b, SAO filter 302c, ALF, and CC-ALF 302d. The one or more input image frames can be represented by a first set of dimensions (e.g., HxWxC1) with a specific height and width for the input image pixels and one or more first channels, where H represents the height of the tensor, W represents the width of the tensor, and C represents the number of channels in the tensor. The dimensions can correspond to the specific height, width, and number of channels present in the tensor of the one or more input image frames. The tensor can be structured in a manner that constitutes the pixel values of the one or more image frames in a multi-dimensional array. Thereafter, the one or more input image frames can be passed through one or more DCNNs to improve the quality of the one or more input image frames. The one or more improved image frames can be represented by a second set of dimensions (e.g., HxWxC2) with a specific height and weighted values for the image pixels of the input image signal and one or more second channels. In one embodiment of the present disclosure, the use of the in-loop filter 302e of the AI foundation for quality improvement in a video codec pipeline allows for a wider range of data deformations to be managed, as the AI foundation module can be trained to remove various types of artifacts (e.g., blocking artifacts, ringing artifacts, etc.). In one embodiment of the present disclosure, the in-loop filter 302e of the AI foundation can be designed to prevent various types of compression artifacts. Decisions from such a filter are made within the loop, which also affects decisions for future image frames. The in-loop filter 302e of the AI foundation can transmit one or more improved image frames to the data distribution model 302f for additional processing, as described in connection with FIG. 6.

[0065] FIG. 6 illustrates an exemplary scenario 600 for generating a first representative data point set according to an embodiment of the present disclosure. The improved image frame is received from the in-loop filter 302e of the AI infrastructure.

[0066] In step 601, the in-loop filter 302e of the AI infrastructure may receive one or more input image frames (e.g., reconstructed image frames, compressed image frames) related to the input video from at least one of the compression module 301, LMCS 302a, DBF 302b, SAO filter 302c, ALF, or CC-ALF 302d. In step 602, when receiving one or more input image frames, the in-loop filter 302e of the AI infrastructure may improve the quality of the one or more input image frames. Thereafter, the in-loop filter 302e of the AI infrastructure may transmit one or more improved image frames to the data distribution model 302f for additional processing, as described below.

[0067] In steps 603-604, the data distribution model 302f can receive (or identify) one or more improved image frames from the in-loop filter 302e of the AI infrastructure and can receive (or identify) user input to determine the number of fragments associated with each improved image frame. For example, the user may specify a preference compression deal in the form of a command line argument.

[0068] In stages 605-606, the data distribution model 302f analyzes content changes within each improved image frame by determining one or more optimal fragmented image frames using one or more distribution mechanisms, where each optimal fragmented image frame may include a group of pixels with unique characteristics. In one embodiment of the present disclosure, the data distribution model 302f can analyze content changes within each improved image frame and generate one or more optimal fragments for each improved image frame. The data distribution model 302f can then perform pixel binning for each improved image frame by using one or more statistical mechanisms. The data distribution model 302f can then fragment each improved image frame based on at least one of the received user input, the analyzed content changes, or the pixel binning. For example, each improved image frame may include four fragments or four sections (e.g., fragment-1, fragment-2, fragment-3, etc.), and each section is indicated by a specific band (e.g., band-1).

[0069] In stage 607, the data distribution model 302f generates a first representative data point set for each of the one or more optimal fragmented image frames, where the first representative data point set can represent a vector of scalar values for each of the one or more optimal fragmented image frames. In one embodiment of the present disclosure, the data distribution model 302f can perform pixel binning for each of the one or more fragments within each optimal fragmented image frame to generate a first representative data point set for each optimal fragmented image frame. The first representative data point set for each optimal fragmented image frame can indicate scalar values for each of the one or more fragments within each optimal fragmented image frame.

[0070] In one embodiment, the data distribution model 302f fragments each improved image frame using, for example, a Gaussian Mixture (GM) model, but is not limited to the Gaussian Mixture (GM) model.

[0071] In one embodiment, each fragmented image frame may capture different types of content within each improved image frame. Thus, the method according to one embodiment of the present disclosure can customize the processing for various data deformations because different types of data deformations can generate different types of internal data distributions leading to different types of fragments. Further, since the disclosed method operates on the fly, it can be generalized for all types of input data (e.g., image frames, input videos, etc.) without the required prior training.

[0072] FIG. 7 illustrates an exemplary scenario 700 for generating a second representative data point set according to one embodiment of the present disclosure.

[0073] In steps 701 - 702, the ground truth data distribution model 302g can receive (or identify) one or more input image frames associated with an input video (e.g., the original video).

[0074] In steps 702 - 703, the ground truth data distribution model 302g can receive (or identify) user input to determine the number of fragments for each input image frame. For example, the user can specify the fragment preference in the form of command - line arguments.

[0075] In steps 704 - 705, the ground truth data distribution model 302g analyzes the content changes within each input image frame by determining one or more optimal fragmented input image frames using one or more distribution mechanisms, where each optimal fragmented input image frame may include a group of pixels with unique characteristics. In one embodiment of the present disclosure, the ground truth data distribution model 302g can analyze the content changes within each input image frame and generate one or more optimal fragments for each input image frame. The ground truth data distribution model 302g can then perform pixel binning for each input image frame by using one or more statistical mechanisms. The ground truth data distribution model 302g can then fragment each input image frame based on the received user input, the analyzed content changes, and the pixel binning. For example, each input image frame may include four fragments or four sections (e.g., fragment - 1, fragment - 2, fragment - 3, etc.), and each section is indicated by a specific band (e.g., band - 1).

[0076] In step 706, the ground truth data distribution model 302g generates a second representative data point set for each of the one or more optimal fragmented input image frames, where the second representative data point set can indicate a vector of scalar values associated with each of the one or more optimal fragmented input image frames. In one embodiment of the present disclosure, the ground truth data distribution model 302g can perform pixel binning for each of the one or more fragments within each optimal fragmented input image frame and generate a second representative data point set associated with each optimal fragmented input image frame. The second representative data point set associated with each optimal fragmented input image frame can indicate scalar values associated with each of the one or more fragments within each optimal fragmented input image frame.

[0077] In one embodiment, the ground truth data distribution model 302g can fragment each input image frame using, for example, a GM model, but is not limited to the GM model.

[0078] In one embodiment, one or more statistical mechanisms used by the data distribution model 302f are the same as one or more statistical mechanisms used by the ground truth data distribution model 302g.

[0079] FIG. 8 illustrates an exemplary scenario 800 for generating a pixel mapping for offset scaling for encoding of an AI infrastructure of media according to one embodiment of the present disclosure.

[0080] In one embodiment of the present disclosure, the pixel mapping module 302i for offset scaling receives (or identifies) modeled data distribution information from the data distribution model 302f and the ground truth data distribution model 302g. One or more distribution mechanisms used by both models 302f, 302g are the same, and subsequently, the pixel mapping module 302i for offset scaling compares the outputs of both models 302f, 302g and determines the difference between the outputs of both models (i.e., the data distribution model 302f and the ground truth data distribution model 302g). Fragment information (e.g., fragment-1) along with the representative value of the domain (e.g., the amplitude of the pixel binning curve as illustrated in FIGS. 6 and 7) may already exist in the modeled data distribution. The pixel mapping module 302i for offset scaling can consider one fragment at a particular time instance and calculate the conversion between the fragment domain representative values of the corresponding segments in the ground truth data (e.g., the input image frame, the fragmented input image frame) and the model output data (e.g., the reconstructed image frame, the enhanced image frame, the fragmented image frame).

[0081] In one embodiment of the present disclosure, the pixel mapping module 302i for offset scaling may receive (or identify) mapping range information (a determination as to whether to apply offset scaling to a specific segment (e.g., fragment-1)) from the RDO control mapping range module 302h. In one embodiment of the present disclosure, the pixel mapping module 302i for offset scaling may calculate a mapping range based on an objective function from the RDO control mapping range module 302h. The mapping range information is determined by two functions as illustrated in Equation 4.

[0082]

Number

[0083] ​For example, graph 801 shows the distribution modeling related to the data distribution model 302f, while graph 802 shows the distribution modeling related to the ground truth data distribution model 302g. Graphs 801 and 802 depict the domain information related to Fragment-1. The pixel mapping module 302i for offset scaling can determine the offset value using mathematical operations (e.g., average operation) for offset scaling. The average value of the data shown on the x-y axis of graph 801 is "2", while the average value of the data shown on the x-y axis of graph 802 is "2.5". As a result, the domain difference average value "0.5" is obtained. To overcome the domain difference average value, one or more optimization operations (e.g., addition, multiplication, etc.) are executed, and the pixel values related to fragment-1 can be scaled to reduce distortion. The above-described process is performed for each fragment in the image frame. As a result, the pixel mapping module 302i for offset scaling can generate an offset map 803. The generated offset map 803 can be transmitted to the in-loop filter output module 302j of the final AI infrastructure for additional processing, as will be described in relation to FIG. 9.

[0084] FIG. 9 illustrates an exemplary scenario 900 for generating offset scaling according to an embodiment of the present disclosure.

[0085] The in-loop filter output module 302j of the final AI infrastructure receives (or identifies) the generated offset map 901 from the pixel mapping module 302i for offset scaling and the improved image frame 902 from the in-loop filter 302e of the AI infrastructure. When receiving this information, the in-loop filter output module 302j of the final AI infrastructure can perform offset scaling on the improved image frame 902 based on the generated offset map 901. As a result, the in-loop filter output module 302j of the final AI infrastructure can generate a content-based scaled output.

[0086] The method according to one embodiment of the present disclosure simply requires that segment-by-segment (or fragment-by-fragment) offsets be encoded into the bitstream information. The method according to one embodiment of the present disclosure does not encode information related to the data distribution model 302f and the ground truth data distribution model 302g. In other words, information regarding ad-hoc modeling need not be provided in the bitstream, and thus, the disclosed method does not significantly increase the rate associated with the encoding process.

[0087] FIGS. 10 and 11 illustrate various ways of modeling a domain according to one embodiment of the present disclosure.

[0088] Referring to FIG. 10, in an exemplary scenario as illustrated in FIG. 10, there is a diagram 1000 showing three principal components at the upper end. The diagram 1000 visualizes pixel points in a three-dimensional space. For example, there are 100 points 1001 having a quantization parameter (QP) value (i.e., 42), 500 points 1002 having a QP value, and all points 1003 having a QP value. To calculate representative points (i.e., the first representative data point set and the second representative data point set) using the principal component analysis (PCA) technique, the pixel mapping module 302i for offset scaling can transform the pixel values using the PCA technique. This transformation allows the transformed values to be visualized as clusters. From these clusters (i.e., 1001, 1002, and 1003), the pixel mapping module 302i for offset scaling can extract cluster centroids that serve as representative points for the particular clusters (i.e., 1001, 1002, and 1003). The representative points can capture the essential characteristics of the clusters (i.e., 1001, 1002, and 1003) and be used for additional analysis or decision-making processes. analysis, PCA) technique, the pixel mapping module 302i for offset scaling can transform the pixel values using the PCA technique. This transformation allows the transformed values to be visualized as clusters. From these clusters (i.e., 1001, 1002, and 1003), the pixel mapping module 302i for offset scaling can extract cluster centroids that serve as representative points for the particular clusters (i.e., 1001, 1002, and 1003). The representative points can capture the essential characteristics of the clusters (i.e., 1001, 1002, and 1003) and be used for additional analysis or decision-making processes.

[0089] Referring to FIG. 11, one or more graphs 1100 show, for example, the mean and standard deviation modeling of the pixel difference (referred to as "diff") between the ground truth data distribution model 302g and the data distribution model 302f according to FIG. 8. When the QP value is 22, graphs 1101 and 1102 display the mean and standard deviation modeling of the pixel difference between the original data points (i.e., the second representative data point set) and the codec regenerated data points (i.e., the first representative data point set). When the QP value is 32, graphs 1103 and 1104 display the mean and standard deviation modeling of the pixel difference between the original data points and the codec regenerated data points. When the QP value is 42, graphs 1104 and 1105 display the mean and standard deviation modeling of the pixel difference between the original data points and the codec regenerated data points.

[0090] FIG. 12 illustrates a block diagram of a VVC decoder pipeline 1200 for AI-based encoding of media according to an embodiment of the present disclosure.

[0091] The VVC decoder pipeline 1200 includes a plurality of modules. The plurality of modules may include a CABAC module 1201, an inverse quantization module 1202, an inverse transform module 1203, LMCS 1204, an in-loop filter 1205, an intra prediction module 1206, an inter prediction module 1207, a CIIP module 1208, LMCS 1209, a DPB 1210, and a reconstructed image module 1211. The CABAC module 1201, the inverse quantization module 1202, the inverse transform module 1203, LMCS 1204, the in-loop filter 1205, the intra prediction module 1206, the inter prediction module 1207, the CIIP module 1208, LMCS 1209, the DPB 1210, and the reconstructed image module 1211 are traditional modules.

[0092] The CABAC module 1201 can decompress the compressed bitstream (i.e., the bitstream information received from the VVC encoder pipeline 300) and reconstruct the original frame (i.e., the input video data). a. Bitstream parsing: The CABAC module 1201 can receive the compressed bitstream and parse it to extract the encoded symbols and related information.

[0093] b. Binary arithmetic decoding: The CABAC module 1201 can decode symbols from the compressed bitstream using binary arithmetic decoding. The CABAC module 1201 can assign probabilities to the decoded symbols and use these probabilities to reconstruct the original symbols associated with the encoded symbols.

[0094] c. Context modeling synchronization: The CABAC module 1201 can use context modeling to provide appropriate synchronization between the VVC encoder pipeline 300 and the VVC decoder pipeline 1200.

[0095] Furthermore, the inverse quantization module 1202 and the inverse transform module 1203 are configured to convert the quantized and transformed video data back to its original form. The LMCS module 1204 is used to adjust the chroma and luma components of the video data to improve the overall image quality.

[0096] Furthermore, the in-loop filter 1205 can receive data from the LMCS module 1204. The in-loop filter 1205 is configured to remove artifacts and noise that may occur during the compression operation. Examples of the in-loop filter 1205 can include LMCS 1205a, DBF 1205b, SAO filter 1205c, ALF, and CC-ALF 1205d. The functions of the various blocks 1205a, 1205b, 1205c, and 1205d are the same as the description in FIG. 2A, and these various blocks are traditional modules.

[0097] Furthermore, the intra prediction module 1206 and the inter prediction module 1207 are configured to predict the values of data based on spatial and temporal correlations. The CIIP module 1208 can combine both the intra prediction technique and the inter prediction technique to further improve the compression efficiency. Finally, the DPB 1210 and the reconstructed image module can work together to store and display the decoded video data. The DPB 1210 stores the previously decoded frames to enable efficient inter-frame prediction, while the reconstructed image module 1211 generates a high-quality display image from the decoded data (e.g., the image frames from the in-loop filter 1205).

[0098] In one or more embodiments of the present disclosure, the in-loop filter 1205 can include an AI-based in-loop filter 1205e, a data distribution model 1205f, a pixel mapping module 1205g for offset scaling, and a final AI-based in-loop filter output module 1205h.

[0099] In one or more embodiments of the present disclosure, the AI-based in-loop filter 1205e receives bitstream information related to the video reconstructed from the VVC encoder pipeline 300 (see FIG. 3), where the bitstream information includes an offset value. The AI-based in-loop filter 1205e can perform various functions related to the AI-based in-loop filter 302e based on the received bitstream information.

[0100] In one or more embodiments of the present disclosure, the AI-based in-loop filter 1205e can receive (or identify) a reconstructed image frame related to the video reconstructed from at least one of LMCS 1204, LMCS 1205a, deblocking filter 1205b, SAO 1205c or ALF, CC-ALF 1205d. The AI-based in-loop filter 1205e can improve the quality of the reconstructed image frame by removing artifacts from the reconstructed image frame.

[0101] In one or more embodiments of the present disclosure, the data distribution model 1205f can generate a model output data distribution related to the reconstructed video. The data distribution model 1205f can generate a model output data distribution related to the improved image frame from the AI-based in-loop filter 1205e.

[0102] In one or more embodiments of the present disclosure, the pixel mapping module 1205g for offset scaling can perform a scaling operation on each pixel value related to each fragmented image frame of the reconstructed video by using the offset value and the generated model output data distribution. In one embodiment of the present disclosure, the pixel mapping module 1205g for offset scaling can identify offset signaling from the bitstream and generate an offset map based on the offset signaling.

[0103] In one or more embodiments of the present disclosure, the in-loop filter output module 1205h of the final AI infrastructure may generate an output video based on the performed scaling operation. In one embodiment of the present disclosure, the in-loop filter output module 1205h of the final AI infrastructure may perform an offset scaling operation on the improved image frame from the in-loop filter 1205e of the AI infrastructure based on the generated offset map (or offset value, offset signaling).

[0104] FIG. 13A illustrates a block diagram of an electronic device for AI infrastructure encoding of media according to one embodiment of the present disclosure. In one or more embodiments of the present disclosure, the electronic device 100 may include one or more modules shown in FIG. 3. Examples of the electronic device 100 may include, but are not limited to, smartphones, tablet computers, personal digital assistants (PDAs), Internet of Things (IoT) devices, and wearable devices.

[0105] In one embodiment of the present disclosure, the electronic device 100 includes a system 101. The system 101 may include a memory 110, a processor 120, and a communication unit 130.

[0106] In one embodiment of the present disclosure, the memory 110 stores instruction words executed by the processor 120 for AI infrastructure encoding of media, as described throughout the present disclosure. The memory 110 may include non-volatile storage elements. Examples of such non-volatile storage elements include magnetic hard disks, optical disks, floppy disks, flash memories, or EPROM (electrically programmable memories) or EEPROM (electrically erasable and programmable It may include forms of [[memories]]. Additionally, in some embodiments, the memory 110 can be regarded as a non-transitory recording medium. The term "non-transitory" may indicate that the recording medium is not embodied as a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted to mean that the memory 110 is non-portable. In some examples, the memory 110 can be configured to store a larger amount of information than a memory. In a specific example, a non-transitory recording medium can store data that changes over time (e.g., in a RAM (Random Access Memory) or a cache). The memory 110 can be an internal storage unit or an external storage unit of the electronic device 100, a cloud storage, or any other type of external storage.

[0107] The processor 120 communicates with the memory 110 and the communication unit 130. The processor 120 is configured to execute the instruction words stored in the memory 110 and to perform various processes for the encoding of the AI base of the media as described throughout this disclosure. The processor 120 includes one or more processors and can be a general-purpose processor, a so-called central processing unit, CPU, an application processor processor, AP, etc., a graphics dedicated processing unit, a so-called graphics processing unit, GPU, a visual processing unit, VPU, and / or an artificial intelligence (Artificial Intelligence, AI) dedicated processor, so to speak, a neural processing unit, NPU.

[0108] The communication unit 130 is configured to communicate internally among the internal hardware components and with external devices (such as servers) through one or more networks (for example, radio technologies). The communication unit 130 includes electronic circuits specialized for standards enabling wired or wireless communication.

[0109] The processor 120 is implemented by a processing circuit such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, and can be optionally driven by firmware. The circuit can be implemented, for example, within one or more semiconductor chips or on a substrate support such as a printed circuit board.

[0110] In one embodiment of the present disclosure, the processor 120 includes an image compression unit 121, an in-loop filter 122 of the AI base, a data distribution model 123, a ground-truth data distribution model 124, an RDO control mapping range module 125, a pixel mapping module 126 for offset scaling, and an optimized media generation unit 127.

[0111] In one embodiment of the present disclosure, the image compression unit 121 is related to the compression module 301, the in-loop filter 122 of the AI base is related to the in-loop filter 302e of the AI base, the data distribution model 123 is related to the data distribution model 302f, the ground-truth data distribution model 124 is related to the ground-truth data distribution model 302g, the RDO control mapping range module 125 is related to the RDO control mapping context module 302h, the pixel mapping module 126 for offset scaling is related to the pixel mapping module 302i for offset scaling, and the optimized media generation unit 127 is related to the output module 302j of the final in-loop filter of the AI base.

[0112] The image compression unit 121 is a traditional module, which can receive an input video. When receiving the input video, the image compression unit 121 can perform one or more predefined compression operations such as conversion, quantization, inverse quantization, and inverse conversion on the received input video. The in-loop filter 122 of the AI base can be configured to apply AI techniques in the loop filtering process of video coding. The in-loop filter 122 of the AI base can improve the performance of the in-loop filters 302a, 302b, 302c, 302d. The in-loop filter 122 of the AI base can utilize a machine running model to learn the characteristics of video content (such as the input video, etc.) and apply an adaptive filtering technique. Instead of relying on predefined rules or heuristics, the in-loop filter 122 of the AI base can analyze and process video content using a neural network or other AI methods. The in-loop filter 122 of the AI base uses an AI model, and the AI model can learn from a large amount of training data and make intelligent decisions regarding the filtering operation. The in-loop filter 122 of the AI base generates one or more reconstructed image frames and transmits one or more reconstructed image frames to the data distribution model 123 for additional processing, which will be described in detail below.

[0113] As described in connection with FIG. 6, the data distribution model 123 can generate a first representative data point set. The data distribution model 123 can perform various steps given below and generate a first representative data point set.

[0114] The data distribution model 123 receives one or more reconstructed image frames (e.g., one or more enhanced image frames) from the in-loop filter 122 of the AI infrastructure, where the one or more reconstructed image frames relate to the reconstructed video. The data distribution model 123 may then receive user input indicating the desired compression level for each of the one or more reconstructed image frames. The data distribution model 123 may then determine the number of segments for each of the one or more reconstructed image frames based on the received user input. The data distribution model 123 may then fragment each of the one or more reconstructed image frames based on the determined number of segments. The data distribution model 123 may then analyze the content changes within each fragmented image frame using one or more distribution mechanisms. The data distribution model 123 may then perform a pixel binning operation on the analyzed content changes using one or more statistical mechanisms to determine one or more optimal fragmented image frames, where each of the one or more optimal fragmented image frames includes a group of pixels with unique characteristics. The data distribution model 123 may then generate a first representative data point set for each of the one or more optimal fragmented image frames.

[0115] The ground truth data distribution model 124 may generate a second representative data point set, as described in connection with FIG. 7. The ground truth data distribution model 124 may perform various steps given below to generate the second representative data point set.

[0116] The ground truth data distribution model 124 may receive one or more input image frames associated with the input video. The ground truth data distribution model 124 may then receive user input indicating a desired compression level for each of the one or more input image frames. The ground truth data distribution model 124 may then determine the number of segments for each of the one or more input image frames based on the received user input. The ground truth data distribution model 124 may then fragment each of the one or more input image frames based on the determined number of segments. The ground truth data distribution model 124 may then analyze the content changes within each fragmented input image frame using one or more distribution mechanisms. The ground truth data distribution model 124 may then perform a pixel binning operation on the analyzed content changes using one or more statistical mechanisms to determine one or more optimal fragmented image frames, where each of the one or more optimal fragmented image frames includes a group of pixels with unique characteristics. The ground truth data distribution model 124 may then generate a second representative data point set for each of the one or more optimal fragmented image frames.

[0117] As described with reference to FIG. 4, the RDO control mapping range module 125 may determine the overall RD cost based on the output of the RD cost module 304 and user input regarding the number of segments. The RDO control mapping range module 125 may further construct an objective function to minimize the overall RD cost.

[0118] The pixel mapping module 126 for offset scaling can determine a data distribution dissimilarity metric between a first representative data point set associated with each fragmented image frame related to the reconstructed video and a second representative data point set associated with each fragmented image frame related to the input video, as described in FIG. 8. The pixel mapping module 126 for offset scaling can then further determine an offset value associated with each fragmented image frame by using the data distribution dissimilarity metric, as described in FIG. 8.

[0119] The optimized media generation unit 127 can perform a scaling operation on each pixel value associated with each fragmented image frame by using the determined offset values, as described in FIG. 9. The optimized media generation unit 127 can further encode the reconstructed video based on the performed scaling operation.

[0120] The functions associated with the various components of the electronic device 100 can be performed through non-volatile memory, volatile memory, and the processor 120. One or more processors control the processing of input data according to predetermined operation rules or AI models stored in the non-volatile memory and volatile memory. The predetermined operation rules or AI models are provided through training or learning. Being provided through learning means that a predetermined operation rule or AI model with desired characteristics is created by applying a learning algorithm to a plurality of learning data. Learning can be performed by the device itself on which the AI according to the embodiment is performed and / or can be embodied through a separate server / system. The learning algorithm is a method of training a predetermined target device (e.g., a robot) using a plurality of learning data in order to induce, allow, or control what the target device determines or predicts. Examples of learning methods include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0121] An AI model can be composed of multiple neural network layers. Each layer has values of multiple weights and performs layer operations through calculations of previous layers and operations of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann Machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks), and deep Q networks.

[0122] Although FIG. 13A illustrates various hardware components of the electronic device 100, it should be understood that one or more embodiments of the present disclosure are not limited thereto. In one or more embodiments of the present disclosure, the electronic device 100 may include fewer or more components. Further, the labels or names of the components are used for illustrative purposes only and do not limit the scope of the invention. One or more components may be combined to perform the same function or substantially the same function for encoding the AI foundation of the media.

[0123] FIG. 13B illustrates a block diagram of an electronic device for AI foundation decoding of media according to an embodiment of the present disclosure. In one or more embodiments of the present disclosure, the electronic device 1300 includes one or more modules shown in FIG. 12. Examples of the electronic device 1300 may include, but are not limited to, smartphones, tablet computers, PDAs, IoT devices, wearable devices, and the like.

[0124] In one embodiment, the electronic device 1300 includes a system 1301. The system 1301 may include a memory 1310, a processor 1320, and a communication unit 1330.

[0125] In one embodiment, the memory 1310 stores instruction words to be executed by the processor 1320 for AI-based decoding of media, as described throughout this disclosure. The memory 1310 may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard disks, optical disks, floppy disks, flash memories, or may include the behavior of EPROM or EEPROM memories. Additionally, the memory 1310 may be regarded as a non-transitory recording medium in some embodiments. The term "non-transitory" may indicate that the recording medium is not embodied as a carrier wave or a propagated signal. However, the term "non-transitory" should not be construed to mean that the memory 1310 is non-portable. In some examples, the memory 1310 may be configured to store a larger amount of information than a memory. In a particular example, the non-transitory recording medium may store data that can change over time (e.g., in a RAM or a cache). The memory 1310 may be an internal storage unit or an external storage unit of the electronic device 1300, a cloud storage, or any other type of external storage.

[0126] The processor 1320 communicates with the memory 1310 and the communication unit 1330. The processor 1320 is configured to execute the instruction words stored in the memory 1310 and to perform various processes for encoding-decoding of the AI-based media, as described throughout this disclosure. The processor 1320 includes one or more processors and may be a general-purpose processor, a so-called central processing unit (CPU), an application processor (AP), etc., a graphics dedicated processing unit, a so-called graphics processing unit (GPU), a vision processing unit (VPU), and / or an artificial intelligence (AI) dedicated processor, a so-called neural processing unit (NPU).

[0127] The communication unit 1330 is configured to communicate internally among internal hardware components and with external devices (e.g., servers) through one or more networks (e.g., radio technologies). The communication unit 1330 includes electronic circuits specialized for standards enabling wired or wireless communication.

[0128] The processor 1320 is implemented by processing circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hard-wired circuits, and can be optionally driven by firmware. The circuit can be implemented, for example, within one or more semiconductor chips or on a substrate support such as a printed circuit board.

[0129] In one embodiment of the present disclosure, the processor 1320 includes an AI-based in-loop filter 1322, a data distribution model 1324, a pixel mapping module 1326 for offset scaling, and an optimized media generation unit 1328.

[0130] In one embodiment of the present disclosure, the AI-based in-loop filter 1322 is related to the AI-based in-loop filter 1205e, the data distribution model 1324 is related to the data distribution model 1205f, the pixel mapping module 1326 for offset scaling is related to the pixel mapping module 1205g for offset scaling, and the optimized media generation unit 1328 is related to the final AI-based in-loop filter output module 1205h.

[0131] The in-loop filter 1322 based on AI can be configured to apply AI techniques in the loop filtering process of video decoding. The in-loop filter 1322 based on AI can improve the performance of the in-loop filters 302a, 302b, 302c, 302d. The in-loop filter 1322 based on AI can utilize a machine running model to learn the characteristics of video content (such as input video, etc.) and apply an adaptive filtering technique. Instead of relying on predefined rules or heuristics, the in-loop filter 1322 based on AI can analyze and process video content using a neural network or other AI methods. The in-loop filter 1322 based on AI uses an AI model, and the AI model can learn from a large amount of training data and make intelligent decisions regarding the filtering operation. The in-loop filter 1322 based on AI generates one or more reconstructed image frames and transmits the one or more reconstructed image frames to the data distribution model 1324 for additional processing, which will be described in detail below.

[0132] As described in FIG. 6, the data distribution model 1324 can generate a first representative data point set. The data distribution model 123 can perform various steps given below to generate a first representative data point set.

[0133] The data distribution model 1324 receives one or more reconstructed image frames (e.g., one or more enhanced image frames) from the in-loop filter 1322 of the AI infrastructure, where the one or more reconstructed image frames pertain to the reconstructed video. The data distribution model 1324 may then receive user input indicating the desired compression level for each of the one or more reconstructed image frames. The data distribution model 1324 may then determine the number of segments for each of the one or more reconstructed image frames based on the received user input. The data distribution model 1324 may then fragment each of the one or more reconstructed image frames based on the determined number of segments. The data distribution model 1324 may then analyze the content changes within each fragmented image frame using one or more distribution mechanisms. The data distribution model 1324 may then perform a pixel binning operation on the analyzed content changes using one or more statistical mechanisms to determine one or more optimal fragmented image frames, where each of the one or more optimal fragmented image frames includes a group of pixels with unique characteristics. The data distribution model 1324 may then generate a first representative data point set for each of the one or more optimal fragmented image frames.

[0134] The pixel mapping module 1326 for offset scaling may perform pixel mapping for scaling operations based on offset information included in the bitstream information from the encoder, as illustrated in FIG. 8.

[0135] The optimized media generation unit 1328 can perform a scaling operation on the image frame reconstructed based on the pixel mapping to generate a scaled image frame. As described in FIG. 9, the optimized media generation unit 1328 can perform a scaling operation on each pixel value associated with each fragmented image frame by using offset information (e.g., values). The optimized media generation unit 1328 can generate an output video based on the scaled image frame.

[0136] Furthermore, the optimized media generation unit 1328 receives bitstream information related to the reconstructed video from an encoder (i.e., the VVC encoder pipeline 300), where the bitstream information includes an offset value. The optimized media generation unit 1328 can generate a model output data distribution related to the reconstructed video by using the offset value with the data distribution model 1205f. The optimized media generation unit 1328 can then perform a scaling operation on each pixel value associated with each fragmented image frame of the reconstructed video by using the offset value and the generated model output data distribution with the pixel mapping module 1205g for offset scaling. The optimized media generation unit 1328 can then generate an output video based on the scaling operation performed by using the final in-loop filtering output module 1205h of the AI infrastructure.

[0137] The functions associated with the various components of the electronic device 100 can be performed through a non-volatile memory, a volatile memory, and the processor 120. One or more processors control the processing of input data according to predetermined operating rules or AI models stored in the non-volatile memory and the volatile memory. The predetermined operating rules or AI models are provided through training or learning. Providing through learning means that a predetermined operating rule or AI model with desired characteristics is created by applying a learning algorithm to a plurality of learning data. Learning can be performed on the device itself where the AI according to the embodiment is executed, and / or can be implemented through a separate server / system. The learning algorithm is a method of training a predetermined target device (e.g., a robot) using a plurality of learning data in order to induce, allow, or control what the target device determines or predicts. Examples of learning methods include, but are not limited to, supervised (teacher) learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0138] The AI model can be composed of a plurality of neural network layers. Each layer has values of a plurality of weight values and performs layer operations through calculations of the previous layer and operations of the plurality of weight values. Examples of neural networks include, but are not limited to, CNN, DNN, RNN, RBM, DBN, BRDNN, GAN, and deep Q-networks.

[0139] Although FIG. 13B illustrates various hardware components of the electronic device 1300, it should be understood that one or more embodiments of the present disclosure are not limited thereto. In one or more embodiments of the present disclosure, the electronic device 100 may further include fewer components or further include more components. Furthermore, the labels or names of the components are used only for illustrative purposes and do not limit the scope of the invention. One or more components can be combined to perform the same function or substantially the same function for encoding-decoding of the AI base of the media.

[0140] FIG. 14 is a flowchart illustrating a method 1400 performed by an electronic device for encoding an AI infrastructure of media according to an embodiment of the present disclosure.

[0141] In step 1410, method 1400 includes compressing an input image associated with an input video. In one embodiment of the present disclosure, method 1400 includes compressing an input video for feeding into an in-loop filter of the AI infrastructure, which may be related to at least one of the compression module 301, in-loop filter 302 (e.g., LMCS 302a, deblocking filter 302b, SAO 302c, ALF, CC-ALF 302d), CABAC module 303, or RD cost module 304 of FIG. 3.

[0142] In step 1420, method 1400 includes generating a reconstructed image frame corresponding to the input image frame using the in-loop filter 302e of the AI infrastructure, which may be related to the output of the in-loop filter 302e of the AI infrastructure of FIG. 3.

[0143] In stage 1430, method 1400 includes generating an offset value based on an input image frame and a reconstructed image frame. In one embodiment of the present disclosure, method 1400 includes generating a model output data distribution related to the reconstructed video, which may be related to data distribution model 302f in FIG. 3. Method 1400 includes various processes for generating a model output data distribution related to the reconstructed video, which are listed below. Method 1400 includes receiving one or more reconstructed image frames from an in-loop filter of an AI platform, where the one or more reconstructed image frames are related to the reconstructed video. The one or more reconstructed image frames are generated by using one or more NN models. Method 1400 further includes receiving user input indicating a desired compression level for each of the one or more reconstructed image frames. Method 1400 further includes determining the number of segments for each of the one or more reconstructed image frames based on the received user input. Method 1400 further includes fragmenting each of the one or more reconstructed image frames based on the determined number of segments. Method 1400 further includes analyzing content changes within each fragmented image frame using one or more distribution mechanisms. Method 1400 further includes performing a pixel binning operation using one or more statistical mechanisms on the content changes analyzed to determine one or more optimal fragmented image frames, where each of the one or more optimal fragmented image frames includes a group of pixels with unique characteristics. Method 1400 further includes generating a first representative data point set for each of the one or more optimal fragmented image frames.

[0144] In one embodiment of the present disclosure, method 1400 includes generating a ground truth data distribution for an input video, which may relate to the ground truth data distribution model 302g of FIG. 3. Method 1400 includes various processes for generating a ground truth data distribution for the input video, which are listed below. Method 1400 includes receiving one or more input image frames associated with the input video. Method 1400 further includes receiving user input indicating a desired compression level for each of the one or more input image frames. Method 1400 further includes determining, based on the received user input, the number of segments for each of the one or more input image frames. Method 1400 further includes fragmenting each of the one or more input image frames based on the determined number of segments. Method 1400 further includes analyzing content changes within each fragmented input image frame using one or more distribution mechanisms. Method 1400 further includes performing a pixel binning operation using one or more statistical mechanisms on the content changes analyzed to determine one or more optimal fragmented image frames, each of the one or more optimal fragmented image frames including a group of pixels having unique characteristics. Method 1400 further includes generating a second representative data point set for each of the one or more optimal fragmented image frames.

[0145] In one embodiment of the present disclosure, method 1400 includes determining an offset value by comparing a model output data distribution and a ground truth data distribution, which may be related to the pixel mapping module 302i for offset scaling in FIG. 3. Method 1400 includes various processes for determining the offset value, which are listed below. Method 1400 includes determining a data distribution dissimilarity metric between a first representative data point set associated with each fragmented image frame related to the reconstructed video and a second representative data point set associated with each fragmented image frame related to the input video. Method 1400 further includes determining an offset value associated with each fragmented image frame by using the data distribution dissimilarity metric.

[0146] In step 1440, method 1400 includes encoding a reconstructed image frame based on the determined offset value. In one embodiment of the present disclosure, method 1400 includes encoding the reconstructed video by using the determined offset value, which may be related to at least one of the in-loop filter output module 302j, the in-loop filter 302 (e.g., LMCS 302a, deblocking filter 302b, SAO 302c, ALF, CC-ALF 302d), the CABAC module 303, or the RD cost module 304 of the final AI base in FIG. 3. Method 1400 includes various processes for encoding the reconstructed video, which are listed hereinafter. Method 1400 includes performing a scaling operation on each pixel value associated with each fragmented image frame by using the determined offset value, where the scaling operation includes at least one of an addition operation, a multiplication operation, a division operation, or an exponential operation. Method 1400 further includes encoding the reconstructed video based on the performed scaling operation.

[0147] In one embodiment of the present disclosure, method 1400 includes transmitting bitstream information associated with the reconstructed video to a decoder, where the bitstream information includes a determined offset value.

[0148] FIG. 15 is a flowchart illustrating a method 1500 for AI-based compression of media according to one embodiment of the present disclosure. Method 1500 may be related to method 1400.

[0149] In step 1510, method 1500 includes determining a desired compression level for the input media, where the desired compression level is indicated by the user.

[0150] In step 1520, method 1500 includes compressing the input media based on the desired compression level, which may be related to steps 1410 and 1420 of FIG. 14.

[0151] In step 1530, method 1500 includes determining an offset value for the compressed media relative to the input media by comparing the compressed media with the input media and the desired compression level for the input media, which may be related to step 1430 of FIG. 14. The offset value is calculated and used at the fragment-by-fragment granularity, where the number of fragments is received as input from the user, and each fragment is calculated during the data distribution modeling operation. Method 1500 includes various processes for determining the offset value, which are listed hereinafter. Method 1500 includes modeling the model output data distribution. Method 1500 further includes modeling the ground truth data distribution. Method 1500 further includes determining a pixel mapping using the modeled model output data distribution, the modeled ground truth data distribution, and a mapping range. The mapping range is a determination of whether to apply offset scaling for a particular fragment based on the RD cost, and the mapping range is determined using the user-defined number of fragments and the codec RD cost.

[0152] In stage 1540, method 1500 includes the stage of obtaining optimized compressed media by adding the determined offset value to the compressed media, which may relate to stage 1440 of FIG. 14.

[0153] FIG. 16 is a flowchart illustrating a method 1600 for AI-based decoding of media according to an embodiment of the present disclosure.

[0154] In stage 1610, method 1600 includes the stage of receiving bitstream information including offset information from an encoder. In one embodiment of the present disclosure, method 1600 includes the stage of receiving bitstream information related to the reconstructed video from an encoder, where the bitstream information includes an offset value.

[0155] In stage 1620, method 1600 includes the stage of generating a reconstructed image frame based on the bitstream information using an AI-based in-loop filter, which may relate to at least one of CABAC 1201, inverse quantization module 1202, inverse transform module 1203, LMCS module 1204, or in-loop filter 1205 of FIG. 12. In one embodiment of the present disclosure, method 1600 includes the model output data distribution related to the reconstructed video by using the offset value, which may relate to stage 1420 of FIG. 14.

[0156] In stage 1630, method 1600 includes performing a scaling operation on the reconfigured image frame based on the offset information to generate a scaled image frame. In one embodiment of the present disclosure, method 1600 includes performing a scaling operation on each pixel value associated with each fragmented image frame of the reconfigured video by using the offset value and the generated model output data distribution, which may be related to at least one of stage 1420 or stage 1430 of FIG. 14. The scaling operation includes at least one of an addition operation, a multiplication operation, a division operation, or an exponentiation operation.

[0157] In stage 1640, method 1600 includes generating an output video based on the scaled image frame, which may be related to at least one of the in-loop filters of FIG. 12 (e.g., LMCS 1205a, deblocking filter 1205b, SAO 1205c, ALF, CC-ALF 1205d), or the reconfigured image module 1211. In one embodiment of the present disclosure, method 1600 may include generating an output video based on the performed scaling operation.

[0158] In one or more embodiments of the present disclosure, the bitstream information may include various tunable parameters. The parameters may include a tunable enhancement flag, a tunable number of classes, and a tunable class offset for offset scaling. For example, the tunable enhancement flag is binary. When the tunable enhancement flag is set to 0, this indicates that the bitstream information does not include syntax elements related to tunable offsets. As a result, these syntax elements cannot be used in the reconstruction process (see FIG. 12). The tunable number of classes is another parameter that can be adjusted based on the user's requirements to determine the number of classes considered during processing (e.g., decoding). Further, the tunable class offset parameter specifies a particular offset applied to each pixel class. These offsets may modify the characteristics of the pixels within each class. By tuning these parameters within the bitstream information, the disclosed method can be applied to different scenarios and user preferences, customize the data, and enable flexible processing.

[0159] In one or more embodiments of the present disclosure, the disclosed method provides a unique strategy for an AI-based model using additional real-time modules that are newer and adjusted to unseen data to reduce data domain gaps, as described in FIGS. 3 through 16. The disclosed method uses pixel mapping and content-based scaling to further fill the data domain gap between newer / unseen data and training data in the in-loop filter of the AI base. As a result, the disclosed method is superior to existing methods that introduce bias and variability inherent in the AI base training model by using well-trained restricted data targeted at specific artifacts. The disclosed method enables content-based scaling to identify pixel mapping for scaling offset values, operates on the fly, requires no pre-training, and can provide more generalization for different types of data. Further, the disclosed method has a very low complexity when compared to existing methods, while existing methods deploy very complex DNN models to achieve a higher gain, but the method is not adaptable to the content base. The disclosed method solves this problem by introducing video data-based statistical regression during the encoding process and using the same regression during the decoding process.

[0160] Unlike existing methods that rely on restricted data targeting specific artifacts, the disclosed method uses content-based scaling to fill data domain gaps and provides more generalization for different types of data. One of the core advantages of the disclosed method is the ability to identify pixel mappings for scaling offsets without requiring pre-training. This allows for immediate adaptation, which is highly important in dynamic environments where data continues to change. Also, the disclosed method has a much lower complexity compared to existing methods that rely on complex DNN models to achieve higher gains. Additionally, the disclosed method resolves bias and variability issues in AI-based trained models by introducing video data-based statistical regression during the encoding and decoding processes. This ensures that the disclosed method is content-based adaptive and provides more accurate predictions. Generally, the disclosed method provides a powerful tool for improving AI-based models and solving the conductive challenges imposed by new and unseen data. An electronic device 100 having the ability to perform content-based scaling, fill data domain gaps, and provide more generalization becomes an essential tool for organizations seeking to utilize AI technology in many operations (e.g., encoding, decoding, etc.).

[0161] In one or more embodiments of the present disclosure, the disclosed method has the potential to improve the efficiency of the encoder and decoder in the electronic device 100, thereby reducing video recording and memory requirements from the electronic device 100. Further, the disclosed method enables excellent video quality even in bandwidth-restricted scenarios. The disclosed method can greatly improve the overall user experience and satisfaction. Such advantages are particularly relevant in today's world where video content consumption is generalized and bandwidth restrictions are prevalent.

[0162] In one or more embodiments of the present disclosure, the disclosed method may improve the BD rate (Bjontegaard-delta rate). The BD rate measures the overall compression efficiency of the entire pipeline. Further, the disclosed method improves visual quality by reducing distortion while keeping the bandwidth requirement constant. Further, the disclosed method may reduce the bitstream or bandwidth requirement while maintaining similar visual quality. Further, the disclosed method may improve visual quality and reduce the bandwidth requirement. These features make the disclosed method a valuable asset in improving video streaming services.

[0163] In one embodiment of the present disclosure, a method for encoding an AI-based media may include compressing an input image frame associated with an input video. In one embodiment of the present disclosure, the method may include generating a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI-based media. In one embodiment of the present disclosure, the method may include determining an offset value based on the input image frame and the reconstructed image frame. In one embodiment of the present disclosure, the method may include encoding the reconstructed image frame based on the determined offset value.

[0164] In one embodiment of the present disclosure, the step of determining an offset value based on the input image frame and the reconstructed image frame may include generating a model output data distribution related to the reconstructed image frame. In one embodiment of the present disclosure, the step of determining an offset value based on the input image frame and the reconstructed image frame may include generating a ground truth data distribution related to the input image frame. In one embodiment of the present disclosure, the step of determining an offset value based on the input image frame and the reconstructed image frame may include determining the offset value by comparing the model output data distribution and the ground truth data distribution.

[0165] In one embodiment of the present disclosure, the step of generating a model output data distribution related to a reconstructed image frame may include the step of identifying the number of fragments related to the reconstructed image frame based on user input. In one embodiment of the present disclosure, the step of generating a model output data distribution related to a reconstructed image frame may include the step of analyzing content changes within the reconstructed image frame using one or more distribution mechanisms. In one embodiment of the present disclosure, the step of generating a model output data distribution related to a reconstructed image frame may include the step of fragmenting the reconstructed image frame based on the number of fragments determined and the content changes analyzed to generate an optimal fragmented image frame. In one embodiment of the present disclosure, the step of generating a model output data distribution related to a reconstructed image frame may include the step of performing a pixel binning operation on one or more fragments of the optimal fragmented image frame using one or more statistical mechanisms. In one embodiment of the present disclosure, the optimal fragmented image frame may include a group of pixels having unique characteristics. In one embodiment of the present disclosure, the step of generating a model output data distribution related to a reconstructed image frame may include the step of generating a first representative data point set related to the optimal fragmented image frame based on the pixel binning operation.

[0166] In one embodiment of the present disclosure, the step of generating the ground truth data distribution related to the input image frame may include the step of identifying the number of fragments related to the input image frame based on user input. In one embodiment of the present disclosure, the step of generating the ground truth data distribution related to the input image frame may include the step of analyzing the content change within the input image frame using one or more distribution mechanisms. In one embodiment of the present disclosure, the step of generating the ground truth data distribution related to the input image frame may include the step of fragmenting the input image frame based on the determined number of fragments and the analyzed content change to generate an optimal fragmented input image frame. In one embodiment of the present disclosure, the step of generating the ground truth data distribution related to the input image frame may include the step of performing a pixel binning operation on one or more fragments of the optimal fragmented input image frame using one or more statistical mechanisms. In one embodiment of the present disclosure, the optimal fragmented input image frame may include a group of pixels with unique characteristics. In one embodiment of the present disclosure, the step of generating the ground truth data distribution related to the input image frame may include the step of generating a second representative data point set related to the optimal fragmented input image frame based on the pixel binning operation.

[0167] In one embodiment of the present disclosure, the step of determining the offset value by comparing the model output data distribution and the ground truth data distribution may include the step of determining a data distribution dissimilarity metric between a first representative data point set related to the fragmented image frame from the reconstructed image frame and a second representative data point set related to the fragmented input image frame from the input image frame. In one embodiment of the present disclosure, the step of determining the offset value by comparing the model output data distribution and the ground truth data distribution may include the step of determining the offset value based on the data distribution dissimilarity metric.

[0168] In one embodiment of the present disclosure, the step of determining the offset value by comparing the model output data distribution and the ground truth data distribution may include the step of determining the pixel mapping in the form of a fragment-by-fragment offset using the model output data distribution, the ground truth data distribution, and the mapping range.

[0169] In one embodiment of the present disclosure, the mapping range is also a determination of whether to apply offset scaling to a specific fragment based on the RD cost. In one embodiment of the present disclosure, the mapping range can be determined based on the number of fragments and the codec RD cost.

[0170] In one embodiment of the present disclosure, the step of encoding the reconstructed image frame may include the step of performing a scaling operation on the reconstructed image frame based on the offset value determined to generate the scaled image frame. In one embodiment of the present disclosure, the scaling operation may include at least one of an addition operation, a multiplication operation, a division operation, or an exponential operation. In one embodiment of the present disclosure, the encoding of the reconstructed image frame may include the step of encoding the scaled image frame.

[0171] In one embodiment of the present disclosure, the reconstructed image frame can be generated by using one or more neural network (NN) models of the in-loop filter of the AI platform.

[0172] In one embodiment of the present disclosure, the method may include the step of transmitting the bitstream information associated with the reconstructed image frame to the decoder. In one embodiment of the present disclosure, the bitstream information may include the determined offset value.

[0173] In one embodiment of the present disclosure, the offset value can be calculated and used at the fragment-by-fragment granularity.

[0174] In one embodiment of the present disclosure, a method for AI-based decoding of media may include receiving bitstream information including offset information from an encoder. In one embodiment of the present disclosure, the method may include generating a reconstructed image frame based on the bitstream information using an in-loop filter of the AI base. In one embodiment of the present disclosure, the method may include performing a scaling operation on the reconstructed image frame based on the offset information to generate a scaled image frame. In one embodiment of the present disclosure, the method may include generating an output video based on the scaled image frame.

[0175] In one embodiment of the present disclosure, the step of performing a scaling operation on the reconstructed image frame may include generating a model output data distribution related to the reconstructed image frame. In one embodiment of the present disclosure, the step of performing a scaling operation on the reconstructed image frame may include performing a pixel mapping for the scaling operation based on the offset information and the model output data distribution. In one embodiment of the present disclosure, the step of performing a scaling operation on the reconstructed image frame may include performing a scaling operation on the reconstructed image frame based on the pixel mapping to generate a scaled image frame.

[0176] In one embodiment of the present disclosure, the scaling operation may include at least one of an addition operation, a multiplication operation, a division operation, or an exponentiation operation.

[0177] A system for encoding an AI infrastructure of media may include a processor operably coupled to a memory and a communication unit. In one embodiment of the present disclosure, the processor may be configured to compress an input image frame associated with an input video. In one embodiment of the present disclosure, the processor may be configured to generate a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI infrastructure. In one embodiment of the present disclosure, the processor may be configured to determine an offset value based on the input image frame and the reconstructed image frame. In one embodiment of the present disclosure, the processor may be configured to encode the reconstructed image frame based on the determined offset value.

[0178] In one embodiment of the present disclosure, the processor may be configured to generate a model output data distribution related to the reconstructed image frame. In one embodiment of the present disclosure, the processor may be configured to generate a ground truth data distribution related to the input image frame. In one embodiment of the present disclosure, the processor may be configured to determine an offset value by comparing the model output data distribution and the ground truth data distribution.

[0179] In one embodiment of the present disclosure, the processor may be configured to identify the number of fragments related to the image frame reconstructed based on user input. In one embodiment of the present disclosure, the processor may be configured to analyze the content change within the reconstructed image frame using one or more distribution mechanisms. In one embodiment of the present disclosure, the processor may be configured to fragment the reconstructed image frame based on the number of fragments determined and the analyzed content change to generate an optimal fragmented image frame. In one embodiment of the present disclosure, the processor may be configured to perform a pixel binning operation on one or more fragments of the optimal fragmented image frame using one or more statistical mechanisms. In one embodiment of the present disclosure, the optimal fragmented image frame may include a group of pixels having unique characteristics. In one embodiment of the present disclosure, the processor may be configured to generate a first representative data point set related to the optimal fragmented image frame based on the pixel binning operation.

[0180] In one embodiment of the present disclosure, the processor may be configured to identify the number of fragments related to the input image frame based on user input. In one embodiment of the present disclosure, the processor may be configured to analyze the content change within the input image frame using one or more distribution mechanisms. In one embodiment of the present disclosure, the processor may be configured to fragment the input image frame based on the number of fragments determined and the analyzed content change to generate an optimal fragmented image frame. In one embodiment of the present disclosure, the processor may be configured to perform a pixel binning operation on one or more fragments of the optimal fragmented input image frame using one or more statistical mechanisms. In one embodiment of the present disclosure, the optimal fragmented input image frame may include a group of pixels having unique characteristics. In one embodiment of the present disclosure, the processor may be configured to generate a second representative data point set related to the optimal fragmented input image frame based on the pixel binning operation.

[0181] In one embodiment of the present disclosure, the processor may be configured to determine a data distribution dissimilarity metric between a first representative data point set related to a fragmented image frame from a reconstructed image frame and a second representative data point set related to a fragmented input image frame from an input image frame. In one embodiment of the present disclosure, the processor may be configured to determine an offset value based on the data distribution dissimilarity metric.

[0182] In one embodiment of the present disclosure, the processor may be configured to determine a pixel mapping in the form of a fragment-by-fragment offset using a model output data distribution, a ground truth data distribution, and a mapping range.

[0183] In one embodiment of the present disclosure, the mapping range is also a decision on whether to apply offset scaling to a specific fragment based on the RD cost. In one embodiment of the present disclosure, the mapping range may be determined based on the number of fragments and the codec RD cost.

[0184] In one embodiment of the present disclosure, the processor may be configured to perform a scaling operation on the reconstructed image frame based on the offset value determined to generate a scaled image frame. In one embodiment of the present disclosure, the scaling operation may include at least one of an addition operation, a multiplication operation, a division operation, or an exponential operation. In one embodiment of the present disclosure, the processor may be configured to encode the scaled image frame.

[0185] In one embodiment of the present disclosure, the reconstructed image frame may be generated by using one or more neural network (NN) models of an in-loop filter of an AI infrastructure.

[0186] In one embodiment of the present disclosure, the processor is configured to transmit bitstream information associated with the reconstructed image frame to a decoder, and in one embodiment of the present disclosure, the bitstream information may include a determined offset value.

[0187] In one embodiment of the present disclosure, the offset value may be calculated and used at a per-fragment granularity.

[0188] A system for AI-based decoding of media may include a processor operably coupled to a memory and a communication unit. In one embodiment of the present disclosure, the processor may be configured to receive bitstream information including offset information from an encoder. In one embodiment of the present disclosure, the processor may be configured to generate a reconstructed image frame based on the bitstream information using an in-loop filter of the AI base. In one embodiment of the present disclosure, the processor may be configured to perform a scaling operation on the reconstructed image frame based on the offset information to generate a scaled image frame. In one embodiment of the present disclosure, the processor may be configured to generate an output video based on the scaled image frame.

[0189] In one embodiment of the present disclosure, the processor may be configured to generate a model output data distribution related to the reconstructed image frame. In one embodiment of the present disclosure, the processor may be configured to perform a pixel mapping for a scaling operation based on the offset information and the model output data distribution. In one embodiment of the present disclosure, the processor may be configured to perform a scaling operation on the reconstructed image frame based on the pixel mapping to generate a scaled image frame.

[0190] In one embodiment of the present disclosure, the scaling operation may include at least one of an addition operation, a multiplication operation, a division operation, or an exponentiation operation.

[0191] According to one embodiment of the present disclosure, one or more operations described as being performed on an image frame are performed on a video including the image frame, and similarly, one or more operations described as being performed on a video can be performed on an image frame included in the video.

[0192] Various actions, acts, blocks, steps, etc. can be performed in the presented order, in a different order from each other, or simultaneously. Further, in some embodiments, some of the actions, acts, blocks, steps, etc. can also be omitted, added, modified, skipped, etc. without departing from the scope of the present invention.

[0193] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The systems, methods, and examples provided in this disclosure are merely illustrative and not limiting.

[0194] Specific language expressions have been used to describe the subject matter of this application, but any limitations resulting therefrom are not intended. As will be apparent to one of ordinary skill in the art, various working modifications can be made to the method to embody the inventive concepts as taught in this disclosure. The drawings and the foregoing description provide multiple examples. One of ordinary skill in the art will understand that one or more of the described elements can be well combined as a single functional element. Alternatively, a particular element can be divided into multiple functional elements. Elements from one embodiment can be added to other embodiments.

[0195] The embodiments disclosed in this application are embodied using at least one hardware device and can perform network management functions to control elements.

[0196] The foregoing description of specific embodiments fully discloses the general nature of the embodiments in the present disclosure such that others, by applying current knowledge, can readily modify and / or adapt such specific embodiments for various applications without departing from the general concept, and thus, such alterations and modifications should be understood to be within the meaning and scope of the equivalents of the disclosed embodiments and are intended to be so understood. It must be understood that the syntax or terminology employed in the present disclosure is for the purpose of description and not of limitation. Therefore, although embodiments in this specification have been described in terms of preferred embodiments, it will be recognized by those of ordinary skill in the relevant art that the embodiments in this specification can be practiced with modifications within the scope of the embodiments as described in this application.

Claims

1. In a method (1400) performed by an electronic device for encoding an artificial intelligence (AI) - based media, compressing an input image frame associated with an input video (1410); generating a reconstructed image frame corresponding to the input image frame using an in - loop filter (302e) of the AI - based platform (1420); determining an offset value based on the input image frame and the reconstructed image frame (1430); encoding the reconstructed image frame based on the determined offset value (1440), the method (1400).

2. The step (1430) of determining the offset value based on the input image frame and the reconstructed image frame includes generating a model output data distribution related to the reconstructed image frame; generating a ground - truth data distribution related to the input image frame; determining the offset value by comparing the model output data distribution and the ground - truth data distribution, the method (1400) according to claim 1.

3. The step of generating the model output data distribution related to the reconstructed image frame includes identifying the number of fragments related to the reconstructed image frame based on user input; analyzing content changes within the reconstructed image frame using one or more distribution mechanisms; fragmenting the reconstructed image frame based on the determined number of fragments and the analyzed content changes to generate an optimal fragmented image frame; performing a pixel binning operation on one or more fragments of the optimal fragmented image frame using one or more statistical mechanisms, wherein the optimal fragmented image frame includes groups of pixels with unique characteristics; generating a first representative data point set related to the optimal fragmented image frame based on the pixel binning operation, the method (1400) according to claim 2.

4. The step of generating the ground - truth data distribution related to the input image frame Identifying the number of fragments related to the input image frame based on user input; Analyzing content changes within the input image frame using one or more distribution mechanisms; Fragmenting the input image frame based on the determined number of fragments and the analyzed content changes to generate an optimal fragmented input image frame; Performing a pixel binning operation on one or more fragments of the optimal fragmented input image frame using one or more statistical mechanisms, wherein the optimal fragmented input image frame includes groups of pixels having unique characteristics; Generating a second representative data point set related to the optimal fragmented input image frame based on the pixel binning operation, the method (1400) according to claim 2 or 3.

5. The step of determining the offset value by comparing the model output data distribution and the ground truth data distribution includes: Determining a data distribution dissimilarity metric between a first representative data point set related to the fragmented image frame from the reconstructed image frame and a second representative data point set related to the fragmented input image frame from the input image frame; Determining the offset value based on the data distribution dissimilarity metric, the method (1400) according to any one of claims 2 to 4.

6. The step of determining the offset value by comparing the model output data distribution and the ground truth data distribution includes: Determining a pixel mapping in the form of a fragment-by-fragment offset using the model output data distribution, the ground truth data distribution, and a mapping range, the method (1400) according to any one of claims 2 to 5.

7. The mapping range is a determination of whether to apply the offset scaling to a specific fragment based on the RD cost, and the mapping range is determined based on the number of fragments and the codec RD cost, the method (1400) according to claim 6.

8. The encoding (1440) of the reconstructed image frame is a step of performing a scaling operation on the reconstructed image frame based on the determined offset value to generate a scaled image frame, wherein the scaling operation includes at least one of an addition operation, a multiplication operation, a division operation, or an exponential operation, and encoding the scaled image frame, the method (1400) according to any one of claims 1 to 7. **Claim 9** The reconstructed image frame is generated by using one or more neural network (NN) models of the in-loop filter of the AI base, the method (1400) according to any one of claims 1 to 8. **Claim 10** transmitting bitstream information related to the reconstructed image frame (the bitstream information includes the determined offset value) to a decoder, the method (1400) according to any one of claims 1 to 9. **Claim 11** The offset value is calculated and used at the fragment-by-fragment granularity, the method (1400) according to any one of claims 1 to 10. **Claim 12** In a method (1600) performed by an electronic device for decoding an artificial intelligence (AI) base of a medium, receiving (1610) bitstream information including offset information from an encoder; generating (1620) a reconstructed image frame based on the bitstream information using an in-loop filter of an AI base; performing (1630) a scaling operation on the reconstructed image frame based on the offset information to generate a scaled image frame; and generating (1640) an output video based on the scaled image frame, the method (1600). **Claim 13** The step of performing (1630) the scaling operation on the reconstructed image frame generating a model output data distribution related to the reconstructed image frame, and Performing a pixel mapping for the scaling operation based on the offset information and the model output data distribution; Performing the scaling operation on the reconstructed image frame based on the pixel mapping to generate the scaled image frame, the method (1600) according to claim 12. **Claim 14** The method (1600) according to claim 12 or 13, wherein the scaling operation includes at least one of an addition operation, a multiplication operation, a division operation, or an exponential operation. **Claim 15** In a system (101) for encoding an artificial intelligence (AI) infrastructure of a media, Including a processor (120) operably coupled to a memory (110) and a communication unit (130), The processor (120) is Compressing an input image frame associated with an input video, Generating a reconstructed image frame corresponding to the input image frame using an in-loop filter of the AI infrastructure, Determining an offset value based on the input image frame and the reconstructed image frame, A system (101) configured to encode the reconstructed image frame based on the determined offset value.