Network-based Image Filtering for Video Encoding

By employing a neural network-based model filtering method that adjusts QpMap values during video encoding, the method addresses the limitations of traditional video coding techniques, achieving enhanced compression efficiency and video quality.

JP7695462B2Active Publication Date: 2025-06-18BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024500126
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-05
Filing Date
2022-07-05
Publication Date
2025-06-18
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

Existing video coding techniques struggle to achieve optimal compression efficiency while maintaining video quality, as they rely on traditional prediction methods that do not fully exploit the redundancy in video data.

Method used

The use of a neural network-based model filtering method during video encoding, where QpMap values are loaded into the neural network to adjust input QP values and filter the input frame, enabling the network to learn and improve image filtering.

Benefits of technology

This approach significantly enhances video coding efficiency by allowing the neural network to adaptively filter video frames, leading to improved compression performance and maintained video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007695462000017
    Figure 0007695462000017
  • Figure 0007695462000018
    Figure 0007695462000018
  • Figure 0007695462000019
    Figure 0007695462000019
Patent Text Reader

Abstract

A method and apparatus for image filtering during video encoding using a neural network is provided, the method includes loading a plurality of quantization parameter (QP) map (QpMap) values ​​in a plurality of QpMap channels into a neural network, obtaining a QP scaling factor by adjusting a plurality of input QP values ​​for an input frame, and adjusting the plurality of QpMap values ​​according to the QP scaling factor for the neural network to learn and filter the input frame to the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the priority of U.S. Provisional Application No. 63 / 218,485, titled "Neural Network based Image filtering for Video Coding", filed on July 5, 2021, the entire content of which is incorporated by reference for all purposes.

[0002] This disclosure relates to video coding, and more particularly, but not limited to, methods and apparatuses for video coding using neural network - based model filtering.

Background Art

[0003] To compress video data, various video coding techniques may be used. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), joint exploration test model (JEM), High - Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, and the like. Video coding generally utilizes prediction methods (e.g., inter - prediction, intra - prediction, etc.) that exploit the redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form using a lower bitrate while avoiding or minimizing the degradation of video quality.

[0004] The first version of the HEVC standard was finalized in October 2013, providing approximately 50% bitrate reduction or equivalent perceptual quality compared to the previous-generation video coding standard H.264 / MPEG AVC. The HEVC standard brings significant coding improvements over its predecessors, but there is evidence that even better coding efficiency than HEVC can be achieved using additional coding tools. Based on this, both VCEG and MPEG initiated work to explore new coding techniques for future video coding standardization. To begin important research on advanced technologies that could enable substantial enhancements in coding efficiency, a Joint Video Exploration Team (JVET) was formed in October 2015 by ITU-T VECG and ISO / IEC MPEG. By integrating some additional coding tools on top of the HEVC Test Model (HM), a reference software called the Joint Exploration Model (JEM) was developed by JVET.

[0005] A joint call for proposals (CfP) for video compression with capabilities beyond HEVC was issued by ITU-T and ISO / IEC. Responses to 23 CfPs were received and evaluated at the 10th JVET meeting, showing an approximately 40% improvement in compression efficiency over HEVC. Based on such evaluation results, JVET launched a new project to develop a new-generation video coding standard named Versatile Video Coding (VVC). To show the reference implementation form of the VVC standard, a reference software codebase called the VVC Test Model (VTM) was established. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] This disclosure provides examples of techniques for improving video coding efficiency by using neural network-based model filtering.

Means for Solving the Problem

[0007] According to a first aspect of the present disclosure, a method for image filtering during video encoding using a neural network is provided. The method includes loading a plurality of QpMap values in one or more quantization parameter (QP) maps (QpMap channels) into the neural network, obtaining a QP scaling factor by adjusting a plurality of input QP values related to the input frame, and adjusting the plurality of QpMap values according to the QP scaling factor so that the neural network learns and filters the input frame to the neural network.

[0008] According to a second aspect of the present disclosure, an apparatus for image filtering during video encoding using a neural network is provided. The apparatus includes one or more processors and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Further, when the one or more processors execute the instructions, they are configured to perform the method according to the first aspect.

[0009] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer-executable instructions that cause one or more computer processors to perform the method according to the second aspect when executed by the one or more computer processors is provided.

[0010] A more detailed description of the examples of the present disclosure will be made by referring to the specific examples illustrated in the accompanying drawings. On the premise that these drawings only depict some examples and are not considered to limit the scope, the examples will be described and explained more specifically and in detail through the use of the accompanying drawings.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 3E

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7A

Figure 7B

Figure 8A

Figure 8B

Figure 8C

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 13

Figure 14

Figure 15

Figure 16A

Figure 16B

Figure 17

Figure 18

Figure 19

Figure 20

DETAILED DESCRIPTION OF THE INVENTION

[0012] Particular embodiments are now referred to in detail, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in the understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative forms may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented in many types of electronic devices having digital video capabilities.

[0013] References throughout this specification to "one embodiment", "an embodiment", "an example", "some embodiments", "some examples", or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are applicable to other embodiments as well, unless expressly specified otherwise.

[0014] Throughout this disclosure, the terms "first", "second", "third", etc. are all used as a nomenclature solely for reference to related elements, such as devices, components, structures, steps, etc., and do not imply any spatial or temporal order unless otherwise explicitly specified. For example, "a first device" and "a second device" may refer to two separately formed devices, or two parts, components, or operational states of the same device, and may optionally be named.

[0015] The terms "module", "sub-module", "circuit", "sub-circuit", "circuitry", "sub-circuitry", "unit", or "sub-unit" may include a memory (shared, dedicated, or grouped) that stores code or instructions executable by one or more processors. A module may include one or more circuits regardless of the presence of stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached, or may or may not be placed adjacent to each other.

[0016] As used herein, the terms "when" or "in case" may be understood to mean "simultaneously with" or "in response to", depending on the context. These terms may not indicate that related limitations or features are conditional or optional when they appear in claims. For example, a method may include the steps of: i) a function or action X' is performed when or in case condition X exists, and ii) a function or action Y' is performed when or in case condition Y exists. The method may be performed with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at different times during multiple executions of the method.

[0017] The unit or module may be executed purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, the unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a specific function.

[0018] FIG. 20 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will be decoded later by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0019] In some implementation forms, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium to enable the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum, or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be beneficial in facilitating communication from the source device 12 to the destination device 14.

[0020] In some other implementations, the encoded video data may be transmitted from the output interface 22 to the storage device 32. Thereafter, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray disk, a digital versatile disk (DVD), a compact disk read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a wireless fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0021] As shown in FIG. 20, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include, for example, a video capture device such as a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a source such as a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As one example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the implementations described in this application may generally be applicable to video encoding and may also be applied to wireless and / or wired applications.

[0022] Captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may further (or alternatively) be stored in the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0023] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and may receive encoded video data via link 16. The encoded video data communicated via link 16 or provided to the storage device 32 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 upon decoding of the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.

[0024] In some implementations, the destination device 14 may include a display device 34, and the display device 34 can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to the user and may include any of various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0025] The video encoder 20 and the video decoder 30 may operate according to a proprietary or industry standard such as VVC, HEVC, MPEG-4, Part10, AVC, or an extended version of such a standard. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally assumed that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is generally assumed that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.

[0026] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device may store software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in their respective devices.

[0027] Like HEVC, VVC is built on a block - based hybrid video encoding framework. FIG. 1 is a block diagram illustrating a block - based video encoder according to some implementation forms of the present disclosure. In encoder 100, the input video signal is processed for each block called a coding unit (CU). Encoder 100 may be a video encoder 20 as shown in FIG. 20. In VTM - 1.0, a CU can be up to 128×128 pixels. However, different from HEVC which divides blocks only based on a quadtree, in VVC, one coding tree unit (CTU) is divided into CUs in order to adapt to various local characteristics based on quadtree / bi - tree / tri - tree. Additionally, the concept of multiple split unit types in HEVC is abolished, that is, the distinction between CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as a basic unit for both prediction and transform without further division. In multiple types of tree structures, a CTU is first divided by a quadtree structure. Then, the leaf nodes of each quadtree can be further divided by bi - tree and tri - tree structures.

[0028] FIGS. 3A - 3E are schematic diagrams illustrating multiple types of tree partition modes according to some implementation forms of the present disclosure. FIGS. 3A - 3E show five partition types including four - way split (FIG. 3A), vertical bi - split (FIG. 3B), horizontal bi - split (FIG. 3C), vertical tri - split (FIG. 3D), and horizontal tri - split (FIG. 3E), respectively.

[0029] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") uses pixels from samples of already encoded neighboring blocks (referred to as reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion compensated prediction") uses pixels reconstructed from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference pictures are supported, one reference picture index is sent additionally, and the reference picture index is used to identify from which reference picture in the reference picture store the temporal prediction signal came.

[0030] After spatial and / or temporal prediction, the intra / inter mode decision circuitry 121 of the encoder 100 selects the best prediction mode, for example, based on a rate distortion optimization method. The block predictor 120 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using the transform circuitry 102 and quantization circuitry 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuitry 116 and inverse transformed by the inverse transform circuitry 118 to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, the reconstructed CU is placed in the reference picture store of the picture buffer 117, and in-loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), may be applied to the reconstructed CU before it is used to encode future video blocks. The encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit 106 so that they are further compressed and packed to form the output video bitstream 114.

[0031] For example, deblocking filters are available in AVC, HEVC, and the current version of VVC. In HEVC, an additional in-loop filter called SAO is defined to further improve the encoding efficiency. In the current version of the VVC standard, yet another in-loop filter called ALF is being actively investigated and is likely to be included in the final standard.

[0032] These in-loop filter operations are optional. Performing these operations helps to improve the encoding efficiency and visual quality. They may further be turned off as a decision made by the encoder 100 to save computational complexity.

[0033] Intra prediction is usually based on non-filtered reconstructed pixels, while inter prediction, it should be noted, is based on filtered reconstructed pixels when these filter options are turned on by the encoder 100.

[0034] FIG. 2 is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section resident in the encoder 100 of FIG. 1. The block-based video decoder 200 may be a video decoder 30 as shown in FIG. 20. In the decoder 200, the incoming video bitstream 201 is first decoded through entropy decoding 202 to derive quantization coefficient levels and prediction-related information. The quantization coefficient levels are then processed through inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residuals. The block predictor mechanism executed by the intra / inter mode selector 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of non-filtered reconstructed pixels is obtained by summing the reconstructed prediction residuals from the inverse transform 206 and the predictive output generated by the block predictor mechanism using the adder 214.

[0035] The reconstructed blocks may further pass through the in-loop filter 209 and are then stored in the picture buffer 213 that functions as a reference picture store. The reconstructed video in the picture buffer 213 is sent to drive a display device and may also be used to predict future video blocks. In the situation where the in-loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0036] This disclosure is for improving the image filtering design of the above-mentioned video coding standards or techniques. The filtering method proposed in this disclosure is based on a neural network and can be applied, for example, as part of in-loop filtering between a deblocking filter and a sample adaptive offset (SAO), or as part of post-loop filtering to improve current video coding techniques, or as part of post-processing filtering after current video coding techniques.

[0037] Neural network techniques, such as fully connected neural networks (FC-NNs), convolutional neural networks (CNNs), and long short-term memory networks (LSTMs), have already achieved remarkable success in many research fields including computer vision and video understanding.

[0038] Fully connected neural network (FC-NN) FIG. 4 illustrates a simple FC-NN consisting of an input layer, an output layer, and a plurality of hidden layers according to some implementation forms of this disclosure. In the k-th layer, the output f k (x k-1 , W k , B k ) is

Equation

Equation

[0039] According to the universal approximation hypothesis and Equation (4), assuming any continuous function g(x) and some ε > 0, there exists a neural network f(x) with a non-linear reasonable choice, such as ReLU, such that ∀x, |g(x) - f(x)| < ε. Therefore, many empirical studies have applied neural networks as approximators to mimic models with hidden variables in order to extract the underlying explainable features. For example, when applied to image recognition, the FC-NN helps researchers build a system that understands not only single pixels but also increasingly deep and complex substructures, such as edges, textures, geometric shapes, and objects.

[0040] Convolutional Neural Network (CNN) FIG. 5A illustrates an FC-NN having two hidden layers according to some implementations of the present disclosure. A CNN is a widely popular neural network architecture for image or video applications, which is very similar to the FC-NN as shown in FIG. 5A and includes weight and bias metrics. A CNN can be regarded as a 3D version of a neural network. FIG. 5B illustrates an example of a CNN in which the dimension of the second hidden layer is [W, H, depth] according to some implementations of the present disclosure. In FIG. 5B, neurons are arranged in a three-dimensional structure (width, height, and depth) to form a CNN, and the second hidden layer is visualized. In this example, the input layer holds an input image or video frame, and thus its width and height are the same as the input data. For application to image or video applications, each neuron in a CNN is a spatial filter element with an extended depth aligned with its input. For example, if there are three color components in the input image, the depth is 3.

[0041] FIG. 6 illustrates an example of applying a spatial filter to an input image according to some implementations of the present disclosure. As shown in FIG. 6, the dimensions of the basic elements in a CNN are defined as [Filter width , Filter height , Input depth , Output depth , and in this example, it is set to [5, 5, 3, 4]. Each spatial filter performs a two-dimensional spatial convolution with weights of 5*5*3 on the input image. The input image may be a 64×64×3 image. Then, four convolution results are output. Thus, when padding the boundary with two additional pixels, the dimension of the filtering result is [64 + 4, 64 + 4, 4].

[0042] Residual Network (ResNet) In image classification, accuracy quickly saturates and deteriorates as the depth of the neural network increases. More specifically, as the gradient gradually vanishes along the deep network and towards the zero gradient at the end, adding more layers to the deep neural network results in a larger training error. Therefore, ResNet, which is composed of residual blocks, resolves the deterioration problem by introducing identity connections.

[0043] FIG. 7A illustrates a ResNet that includes, as an element of the ResNet, a residual block in which the input of the residual block is added element - by - element by an identity connection according to some implementations of the present disclosure. As shown in FIG. 7A, the basic module of the ResNet consists of a residual block and an identity connection. According to the universal approximation hypothesis, assuming an input x, the weighted layer with an activation function within the residual block approximates a hidden function F(x), rather than an output H(x)=F(x)+x.

[0044] By stacking non - linear multi - layer neural networks, the residual blocks explore features that represent local characteristics of the input image. Without introducing additional parameters or computational complexity, the identity connection enables the training of deep learning networks by skipping one or more non - linear weighted layers as shown in FIG. 7A. When skipping the weighted layer, the differential output of the residual layer can be described as

Equation

[0045] TIFF0007695462000006.tif48170

[0046] Variant forms of ResNet In FIGS. 8A-8B, several variant forms of ResNet were proposed to improve the restored image quality of single-image super-resolution (SISR) and enhance the accuracy of image classification. FIG. 8A illustrates an example of ResNet including a plurality of residual blocks with overall identity connections according to some implementation forms of the present disclosure. In FIG. 8A, a variant form of ResNet is proposed to enhance the visual quality of the upsampled image. Specifically, an overall identity connection from the input of the first residual block to the output of the last residual block is applied to facilitate the integration of the training procedure.

[0047] FIG. 8B illustrates another example of ResNet with a plurality of residual blocks stacked to further improve the video encoding efficiency according to some implementation forms of the present disclosure. Each residual block directly propagates its own input to the next unit by means of a concatenation operation. In other words, since multi-level information can flow through an identical connection, each intermediate block can receive multi-layered information from the previous unit. The parameters of each residual block in FIG. 8B linearly increase with the number of layers by means of the concatenation operation.

[0048] In FIGS. 8A-8B, before the residual information can be propagated to subsequent modules, the residual features must pass through one or several modules. Due to the identity connection, these residual features can be quickly combined with the identity features at a specific layer and the propagation to subsequent modules can be stopped. Therefore, the residual features in the previous two variant forms are locally limited and the performance deteriorates.

[0049] FIG. 8C illustrates another example of a ResNet that addresses single-image super-resolution (SISR) by aggregating the outputs of residual blocks, according to some implementations of the present disclosure. In FIG. 8C, the output of the last residual block is concatenated with all of the outputs of the previous three modules. The concatenated hierarchical features are fused by a convolution operation before being applied to an element-wise addition with the input of the first residual block. Unlike the first two variant forms, the aggregated ResNet enables non-local features to be applied to the last residual module so that hierarchical information can be propagated to subsequent blocks, achieving feature representation in a more discriminative manner.

[0050] In the present disclosure, methods and apparatuses for neural-network-based image filtering are proposed to further improve the encoding efficiency of current hybrid video encoding. The proposed methods and apparatuses may be applied, for example, as part of in-loop filtering between a deblocking filter and a sample adaptive offset (SAO), as shown in FIG. 2, or as post-loop filtering to improve current video encoding techniques, or as post-processing filtering after current video encoding techniques.

[0051] FIG. 9 illustrates a typical neural network-based model for performing image filtering for video encoding, according to some implementations of the present disclosure. The YUV components may be provided in parallel to the neural network model. This parallel input of the YUV components not only reduces the processing delay, but may also be beneficial for the neural network model to learn the correlation between the arranged YUV information, such as, for example, cross-component filtering and / or luma guided chroma filtering. The on / off control of this neural network model-based filter may be performed at the coding tree unit (CTU) level for a reasonable trade-off between control granularity and signaling overhead. The on / off control of the neural network-based filter for the YUV components may be performed at the same or different granularities. For example, the on / off control of this neural network model-based filter for the Y component may be performed at the CTU level, while the on / off control for the U and V components may be performed at the frame level, for example, to save the CTU-level flag signaling overhead.

[0052] Feature map resolution alignment When CTU-level YUV information is provided to the neural network model filter as shown in FIG. 9, the resolutions of the YUV CTU batches may or may not be the same. For example, when the encoded video content is YUV420, the resolutions of the three arranged YUV patches may not be the same. In this case, resolution alignment is required. For easier illustration, it is assumed in the present disclosure that the video content is YUV420. The proposed method may be easily extended for different content formats, such as, for example, YUV422, YUV444.

[0053] In some examples, the resolution alignment may be performed before the YUV patches enter the neural network.

[0054] In some examples, one 128×128 Y patch may be downsampled to one 64×64 patch or four 64×64 patches. When four 64×64 patches are generated, all the information of the original 128×128 patch may be retained and distributed among the four patches. The method used for information distribution of the original 128×128 patch may be based on splitting. For example, one 64×64 patch may be from the upper left of the original 128×128 patch, and another 64×64 patch may be from the upper right of the original 128×128 patch. Alternatively, the method used for information distribution of the original 128×128 patch may be based on interleaving. For example, all four adjacent samples of the original 128×128 patch may be evenly distributed among the four 64×64 patches.

[0055] In some examples, one 64×64 U or V patch may be upsampled to one 128×128 patch.

[0056] In some examples, resolution alignment may be performed after the YUV patches enter the neural network. In one example, the Y input resolution may be decreased to match the UV input. One way to achieve this is to use a convolutional layer with a stride size twice that of the UV input. In this example, at the end of the neural network, a resolution increase layer is required to scale up the Y content so that the output of the model has the same resolution as the input. One way to achieve this is to use a pixel shuffle layer to scale up the Y resolution. In another example, the UV input resolution may be increased to match the Y input. One way to achieve this is to use a pixel shuffle layer to scale up the UV at the beginning of the neural network and then scale it down at the end of the neural network.

[0057] Feature Map Resolution Control The feature map resolution affects proportionally the neural network processing overhead, but may not affect proportionally the performance of the neural network. To control the computational complexity of model filtering, different solutions may become available, such as, for example, the number of residual blocks, and the number of input and output channels of the convolutional layers per residual block. Controlling the resolution of the feature map in the convolutional layer is another effective option for controlling the computational complexity.

[0058] FIG. 10 illustrates region-based feature map resolution control according to some implementations of the present disclosure. As shown in FIG. 10, three regions may be used to adjust the feature map resolution for computational complexity control (which may be referred to as region-based feature map resolution control). In region 1, the resolution of the input YUV patch is determined and the corresponding upscaling / downscaling operations are performed. For example, an up / downsampling method was introduced in "feature map resolution alignment". An example is shown in FIG. 15.

[0059] FIG. 15 illustrates luma downsampling in region 1 of a neural network according to some implementations of the present disclosure. As shown in FIG. 15, in region 1 before entering the neural network, the original Y patch is downsampled to four downsampled Y patches. For example, one 128×128 Y patch may be downsampled to four 64×64 Y patches. After the processing of the neural network is complete, the reverse operation, such as upsampling, is performed in region 1. As shown in FIG. 15, the four downsampled Y patches output by the neural network are upsampled to one original Y patch. For example, four 64×64 Y patches may be upsampled to one 128×128 Y patch.

[0060] In Region 2, the resolution of the input YUV patch is determined, and the corresponding upscaling / downscaling operation is performed before YUV concatenation. Since this region is at the beginning of the neural network, if a downscaling operation is performed, the input information may be significantly lost, and the overall performance after model training may be impaired. Two examples are shown in FIGS. 13-14 respectively.

[0061] FIG. 13 illustrates chroma upsampling in Region 2 of a neural network according to some implementations of the present disclosure. As shown in FIG. 13, the UV patch is upscaled by the corresponding convolutional block or layer in Region 2 of the neural network after entering the neural network. As shown in FIG. 13, an inverse operation, such as scaling-down, is performed on the corresponding UV patch output by the last residual block in Region 3.

[0062] FIG. 14 illustrates luma downsampling in Region 2 of a neural network according to some implementations of the present disclosure. As shown in FIG. 14, the Y patch is downscaled by the corresponding convolutional block or layer in Region 2 of the neural network after entering the neural network. As shown in FIG. 14, an inverse operation, such as scaling-up, is performed on the corresponding Y patch output by the last residual block in Region 3.

[0063] In Region 3, the resolution of the input YUV patch may be scaled up / down in one of the initial residual blocks, and in subsequent residual blocks, an inverse operation, such as scale down / up, may be performed. Since this region is after YUV concatenation, when a scale-down operation is performed, most of the input information has already been captured or learned in the initial convolutional layer with sufficient depth for information learning, so the input information is not lost as significantly as in Region 2. For example, after Region 2, three channels of YUV content with UV scaled up to 128×128 are generated. The Y input information may already be learned / extracted and dispersed / duplicated in the initial convolutional layer before concatenation. Alternatively, since the first residual block may have sufficient channels to learn / extract the Y input information features, a scale-down operation may be performed after the first residual block.

[0064] QP-independent neural network model To facilitate a simpler deployment of the proposed neural network model filtering, it is desirable to remove the input quantization parameter (QP) dependency from the neural network model. Thus, a single neural network model may be used for image filtering independent of the input QP used for video encoding.

[0065] FIG. 11 illustrates a typical QP-independent neural network model according to some implementation forms of the present disclosure. In the case of a typical video encoding system, QP values are used to calculate quantization step sizes for predictive residual quantization / dequantization. Therefore, different QP values represent different levels of video quality. To handle different video frames with different input QPs and qualities, a QpMap is provided to the neural network. The QpMap adds another dimension of information that the neural network should learn to adaptively filter the provided YUV input that may include different levels of video quality (e.g., input Qp). In some common video encoding standards such as HEVC, VVC, or AVS, when the input Qp value is converted to the Qp step size for predictive residual quantization, a predefined relationship (e.g., Q step =2 (QP-4) / 6 ) is typically used. As a simpler illustration, a QpMap having or including an input Qp value or a Qp step size value is used to introduce the proposed concept as follows.

[0066] Dynamic range control of QpMap values FIG. 11 illustrates a typical QP-independent neural network-based model for performing image filtering for video encoding according to some implementation forms of the present disclosure. The QpMap is concatenated with the YUV input channels. Each QpMap channel may have the same value at different coordinates. Also, each QpMap channel may have the same resolution as the associated input channel. That is, the QpMap channel for the Y input has the same resolution as the Y input channel, and each value of the QpMap indicates that all samples within the Y input channel have the same Qp value.

[0067] To generate QpMap values, the input QP value for each video frame / image may be directly used. Alternatively, the input QP value for each video frame / image is, for example, Q step =2 (QP-4) / 6It may first be converted to a Qp step size, for example. In some examples of the present disclosure, the Qp step size Q step is sometimes also referred to as QP step .

[0068] When a QpMap value is generated from an input QP or Qp step size, it is desirable for the dynamic range of the QpMap value to be reasonable in the following three senses.

[0069] First, the range should be large enough so that different QpMap values can be easily used to represent / differentiate different input Qp or Qp step sizes. In other words, assuming two input Qp values, the corresponding QpMap values should not be close to each other.

[0070] Second, the range should be well-balanced so that different QpMap values can be evenly distributed at different positions within the range.

[0071] Third, the range should match the dynamic range of the associated YUV sample values. For example, when the YUV sample values are normalized to [0,1] by dividing by P max and P max = 2 bitdepth , the QpMap values should similarly be normalized within a similar range.

[0072] Therefore, when the QpMap is mapped to a dynamic range, it is proposed not to use the maximum or minimum input Qp or Qp step size as a divisor factor, because otherwise the division will push the generated QpMap values to one side of the dynamic range, which is equivalent to reducing the dynamic range of the QpMap values.

[0073] For example, when using a maximum QpStep size of 912 (corresponding to a maximum input Qp of 63), the theoretical dynamic range is (0,1), but when the input Qp step size is typically less than 45 (corresponding to an input Qp of 37), the effective dynamic range is only (0,0.05), which means that in most cases, the QpMap value is close to 0.

[0074] Instead, it is proposed to perform normalization using the central / intermediate input Qp or Qp step size, so that the generated QpMap values can be distributed on either side of the dynamic range, for example, in the range [0.5,1.5]. One exemplary selected central / intermediate input Qp value is Qp32, and then, according to the QP to Qstep (Qp step size) equation (e.g., Q step =2 (QP-4) / 6 ), the converted Qp step size is approximately 25.5. Therefore, any input Qp value is first converted to the corresponding Qp step size, and then followed by division by the selected Qp step size of 25.5.

[0075] This selected input Qp or Qp step size for the normalization division may be flexibly determined based on the actual input Qp range. For example, in the case of an actual Qp range of [22,42], Qp37 or its corresponding Qp step size may be selected as the division factor so that the normalized value of the lower Qp of the actual Qp range does not get too close to zero, while the normalized value of the higher Qp does not exceed 1.0 by much. Alternatively, when the maximum value is not much larger than the minimum value (e.g., within a factor of 2), the maximum value of the actual input QP range (e.g., 42 in the case of an actual Qp range of [22,42]) or its corresponding Qp step size may be selected as the division factor.

[0076] Prediction-based adjustment of QpMap values As described above, the QpMap value may be directly generated by the input Qp value, or may be generated by the Qp step size value according to the mapping relationship between the input Qp value and the Qp step size. For the sake of a simpler illustration, the following description assumes that the QpMap value is directly generated by the input Qp value. When the QpMap value is generated by the Qp step size value, the proposed concept / method may be extended in the same way.

[0077] In the case of an inter-predicted video frame / image, most blocks / CTUs within the frame / image may be inter-predicted in a state where the residual is small or non-existent, such as in skip mode. In this case, the corresponding reference frame / image should determine the valid input Qp value.

[0078] In some examples, the input Qp value of the corresponding reference frame / image may be saved and retrieved when the current image is reconstructed during the motion compensation process. The input QP value for each current frame is known. However, when this frame is not the current one and this frame is the reference frame of another frame, the input Qp value of this frame becomes unknown. Therefore, the Qp value must be saved in the future in order to obtain it.

[0079] In some examples, the input Qp value of the corresponding reference frame / image may be derived by subtracting a specific value from the QP value of the currently inter-coded frame, and the specific value may be obtained by checking the temporal layer index of the currently inter-coded frame.

[0080] In some other examples, when the reference frame / image is in a chain of reference images (the reference frame / image of the reference frame / image), this information may be inherited or carried over from the signaling.

[0081] For a simple solution, in the case of inter-predicted video frames / images, the effective Qp step size may be derived from the Qp step size of the current frame by a constant scaling factor such as 0.5, which corresponds to an input Qp difference with a value of 6. This scaling operation is an approximation to the reference frame / image input Qp or Qp step size.

[0082] For a typical mapping relationship between the Qp step size and the input Qp (e.g., Q step =2 (QP-4) / 6 ), the scaling operation of the Qp step size is equivalent to the subtraction / addition operation of the input Qp value. That is, the scaling operation of the Qp step size can be implemented by applying the subtraction / addition operation of the input Qp value.

[0083] Depending on the trade-off between the signaling overhead and the prediction-based adjustment accuracy, the scaling of the Qp step size, or the subtraction / addition of the input Qp value, may be represented with different accuracies or / and different granularities. As a simpler illustration, the proposed concept / method below assumes that the scaling of the Qp step size is used for the prediction-based adjustment of the QpMap value. When the subtraction / addition of the input Qp value is used for the prediction-based adjustment of the QpMap value, the proposed concept / method may be directly extended.

[0084] In one or more examples, the scaling of the Qp step size may also be applied to the intra-predicted frame / image to compensate for the inaccuracy of a constant scaling factor used for subsequent inter-predicted frames.

[0085] In another example, the Qp scaling factor may be flexibly derived using the following different methods.

[0086] In the first method, the Qp scaling factor may be selected by the encoder from a set of values. The set of scaling factors may be sequence-based or picture / slice-based, which means that the set may be encoded within the picture header or the sequence parameter set. The index of the selected Qp scaling factor within the set of scaling factors may be selected on the encoder side based on a rate-distortion optimization algorithm at different granularities, such as picture-level index selection, CTU-level index selection, block-level selection, etc., for a good trade-off between picture quality and signaling overhead (e.g., a picture may be distributed to different blocks based on a quadtree partition).

[0087] In the second method, the Qp scaling factor may be converted to a Qp offset / regulation. The Qp offset / regulation may be applied to the input Qp value or the Qp step size before calculating the QpMap value.

[0088] In one example, the adjusted input Qp value may be expressed as Q p_new =Q p_old -Q p_offset_stepsize ×(lower_bound-offset_index), where Q p_old is the original Qp value of the current slice or CTU, Q p_new is the new Qp value after adjustment, Q p_offset_stepsize is the step size for each Qp adjustment, lower_bound is an integer value that determines the maximum Qp reduction, and offset_index is the index value to be signaled (e.g., a value in the range [0,3]). The offset_index is determined on the encoder side and parsed / used on the decoder side. Note that Qp_offset_stepsize and lower_bound are predefined constant values or are signaled in a similar manner.

[0089] For example, Qp_offset_stepsize may be a constant value such as 4, lower_bound is 2, and when the signaled offset_index is 1, the decoder may adjust the current Qp value of 32 to 28, where 28 = 32 - 4*(2 - 1).

[0090] In the third method, for the current CTU / picture / block, the Qp scaling factor may be inherited from an adjacent CTU / picture / block (e.g., the left or upper CTU / block in the spatial domain, or the reference block / CTU in the temporal domain) to save signaling overhead.

[0091] In the fourth method, instead of signaling or inheritance, the Qp scaling factor may be calculated on the decoder side based on the Qp difference between the current CTU / block and the reference CTU / block. If there are multiple reference CTU / blocks corresponding to the current CTU / block, the average value of the Qp scaling factors may be calculated. When the reference CTU / block is included in the reference chain, the reference depth may be limited, and the Qp scaling factor may be calculated according to the parent reference CTU / block at most limited reference depths.

[0092] In the fifth method, for lower complexity, the Qp scaling factors of different components may be signaled / selected / calculated together. Alternatively, the Qp scaling factors of different components may be signaled / selected / calculated separately. Alternatively, the Qp scaling factors of luma and chroma may be signaled / selected / calculated separately.

[0093] In the sixth method, any combination of the above methods may be used as a hybrid method.

[0094] Furthermore, in the second method described above, one embodiment of the derivation process can be introduced as follows.

[0095] Generally, the QpMap value (QM CH (x, y)) at the position (x, y) of channel CH can be calculated as follows.

Number

Number

[0096] In VVC, the relationship between QP and quantization step Qstep is given by QP step = 2 (QP-4) / 6 . Therefore, equation (7) can be rewritten as

Number

[0097] In one embodiment, for equation (8), QP is the input QP of the current frame and is channel-dependent. QP max is the maximum allowable input QP equal to 63 in VVC. Using equations (7) and (8), equation (6) can be rewritten as

Number

[0098] From equation (9), it can be seen that the entire term is either channel-dependent or a constant value. Therefore, for each channel of the quality map, the values at different positions are the same.

[0099] As a simpler implementation form, two modifications can be further made to equation (9). The first modification is to use a constant value QP selected instead of the actual value 63 in VVC as QP max . The motivation is to avoid the vanishing gradient problem caused by very small quality map values. QP selectedNote that this is equivalent to controlling the dynamic range of the QpMap value, which has already been introduced in the section "Dynamic Range Control of QpMap Value". The second modification is to signal the QP offset value instead of the scaling factor α CH . In this case, Equation (9) can be rewritten as

Equation

[0100] Comparing Equation (10) with Equation (9), this is equivalent to representing α offset in a closed-form expression of the newly signaled QP offset value QP CH . In one embodiment, since the corresponding reference block is generated from an independently selected temporal reference picture, QP offset needs to be determined for each input video block. In another embodiment, QP offset is determined and signaled at the frame level to save signaling bits, which means that all CTUs within the same video frame share the same value of QP offset .

[0101] In one example, QP offset is implemented as a look-up table (LUT). As shown in Table 1, instead of directly signaling the QP offset value, a predefined codeword may be defined. Based on the codeword received in the bitstream, the actual QP offset value used for Equation (10) can be retrieved from the LUT shown in Table 1. For example, when the encoded word "01" is received on the decoder side, in the example shown in Table 1, the QP offset value 8 can be derived. Note that in Table 1, the mapped QP offset has a fixed step size at 4, and another example of LUT with a step size at 5 may be defined in Table 2.

[0102] Table 1 shows an exemplary LUT used to signal the scaling factor.

[0103] [Table 1]

[0104] Table 2 shows another exemplary LUT used to signal the scaling factor.

[0105] [Table 2]

[0106] QP offset The LUT-based implementation form of QP may be defined at different granularities, such as sequence level, frame level, etc. For different LUTs, the LUT difference may be signaled, or the LUT step size difference may be signaled. For example, the LUT step size difference between Table 1 and Table 2 is equal to 1 (5 - 4 = 1).

[0107] Scaling of sample values based on QpMap values The QP-independent neural network model may not explicitly include the QpMap channel in the network. For example, as shown in FIG. 11, the QpMap values generated for each YUV channel are concatenated with the YUV channel. Alternatively, the QpMap values fed into the network may be used to directly scale the sample values for each YUV channel. In this way, the QpMap channel is not concatenated with the YUV channel, which represents the implicit use of QpMap in the network.

[0108] When the QpMap channel is input into the network and concatenated with the YUV channel, similar to what is shown in FIG. 11, the scaling of the sample values, i.e., the scaling of the sample values in the YUV channel, may be directly performed by element-wise multiplication. For example, each element of the QpMap channel for the Y component is multiplied by the corresponding element of the Y channel, and each element of the QpMap channel for the U or V component is multiplied by the corresponding element of the U or V channel. Note that the resolution of the QpMap channel may already be aligned with the resolution of the corresponding component channel.

[0109] In another example, the element-wise scaling may also be performed for each residual block. FIG. 17 illustrates an example of the element-wise scaling performed for each residual block according to some implementations of the present disclosure. In FIG. 17, the QpMap channel is first concatenated with the YUV channel after the YUV resolutions are aligned. Then, the QpMap channel is used not only as the input feature map for the first residual block but also as the scaling factor of the sample values for each residual block.

[0110] Note that the above two sample scaling mechanisms may be used exclusively or combined. In other words, the sample scaling directly applied to the YUV samples before concatenation such as in FIG. 11, and the sample scaling applied for each residual block such as in FIG. 17, may be used both in the same neural network or separately in different neural networks.

[0111] In some examples regarding the implicit use of QpMap, the QpMap data may not need to be fed into the network, and the scaling of the sample values for each YUV channel is performed before the neural network.

[0112] Interaction between Neural Network-based Model Filtering and Other In-loop Filters When a QpMap channel is provided to a neural network to filter video content having different qualities, the QpMap channel may include Qp information from one or more components.

[0113] FIG. 12A illustrates an example of a layout arranging a QpMap channel and YUV channels according to some implementations of the present disclosure. FIG. 12B illustrates another example of a layout arranging a QpMap channel and YUV channels according to some implementations of the present disclosure. In FIGS. 12A-12B, a block named Map Y indicates a QpMap channel for the Y channel, a block named Map U indicates a QpMap channel for the U channel, and a block named Map V indicates a QpMap channel for the V channel. Blocks Y, U, and V indicate the Y channel, the U channel, and the V channel, respectively.

[0114] Assuming YUV420 content, the UV components may first be upsampled, and then the YUV may be sandwiched between and arranged with the corresponding QpMap channels as shown in FIG. 12A, or the Y channel may first be downsampled into four smaller Y channels, and then the YUV may be sandwiched between and arranged with the QpMap channels as shown in FIG. 12B. In some examples, the upsampling or downsampling is performed, for example, within the network in regions 2 and 3 of FIG. 10, or outside the network in region 1 of FIG. 10, for example.

[0115] FIG. 16A illustrates another example of a layout arranging a QpMap channel and YUV channels according to some implementations of the present disclosure. FIG. 16B illustrates another example of a layout arranging a QpMap channel and YUV channels according to some implementations of the present disclosure. In FIGS. 16A-16B, a block named Map Y indicates a QpMap channel for the Y channel, a block named Map U indicates a QpMap channel for the U channel, and a block named Map V indicates a QpMap channel for the V channel. Blocks Y, U, and V indicate the Y channel, the U channel, and the V channel, respectively. As shown in FIGS. 16A-16B, multiple QpMap channels may first be connected internally and then connected to the YUV channels.

[0116] When only one or more QpMap channels for one component are provided to the neural network, the one or more QpMap channels may be placed on one side of the YUV channels, which indicates that the YUV channels are connected before the addition of the QpMap channels so that the YUV channels are arranged adjacent to each other.

[0117] In another example, in addition to QpMpa channels from different components, additional QpMap channels for different types of training data may be required. For example, when the training data is cropped from an I-frame or a B-frame or a P-frame, a QpMap including frame type information may be generated and connected. An I-frame is an intra-coded frame, a B-frame is a bi-directional prediction frame, and a P-frame is a prediction frame.

[0118] Filtering offset or scaling of the output of the neural network model-based filter For generalization, an integrated neural network model-based filter may be used for different video contents with different levels of quality, motion, and lighting environment. The output of the neural network model-based filter may be slightly adjusted in the form of an offset or scaling on the encoder side for better encoding efficiency.

[0119] The filtering offset or scaling value may be adaptively selected by the encoder from a set of values. The offset or scaling set may be sequence-based or picture / slice-based, which means that the set may be encoded within the picture / slice header or sequence parameter set. The index of the selected offset or scaling value within the set may be selected on the encoder side based on a rate-distortion optimization algorithm at different granularities, such as picture level index selection, CTU level index selection, block level selection, etc., for a good trade-off between picture quality and signaling overhead. For example, a picture may be distributed to different blocks based on a quadtree partition.

[0120] The selection of the adaptive filtering offset or scaling value may be based on a specific classification algorithm, such as content smoothness or histogram of oriented gradients. The adaptive filtering offset or scaling value for each category is calculated and selected in the encoder and explicitly signaled to the decoder to effectively reduce sample distortion. On the other hand, the classification for each sample is performed in both the encoder and the decoder to significantly save side information.

[0121] The selection of the adaptive filtering offset or scaling value may be done together or separately for different components. For example, YUV may have different adaptive filtering offset or scaling values.

[0122] Training data generation and training process When a neural network-based filter model is trained, the preparation of the training data and the training process may be performed in different ways.

[0123] In some examples, the model may be trained based on a dataset having only still images. The dataset may be encoded with all I-frames from a video encoding tool where a neural network-based filter is used.

[0124] In some examples, the model may be trained based on a two-path process. In the first path, the dataset may be encoded with all I-frames, and model A may be trained based on all the I-frames. In the second path, the same dataset or a new dataset may be encoded with a combination of I, B, and P frames having different ratios (the ratio of the number of I, B, and P frames included). In some examples, the generated I / B / P frames are encoded by applying model A trained in the first path. Based on the newly generated I / B / P frames, a new model B may be trained.

[0125] When model B is trained, model A may be loaded as a pre-trained model such that model B is a refined model starting from model A. In another example, another model different from model A may be loaded as a pre-trained point.

[0126] Alternatively, model B may be trained from scratch.

[0127] In some examples, the model may be a trained multi-path with three or more paths. In the first path, model A may be trained based on I-frames. In the second path, model B may be trained or refined based on model A based on the combination of I / B / P frames when model A is applied to the encoder. Note that the selected combination of B / P frames in this second training path may be from only the lower temporal layers. In the third path or further paths, model C may be trained or refined based on model B based on higher temporal layers of B / P frames. When higher temporal layers of B / P frames are generated and selected, model B or / and model A may be applied on the encoder side.

[0128] Before network training, training data must be generated. In this multi-path method, the training data is generated by three paths including the following. The first path is to generate only the I-frames used to train model A. When model A is ready, the encoder may or may not load model A and generate the low temporal layer B / P frames called the second path. These generated low temporal layer B / P frames are used to train model B by new training or refined based on model A.

[0129] Furthermore, when model B is ready, the encoder may or may not load models A and B and generate the high temporal layer B / P frames called the third path. These generated high temporal layer B / P frames are used to train model C by new training or refined based on model A or / and B.

[0130] Interaction between neural network-based model filtering and other in-loop filters When neural network-based model filtering is signaled to be turned on at the CTU level or frame level, deblocking filtering may be skipped to avoid unnecessary calculations or excessive smoothing. Alternatively, deblocking filtering may still be performed for visual quality.

[0131] When neural network-based model filtering is signaled to be turned on at the CTU level or frame level, some other in-loop filters such as ALF, Cross Component Adaptive Loop Filter (CCALF), and SAO may be turned off.

[0132] When neural network-based model filtering is signaled to be turned on at the CTU level or frame level, other in-loop filters may be selectively turned on or off at the CTU level or frame level. For example, when an intra-frame or intra-frame CTU is enabled for neural network-based model filtering, deblocking filtering, and / or other in-loop filters such as ALF, and / or CCALF, and / or SAO are disabled for the current intra-frame, or the current intra-frame CTU.

[0133] FIG. 18 is a block diagram illustrating an apparatus for image filtering during video encoding using a neural network according to some implementations of the present disclosure. The apparatus 1800 may be a terminal such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.

[0134] As shown in FIG. 18, the apparatus 1800 may include one or more of components such as a processing component 1802, a memory 1804, a power component 1706, a multimedia component 1808, an audio component 1810, an input / output (I / O) interface 1812, a sensor component 1814, and a communication component 1816.

[0135] The processing component 1802 generally controls all operations of the apparatus 1800, such as operations related to the display, calls, data communications, camera operations, and recording operations. The processing component 1802 may include one or more processors 1820 for executing instructions to complete all or part of the steps of the above methods. Further, the processing component 1802 may include one or more modules to facilitate the interaction between the processing component 1802 and other components. For example, the processing component 1802 may include a multimedia module to facilitate the interaction between the multimedia component 1808 and the processing component 1802.

[0136] The memory 1804 is configured to store different types of data to support the operation of the apparatus 1800. Examples of such data include instructions for any application or method operating on the apparatus 1800, contact data, phone book data, messages, pictures, videos, etc. The memory 1804 may be implemented by any type of volatile or non-volatile storage device, or a combination thereof, and the memory 1804 may be a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or a compact disk.

[0137] The power supply component 1806 supplies power to different components of the device 1800. The power supply component 1806 may include a power supply management system, one or more power supplies, and other components associated with the generation, management, and distribution of power for the device 1800.

[0138] The multimedia component 1808 includes a screen that provides an output interface between the device 1800 and the user. In some examples, the screen may include a liquid crystal display (LCD) and a touch panel (TP). When the screen includes a touch panel, the screen may be implemented as a touch screen that receives input signals from the user. The touch panel may include one or more touch sensors for detecting touches, slides, and gestures on the touch panel. The touch sensors may not only detect the boundaries of a touch or slide action, but may also further detect the duration and pressure associated with the touch or slide operation. In some examples, the multimedia component 1808 may include a front camera and / or a rear camera. When the device 1800 is in an operating mode such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data.

[0139] The audio component 1810 is configured to output and / or input audio signals. For example, the audio component 1810 includes a microphone (MIC). When the device 1800 is in an operating mode such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals may be further stored in the memory 1804 or sent via the communication component 1816. In some examples, the audio component 1810 further includes a speaker for outputting audio signals.

[0140] The I / O interface 1812 provides an interface between the processing component 1802 and the peripheral interface module. The above-mentioned peripheral interface module may be, for example, a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0141] The sensor component 1814 includes one or more sensors for providing state evaluation of the device 1800 in different modes. For example, the sensor component 1814 may detect the on / off state of the device 1800 and the relative location of components. For example, the components may be the display and keypad of the device 1800. The sensor component 1814 may also detect a change in the position of the device 1800 or a component of the device 1800, the presence or absence of user contact with the device 1800, the orientation or acceleration / deceleration of the device 1800, and a change in the temperature of the device 1800. The sensor component 1814 may include a proximity sensor configured to detect the presence of nearby objects even without physical contact. The sensor component 1814 may further include an optical sensor such as a CMOS or CCD image sensor used in an imaging application. In some examples, the sensor component 1814 may further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0142] The communication component 1816 is configured to facilitate wired or wireless communication between the device 1800 and other devices. The device 1800 may access a wireless network based on a communication standard such as Wi-Fi, 4G, or a combination thereof. In an example, the communication component 1816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an example, the communication component 1816 may further include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0143] In an example, the device 1800 may be implemented by one or more of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic elements for performing the above methods. The non-transitory computer-readable storage medium may be, for example, a Hard Disk Drive (HDD), a Solid State Drive (SSD), a flash memory, a hybrid drive or a Solid State Hybrid Drive (SSHD), a Read Only Memory (ROM), a Compact Disk Read Only Memory (CD-ROM), a magnetic tape, a floppy disk, etc.

[0144] FIG. 19 is a flowchart illustrating a process for image filtering during video encoding using a neural network according to some implementations of the present disclosure.

[0145] In step 1901, the processor 1820 loads a plurality of QpMap values in one or more QpMap channels into the neural network. As shown in FIG. 11, each of the plurality of QpMap channels is combined with a corresponding YUV channel after three convolutional blocks or layers before the concatenation block.

[0146] For example, since the QpMap may have different QP values for YUV, it may have three channels for YUV, namely QP - Y, QP - U, and QP - V, respectively.

[0147] In step 1902, the processor 1820 obtains a QP scaling factor by adjusting a plurality of input QP values regarding the input frame. For example, as discussed in the second method of the section "Prediction - based adjustment of QpMap values", the Qp offset / adjustment may be applied to the input Qp value or Qp · step size before calculating the QpMap value. In the present disclosure, the abbreviation "QP" is the same as "Qp".

[0148] In step 1903, the processor 1820 adjusts a plurality of QpMap values according to the QP scaling factor for the neural network to learn and filter the input frame to the neural network.

[0149] In some examples, the processor 1820 may obtain a QP offset based on a QP offset step size, a lower limit, and an offset index. The QP offset step size may be a step size for adjusting each input QP value, the lower limit may be an integer value for determining the maximum QP value reduction, and the offset index may be a signaled index value. Further, the processor 1820 may subtract the QP offset from the input QP value. For example, as discussed in the second method of the section "Prediction - based adjustment of QpMap values", the adjusted input Qp value is Q p_new =Q p_old -Q p_offset_stepsizeIt may be expressed as ×(lower_bound - offset_index).

[0150] In some examples, the encoder may signal a QP offset step size, a lower bound, and an offset index. The offset index may be an integer from 0 to 3, and the QP offset step size and the lower bound may each be a predefined constant value.

[0151] In some examples, the encoder may pre - define a QP offset step size and a lower bound. The encoder may further signal an offset index, which may be an integer from 0 to 3. For example, the QP offset step size and the lower bound may be predefined as constant values. In this case, the QP offset step size and the lower bound do not need to be signaled. When the QP offset step size and the lower bound are not constant values, signaling is required.

[0152] In some examples, the processor 1820 may obtain a QP scaling factor by obtaining a QP offset, and may adjust a plurality of QpMap values by subtracting the QP offset from the input QP value.

[0153] In some examples, the QP scaling factor may be obtained using an equation such as those shown in equations (9) and (10) TIFF0007695462000015.tif15170, where α CH represents the QP scaling factor and QP offset represents the QP offset.

[0154] In some examples, the QP offset may be determined for each input video block.

[0155] In some examples, the QP offset may be signaled at the frame level.

[0156] In some examples, the processor 1820 may pre-define a LUT, which may include a plurality of codewords and a plurality of QP offsets corresponding to the plurality of codewords as shown in Table 1 or 2.

[0157] In some examples, the processor 1820 may further signal a codeword so that the decoder fetches a QP offset corresponding to the codeword based on the LUT.

[0158] In some examples, the processor 1820 may pre-define a plurality of LUTs with different granularities and signal a codeword so that the decoder fetches a QP offset corresponding to the codeword based on the plurality of LUTs with different granularities.

[0159] For example, the LUT may be pre-defined at the frame level such that all video blocks within the same frame have the same LUT and the encoder only needs to signal the codeword. The decoder may receive the signaled codeword and fetch the corresponding QP offset based on the received codeword and the same LUT.

[0160] In some examples, the processor 1820 may pre-define a plurality of LUTs with different granularities and signal the codeword and the LUT step size difference so that the decoder fetches a QP offset corresponding to the codeword based on the codeword and the LUT step size difference. As shown in Table 1 and Table 2, the LUT step size difference is the difference between a first step size and a second step size. The first step size is the difference between two adjacent QP offsets in the first LUT, and the second step size is the difference between two adjacent QP offsets in the second LUT.

[0161] For example, as shown in Tables 1 to 2, the first step size in Table 1 is 4, the second step size in Table 2 is 5, and the LUT step size difference is 1 (=5 - 4). In some examples, the encoder does not need to send all the LUTs to the decoder such that the decoder can find the corresponding QP offset based on the LUT step size difference and the LUT previously stored in the decoder, but only needs to send the LUT step size difference.

[0162] In some other examples, a non - transitory computer - readable storage medium 1804 storing instructions is provided. When the instructions are executed by one or more processors 1820, the instructions cause the processor to perform in any manner as described in FIG. 19 and above.

[0163] The description of the present disclosure is presented for purposes of illustration and is not intended to be exhaustive or to limit the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.

[0164] The examples were chosen and described in order to explain the principles of the present disclosure and to enable those skilled in the art to understand the disclosure for various implementations and to best utilize the principles and various implementations with various modifications as are suited to the particular use anticipated. Accordingly, it is to be understood that the scope of the present disclosure should not be limited to the specific examples of the disclosed implementations, and that modifications and other implementations are intended to be included within the scope of the present disclosure.

Claims

1. A method for image filtering during video encoding, loading a plurality of QpMap values in one or more quantization parameter (QP) map (QpMap) channels into a neural network for image filtering; obtaining a QP scaling factor by adjusting a plurality of input QP values related to an input frame; adjusting the plurality of QpMap values according to the QP scaling factor so that the neural network learns and filters the input frame; comprising wherein adjusting the plurality of input QP values related to the input frame is to obtain a QP offset based on a QP offset step size, a lower limit, and an offset index, wherein the QP offset step size is a step size for adjusting each input QP value, the lower limit is an integer value for determining the maximum QP value reduction, and the offset index is a signaled index value; and subtracting the QP offset from the input QP value.

2. further comprising signaling, by an encoder, the QP offset step size, the lower limit, and the offset index, wherein the offset index is an integer from 0 to 3, the method according to claim 1.

3. predefining, by an encoder, the QP offset step size and the lower limit; further comprising signaling, by the encoder, the offset index, wherein the offset index is an integer from 0 to 3, the method according to claim 1.

4. obtaining the QP scaling factor includes obtaining a QP offset, The method according to claim 1, wherein adjusting the plurality of QpMap values includes subtracting the QP offset from the input QP value.

5. Obtaining the QP scaling factor includes obtaining a QP offset, and 【Equation 1】 obtaining the QP scaling factor using an operation such as α CH represents the QP scaling factor, and QP offset represents the QP offset, the method according to claim 1.

6. The method according to claim 5, further comprising determining the QP offset for each input video block.

7. The method according to claim 5, further comprising signaling the QP offset at the frame level by an encoder.

8. Defining in advance a look-up table (LUT) by an encoder, the LUT including a plurality of codewords and a plurality of QP offsets corresponding to the plurality of codewords, and signaling the codewords by the encoder such that a decoder retrieves a QP offset corresponding to a codeword based on the LUT, and The method according to claim 4, further comprising.

9. Defining in advance by an encoder a plurality of look-up tables (LUTs) at different granularities, and signaling the codewords by the encoder such that a decoder retrieves a QP offset corresponding to a codeword based on the plurality of LUTs at different granularities, and The method according to claim 8, further comprising.

10. By an encoder, a plurality of look-up tables (LUTs) are predefined at different granularities, and the encoder signals the codeword and the LUT step size difference such that a decoder extracts a QP offset corresponding to the codeword based on the codeword and the LUT step size difference, where the LUT step size difference is a difference between a first step size and a second step size, the first step size is a difference between two adjacent QP offsets in a first LUT, and the second step size is a difference between two adjacent QP offsets in a second LUT, and The method according to claim 8, further comprising. **Claim 11** An apparatus for image filtering during video encoding using a neural network, comprising one or more processors, and a memory coupled to the one or more processors and configured to store instructions and a bitstream executable by the one or more processors. When the one or more processors execute the instructions, they are configured to perform the method according to any one of claims 1 to 10 to generate the bitstream. Apparatus. **Claim 12** A non-transitory computer-readable storage medium storing computer-executable instructions and a bitstream, wherein when the computer-executable instructions are executed by one or more computer processors, the one or more computer processors are caused to store the bitstream and perform the method according to any one of claims 1 to 10 to generate the bitstream. **Claim 13** A method for storing a bitstream, the bitstream being generated by a video encoding method, wherein the video encoding method Loading a plurality of QpMap values in one or more quantization parameter (QP) map (QpMap) channels into a neural network for image filtering in video encoding, Obtaining a QP scaling factor by adjusting a plurality of input QP values for an input frame, Adjusting the plurality of QpMap values according to the QP scaling factor for the neural network to learn and filter the input frame into the neural network, including: Adjusting the plurality of input QP values for the input frame is Obtaining a QP offset based on a QP offset step size, a lower limit, and an offset index, where the QP offset step size is a step size for adjusting each input QP value, the lower limit is an integer value for determining a maximum QP value reduction, and the offset index is a signaled index value; Subtracting the QP offset from the input QP value, including a method.

Citation Information

Patent Citations

  • Image filter device, image decoding device, and image encoding device

    JP2019201255A

  • Methods and Apparatuses of Video Encoding or Decoding with Adaptive Quantization of Video Data

    US20200404275A1

  • Apparatus and method for applying artificial neural network to image encoding or decoding

    US20210021823A1