On architectures for neural network filters

Neural network filters are integrated into video codecs to enhance image and video quality for machine analysis, addressing the limitations of conventional codecs by improving distortion reduction and quality for machine vision tasks.

WO2025149917A1PCT designated stage expired Publication Date: 2025-07-17NOKIA TECHNOLOGIES OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050203
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2025-01-08
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Conventional video codecs struggle to optimize image and video compression for machine analysis tasks, as they are primarily designed for human perceptual quality, neglecting the specific requirements of machine consumption.

Method used

Implementing neural network filters within the encoding and decoding loops of video codecs to enhance image and video quality for machine analysis, utilizing architectures that include convolutional layers, summary-based modulation, and auxiliary neural networks to blend signals, ensuring effective processing of luma and chroma components.

Benefits of technology

The proposed solution improves the performance of machine analysis tasks by reducing distortion and enhancing the quality of compressed data, thereby improving the effectiveness of machine vision applications such as object detection and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050203_17072025_PF_FP_ABST
    Figure IB2025050203_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An example apparatus includes: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a luma of at least one image; determine a chroma of the at least one image; and filter, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.
Need to check novelty before this filing date? Find Prior Art

Description

ON ARCHITECTURES FOR NEURAL NETWORK FILTERSTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to multimedia transport and, more particularly, to architectures for neural network filters.BACKGROUND

[0002] It is known to perform data compression and data decompression in a multimedia system.BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0004] FIG. 1 shows a system pipeline for VCM.

[0005] FIG. 2 shows an example architecture of an NN filter.

[0006] FIG. 3 shows an example architecture of an NN filter.

[0007] FIG. 4 shows an example SBM layer.

[0008] FIG. 5 shows an example architecture of an NN filter.

[0009] FIG. 6 shows an example architecture of an NN filter.

[0010] FIG. 7 shows an example where the output of the auxNN represents the blended output.

[0011] FIG. 8 shows an example where the output of the auxNN represents one blending map, or two or more blending maps.

[0012] FIG. 9 shows an example where the luma backbone blocks in the general architecture of the NN filter are densely connected.

[0013] FIG. 10 shows an example architecture of an NN filter.

[0014] FIG. 11 is a block diagram illustrating a system in accordance with an example.

[0015] FIG. 12 is an example apparatus configured to implement the examples described herein.

[0016] FIG. 13 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0017] FIG. 14 is an example method, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0018] Fundamentals of neural networks

[0019] A neural network (NN) is a computation graph consisting of several layers of computation. Each layer consists of one or more units, where each unit performs an elementary computation. A unit is connected to one or more other units, and the connection may be associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, i.e., values which can be learned from training data. There may be other learnable parameters, such as those of batch-normalization layers.

[0020] Neural networks are being utilized in an ever-increasing number of applications for many different types of device, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, etc.

[0021] One of the properties of neural nets (and other machine learning tools) is that they are able to learn properties of input data or learn to perform tasks given input data, based on a learning algorithm, or training algorithm. The learning algorithm may comprise (but may not be limited to) one or more of a supervised learning algorithm, a self-supervised learning algorithm, an unsupervised learning algorithm. Examples ofproperties of input data that may be learned comprise categories, correlations, causation relations, etc. Examples of tasks comprise classification or categorization, regression, prediction, etc.

[0022] In general, the training algorithm consists of changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network can be used to derive a class or category index which indicates the class or category that the object in the input image belongs to. Training usually happens by minimizing or decreasing the output’s error, also referred to as the loss. Examples of losses are mean squared error (MSE), cross-entropy, etc. The training algorithm may comprise an iterative process, where at each iteration the algorithm may modify the weights of the neural net to make a gradual improvement of the network’s output, i.e., to gradually decrease the loss.

[0023] The terms “model”, “neural network”, “neural net”, “network”, “NN” are used herein interchangeably, and also the weights of a neural network are sometimes referred to as learnable parameters or simply as parameters.

[0024] Training a neural network may be regarded as an optimization process, where the goal is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, i.e., data which was not used for training the model. This is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, i.e., to modify its learnable parameters in order to minimize the loss. The validation set is used for checking the performance of the network on data which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set are monitored during the training process to understand the following things:

[0025] —If the network is learning at all - in this case, the training set error should decrease, otherwise the model is in the regime of underfitting.

[0026] —If the network is learning to generalize - in this case, also the validation set error needs to decrease and to be not too much higher than the training set error. If the training set error is low, but the validation set error is much higher than the training set error, or it does not decrease, or it even increases, the model is in the regime of overfitting. This means that the model has just memorized the training set’s properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters.

[0027] Neural networks have been used for compressing and de-compressing data such as images, i.e., in an image codec. One common architecture for realizing one component of an image codec is the auto-encoder, which is a neural network consisting of two parts: a neural encoder and a neural decoder. The neural encoder takes as input an image and produces a code which requires less bits than the input image. This code may be obtained by applying a binarization or quantization process to the output of the encoder. The neural decoder takes in this code and reconstructs the image which was input to the encoder.

[0028] Such neural encoder and neural decoder are usually trained to minimize a combination of bitrate and distortion, where the distortion may be based on one or more of the following metrics: Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), or similar.

[0029] Fundamentals of video / image coding

[0030] A video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).

[0031] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closelyto the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).

[0032] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).

[0033] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred to in the prediction process. Inter-layer or inter- view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter- view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.

[0034] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0035] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy- coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.

[0036] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.

[0037] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or colocated blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporalreference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0038] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.

[0039] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor I to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + .R where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0040] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications contain both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit contains one or more SEI messages, which are not required for the decoding of output pictures but may assistin related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC. and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.

[0041] Background information on Video Coding for Machines (VCM)

[0042] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, i.e. consuming / watching the decoded image. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (i.e., autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. Example use cases and applications are self-driving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. This may raise the following question: when decoded data is consumed by machines, shouldn’t the aim be at a different quality metric -other than human perceptual quality- when considering media compression in inter-machine communications? Also, dedicated algorithms for compressing and decompressing data for machine consumption are likely to be different than those for compressing anddecompressing data for human consumption. The set of tools and concepts for compressing and decompressing data for machine consumption is referred to here as Video Coding for Machines.

[0043] It is likely that the receiver-side device has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is for example determined by an orchestrator sub-system. The multiple machines may be used for example in succession, based on the output of the previously used machine, and / or in parallel. For example, a video which was compressed and then decompressed may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.

[0044] Also, please notice that the term “receiver-side” or “decoder-side” is used herein to refer to the physical or abstract entity or device which contains one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.

[0045] The encoded video data may be stored into a memory device, for example as a file. The stored file may later be provided to another device.

[0046] Alternatively, the encoded video data may be streamed from one device to another.

[0047] FIG. 1 is a general illustration of the pipeline 100 of Video Coding for Machines. A VCM encoder 104 encodes the input video 102 into a bitstream 106. A bitrate 110 may be computed 108 from the bitstream 106 in order to evaluate the size of the bitstream 106. A VCM decoder 112 decodes the bitstream 106 output by the VCM encoder 104. The output of the VCM decoder 112 is referred to in FIG. 1 as “Decoded data for machines” 114. This data 114 may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline 100, this data 114 may not have the same or similar characteristics as the original video 102 which was input to the VCM encoder 104. For example, this data 114 may not be easilyunderstandable by a human by simply rendering the data onto a screen. The output 114 of VCM decoder 112 is then input to one or more task neural network (116-1, 116-2, 116-3, 116-X). In FIG. 1, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs (116-1, 116-2, 116-3), and a non-specified one (Task-NN X 116-X). The goal of VCM is to obtain a low bitrate while guaranteeing that the task-NNs (116-1, 116-2, 116-3, 116-X) still perform well in terms of the evaluation metric associated to each task.

[0048] As shown in FIG 1, the pipeline 100 includes evaluating the performance (118-1) of Task-NN 1 (116-1) used for object detection to determine a Task-NN 1 performance (120-1), evaluating the performance (118-2) of Task-NN 2 (116-2) used for object segmentation to determine a Task-NN 2 performance (120-2), evaluating the performance (118-3) of Task-NN 3 (116-3) used for object tracking to determine a Task-NN 3 performance (120-3), and evaluating the performance (118-X) of Task-NN X (116-X) used for another task to determine a Task-NN X performance (120-X).

[0049] When a conventional video encoder, such as a H.266 / VVC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks:

[0050] —One or more regions of interest (ROIs) may be detected. An ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways: The quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise; The video is preprocessed to contain only the ROIs, while the other areas are replaced by one or more constant values or removed; A grid is formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that contain no ROIs are downsampled as preprocessing to encoding.

[0051] —Quantization parameter of the highest temporal sublayer(s) is increased (i.e.coarser quantization is used) when compared to practices for human watchable video.

[0052] —The original video is temporally downsampled as preprocessing prior to encoding. A frame rate upsampling method may be used as postprocessing subsequent to decoding, if machine analysis at the original frame rate is desired.

[0053] —A filter is used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.

[0054] Background on neural network based filtering

[0055] In some video codecs, a neural network may be used as filter in the encoding and decoding loop (also referred to simply as coding loop), and it may be referred to as neural network loop filter, or neural network in-loop filter. The NN loop filter may replace all other loop filters of an existing video codec, or may represent an additional loop filter with respect to the already present loop filters in an existing video codec.

[0056] In the context of image and video enhancement, a neural network may be used as post-processing filter, for example applied to the output of an image or video decoder in order to remove or reduce coding artifacts.

[0057] For simplicity, a neural network filter or an NN filter is referred to as a filter that comprises one or more neural networks and is used either as a loop filter in the coding loop or as a post-processing filter.

[0058] The following example system is used in several embodiments to illustrate or describe the idea. The example system comprises a codec that comprises one or more NN loop filters. For example, the codec could comprise a modified version of a VVC / H.266 compliant codec (e.g., a VVC / H.266 compliant codec that has been modified so that it would comprise one or more NN loop filters, where the modified version of a VVC / H.266 compliant codec may not anymore be a VVC / H.266 compliant codec). The input to the one or more NN loop filters may comprise at least a reconstructed block or frames (simply referred to as reconstruction) or data derived from a reconstructed block or frame (e.g., the output of a conventional loop filter). The reconstruction may be obtained based on predicting a block or frame (e.g., by means ofintra-frame prediction or inter-frame prediction) and performing residual compensation. The one or more NN loop filters (may be referred to simply as NN filters in some of the embodiments) may enhance the quality of at least one of their input, so that a rate-distortion loss is decreased. The rate may indicate a bitrate (e.g., an estimated bitrate or a real bitrate) of the encoded video. The distortion may indicate a pixel fidelity distortion such as one or more of the following: Mean-squared error (MSE); Mean absolute error (MAE); Mean Average Precision (mAP) computed based on the output of a task NN (such as an object detection NN) when the input is the output of the postprocessing NN; Other machine task-related metric, for tasks such as object tracking, video activity classification, video anomaly detection, etc.

[0059] The enhancement may result into a coding gain, which can be expressed for example in terms of BD-rate or BD-PSNR.

[0060] However at least some of the embodiments described herein are applicable to a NN filter which is not a loop filter of a codec. For example, the NN filter may be a NN post-processing filter, whose input may comprise one or more outputs of a video codec. In this case, the filter may be used only for increasing a quality metric of at least one of its inputs, where the quality metric may be, for example, peak signal-to-noise ratio (PSNR), mAP for object detection, MOTA for object tracking, etc.

[0061] The examples described herein address the general problem of filtering an input data item, such as an image or a video frame, for one or more purposes including: enhancing the visual quality, enhancing machine analysis results, etc.

[0062] More specifically, the examples described herein are directed to the architecture of a NN filter.

[0063] General information

[0064] For the sake of simplicity, at least some embodiments are described herein as applied to a filter. A filter takes as input at least one or more first images to be filtered and outputs at least one or more second images, where the one or more second images are the filtered version of the one or more first images. In one example, the filter takesas input one image and outputs one image. In another example, the filter takes as input more than one images and outputs one image. In another example, the filter takes as input more than one image and outputs more than one image.

[0065] It is to be understood that a filter may take as input also other data (also referred to as auxiliary data) than the data that is to be filtered, such as data that can aid the filter to perform a better filtering than if no auxiliary data was provided as input. In one example, the auxiliary data comprises information about prediction data, and / or information about the picture type, and / or information about the slice type, and / or information about a Quantization Parameter (QP) used for encoding, and / or information about boundary strength, etc. In one example, the filter takes as input one image and other data associated to that image, such as information about the quantization parameter (QP) used for quantizing and / or dequantizing that image, and outputs one image.

[0066] A filter may be a neural network based filter, or may be another type of filter. However, several embodiments describe training aspects which may be applicable to machine learning based filters such as neural network based filters.

[0067] A filter may be, for example, an in-loop filter that is used in the coding loop of a codec, or a post-processing filter applied on the data decoded by the codec.

[0068] Even if at least some of the embodiments are described with reference to a filter, those embodiments may be applied also to other operations than just filters, such as an operation performing intra-frame prediction, or an operation performing interframe prediction, or an operation performing frame-rate upsampling, or an operation performing encoding and / or decoding (e.g., an end-to-end learned codec).

[0069] While at least some embodiments are described such that the input and output data are in the form of images or (video) frames or pictures, those embodiments may be applicable also to other types of data, such as audio frames. Furthermore, while at least some embodiments are described by considering a full image, those embodiments may be applicable also to one or more blocks or portions of an image.

[0070] General architecture considered by the examples described herein

[0071] Considered herein is an example architecture of a NN filter in order to describe one or more of the herein described embodiments. An illustration of such architecture 200 is shown in FIG. 2.

[0072] In FIG. 2, “luma” 202 and “chroma” 204 refer to the reconstructed luma and chroma that are to be enhanced by the NN filter, and may represent an intermediate result of an encoding or decoding operation. For example, they may represent the result of combining a predicted block with a decoded residual. The luma 202 and chroma 204 are concatenated to form the reconstruction Rec 206. The luma 202 and chroma 204 are assumed to have the same size in terms of height and width. In case the format of data is YUV 420 (201), or anyway a format such that the chroma 204 has smaller resolution than the luma 202, such as half height and half width with respect to the height and width, respectively, of luma 202, then the chroma 204 may be upsampled to match the height and width of the luma 202 (and this upsampling operation is denoted as “Up2” 203 in FIG. 2, referring to upsampling by 2 in both the height dimension and the width dimension). As used herein, chroma or input chroma are referred to as the chroma 204 which may have been upsampled in order to match the height and width of luma 202, and original chroma or original input chroma are referred to as the chroma before any upsampling that was performed in order to match the height and width of luma 202.

[0073] The terms Rec 206, Pred 208, BS 210, BaseQP 212, SliceQP 214, IPB 216 represent the inputs to the NN filter and are usually in the format of a tensor of shape BxCxHxW, where B indicates a batch size, C indicated a number of channels, H and W indicate a height and width, respectively. The square brackets and the number within them (e.g., Rec[3]), indicate the number of channels of the associated tensor. For example, Rec[3] indicates that the input tensor Rec 206 contains 3 channels, thus has shape Bx3xHxW, where the 3 channels may represent the luma channel, the Blue- Yellow Chrominance (Cb) channel and the Red-Green Chrominance (Cr) channel. The Cb channel and the Cr channel may be collectively referred to as chroma.

[0074] Pred (208) stands for prediction. BS (210) stands for boundary strength,BaseQP (212) stands for the sequence-level quantization parameter (QP), SliceQP (214) stands for the slice-level QP, IPB (216) stands for the type of slice or type of picture.

[0075] Rec 206 may be referred to as main input, or data to be filtered, whereas Pred 208, BS 210, BaseQP 212, SliceQP 214 and IPB 216 may be referred to as auxiliary input, or auxiliary data, or data not to be filtered.

[0076] Each block in FIG. 2 represents an operation, such as a NN layer or a combination of NN layers. A block “ConvKlxK2,Z”, where KI, K2, Z may be denoted differently for different blocks or layers in FIG. 2, indicates a convolutional layer with kernel size KlxK2 and number of kernels equal to Z. When present, the term “+ PReEU” indicates that a layer is followed by a Parametric Rectified Einear Unit (PReLU). When present, the term “s=2” indicates that a convolutional layer has stride equal to 2; when not present, the convolutional layer has stride equal to 1. “Split” (220) refers to an operation that splits a tensor across the channel dimension. “Luma backbone block” (such as 222-1) and “Chroma backbone block” (such as 226-1) indicate backbone blocks used for filtering or processing the luma channel and the chroma channel, respectively; the architecture of a backbone block 230 is also shown in the figure and comprises several layers and operations. “SepConv3x3” indicates a block (241, 242, 243) that comprises a separable convolution; an illustration of the SepConv3x3 block 240 is also shown in FIG. 2 and it comprises several layers. “PixelShuffle” refers to an operation 244 that rearranges elements in a tensor of shape Bx(C*r*r)xHxW to a tensor of shape BxCx(H*r)x(W*r), where r is an upscaling or upsampling factor. “Downsample by 2” indicates an operation 246 that downsamples the input by a factor of 2. In particular, in FIG. 2, the input chroma 204 is downsampled by 2. “LumaOut” (250) and “ChromaOut” (252) represent the filtered luma and the filtered chroma, respectively, i.e., the final outputs from the NN filter.

[0077] For the sake of simplicity, the NN architecture 200 is figuratively organized into the following sections: head 254, fuse 256, transition 258, luma backbone 260, chroma backbone 262, luma tail 264, chroma tail 266. However, it is to be noted that other organizations of the NN into subsets or blocks or sections may be possible.

[0078] All the input tensors are input to respective convolutional layers (271, 272, 273, 274, 275, 276) that are part of the “head” section 254 of the NN filter. The outputs of those convolutional layers are tensors, referred to as head tensors. As part of the operations of the “fuse” section 256 of the NN, the head tensors are concatenated 278 into a single tensor across the channel dimension. The concatenated tensor is input to a convolutional layer, followed by a non-linear activation function PreLU (279). The output of the fuse section 256 is input to the “transition” section 258 of the NN, which comprises a convolutional layer with stride equal to 2, followed by a PReLU activation function (280). The output of the transition section is a tensor of shape Bx(2*C)x(H / 2)x(W / 2), and it is split 220 into two sub-tensors, where each subtensor is of shape BxCx(H / 2)x(W / 2). A first subtensor is used to filter the luma and a second subtensor is used to filter the chroma. The first subtensor is input to the “luma backbone” section 260, and the second subtensor is input to the “chroma backbone” section 262. The luma backbone section 260 comprises Ny luma backbone blocks (including luma backbone block 1 222-1 and luma backbone block Ny 222-Ny), and the chroma backbone section 262 comprises Nuv chroma backbone blocks (including chroma backbone block 1 226-1 and chroma backbone block Nuv 226-Nuv). The output of a backbone block is input to the next backbone block, until the last backbone block in the section. The output of the last luma backbone block 222-Ny is input to the “luma tail” section 264, that comprises a SepConv3x3 block, followed by a PReLU operation (241), a Conv3x3 layer 281 and a PixelShuffle operation 244. The output of the PixelShuffle operation 244 is added 282 to the input luma 202, in order to obtain LumaOut 250. The output of the last chroma backbone block 226-Nuv is input to the “chroma tail” section 266, that comprises a SepConv3x3 block, followed by a PReLU operation (242), and a Conv3x3 layer 283. The output of the chroma tail section 266 is added 284 to the input chroma 204, in order to obtain ChromaOut 252.

[0079] In FIG. 2, also examples of the values of various hyper-parameters (collectively 290) of the NN are indicated, such as the number of channels of convolutional layers, DI, D2, D3, D4, D5, D6, C, Cl, the number of luma and chroma backbone blocks Ny and Nuv.

[0080] The example backbone block 230 includes an input 291, a Convlxl layer,followed by a PReLU operation (292), a Convlxl layer 293, a SepConv3x3 layer 294, an addition operation 285, and an output 295. The example SepConv3x3 layer 240 includes an input 296, a Conv3xl layer 297, a Convlx3 layer 298, and an output 286.

[0081] Embodiments on luma and chroma resolutions

[0082] Considered herein is the case where the data format is YUV 4:2:0, i.e., the height and width of the chroma components are half with respect to the height and width of the luma component, and where the input chroma that is provided to the filter as part of the Rec tensor has been upsampled in order to have same height and width of luma.

[0083] When an operation gets as input a first tensor (or component, or channel) and outputs a second tensor, the “effective height and width” (or, similarly, “effective resolution”) of the second tensor is referred to as being the same or substantially the same as the height and width of the first tensor, if the operation does not introduce or remove any or significant information.

[0084] In one example, the operation is a spatial upsampling operation that is based on the “nearest” or “nearest neighbour” algorithm; this operation re-uses some of the pixels already present in the input tensor in order to output another tensor with higher resolution. The output tensor has same effective resolution as the input tensor. If the output (upsampled) tensor is then downsampled by a downsampling operation, the effective resolution of the downsampled operation is still the same as the effective resolution of the upsampled tensor and of the input tensor to the upsampling operation, even when there are other operations in between the upsampling and downsampling which do not affect the spatial resolution.

[0085] One embodiment herein relates to processing, in at least a portion of the architecture of the NN filter, the chroma and luma components in such a way that the relation or ratio of their effective height and width is same as in the original input data format.

[0086] For example, if the original data format is YUV 4:2:0, the ratio of height of chroma and height of luma is half, and the ratio of width of chroma and width of lumais half, thus the chroma and luma are processed, at least in part of the NN, so that the ratio of their effective heights is half and the ratio of their effective widths is half.

[0087] By considering the example architecture in FIG. 2, even after the upsampling operation 203 performed on the input chroma 204 (denoted as “Up2” in the figure), the ratio of effective heights and the ratio of effective widths of chroma and luma are half in Rec tensor 206 and continue to be half until the “transition” section 258, where the whole input tensor (that comprises information about both luma and chroma) is downsampled by 2 by means of a convolutional layer with stride equal to 2 (280) and then the output tensor is split 220 into two sub-tensors that are input to two respective paths (the luma backbone 260 and the chroma backbone 262). The two sub-tensors may also be referred to as split tensors. The two sub-tensors, representing (processed) luma and chroma, have now different ratios of effective heights and widths. In particular, the effective height and width of chroma are the same as the height and width of the input chroma, and the effective height and width of luma are half of the height and width of the input luma, thus the ratio changed from half (at the input) to 1 when considering the sub-tensors.

[0088] Referring to FIG. 3, one embodiment described herein relates to decomposing the convolutional layer in the “transition” section 258 of FIG. 2 into two convolutional layers (381, 382), where both the two convolutional layers (381, 382) get as input the output 357 of the “fuse” section 356, and where a first convolutional layer 381 of the two convolutional layers has stride equal to 2 and a second convolutional layer 382 of the two convolutional layers has stride equal to 4, and where the output 383 of the first convolutional layer 381 represents (processed) luma and the output 384 of the second convolutional layer 382 represents (processed) chroma.

[0089] In an alternative embodiment, both the first and second convolutional layers 381 and 382 have stride equal to 2, and a third convolutional layer with stride equal to 2 or a downsampling operation with downsampling factor equal to 2 is applied to the output of the second convolutional layer 382. Then, the output of the third convolutional layer represents the (processed) chroma tensor 384.

[0090] In one embodiment, the ChromaOut tensor (252, 352) may be upsampled, forexample by a PixelShuffle operation.

[0091] Thus FIG. 3 illustrates an example of some of these embodiments (of architecture 300), where the transition section 358 has been modified so that there are two convolutional layers (381, 382) instead of one (e.g. convolutional layer 280), and where the two convolutional layers (381, 382) perform a downsampling operation with different downsampling factors so that their output tensors (383, 384) have the same ratio of effective resolutions.

[0092] Embodiments on summary-based modulation

[0093] In one embodiment, a new NN layer for the NN filter is introduced, implemented, and / or used, and is referred to as summary-based modulation (SBM) layer.

[0094] In one embodiment, an input tensor to an SBM layer is first summarized by means of a summarization operation, obtaining a summary tensor, then the summary tensor may be processed by zero, one or more NN layers, obtaining a processed summary tensor, then the processed summary tensor is combined with the input tensor (or with a tensor derived from the input tensor) by means of a combination operation.

[0095] In one embodiment, the summarization operation reduces the height and width of the input tensor to a smaller height and width. In one example, the input tensor has shape BxCxHxW, the summarized tensor has shape BxCxlxl, i.e., the summarized tensor has height equal to 1 and width equal to 1.

[0096] In another embodiment, the summarization operation reduces the number of channels of the input tensor to a smaller number of channels. In one example, the input tensor has shape BxCxHxW, the summarized tensor has shape BxlxHxW.

[0097] In one embodiment, the summarization operation is a mean or average operation. In one example, the summarization operation is a mean computed over the elements on the height and width axes and separately for each sample in the batch axis and separately for each channel. In another example, the summarization operation is a mean computed over all channels separately for each sample on the batch axis andseparately for each spatial element.

[0098] In another embodiment, the summarization operation is a sum operation.

[0099] In another embodiment, the summarization operation is a median operation.

[0100] In another embodiment, the summarization operation is a mode operation (i.e., computing the most frequent value).

[0101] In another embodiment, the summarization operation is a minimum operation (i.e., computing the minimum value).

[0102] In another embodiment, the summarization operation is a maximum operation (i.e., computing the maximum value).

[0103] In one embodiment, when the tensors to be combined by the combination operation have different shape, a broadcasting operation in one or more dimensions or axes may be performed on one or more of the tensors to be combined.

[0104] In one embodiment, the combination operation may be an element-wise sum.

[0105] In one example, the summarized tensor (and the processed summarized tensor) has shape BxCxlxl, and is added to the input tensor by broadcasting the same value to all spatial positions.

[0106] In one embodiment, the combination operation may be an element-wise multiplication.

[0107] FIG. 4 illustrates an example of SBM layer 400, where Input 410 represents the input tensor, Output 450 represents the output tensor, Summarization 420 represents a summarization operation and Conv 430 represents a convolutional layer. The processed summarized tensor 435 is multiplied with the input tensor 410 by means of the Multiply block 440.

[0108] In one embodiment, the SBM layer 400 is used before one or more convolutional layers of the architecture (200, 300) of the NN filter.

[0109] Embodiments on learned blending

[0110] Referring to FIG. 7, in one embodiment, an auxiliary neural network (auxNN) 708 is used within architecture 700 to blend two or more signals or tensors, where at least one of the two or more signals or tensors is an output 706 of an NN filter 704. In the example shown in FIG. 7, the other one of the blended two or more signals or tensors is input data 702. The auxNN 708 produces blended data 710. The blended data 710 may represent the final output of the filtering process.

[0111] In one embodiment, the two or more signals or tensors comprise an input to the NN filter and an output from the NN filter.

[0112] In one example, one of the two or more tensors comprises the luma and chroma that are input to the NN filter and another of the two or more tensors comprises the filtered luma and chroma that are output from the NN filter.

[0113] In one embodiment, an input to the auxNN comprises one or more input to the NN filter.

[0114] In one example, the input to the auxNN is same as the inputs to the NN filter, i.e., the main input Rec and all other auxiliary inputs.

[0115] In one embodiment, an input to the auxNN may be an output from the NN filter.

[0116] In one embodiment, the input to the auxNN comprises all the inputs to the NN filter and also the output from the NN filter.

[0117] In one embodiment, the auxNN has a similar architecture as the NN filter.

[0118] In one embodiment, the input Rec to the auxNN is added to an output of the last NN layer or block of the auxNN. FIG. 5 illustrates an example of an auxNN (within architecture 500) where the auxNN takes as input all the inputs to the NN filter and also the output of the NN filter. The output of the NN filter is denoted as FilterOut 518.

[0119] In an alternative embodiment, when an input to the auxNN is an output of theNN filter, the output of the NN filter is added to an output of the last NN layer or block of the auxNN. FIG. 6 illustrates an example of an auxNN (of architecture 600) where the auxNN takes as input all the inputs to the NN filter and also the output of the NN filter. The output of the NN filter is denoted as FilterOut 604. The components of the FilterOut tensor are added to the output of the last blocks.

[0120] In one embodiment, the auxNN is trained jointly with the NN filter.

[0121] In another embodiment, the auxNN and the NN filter are trained separately.

[0122] In one example, the NN filter is trained first, then the auxNN is trained by keeping the NN filter frozen.

[0123] In another example, the NN filter is trained first, then the auxNN is trained and the NN filter is finetuned jointly with the training of the auxNN.

[0124] In one embodiment, referring to FIG. 7, the output 710 of the auxNN 708 represents the blended output 710, such as the blending luma and chroma. Thus, the auxNN 708 represents the blending operation applied to the two or more signals (for example input data 702, filtered data 706) to be blended. FIG. 7 illustrates an example of this embodiment.

[0125] In another embodiment, referring to FIG. 8, the output 812 of the auxNN 808 of architecture 800 represents one blending map 812 or two or more blending maps 812, where the blending map 812 or blending maps 812 may be used for blending the two or more tensors (such as input data 802 and filtered data 806) by means of a blending operation 810. FIG. 8 illustrates an example of this embodiment. In FIG. 8, NN filter 804 uses input data 802 to generate filtered data 806.

[0126] In one embodiment, the blending operation may be a linear combination, or a weighted combination, where the coefficients or weights are derived based on the blending map.

[0127] In another embodiment, the blending operation may be a learned module, such as another neural network that gets as input the two or more tensors to be blended andthe blending map.

[0128] In one embodiment, an input to the auxNN may be derived from a signal or data that is signaled from an encoder.

[0129] Embodiments on dense connections

[0130] In one embodiment, three or more NN layers of the NN filter are densely connected. By densely connected it is meant that the input to a certain layer comprises or is derived from the input and / or output of some or all the previous layers (in a processing order) that are comprised in the three or more NN layers.

[0131] In one embodiment, an input to a certain layer may be derived from the input and / or output of some or all the previous layers based at least on one or more NN layers.

[0132] In one example, the luma backbone blocks (such as luma backbone block 1 222-1, luma backbone block Ny 222-Ny) in the general architecture of the NN filter are densely connected. FIG. 9 illustrates this example, where there are three backbone blocks denoted as Backbone block 1 904, Backbone block 2 912, Backbone block 3 920; the input to these blocks are Inpl 902, Inp2 910, Inp3 918, and the outputs from these blocks are Outl 906, Out2 914, Out3 922; the Combine Inputs blocks (908, 916) combine two or more inputs in order to generate an input for the next Backbone block. The Combine inputs block (for example combine inputs block 908 and / or combine inputs block 916) may comprise one or more NN layers.

[0133] Embodiments on conditioning between components

[0134] In one embodiment, one or more intermediate or final tensors from which a filtered luma tensor may be derived are used as input to one or more NN layers whose outputs may be used to derive a filtered chroma tensor.

[0135] In one embodiment, the tensor representing the filtered luma, or a signal derived from the tensor representing the filtered luma, is used as one of the inputs to one or more NN layers that are used for filtering chroma.

[0136] In one example, referring to the example of the general architecture of the NNfilter, the output of the last convolutional layer of the luma tail (before PixelShuffle) is used as an input to all the Chroma backbone blocks. In another example, the output of the last convolutional layer of the luma tail is used as an input to the convolutional layers in the Chroma tail. In yet another example, the output of the last convolutional layer of the luma tail is input to one or more blocks, where each of the one or more blocks comprises one or more NN layers, and the outputs of the one or more blocks are used as an input to respective one or more layers or blocks of the Chroma backbone or the Chroma tail.

[0137] Similar embodiments could be derived for the other case of conditioning from chroma to luma, instead of luma to chroma.

[0138] FIG. 10 illustrates an example of these embodiments (within architecture 1000), where a tensor 1065 from the Luma tail 1064 is used as an input to a block 1070 denoted as “Luma2chroma”, which comprises one or more NN layers. The output 1071 of the “Luma2chroma” block 1070 is used as one of the inputs to the blocks and NN layers in the Chroma backbone 1062 and Chroma tail 1066. For example, an input to a block or NN layer in the Chroma backbone 1062 or Chroma tail 1066 may be built by means of concatenating an output 1071 of the “Luma2chroma” block 1070 with the output of a previous block or layer in the Chroma Backbone 1062 or Chroma tail 1066.

[0139] As shown in FIG. 10, the output 1071 of the Luma2chroma block 1070 is used as input to the chroma backbone block 1 1026-1 and to the chroma backbone block Nuv 1026-Nuv of the chroma backbone 1062, and the output 1071 of the Luma2chroma block 1070 is used as input to the sepConv3x3,C + PreLU block 1042 and to the Conv3x3,2 block 1083 of the chroma tail 1066.

[0140] Notes

[0141] NOTE: it is to be understood that, as used herein, the terms “picture”, "image", and "frame" may be used interchangeably. In some cases, also the term “block” and “picture” may be used interchangeably, as a block can be considered to comprise part of a picture.

[0142] NOTE: in some embodiments, the term “block” is used to describe a component of a block diagram in a figure; in some other embodiments, the term “block” is used to describe part of a picture. The actual meaning of the term “block” should be clear from the context to a person skilled in the art.

[0143] NOTE: it is to be understood that, as used herein, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably in at least some embodiments.

[0144] NOTE: it is to be understood that, as used herein, the terms “machine consumption” and “machine analysis” may be used interchangeably.

[0145] NOTE: it is to be understood that, as used herein, the terms “post-filter”, "postprocessing filter" and “postprocessing filter" may be used interchangeably.

[0146] FIG. 11 is a block diagram illustrating a system 1100 in accordance with an example. In the example, the encoder 1130 is used to encode video from the scene 1115, and the encoder 1130 is implemented in a transmitting apparatus 1180. The encoder 1130 produces a bitstream 1110 comprising signaling that is received by the receiving apparatus 1182, which implements a decoder 1140. The encoder 1130 sends the bitstream 1110 that comprises the herein described signaling. The decoder 1140 forms the video for the scene 1115-1, and the receiving apparatus 1182 would present this to the user, e.g., via a smartphone, television, or projector among many other options.

[0147] In some examples, the transmitting apparatus 1180 and the receiving apparatus 1182 are at least partially within a common apparatus, and for example are located within a common housing 1150. In other examples the transmitting apparatus 1180 and the receiving apparatus 1182 are at least partially not within a common apparatus and have at least partially different housings. Therefore in some examples, the encoder 1130 and the decoder 1140 are at least partially within a common apparatus, and for example are located within a common housing 1150. For example the common apparatus comprising the encoder 1130 and decoder 1140 implements a codec. In other examples the encoder 1130 and the decoder 1140 are at least partially not within acommon apparatus and have at least partially different housings, but when together still implement a codec.

[0148] 3D media from the capture (e.g., volumetric capture) at a viewpoint 1112 of the scene 1115, which includes a person 1113) is converted via projection to a series of 2D representations with occupancy, geometry, and attributes. Additional atlas information is also included in the bitstream to enable inverse reconstruction. For decoding, the received bitstream 1110 is separated into its components with atlas information; occupancy, geometry, and attribute 2D representations. A 3D reconstruction is performed to reconstruct the scene 1115-1 created looking at the viewpoint 1112-1 with a “reconstructed” person 1113-1. The “-1” are used to indicate that these are reconstructions of the original. As indicated at 1120, the decoder 1140 performs an action or actions based on the received signaling.

[0149] FIG. 12 is an example apparatus 1200, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 1200 comprises at least one processor 1202 (e.g., an FPGA and / or CPU), one or more memories 1204 including computer program code 1205, the computer program code 1205 having instructions to carry out the methods described herein, wherein the at least one memory 1204 and the computer program code 1205 are configured to, with the at least one processor 1202, cause the apparatus 1200 to implement circuitry, a process, component, module, or function (implemented with control module 1206) to implement the examples described herein, including encapsulating and streaming attenuation maps for green metadata. Filter 1230 of the control module 1206 implements the embodiments described herein related to architectures for neural network filters. The memory 1204 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).

[0150] The apparatus 1200 includes a display and / or I / O interface 1208, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one ormore sensors, etc. The apparatus 1200 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 1210. The communication FF(s) 1210 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 1224. The communication FF(s) 1210 may comprise one or more transmitters or one or more receivers.

[0151] The transceiver 1216 comprises one or more transmitters 1218 and one or more receivers 1220. The transceiver 1216 and / or communication I / F(s) 1210 may comprise standard well-known components such as an amplifier, filter, frequencyconverter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 1214 used for communication over wireless link 1226.

[0152] The control module 1206 of the apparatus 1200 comprises one of or both parts 1206-1 and / or 1206-2, which may be implemented in a number of ways. The control module 1206 may be implemented in hardware as control module 1206-1, such as being implemented as part of the one or more processors 1202. The control module 1206-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 1206 may be implemented as control module 1206-2, which is implemented as computer program code (having corresponding instructions) 1205 and is executed by the one or more processors 1202. For instance, the one or more memories 1204 store instructions that, when executed by the one or more processors 1202, cause the apparatus 1200 to perform one or more of the operations as described herein. Furthermore, the one or more processors 1202, one or more memories 1204, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.

[0153] The apparatus 1200 to implement the functionality of control 1206 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 1200 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 1200 may be part of a self-organizing / optimizing network (SON) node or other node, such as a node in a cloud.

[0154] The apparatus 1200 may also be distributed throughout the network includingwithin and between apparatus 1200 and any network element (such as a base station and / or terminal device and / or user equipment).

[0155] Interface 1212 enables data communication and signaling between the various items of apparatus 1200, as shown in FIG. 12. For example, the interface 1212 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 1205, including control 1206 may comprise object- oriented software configured to pass data or messages between objects within computer program code 1205. The apparatus 1200 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 1200 may at least partially reside in a common housing 1228, or a subset of the various components of apparatus 1200 may at least partially be located in different housings, which different housings may include housing 1228.

[0156] FIG. 13 shows a schematic representation of non-volatile memory media 1300a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 1300b (e.g. universal serial bus (USB) memory stick) and 1300c (e.g. cloud storage for downloading instructions and / or parameters 1302 or receiving emailed instructions and / or parameters 1302) storing instructions and / or parameters 1302 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein. Instructions and / or parameters 1302 may represent or correspond to a non-transitory computer readable medium.

[0157] FIG. 14 is an example method 1400, based on the example embodiments described herein. At 1410, the method includes determining a luma of at least one image. At 1420, the method includes determining a chroma of the at least one image. At 1430, the method includes filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output. Method 1400 may be performed with architecture 100, architecture 200, architecture 300, architecture 500, architecture 600, architecture 700, architecture 800, architecture 1000, transmitting apparatus 1180 with encoder 1130, receiving apparatus 1182 with decoder 1140, or apparatus 1200.

[0158] The following examples are provided and described herein.

[0159] Example 1. An apparatus including: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a luma of at least one image; determine a chroma of the at least one image; and filter, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

[0160] Example 2. The apparatus of example 1, wherein an input to the neural network filter comprises a concatenation of the luma and the chroma.

[0161] Example 3. The apparatus of any of examples 1 to 2, wherein an input to the neural network filter comprises one or more of: a prediction tensor, or a boundary strength, or a sequence-level a quantization parameter, or a slice-level quantization parameter, or a type of slice, or a type of picture.

[0162] Example 4. The apparatus of any of examples 1 to 3, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: upsample the chroma to match a height and width of the luma, in response to the chroma having a smaller resolution than the luma.

[0163] Example 5. The apparatus of any of examples 1 to 4, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: produce, using a first convolutional layer of the neural network filter, a first tensor; and produce, using a second convolutional layer of the neural network filter, a second tensor.

[0164] Example 6. The apparatus of example 5, wherein: the first convolutional layer of the neural network filter used to produce the first tensor is within a head section of the neural network filter; and the second convolutional layer of the neural network filter used to produce the second tensor is within the head section of the neural network filter.

[0165] Example 7. The apparatus of any of examples 5 to 6, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: concatenate the first tensor and the second tensor across a channel dimension to generate a concatenated tensor; process the concatenated tensor with a convolutional layer togenerate a convolutional layer output; and process the convolutional layer output with a parametric rectified linear unit operation to generate a fused output of the neural network filter.

[0166] Example 8. The apparatus of example 7, wherein: the first tensor and the second tensor are concatenated across the channel dimension to generate the concatenated tensor within a fuse section of the neural network filter; the concatenated tensor is processed with the convolutional layer to generate the convolutional layer output within the fuse section of the neural network filter; and the convolutional layer output is processed with the parametric rectified linear unit operation to generate the fused output of the neural network filter within the fuse section of the neural network filter.

[0167] Example 9. The apparatus of any of examples 7 to 8, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the fused output of the neural network filter with at least one convolutional layer and at least one parametric rectified linear unit operation, to generate a transition output of the neural network filter.

[0168] Example 10. The apparatus of example 9, wherein the at least one convolutional layer comprises a stride of two.

[0169] Example 11. The apparatus of any of examples 9 to 10, wherein: the fused output of the neural network filter is processed with a convolutional layer comprising a stride of two and a parametric rectified linear unit operation, to produce a tensor used to filter the luma; and the fused output of the neural network filter is processed with a convolutional layer comprising a stride of four and a parametric rectified linear unit operation, to produce a tensor used to filter the chroma; and

[0170] Example 12. The apparatus of any of examples 9 to 11, wherein the fused output is processed with the at least one convolutional layer and the at least one parametric rectified linear unit operation, to generate the transition output of the neural network filter within a transition section of the neural network filter.

[0171] Example 13. The apparatus of any of examples 9 to 12, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: split the transition output of the neural network filter into a tensor used to filter the luma, and a tensor used to filter the chroma.

[0172] Example 14. The apparatus of example 11 or 13, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process, using a number of luma blocks of the neural network filter, the tensor used to filter the luma; wherein an output of a luma block is input to a next luma block; wherein the luma block is part of the luma blocks, and the next luma block is part of the luma blocks; and process, using a number of chroma blocks of the neural network filter, the tensor used to filter the chroma; wherein an output of a chroma block is input to a next chroma block; wherein the chroma block is part of the chroma blocks, and the next chroma block is part of the chroma blocks.

[0173] Example 15. The apparatus of example 14, wherein: the tensor used to filter the luma is processed using the number of luma blocks of the neural network filter within a luma backbone section of the neural network filter; the luma block comprises a luma backbone block and the next luma block comprises a luma backbone block; the tensor used to filter the chroma is processed using the number of chroma blocks of the neural network filter with a chroma backbone section of the neural network filter; and the chroma block comprises a chroma backbone block and the next chroma block comprises a chroma backbone block.

[0174] Example 16. The apparatus of any of examples 14 to 15, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process an output of a last luma block of the luma blocks with a separable convolutional layer, a parametric rectified linear unit operation, a convolutional layer, and a pixel shuffle operation; and process an output of a last chroma block of the chroma blocks with a separable convolutional layer, a parametric rectified linear unit operation, and a convolutional layer to generate a chroma tail output .

[0175] Example 17. The apparatus of example 16, wherein: the output of the last luma block of the luma blocks is processed with the separable convolution layer, theparametric rectified linear unit operation, the convolutional layer, and the pixel shuffle operation within a luma tail section of the neural network filter; and the output of the last chroma block of the chroma blocks is processed with the separable convolutional layer, the parametric rectified linear unit operation, and the convolutional layer to generate the chroma tail output within a chroma tail section of the neural network filter.

[0176] Example 18. The apparatus of any of examples 16 to 17, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: add an output of the pixel shuffle operation to the luma to generate the luma output; and add the chroma tail output to the chroma to generate the chroma output.

[0177] Example 19. The apparatus of any of examples 14 to 18, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the tensor used to filter the luma with a first luma block to generate a first output, the first luma block being part of the luma blocks; combine the tensor used to filter the luma with the first output to generate an input to a second luma block used to generate a second output, the second luma block being part of the luma blocks; and combine the tensor used to filter the luma with the first output and the second output to generate an input to a third luma block used to generate a third output used to generate the luma output, the third luma block being part of the luma blocks.

[0178] Example 20. The apparatus of any of examples 14 to 19, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the tensor used to filter the chroma with a first chroma block to generate a first output, the first chroma block being part of the chroma blocks; combine the tensor used to filter the chroma with the first output to generate an input to a second chroma block used to generate a second output, the second chroma block being part of the chroma blocks; and combine the tensor used to filter the chroma with the first output and the second output to generate an input to a third chroma block used to generate a third output used to generate the chroma output, the third chroma block being part of the chroma blocks.Example 21. The apparatus of any of examples 16 to 20, wherein the instructions, when executed by the at least one processor, cause the apparatus at leastto: process, using at least one neural network layer, an output of one or more luma blocks, to generate an input to one or more of: one or more of the chroma blocks, a separable convolutional layer used to process the output of the last chroma block of the chroma blocks, or the parametric rectified linear unit operation used to process the output of the last chroma block of the chroma blocks.

[0179] Example 22. The apparatus of any of examples 1 to 21, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process, using a summary-based modulation layer of the neural network filter, an input tensor with a summarization operation and a convolutional layer to produce a processed summarized tensor; and multiply, using the summary-based modulation layer of the neural network filter, the processed summarized tensor with the input tensor to produce an output tensor.

[0180] Example 23. The apparatus of any of examples 1 to 22, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: blend, using an auxiliary neural network of the neural network filter, two or more signals or tensors.

[0181] Example 24. The apparatus of example 23, wherein at least one of the two or more signals or tensors is an output of the neural network filter.

[0182] Example 25. The apparatus of any of examples 23 to 24, wherein the two or more signals or tensors comprise an input to the neural network filter and an output of the neural network filter.

[0183] Example 26. The apparatus of any of examples 23 to 25, wherein an output of the auxiliary neural network of the neural network filter comprises a signal or tensor that represents a blending or combination of the two or more signals or tensors.

[0184] Example 27. The apparatus of any of examples 1 to 26, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: produce, using an auxiliary neural network, a blending map using an input to the neural network filter and an output of the neural network filter.

[0185] Example 28. The apparatus of example 27, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: use the blending map to blend the input to the neural network filter and the output of the neural network filter.

[0186] Example 29. The apparatus of any of examples 1 to 28, wherein the apparatus comprises an encoder, or the encoder comprises the apparatus.

[0187] Example 30. The apparatus of any of examples 1 to 29, wherein the apparatus comprises a decoder, or the decoder comprises the apparatus.

[0188] Example 31. A method including: determining a luma of at least one image; determining a chroma of the at least one image; and filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

[0189] Example 32. The method of example 31, wherein an input to the neural network filter comprises a concatenation of the luma and the chroma.

[0190] Example 33. The method of any of examples 31 to 32, wherein an input to the neural network filter comprises one or more of: a prediction tensor, or a boundary strength, or a sequence-level a quantization parameter, or a slice-level quantization parameter, or a type of slice, or a type of picture.

[0191] Example 34. The method of any of examples 31 to 33 further comprising: upsampling the chroma to match a height and width of the luma, in response to the chroma having a smaller resolution than the luma.

[0192] Example 35. The method of any of examples 31 to 34 further comprising: producing, using a first convolutional layer of the neural network filter, a first tensor; and producing, using a second convolutional layer of the neural network filter, a second tensor.

[0193] Example 36. The method of example 35, wherein: the first convolutional layer of the neural network filter used to produce the first tensor is within a head section of the neural network filter; and the second convolutional layer of the neural network filterused to produce the second tensor is within the head section of the neural network filter.

[0194] Example 37. The method of any of examples 35 to 36 further comprising: concatenating the first tensor and the second tensor across a channel dimension to generate a concatenated tensor; processing the concatenated tensor with a convolutional layer to generate a convolutional layer output; and processing the convolutional layer output with a parametric rectified linear unit operation to generate a fused output of the neural network filter.

[0195] Example 38. The method of example 37, wherein: the first tensor and the second tensor are concatenated across the channel dimension to generate the concatenated tensor within a fuse section of the neural network filter; the concatenated tensor is processed with the convolutional layer to generate the convolutional layer output within the fuse section of the neural network filter; and the convolutional layer output is processed with the parametric rectified linear unit operation to generate the fused output of the neural network filter within the fuse section of the neural network filter.

[0196] Example 39. The method of any of examples 37 to 38 further comprising: processing the fused output of the neural network filter with at least one convolutional layer and at least one parametric rectified linear unit operation, to generate a transition output of the neural network filter.

[0197] Example 40. The method of example 39, wherein the at least one convolutional layer comprises a stride of two.

[0198] Example 41. The method of any of examples 39 to 40, wherein: the fused output of the neural network filter is processed with a convolutional layer comprising a stride of two and a parametric rectified linear unit operation, to produce a tensor used to filter the luma; and the fused output of the neural network filter is processed with a convolutional layer comprising a stride of four and a parametric rectified linear unit operation, to produce a tensor used to filter the chroma.

[0199] Example 42. The method of any of examples 39 to 41, wherein the fusedoutput is processed with the at least one convolutional layer and the at least one parametric rectified linear unit operation, to generate the transition output of the neural network filter within a transition section of the neural network filter.

[0200] Example 43. The method of any of examples 39 to 42 further comprising: splitting the transition output of the neural network filter into a tensor used to filter the luma, and a tensor used to filter the chroma.

[0201] Example 44. The method of example 41 or 43 further comprising: processing, using a number of luma blocks of the neural network filter, the tensor used to filter the luma; wherein an output of a luma block is input to a next luma block; wherein the luma block is part of the luma blocks, and the next luma block is part of the luma blocks; and processing, using a number of chroma blocks of the neural network filter, the tensor used to filter the chroma; wherein an output of a chroma block is input to a next chroma block; wherein the chroma block is part of the chroma blocks, and the next chroma block is part of the chroma blocks.

[0202] Example 45. The method of example 44, wherein: the tensor used to filter the luma is processed using the number of luma blocks of the neural network filter within a luma backbone section of the neural network filter; the luma block comprises a luma backbone block and the next luma block comprises a luma backbone block; the tensor used to filter the chroma is processed using the number of chroma blocks of the neural network filter with a chroma backbone section of the neural network filter; and the chroma block comprises a chroma backbone block and the next chroma block comprises a chroma backbone block.

[0203] Example 46. The method of any of examples 44 to 45 further comprising: processing an output of a last luma block of the luma blocks with a separable convolutional layer, a parametric rectified linear unit operation, a convolutional layer, and a pixel shuffle operation; and processing an output of a last chroma block of the chroma blocks with a separable convolutional layer, a parametric rectified linear unit operation, and a convolutional layer to generate a chroma tail output.

[0204] Example 47. The method of example 46, wherein: the output of the last lumablock of the luma blocks is processed with the separable convolution layer, the parametric rectified linear unit operation, the convolutional layer, and the pixel shuffle operation within a luma tail section of the neural network filter; and the output of the last chroma block of the chroma blocks is processed with the separable convolutional layer, the parametric rectified linear unit operation, and the convolutional layer to generate the chroma tail output within a chroma tail section of the neural network filter.

[0205] Example 48. The method of any of examples 46 to 47 further comprising: adding an output of the pixel shuffle operation to the luma to generate the luma output; and adding the chroma tail output to the chroma to generate the chroma output.

[0206] Example 49. The method of any of examples 44 to 48 further comprising: processing the tensor used to filter the luma with a first luma block to generate a first output, the first luma block being part of the luma blocks; combining the tensor used to filter the luma with the first output to generate an input to a second luma block used to generate a second output, the second luma block being part of the luma blocks; and combining the tensor used to filter the luma with the first output and the second output to generate an input to a third luma block used to generate a third output used to generate the luma output, the third luma block being part of the luma blocks.

[0207] Example 50. The method of any of examples 44 to 49 further comprising: processing the tensor used to filter the chroma with a first chroma block to generate a first output, the first chroma block being part of the chroma blocks; combining the tensor used to filter the chroma with the first output to generate an input to a second chroma block used to generate a second output, the second chroma block being part of the chroma blocks; and combining the tensor used to filter the chroma with the first output and the second output to generate an input to a third chroma block used to generate a third output used to generate the chroma output, the third chroma block being part of the chroma blocks.

[0208] Example 51. The method of any of examples 46 to 50 further comprising: processing, using at least one neural network layer, an output of one or more luma blocks, to generate an input to one or more of: one or more of the chroma blocks, aseparable convolutional layer used to process the output of the last chroma block of the chroma blocks, or the parametric rectified linear unit operation used to process the output of the last chroma block of the chroma blocks.

[0209] Example 52. The method of any of examples 31 to 51 further comprising: processing, using a summary-based modulation layer of the neural network filter, an input tensor with a summarization operation and a convolutional layer to produce a processed summarized tensor; and multiplying, using the summary-based modulation layer of the neural network filter, the processed summarized tensor with the input tensor to produce an output tensor.

[0210] Example 53. The method of any of examples 31 to 52 further comprising: blending, using an auxiliary neural network of the neural network filter, two or more signals or tensors.

[0211] Example 54. The method of example 53, wherein at least one of the two or more signals or tensors is an output of the neural network filter.

[0212] Example 55. The method of any of examples 53 to 54, wherein the two or more signals or tensors comprise an input to the neural network filter and an output of the neural network filter.

[0213] Example 56. The method of any of examples 53 to 55, wherein an output of the auxiliary neural network of the neural network filter comprises a signal or tensor that represents a blending or combination of the two or more signals or tensors.

[0214] Example 57. The method of any of examples 31 to 56 further comprising: producing, using an auxiliary neural network, a blending map using an input to the neural network filter and an output of the neural network filter.

[0215] Example 58. The method of example 57 further comprising: using the blending map to blend the input to the neural network filter and the output of the neural network filter.

[0216] Example 59. An apparatus including: means for determining a luma of at leastone image; means for determining a chroma of the at least one image; and means for filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

[0217] Example 60. The apparatus of claim 59, wherein the apparatus further comprises means for performing the methods as claimed in any of the claims 32 to 58.

[0218] Example 61. A computer readable medium including instructions stored thereon for performing at least the following: determining a luma of at least one image; determining a chroma of the at least one image; and filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

[0219] Example 62. The computer readable medium comprising of claim 61 , wherein computer readable medium further comprises instructions for performing the methods as claimed in any of the claims 32 to 58.

[0220] Example 63. The computer readable medium of any of the claims 61 or 62, wherein the computer readable medium comprises a non-transitory computer readable medium.

[0221] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi- processor architectures and sequential / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGAs), application specific circuits (ASICs), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, etc.

[0222] As used herein, the term ‘circuitry’, ‘circuit’ and variants may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions ofprocessor(s) / software including digital signal processor(s), software, and one or more memories that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device. Circuitry or circuit may also be used to mean a function or a process used to execute a method.

[0223] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.

[0224] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen, dash (-), or number, and may be case insensitive):2D two-dimensional3D three-dimensionalASIC application specific integrated circuit aux auxiliaryA VC advanced video codingB batch sizeBaseQP sequence-level quantization parameterBD bit distortionB -frame bidirectional predicted pictureBS boundary strengthC number of channelsCb blue-yellow chrominanceConv convolutionalCPU central processing unitCr red-green chrominanceCTU coding tree unitDCT discrete cosine transformFPGA field programmable gate arrayH heightH.2xx family of video coding standards in the domain of the ITU-T(e.g. H.263, H.264, H.265, H.266, H.274)HE VC high efficiency video codingIBC intra block copyI-frame intra-coded picturePF interfaceI / O input / outputIPB 1-frame, P-frame, B-frameITU International Telecommunication UnionITU-T ITU Telecommunication Standardization SectorMAE mean absolute error mAP mean average precisionMOTA multiple object tracking accuracyMSE mean squared errorNAL network abstraction layerNN neural networkNNPF neural-network-based post filteringN / W networkP-frame prediction picturePReLU parametric rectified linear unitPred predictionPSNR peak signal-to-noise ratioQP quantization parameterRAM random access memoryRec reconstructionROI region of interestROM read only memory s strideSBM summary-based modulationSEI supplemental enhancement informationSep separableSliceQP slice-level QPSON self-organizing / optimizing networkTV televisionUI user interfaceUp upsamplingUSB universal serial busVCM video coding for machinesVSEI versatile supplemental enhancement informationVVC versatile video codingW widthYUV color model having a luma component (Y), a chroma component(U), and a chroma component (V)

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: determine a luma of at least one image; determine a chroma of the at least one image; and filter, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

2. The apparatus of claim 1, wherein an input to the neural network filter comprises a concatenation of the luma and the chroma.

3. The apparatus of any of claims 1 to 2, wherein an input to the neural network filter comprises one or more of: a prediction tensor, or a boundary strength, or a sequence-level a quantization parameter, or a slice-level quantization parameter, or a type of slice, or a type of picture.

4. The apparatus of any of claims 1 to 3, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to:upsample the chroma to match a height and width of the luma, in response to the chroma having a smaller resolution than the luma.

5. The apparatus of any of claims 1 to 4, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: produce, using a first convolutional layer of the neural network filter, a first tensor; and produce, using a second convolutional layer of the neural network filter, a second tensor.

6. The apparatus of claim 5, wherein: the first convolutional layer of the neural network filter used to produce the first tensor is within a head section of the neural network filter; and the second convolutional layer of the neural network filter used to produce the second tensor is within the head section of the neural network filter.

7. The apparatus of any of claims 5 to 6, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: concatenate the first tensor and the second tensor across a channel dimension to generate a concatenated tensor; process the concatenated tensor with a convolutional layer to generate a convolutional layer output; and process the convolutional layer output with a parametric rectified linear unit operation to generate a fused output of the neural network filter.

8. The apparatus of claim 7, wherein: the first tensor and the second tensor are concatenated across the channel dimension to generate the concatenated tensor within a fuse section of the neural network filter;the concatenated tensor is processed with the convolutional layer to generate the convolutional layer output within the fuse section of the neural network filter; and the convolutional layer output is processed with the parametric rectified linear unit operation to generate the fused output of the neural network filter within the fuse section of the neural network filter.

9. The apparatus of any of claims 7 to 8, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the fused output of the neural network filter with at least one convolutional layer and the at least one parametric rectified linear unit operation, to generate a transition output of the neural network filter.

10. The apparatus of claim 9, wherein the at least one convolutional layer comprises a stride of two.

11. The apparatus of any of claims 9 to 10, wherein: the fused output of the neural network filter is processed with a convolutional layer comprising a stride of two and the parametric rectified linear unit operation, to produce a tensor used to filter the luma; and the fused output of the neural network filter is processed with a convolutional layer comprising a stride of four and the parametric rectified linear unit operation, to produce a tensor used to filter the chroma.

12. The apparatus of any of claims 9 to 11, wherein the fused output is processed with the at least one convolutional layer and the at least one parametric rectified linear unit operation, to generate the transition output of the neural network filter within a transition section of the neural network filter.

13. The apparatus of any of claims 9 to 12, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: split the transition output of the neural network filter into a tensor used to filterthe luma, and a tensor used to filter the chroma.

14. The apparatus of claim 11 or 13, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process, using a number of luma blocks of the neural network filter, the tensor used to filter the luma; wherein an output of a luma block is input to a next luma block; wherein the luma block is part of the luma blocks, and the next luma block is part of the luma blocks; and process, using a number of chroma blocks of the neural network filter, the tensor used to filter the chroma; wherein an output of a chroma block is input to a next chroma block; wherein the chroma block is part of chroma blocks, and the next chroma block is part of the chroma blocks.

15. The apparatus of claim 14, wherein: the tensor used to filter the luma is processed using the number of luma blocks of the neural network filter within a luma backbone section of the neural network filter; the luma block comprises a luma backbone block and the next luma block comprises a luma backbone block; the tensor used to filter the chroma is processed using the number of chroma blocks of the neural network filter with a chroma backbone section of the neural network filter; and the chroma block comprises a chroma backbone block and the next chroma block comprises a chroma backbone block.

16. The apparatus of any of claims 14 to 15, wherein the instructions, when executedby the at least one processor, cause the apparatus at least to: process an output of a last luma block of the luma blocks with a separable convolutional layer, the parametric rectified linear unit operation, a convolutional layer, and a pixel shuffle operation; and process an output of a last chroma block of the chroma blocks with a separable convolutional layer, the parametric rectified linear unit operation, and the convolutional layer to generate the chroma tail output.

17. The apparatus of claim 16, wherein: the output of the last luma block of the luma blocks is processed with the separable convolution layer, the parametric rectified linear unit operation, the convolutional layer, and the pixel shuffle operation within a luma tail section of the neural network filter; and the output of the last chroma block of the chroma blocks is processed with the separable convolutional layer, the parametric rectified linear unit operation, and the convolutional layer to generate the chroma tail output within a chroma tail section of the neural network filter.

18. The apparatus of any of claims 16 to 17, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: add an output of the pixel shuffle operation to the luma to generate the luma output; and add the chroma tail output to the chroma to generate the chroma output.

19. The apparatus of any of claims 14 to 18, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the tensor used to filter the luma with a first luma block to generate a first output, the first luma block being part of the luma blocks;combine the tensor used to filter the luma with the first output to generate an input to a second luma block used to generate a second output, the second luma block being part of the luma blocks; and combine the tensor used to filter the luma with the first output and the second output to generate an input to a third luma block used to generate a third output used to generate the luma output, the third luma block being part of the luma blocks.

20. The apparatus of any of claims 14 to 19, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process the tensor used to filter the chroma with a first chroma block to generate a first output, the first chroma block being part of the chroma blocks; combine the tensor used to filter the chroma with the first output to generate an input to a second chroma block used to generate a second output, the second chroma block being part of the chroma blocks; and combine the tensor used to filter the chroma with the first output and the second output to generate an input to a third chroma block used to generate a third output used to generate the chroma output, the third chroma block being part of the chroma blocks.

21. The apparatus of any of claims 16 to 20, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process, using at least one neural network layer, an output of one or more luma blocks, to generate an input to one or more of: one or more of the chroma blocks, a separable convolutional layer used to process the output of the last chroma block of the chroma blocks, or the parametric rectified linear unit operation used to process the output of the last chroma block of the chroma blocks.

22. The apparatus of any of claims 1 to 21, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: process, using a summary-based modulation layer of the neural network filter, an input tensor with a summarization operation and a convolutional layer to produce aprocessed summarized tensor; and multiply, using the summary-based modulation layer of the neural network filter, the processed summarized tensor with the input tensor to produce an output tensor.

23. The apparatus of any of claims 1 to 22, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: blend, using an auxiliary neural network of the neural network filter, two or more signals or tensors.

24. The apparatus of claim 23, wherein at least one of the two or more signals or tensors is an output of the neural network filter.

25. The apparatus of any of claims 23 to 24, wherein the two or more signals or tensors comprise an input to the neural network filter and an output of the neural network filter.

26. The apparatus of any of claims 23 to 25, wherein an output of the auxiliary neural network of the neural network filter comprises a signal or tensor that represents a blending or combination of the two or more signals or tensors.

27. The apparatus of any of claims 1 to 26, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: produce, using an auxiliary neural network, a blending map using an input to the neural network filter and an output of the neural network filter.

28. The apparatus of claim 27, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: use the blending map to blend the input to the neural network filter and the output of the neural network filter.

29. The apparatus of any of claims 1 to 28, wherein the apparatus comprises an encoder, or the encoder comprises the apparatus.

30. The apparatus of any of claims 1 to 29, wherein the apparatus comprises a decoder, or the decoder comprises the apparatus.

31. A method comprising: determining a luma of at least one image; determining a chroma of the at least one image; and filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

32. The method of claim 31 , wherein an input to the neural network filter comprises a concatenation of the luma and the chroma.

33. The method of any of claims 31 to 32, wherein an input to the neural network filter comprises one or more of: a prediction tensor, or a boundary strength, or a sequence-level a quantization parameter, or a slice-level quantization parameter, or a type of slice, or a type of picture.

34. The method of any of claims 31 to 33 further comprising: upsampling the chroma to match a height and width of the luma, in response to the chroma having a smaller resolution than the luma.

35. The method of any of claims 31 to 34 further comprising: producing, using a first convolutional layer of the neural network filter, a first tensor; andproducing, using a second convolutional layer of the neural network filter, a second tensor.

36. The method of claim 35, wherein: the first convolutional layer of the neural network filter used to produce the first tensor is within a head section of the neural network filter; and the second convolutional layer of the neural network filter used to produce the second tensor is within the head section of the neural network filter.

37. The method of any of claims 35 to 36 further comprising: concatenating the first tensor and the second tensor across a channel dimension to generate a concatenated tensor; processing the concatenated tensor with a convolutional layer to generate a convolutional layer output; and processing the convolutional layer output with a parametric rectified linear unit operation to generate a fused output of the neural network filter.

38. The method of claim 37, wherein: the first tensor and the second tensor are concatenated across the channel dimension to generate the concatenated tensor within a fuse section of the neural network filter; the concatenated tensor is processed with the convolutional layer to generate the convolutional layer output within the fuse section of the neural network filter; and the convolutional layer output is processed with the parametric rectified linear unit operation to generate the fused output of the neural network filter within the fuse section of the neural network filter.

39. The method of any of claims 37 to 38 further comprising:processing the fused output of the neural network filter with at least one convolutional layer and at least one parametric rectified linear unit operation, to generate a transition output of the neural network filter.

40. The method of claim 39, wherein the at least one convolutional layer comprises a stride of two.

41. The method of any of claims 39 to 40, wherein: the fused output of the neural network filter is processed with a convolutional layer comprising a stride of two and a parametric rectified linear unit operation, to produce a tensor used to filter the luma; and the fused output of the neural network filter is processed with a convolutional layer comprising a stride of four and a parametric rectified linear unit operation, to produce a tensor used to filter the chroma.

42. The method of any of claims 39 to 41, wherein the fused output is processed with the at least one convolutional layer and the at least one parametric rectified linear unit operation, to generate the transition output of the neural network filter within a transition section of the neural network filter.

43. The method of any of claims 39 to 42 further comprising: splitting the transition output of the neural network filter into a tensor used to filter the luma, and a tensor used to filter the chroma.

44. The method of claim 41 or 43 further comprising: processing, using a number of luma blocks of the neural network filter, the tensor used to filter the luma; wherein an output of a luma block is input to a next luma block; wherein the luma block is part of the luma blocks, and the next luma block is part of the luma blocks; andprocessing, using a number of chroma blocks of the neural network filter, the tensor used to filter the chroma; wherein an output of a chroma block is input to a next chroma block; wherein the chroma block is part of the chroma blocks, and the next chroma block is part of the chroma blocks.

45. The method of claim 44, wherein: the tensor used to filter the luma is processed using the number of luma blocks of the neural network filter within a luma backbone section of the neural network filter; the luma block comprises a luma backbone block and the next luma block comprises a luma backbone block; the tensor used to filter the chroma is processed using the number of chroma blocks of the neural network filter with a chroma backbone section of the neural network filter; and the chroma block comprises a chroma backbone block and the next chroma block comprises a chroma backbone block.

46. The method of any of claims 44 to 45 further comprising: processing an output of a last luma block of the luma blocks with a separable convolutional layer, a parametric rectified linear unit operation, a convolutional layer, and a pixel shuffle operation; and processing an output of a last chroma block of the chroma blocks with a separable convolutional layer, a parametric rectified linear unit operation, and a convolutional layer to generate a chroma tail output.

47. The method of claim 46, wherein: the output of the last luma block of the luma blocks is processed with the separable convolution layer, the parametric rectified linear unit operation, theconvolutional layer, and the pixel shuffle operation within a luma tail section of the neural network filter; and the output of the last chroma block of the chroma blocks is processed with the separable convolutional layer, the parametric rectified linear unit operation, and the convolutional layer to generate the chroma tail output within a chroma tail section of the neural network filter.

48. The method of any of claims 46 to 47 further comprising: adding an output of the pixel shuffle operation to the luma to generate the luma output; and adding the chroma tail output to the chroma to generate the chroma output.

49. The method of any of claims 44 to 48 further comprising: processing the tensor used to filter the luma with a first luma block to generate a first output, the first luma block being part of the luma blocks; combining the tensor used to filter the luma with the first output to generate an input to a second luma block used to generate a second output, the second luma block being part of the luma blocks; and combining the tensor used to filter the luma with the first output and the second output to generate an input to a third luma block used to generate a third output used to generate the luma output, the third luma block being part of the luma blocks.

50. The method of any of claims 44 to 49 further comprising: processing the tensor used to filter the chroma with a first chroma block to generate a first output, the first chroma block being part of the chroma blocks; combining the tensor used to filter the chroma with the first output to generate an input to a second chroma block used to generate a second output, the second chroma block being part of the chroma blocks; andcombining the tensor used to filter the chroma with the first output and the second output to generate an input to a third chroma block used to generate a third output used to generate the chroma output, the third chroma block being part of the chroma blocks.

51. The method of any of claims 46 to 50 further comprising: processing, using at least one neural network layer, an output of one or more luma blocks, to generate an input to one or more of: one or more of the chroma blocks, a separable convolutional layer used to process the output of the last chroma block of the chroma blocks, or the parametric rectified linear unit operation used to process the output of the last chroma block of the chroma blocks.

52. The method of any of claims 31 to 51 further comprising: processing, using a summary-based modulation layer of the neural network filter, an input tensor with a summarization operation and a convolutional layer to produce a processed summarized tensor; and multiplying, using the summary-based modulation layer of the neural network filter, the processed summarized tensor with the input tensor to produce an output tensor.

53. The method of any of claims 31 to 52 further comprising: blending, using an auxiliary neural network of the neural network filter, two or more signals or tensors.

54. The method of claim 53, wherein at least one of the two or more signals or tensors is an output of the neural network filter.

55. The method of any of claims 53 to 54, wherein the two or more signals or tensors comprise an input to the neural network filter and an output of the neural network filter.

56. The method of any of claims 53 to 55, wherein an output of the auxiliary neural network of the neural network filter comprises a signal or tensor that represents ablending or combination of the two or more signals or tensors.

57. The method of any of claims 31 to 56 further comprising: producing, using an auxiliary neural network, a blending map using an input to the neural network filter and an output of the neural network filter.

58. The method of claim 57 further comprising: using the blending map to blend the input to the neural network filter and the output of the neural network filter.

59. An apparatus comprising: means for determining a luma of at least one image; means for determining a chroma of the at least one image; and means for filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

60. The apparatus of claim 59, wherein the apparatus further comprises means for performing the methods as claimed in any of the claims 32 to 58.

61. A computer readable medium comprising instructions stored thereon for performing at least the following: determining a luma of at least one image; determining a chroma of the at least one image; and filtering, using a neural network filter, the luma and the chroma to generate a luma output and a chroma output.

62. The computer readable medium comprising of claim 61 , wherein computer readable medium further comprises instructions for performing the methods as claimed in any of the claims 32 to 58.

63. The computer readable medium of any of the claims 61 or 62, wherein the computer readable medium comprises a non-transitory computer readable medium.

Citation Information

Patent Citations

  • Techniques for convolutional neural network-based multi-exposure fusion of multiple image frames and for deblurring multiple image frames

    US20200265567A1

  • Method and device for video coding using improved inloop filter for chroma component

    WO2023200135A1