Method for sharing neural network inference information in video compression
By using metadata to define an acceptable margin around patches for neural network-based image processing tools, the method addresses the challenges of replicability and flexibility in video encoding and decoding, ensuring efficient and accurate processing of video streams.
Patent Information
- Application Number
- JP2024563984
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-04-12
- Publication Date
- 2025-05-27
AI Technical Summary
Existing video encoding and decoding technologies face challenges in ensuring replicability and flexibility of neural network-based inference processes across different hardware and software constraints, particularly in processing video streams efficiently.
The method involves obtaining a video stream and associated metadata that represents an acceptable margin around a patch for the inference process of a neural network-based image processing tool. This metadata includes syntax elements that define the receptive field of the neural network, allowing for flexible processing of patches larger than the original patch size considered in the network's definition.
The proposed solution ensures replicability and flexibility in the inference process, enabling efficient processing of video streams across different hardware and software constraints while maintaining accurate output results.
Smart Images

Figure 2025516252000001_ABST
Abstract
Description
Technical Field
[0001] At least one of the embodiments generally relates to methods and devices for encoding and decoding video data using a neural network, and more particularly to a method that enables sharing of neural network information that allows for a flexible inference process on the decoder side.
Background Art
[0002] To achieve high compression efficiency, video encoding schemes typically employ prediction and transformation that exploit the spatial and temporal redundancy of video content. During encoding, pictures of video content are divided into blocks of samples (i.e., pixels), and these blocks are then further divided into one or more sub-blocks hereinafter referred to as original sub-blocks. Intra prediction or inter prediction is then applied to each sub-block to utilize intra-picture or inter-picture correlation. Regardless of the prediction method (intra or inter) used, a predictor sub-block is determined for each original sub-block. The sub-block representing the difference between the original sub-block and the predictor sub-block is then often referred to as a prediction error sub-block, a prediction residual sub-block, or simply a residual sub-block, and is transformed, quantized, and entropy encoded to generate an encoded video stream. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to transformation, quantization, and entropy encoding.
[0003] In recently considered video codec solutions, neural network-based processing has been proposed, for example, in the post-filtering stage or for block prediction. Before being actually used, the neural network needs to be trained so that it can provide accurate results. Training of the neural network is a computationally intensive process that generally requires comparing the output data provided by the neural network with the expected values of these output data for a large number of input data. Once trained, the neural network can apply what it has learned to the input data, even if these input data were not considered during the training process. The process of applying a trained neural network to input data to obtain output data is called inference.
[0004] The process applied on the encoder side should be replicable identically on the decoder side in order to ensure that there is no drift between the encoder and the decoder, which is well known in the video compression area. The same also applies to the neural network (NN)-based process applied in the prediction loops of the encoder and the decoder. This requirement of replicability means that the output data inferred on the decoder side should be the same as the output data inferred on the encoder side. Furthermore, if possible, it is generally expected that two decoders with different implementations provide systematically the same results. However, it is common to have encoders and decoders designed with different software or hardware constraints. For example, the encoder can store more data in memory than the decoder. The decoder, for which the processing speed is generally an important issue, can process more data in parallel than an encoder without the same processing speed issue. A decoder implemented in a smartphone generally does not have the same hardware constraints as a decoder implemented in a PC. In this context, in the design of the inference process of the NN-based process, as much flexibility as possible should be left to the developers of the encoder and the decoder, while, first, ensuring the replicability of the output provided to the decoder side by the NN-based in-loop process on the encoder side, and second, ensuring that two decoders with different implementations provide the same results.
[0005] It is desirable to propose a solution that makes it possible to overcome the above problems. In particular, it is desirable to propose a solution that makes it possible to ensure the flexibility of the inference process of the NN inference process. SUMMARY OF THE INVENTION
[0006] In a first aspect, one or more of the present embodiments provide a method, the method comprising: obtaining a video stream; Obtaining metadata associated with a video stream that represents an acceptable margin around a patch for an inference process of a neural network-based image processing tool, including applying a neural network-based image processing tool to decode the video stream.
[0007] In an embodiment, the acceptable margin depends on the receptive field that depends on the neural network used in the neural network-based image processing tool.
[0008] In an embodiment, the metadata includes at least one syntax element that represents a receptive field that depends on the neural network used in the neural network-based image processing tool.
[0009] In an embodiment, the at least one syntax element includes a first syntax element that defines the receptive field vertically and a second syntax element that defines the receptive field horizontally.
[0010] In an embodiment, the ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network and the margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool.
[0011] In an embodiment, the ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified in the metadata by the syntax element.
[0012] In an embodiment, the metadata includes at least one syntax element representing a value for a receptive field that depends on a neural network, or at least one offset added to a margin around a patch considered in the definition of a neural network used in a neural network-based image processing tool, the offset being used in response to a current patch being processed by an inference process and having a size smaller than a patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch.
[0013] In an embodiment, the metadata includes at least one syntax element representing the position of an output patch of an inference process of a neural network-based image processing tool in an output tensor generated by the inference process.
[0014] In a second aspect, one or more of the present embodiments provide a method that includes acquiring a video stream, and signaling information representing an acceptable margin around a patch for an inference process of a neural network-based image processing tool in a form of metadata associated with the video stream.
[0015] In an embodiment, the acceptable margin depends on a receptive field that depends on a neural network used in a neural network-based image processing tool.
[0016] In an embodiment, the metadata includes at least one syntax element representing a receptive field that depends on a neural network used on a neural network-based image processing tool.
[0017] In an embodiment, the at least one syntax element includes a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally.
[0018] In an embodiment, the ability of an inference process to process a patch larger than a patch size considered in the definition of a neural network used in a neural network-based image processing tool is determined by comparing at least one value representing a receptive field that depends on the neural network with a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool.
[0019] In an embodiment, the ability of an inference process to process a patch larger than a patch size considered in the definition of a neural network used in a neural network-based image processing tool is specified in metadata by a syntax element.
[0020] In an embodiment, the metadata includes at least one syntax element representing at least one offset added to a value representing a receptive field that depends on the neural network or to a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool, the offset being used in response to the current patch being processed by the inference process and having a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch.
[0021] In an embodiment, the metadata includes at least one syntax element representing the position of an output patch of the inference process of the neural network-based image processing tool in an output tensor generated by the inference process.
[0022] In an embodiment, a video stream is obtained by applying a video compression process to an original video, the video compression process including a neural network-based image processing tool within a prediction loop of the video compression process or the neural network-based image processing tool being a post-processing tool.
[0023] In a third aspect, one or more of the present embodiments provide a signal associated with a video stream that represents an acceptable margin around a patch for an inference process of a neural network-based image processing tool.
[0024] In a fourth aspect, one or more of the present embodiments provide a computer program including program code instructions for implementing the method of the first aspect or the second aspect.
[0025] In a fifth aspect, one or more of the present embodiments provide a non-transitory information storage medium storing program code instructions for implementing the method of the first aspect or the second aspect.
[0026] In a sixth aspect, one or more of the present embodiments provide a device including an electronic circuit, the electronic circuit being configured to acquire a video stream, acquire metadata associated with the video stream that represents an acceptable margin around a patch for an inference process of a neural network-based image processing tool, and apply a neural network-based image processing tool to decode the video stream.
[0027] In an embodiment, the acceptable margin depends on the receptive field that depends on the neural network used in the neural network-based image processing tool.
[0028] In an embodiment, the metadata includes at least one syntax element representing a receptive field that depends on the neural network used in the neural network-based image processing tool.
[0029] In an embodiment, the at least one syntax element includes a first syntax element that vertically defines the receptive field and a second syntax element that horizontally defines the receptive field.
[0030] In an embodiment, the ability of an inference process to process patches larger than a patch size considered in the definition of a neural network used in a neural network-based image processing tool is determined by comparing at least one value representing a receptive field that depends on the neural network with a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool.
[0031] In an embodiment, the ability of an inference process to process patches larger than a patch size considered in the definition of a neural network used in a neural network-based image processing tool is specified in metadata by a syntax element.
[0032] In an embodiment, the metadata includes at least one syntax element representing at least one offset added to a value representing a receptive field that depends on the neural network or to a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool, the offset being used in response to the current patch being processed by the inference process and having a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch.
[0033] In an embodiment, the metadata includes at least one syntax element representing the position of an output patch of the inference process of the neural network-based image processing tool in an output tensor generated by the inference process.
[0034] In a seventh aspect, one or more of the present embodiments provide a device comprising an electronic circuit, the electronic circuit being acquiring a video stream, and Signaling information representing an acceptable margin around a patch for an inference process of a neural network-based image processing tool in the form of metadata associated with a video stream.
[0035] In an embodiment, the acceptable margin depends on the receptive field that depends on the neural network used in the neural network-based image processing tool.
[0036] In an embodiment, the metadata includes at least one syntax element representing a receptive field that depends on the neural network used on the neural network-based image processing tool.
[0037] In an embodiment, the at least one syntax element includes a first syntax element that defines the receptive field vertically and a second syntax element that defines the receptive field horizontally.
[0038] In an embodiment, the ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network and the margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool.
[0039] In an embodiment, the ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified within the metadata by the syntax element.
[0040] In an embodiment, the metadata includes at least one syntax element representing a value representing a receptive field that depends on a neural network, or at least one offset added to a margin around a patch considered in the definition of a neural network used in a neural network-based image processing tool, the offset being used in response to a current patch being processed by an inference process and having a size smaller than a patch size considered in the definition of a neural network used in a neural network-based image processing tool based on the position of the current patch.
[0041] In an embodiment, the metadata includes at least one syntax element representing the position of an output patch of an inference process of a neural network-based image processing tool in an output tensor generated by the inference process.
[0042] In an embodiment, a video stream is obtained by applying a video compression process to an original video, the video compression process including a neural network-based image processing tool within a prediction loop of the video compression process, or the neural network-based image processing tool being a post-processing tool.
Brief Description of the Drawings
[0043]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14A
Figure 14B
Best Mode for Carrying Out the Invention
[0044] The following examples of embodiments are described in the context of a video format similar to Versatile Video Coding (VVC) being developed by a joint team of ITU-T and ISO / IEC experts known as the Joint Video Experts Team (JVET). However, these embodiments are not limited to video encoding / decoding methods corresponding to VVC. These embodiments are applicable, in particular, to various video formats including, for example, HEVC (ISO / IEC 23008-2-MPEG-H Part 2, High Efficiency Video Coding / ITU-T H.265), AVC ((ISO / CEI 14496-10), EVC (Essential Video Coding / MPEG-5), AV1, AV2, and VP9.
[0045] FIG. 1 schematically shows an example of the context in which an embodiment is implemented.
[0046] In FIG. 1, a system 11, which can be any device capable of delivering a camera, a storage device, a computer, a server, or a video stream, uses a communication channel 12 to transmit a video stream to a system 13. The video stream is encoded by the system 11 and transmitted, or received and / or stored by the system 11 and then transmitted. The communication channel 12 is a wired (e.g., Internet or Ethernet) or wireless (e.g., WiFi, 3G, 4G, or 5G) network link.
[0047] A system 13, which can be, for example, a set-top box, receives and decodes the video stream to generate a sequence of decoded pictures.
[0048] Next, the obtained sequence of decoded pictures is transmitted to a display system 15 using a communication channel 14, which can be a wired or wireless network. Then, the display system 15 displays this picture.
[0049] In an embodiment, system 13 is included in display system 15. In that case, system 13 and display 15 are included in, for example, a TV, a computer, a tablet, a smartphone, a head-mounted display, and the like.
[0050] Figures 2, 3, and 4 introduce examples of video formats.
[0051] Figure 2 shows an example of the division that the picture of pixel 21 of the original video sequence 20 undergoes. Here, the pixel is considered to be composed of three components, namely a luminance component and two chrominance components. However, other types of pixels can contain fewer or more components, such as only a luminance component or additional depth or transparency components.
[0052] The picture is divided into a plurality of coded entities. First, as represented by reference numeral 23 in Figure 2, the picture is divided into a grid of blocks called coding tree units (CTUs). A CTU consists of an N×N block of luminance samples and two corresponding blocks of chrominance samples. N is generally a power of 2 having a maximum value of, for example, "128". Second, the picture is divided into one or more groups of CTUs. For example, the picture can be divided into one or more tile rows and tile columns, where a tile is a sequence of CTUs covering a rectangular region of the picture. In some cases, a tile can be divided into one or more bricks, each of which consists of at least one CTU row within the tile. Above the concepts of tiles and bricks, there exists another coded entity called a slice that can include at least one tile of the picture or at least one brick of a tile.
[0053] In the example of Figure 2, as represented by reference numeral 22, picture 21 is divided into three slices S1, S2, and S3 in a raster scan slice mode, each of which contains a plurality of tiles (not shown), and each tile contains only one brick.
[0054] As represented by reference numeral 24 in FIG. 2, the CTU can be divided into a hierarchical tree form of one or more sub-blocks called coding units (CUs). The CTU is the root of the hierarchical tree (i.e., the parent node) and can be divided into a plurality of CUs (i.e., child nodes). Each CU becomes a leaf of the hierarchical tree if it is not further divided into smaller CUs, and becomes the parent node of smaller CUs (i.e., child nodes) if it is further divided.
[0055] In the example of FIG. 2, the CTU24 is first divided into "4" square CUs using a quadtree type of division. The upper left CU is a leaf of the hierarchical tree because it is not further divided, i.e., it is not the parent node of other CUs. The upper right CU is further divided into "4" smaller square CUs using a quadtree type of division. The lower right CU is vertically divided into "2" rectangular CUs using a binary tree type of division. The lower left CU is vertically divided into "3" rectangular CUs using a ternary tree type of division.
[0056] During the encoding of a picture, the division is adaptive, and each CTU is divided so as to optimize the compression efficiency based on the CTU.
[0057] In HEVC, the concepts of prediction unit (PU) and transform unit (TU) emerged. In fact, in HEVC, the encoding entities used for prediction (i.e., PU) and transform (i.e., TU) can be sub-divisions of the CU. For example, as shown in FIG. 2, a CU of size 2N×2N can be divided into a PU2411 of size N×2N or a PU2411 of size 2N×N. In addition, this CU can be divided into "4" TUs2412 of size N×N, or
[0058]
Number
[0059] In VVC, it should be noted that, except for some specific cases, the frontiers of TUs and PUs are aligned with the frontier of the CU. Therefore, a CU generally includes one TU and one PU.
[0060] In this application, the terms "block" or "picture block" can be used to refer to any one of CTU, CU, PU, and TU. Further, the terms "block" or "picture block" can be used to refer to macroblocks, partitions, and sub-blocks as specified in H.264 / AVC or other video coding standards, and more generally, can be used to refer to an array of samples of multiple sizes.
[0061] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", "sub-picture", "slice", and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstructed" is used on the encoder side, while the term "decoded" is used on the decoder side.
[0062] FIG. 3 schematically shows a method for encoding a video stream executed by an encoding module. Although variations of this method for encoding are contemplated, hereinafter, for the purpose of clarity, the method for encoding in FIG. 3 will be described without explaining all the contemplated variations.
[0063] Before being symbolized, the current original picture of the original video sequence can go through preprocessing. For example, in step 301, color conversion is applied to the current original picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or remapping is applied to the components of the current original picture to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization for one of the color components). The picture obtained by preprocessing is hereinafter referred to as the preprocessed picture.
[0064] Encoding the preprocessed picture starts with the partitioning of the preprocessed picture during step 302, as described in relation to Figure 2. Thus, the preprocessed picture is partitioned into CTUs, CUs, PUs, TUs, etc. For each block, the encoding module determines the encoding mode between intra prediction and inter prediction.
[0065] Intra prediction consists of predicting the pixels of the current block from a predicted block derived from the pixels of the reconstructed blocks located in the causal neighborhood of the current block being encoded during step 303 according to the intra prediction method. The result of intra prediction is a prediction direction indicating which pixels of the neighboring blocks are used, and a residual block resulting from the calculation of the difference between the current block and the predicted block.
[0066] Inter prediction consists of predicting the pixels of a current block from a block of pixels (referred to as a reference block) of a picture before or after the current picture (this picture is referred to as a reference picture). During the encoding of the current block by the inter prediction method, according to a similarity criterion, the block of the reference picture that is closest to the current block is determined by the motion estimation step 304. During step 304, a motion vector indicating the position of the reference block within the reference picture is determined. This motion vector is used during the motion compensation step 305, during which the residual block is calculated in the form of the difference between the current block and the reference block. In the first video compression standard, the one-way inter prediction mode described above was the only available inter mode. As video compression standards have evolved, the family of inter modes has grown significantly and now includes many different inter modes.
[0067] During the selection step 306, according to the rate / distortion optimization criterion (i.e., the RDO criterion), the prediction mode that optimizes the compression performance is selected by the encoding module from among the tested prediction modes (intra prediction mode, inter prediction mode).
[0068] Once the prediction mode is selected, the residual block is transformed during step 307. Next, during step 309, the transformed block is quantized.
[0069] Note that the symbolization module can skip the conversion and directly apply quantization to the unconverted residual signal. When the current block is encoded according to the intra prediction mode, the prediction direction and the transformed and quantized residual block are encoded by the entropy encoder during step 310. When the current block is encoded according to inter prediction, if appropriate, the motion vector of the block is predicted from a predicted vector selected from a set of motion vector predictors derived from the reconstructed blocks located spatially and temporally close to the block to be encoded. Next, the motion information is encoded by the entropy encoder during step 310 in the form of a motion residual and in the form of an index for identifying the predicted vector. The transformed and quantized residual block is encoded by the entropy encoder during step 310.
[0070] Note that the symbolization module can bypass both the conversion and quantization, that is, the entropy encoding is applied to the residual without applying the conversion process or the quantization process. The result of the entropy encoding is inserted into the encoded video stream 311.
[0071] Metadata such as SEI (supplemental enhancement information) messages can be added to the encoded video stream 311. For example, an SEI message defined in a standard such as AVC, HEVC, or VVC is a data container associated with the video stream and containing metadata that provides information about the video stream.
[0072] After quantization step 309, the current block is reconstructed so that the pixels corresponding to that block can be used for future prediction. This reconstruction stage is also called the prediction loop. Thus, inverse quantization is applied to the residual block that was transformed and quantized during step 312, and inverse transformation is applied during step 313. Depending on the prediction mode used for the block obtained during step 314, the predicted block of the block is reconstructed. When the current block is encoded according to the inter prediction mode, the encoding module applies motion compensation using the motion vector of the current block to identify the reference block of the current block, if appropriate, during step 316. When the current block is encoded according to the intra prediction mode, the prediction direction corresponding to the current block is used to reconstruct the predicted block of the current block during step 315. The predicted block and the reconstructed residual block are added to obtain the reconstructed current block.
[0073] After reconstruction, during step 317, in-loop filtering intended to reduce encoding artifacts is applied to the reconstructed block. This filtering is performed in the prediction loop because the decoder obtains the same reference picture as the encoder and thus avoids drift between the encoding process and the decoding process, and is therefore called in-loop filtering. In-loop filtering tools include deblocking filtering, SAO (Sample adaptive Offset), and ALF (Adaptive Loop Filtering).
[0074] Once reconstructed, the block is inserted into the reconstructed picture stored in the memory 319 of the reconstructed pictures, generally called the Decoded Picture Buffer (DPB), during step 318. The reconstructed picture stored in this way can then function as a reference picture for other pictures to be encoded.
[0075] Figure 4 schematically shows a method for decoding an encoded video stream 311 encoded according to the method described in relation to Figure 3, which is executed by a decoding module. Although variations of this decoding method are contemplated, for the sake of clarity, the decoding method of Figure 4 will be described below without explaining all the variations that are expected.
[0076] Decoding is performed block by block. For the current block, decoding begins with entropy decoding the current block during step 410. Entropy decoding enables obtaining at least the prediction mode of the block.
[0077] If the block is encoded according to the inter prediction mode, entropy decoding enables obtaining, if appropriate, the prediction vector index, the motion residual, and the residual block. During step 408, the motion vector is reconstructed for the current block using the prediction vector index and the motion residual.
[0078] If the block is encoded according to the intra prediction mode, entropy decoding enables obtaining the prediction direction and the residual block. Steps 412, 413, 414, 415, 416, and 417 performed by the decoding module are identical in all respects to steps 312, 313, 314, 315, 316, and 317 performed by the encoding module, respectively.
[0079] In step 418, the decoded block is stored in the decoded picture, and the decoded picture is stored in the DPB 419. When the decoding module decodes a given picture, the pictures stored in the DPB 419 are identical to the pictures stored in the DPB 319 by the encoding module during the encoding of the aforementioned given picture. The decoded picture can also be output by the decoding module, for example, for display.
[0080] The post - processing step 421 can include inverse color conversion (e.g., conversion from YCbCr4:2:0 to RGB4:4:4), inverse mapping that performs the reverse of the remapping process executed in the pre - processing of step 301, and post - filtering for improving the reconstructed picture based on, for example, filter parameters provided in the SEI message.
[0081] In recently considered video coding solutions, NN - based processing has been proposed, for example, for post - filtering or block prediction (intra and inter). Further, solutions have been proposed that enable the exchange of information representing the NN between the encoder and the decoder. For example, such information is transmitted as side information in the form of an SEI message.
[0082] FIG. 6 schematically shows an example of an NN - based process.
[0083] In FIG. 6, a block 51 of the original picture 50 is processed. Before the processing, the block 51 is enlarged by a certain margin. The enlarged block is then input to the NN 52. The output of the NN is a block 53 having the same size as the block 51. Next, the block 53 is inserted into the picture 54 so as to be used in the encoding or decoding process or for display without being used in the encoding or decoding process. In some variations, the block 53 has the same size as the enlarged block. In that case, the margin is used for blending at the block boundary to reduce blocking artifacts.
[0084] Table TAB1 provides an example of an SEI message nnr_post_filter that enables the transport of information representing the NN post - filtering process.
[0085] [Table 1]
[0086] The semantics of the syntax elements in Table TAB1 are briefly described below (readers can refer to the document JVET_Z0052: Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 26th Meeting, by teleconference, 20-29 April 2022, Hannuksela et al. for detailed semantics). · When nnrpf_constant_patch_size_flag is "0", it is specified that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnrpf_patch_size_minus1. When nnrpf_constant_patch_size_flag is "1", it is specified that the post-processing filter accepts exactly the patch size indicated by nnrpf_patch_size_minus1 as input. · nnrpf_patch_size_minus1 + 1 specifies the horizontal and vertical sample counts of the patch size of the post-processing filter. The value of nnrpf_patch_size_minus1 ranges from "0" to "32766" inclusive. · nnrpf_overlap specifies the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnrpf_overlap ranges from "0" to "16383" inclusive. · nnrpf_component_last_flag indicates whether the channel is shown as the second or last dimension in the tensor in the input to the NN and the tensor in the output of the NN. · nnrpfwidth_pic_width_in_luma_samples and nnrpf_pic_height_in_luma_samples respectively specify the width and height of the luma sample array of the picture obtained by applying the post-processing filter identified by nnrpf_id to the cropped decoded output picture. ·The nnrpf_id includes an identification number that can be used to identify a post-processing filter.
[0087] In this semantics, a patch is a tensor that includes picture data related to a block or the entire picture that can include, for example, picture samples (e.g., block samples + samples in the vicinity of the block), quantization parameters, motion information, etc.
[0088] The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows. ·inpPatchWidth = nnrpf_patch_size_minus1 + 1, ·inpPatchHeight = nnrpf_patch_size_minus1 + 1, ·outPatchWidth = (nnrpf_pic_width_in_luma_samples * inpPatchWidth) / InpPicWidthInLumaSamples, ·outPatchHeight = (nnrpf_pic_height_in_luma_samples * inpPatchHeight) / InpPicHeightInLumaSamples, ·horCScaling = InpSubWidthC / outSubWidthC, ·verCScaling = InpSubHeightC / outSubHeightC, ·outPatchCWidth = outPatchWidth * horCScaling, ·outPatchCHeight = outPatchHeight * verCScaling, ·overlapSize = nnrpf_overlap.
[0089] InpPicHeightInLumaSamples represents the input picture height in luma samples.
[0090] InpPicWidthtInLumaSamples represents the input picture width in luma samples.
[0091] Using the above variables, the following pseudo - code shows an example of the derivation process of the input tensor inputTensor when the patch size (inpPatchWidth, inpPatchHeight) and the overlap size overflowSize are given.
[0092] [Table 2]
[0093] Clip3(a, b, c) represents a function that keeps the value of variable b between a and c.
[0094] (cTop, cLeft) represents the coordinates of the top - left pixel of the current block in the luma channel of the current picture. Note that the convention in video compression is used. This means that the column coordinate cTop is placed first in the coordinate tuple.
[0095] The inputTensor represents the tensor input to the NN. Note that for indexing of the inputTensor, the Machine-Learning convention is used. This means that the row coordinate comes before the column coordinate. When nnrpf_component_last_flag is equal to "0", i.e., when the channel is shown as the second dimension, the row and column are shown as the third and fourth dimensions. When nnrpf_component_last_flag is equal to "1", i.e., when the channel is shown as the fourth dimension within the inputTensor, the row and column are shown as the second and third dimensions within the inputTensor. inpY represents the luminance channel of the current picture. CroppedYPic[y][x] represents the cropped and decoded output picture.
[0096] The following pseudo-code shows an example of the process of deriving new filtered samples when the output tensor outputTensor is given.
[0097]
Table 3
[0098] In this pseudo-code, FilteredYPic represents an array of luminance-filtered samples.
[0099] OutY and OutC represent functions that convert the luminance sample value and chrominance sample value output in post-processing into integer values of bit depths BitDepthY and BitDepthC respectively. OutY and OutC are specified as follows. OutY(x)=Clip3(0,(1<<BitDepthY)-1,Round(x * ((1<<BitDepthY)-1))) OutC(x)=Clip3(0,(1<<BitDepthC)-1,Round(x * ((1<<BitDepthC)-1)))
[0100] Here, Round(a) rounds the value a to the nearest integer value.
[0101] Typically, the SEI message nnr_post_filter describes the following. · Corresponds to the patch size and the output size. · The input size is equal to the patch size plus some overlapping size overlapSize, and the overlapping size is the same size on all boundaries of the block.
[0102] To enable the decoder to execute an accurate and reproducible inference process, some information is required to describe the input and output of such an NN.
[0103] Figure 7A schematically represents a block-based inference process and a sub-block-based inference process. In Figure 7A, the left part shows a block B with a margin of d samples, which is used by default in the inference process. For example, it is expected that inpPatchWidth and inpPatchHeight are related to the block width (w) and height (h), and overlapSize is the additional margin d.
[0104] On the right side of Figure 7A, an example is shown where the same block B is divided into "4" sub-blocks for the inference process. For example, while the original NN was acting on the complete block B, the memory constraints on the decoder side do not allow the processing of the complete block B, but require the use of "4" separate inference processes to process smaller blocks such as sub-blocks B1 to B4.
[0105] Figure 7B schematically represents a block-based inference process and a super-block-based inference process.
[0106] Figure 7B shows the inverse problem of Figure 7A. While the original NN was acting on block B, the decoder can process more data in parallel and process larger blocks. In Figure 7B, the processing area of the inference process is composed of four blocks B1 to B4, forming a superblock that is processed at once.
[0107] As can be seen, flexibility is required in the inference process. However, in the current signaling proposed in nnr_post_filter, in order to enable such flexibility, especially to enable the processing of sub-blocks and super-blocks during the inference process while the original NN is defined using blocks, some information is lacking. The various embodiments described below propose solutions to this problem.
[0108] In particular, the various embodiments introduce information that enables the decoder to signal possible block size variations in which the inference process can operate.
[0109] Figures 5A, 5B, and 5C illustrate examples of devices, apparatuses, and / or systems that enable the implementation of various embodiments.
[0110] Figure 5A schematically shows an example of the hardware architecture of a processing module 500 that can implement an encoding module or a decoding module that can respectively implement the encoding method of Figure 3 and the decoding method of Figure 4 modified according to different aspects and embodiments. The encoding module is included in system 11, for example, when this system is responsible for encoding a video stream. The decoding module is included in system 13, for example.
[0111] The processing module 500 includes, as a non-limiting example, a processor or central processing unit (CPU) 5000 including one or more microprocessors, general-purpose computers, dedicated computers, and processors based on multi-core architectures connected by a communication bus 5005, a random access memory (RAM) 5001, a read-only memory (ROM) 5002, an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash, a magnetic disk drive, and / or an optical disk drive, or a storage unit 5003 that can include, but is not limited to, non-volatile memory and / or volatile memory such as a secure digital (SD) card reader and / or a hard disc drive (HDD), and at least one communication interface 5004 for exchanging data with other modules, devices, or systems. The communication interface 5004 can include, but is not limited to, a transceiver configured to transmit and receive data via a communication channel. The communication interface 5004 can include, but is not limited to, a modem or a network card.
[0112] When the processing module 500 implements a decoding module, the communication interface 5004 enables, for example, the processing module 500 to receive an encoded video stream and provide a sequence of decoded pictures. When the processing module 500 implements an encoding module, the communication interface 5004 enables, for example, the processing module 500 to receive and encode a sequence of original picture data and provide an encoded video stream.
[0113] The processor 5000 can execute instructions loaded from the ROM 5002, an external memory (not shown), a storage medium, or a communication network into the RAM 5001. When the processing module 500 is powered on, the processor 5000 can read instructions from the RAM 5001 and execute them. These instructions form, for example, a computer program that causes the processor 5000 to implement the decoding method described in connection with FIG. 4, and / or the encoding method described in connection with FIG. 3, and the methods described in connection with FIGS. 14A or 14B. The decoding method and the encoding method include various aspects and embodiments described hereinafter in this specification.
[0114] All or part of the algorithms and steps of the methods of FIGS. 3, 4, 14A, and 14B may be implemented in software form by the execution of an instruction set by a programmable machine such as a digital signal processor (DSP) or a microcontroller, or in hardware form by a machine or a dedicated component such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).
[0115] As can be seen, a microprocessor, general-purpose computer, dedicated computer, processor based on or not based on a multi-core architecture, DSP, microcontroller, FPGA, and ASIC are electronic circuits adapted to at least partially implement the methods of FIGS. 3, 4, 14A, and 14B.
[0116] FIG. 5C shows a block diagram of an example of a system 13 in which various aspects and embodiments are implemented. The system 13 can be embodied as a device including various components described hereinafter and is configured to execute one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and head-mounted displays. The elements of the system 13 can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components, either alone or in combination. For example, in at least one embodiment, the system 13 includes one processing module 500 that implements a decoding module. In various embodiments, the system 13 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, the system 13 is configured to implement one or more of the aspects described herein.
[0117] Inputs to processing module 500 can be provided via various input modules as shown in block 531. Such input modules include, but are not limited to, (i) a radio frequency (RF) module that receives RF signals wirelessly transmitted from a broadcast station, for example, (ii) a component (COMP) input module (or a set of COMP input modules), (iii) a Universal Serial Bus (USB) input module, and / or (iv) a High Definition Multimedia Interface (HDMI) input module. Other examples include composite video, although not shown in FIG. 5C.
[0118] In various embodiments, the input module of block 531 has respective input processing elements known in the art. For example, an RF module may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in certain embodiments, band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF module of various embodiments may include one or more elements that perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF portion may include a tuner that performs various of these functions, including for example down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one embodiment of a set-top box, the RF module and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. In various embodiments, the order of the above-described (and other) elements is rearranged, some of these elements are omitted, and / or other elements that perform similar or different functions are added. Adding elements can include, for example, inserting elements between existing elements such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF module includes an antenna.
[0119] Furthermore, the USB module and / or the HDMI module can each include an interface processor for connecting the system 13 to other electronic devices via a USB connection and / or an HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or, if necessary, within the processing module 500. Similarly, aspects of USB or HDMI interface processing can be implemented, if necessary, within a separate interface IC or within the processing module 500. The demodulated, error-corrected, and demultiplexed streams are provided to the processing module 500.
[0120] The various elements of the system 13 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and data can be transmitted between them using an internal bus known in the art, including a suitable connection arrangement, such as an inter-integrated circuit (I2C) bus, wiring, and a printed circuit board. For example, in the system 13, the processing module 500 is interconnected to the other elements of the system 13 by a bus 5005.
[0121] The communication interface 5004 of the processing module 500 enables the system 13 to communicate over the communication channel 12. As already mentioned above, the communication channel 12 can be implemented, for example, within a wired and / or wireless medium.
[0122] In various embodiments, data is streamed to, or otherwise provided to, system 13 using a Wi-Fi network, such as a wireless network like IEEE 802.11 (where IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of such embodiments are received via communication channel 12 and communication interface 5004 that are adapted for Wi-Fi communication. Typically, the communication channel 12 of such embodiments is connected to an access point, or router, that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the RF connection of input block 531 is used to provide streaming data to system 13. As indicated above, various embodiments provide data in a non-streaming fashion. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0123] System 13 can provide output signals to various output devices including display system 15, speaker 535, and other peripheral devices 536. The display system 15 of various embodiments includes, for example, one or more of a touch screen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 15 can be for a television, a tablet, a laptop, a mobile phone, a head-mounted display, or other devices. The display system 15 can also be integrated with other components (such as in a smartphone) or separate (such as an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 536 include one or more of a stand-alone digital video disc (or digital versatile disc, both terms abbreviated as DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 536 that provide functions based on the output of system 13. For example, a disc player performs the function of playing back the output of system 13.
[0124] In various embodiments, the control signal is communicated between the system 13 and the display system 15, the speaker 535, or other peripheral device 536 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices, with or without user intervention. The output devices can be communicatively coupled to the system 13 via dedicated connections through their respective interfaces 532, 533, and 534. Alternatively, the output devices can be connected to the system 13 using the communication channel 12 via the communication interface 5004, or using a dedicated communication channel corresponding to the communication channel 12 of FIG. 5C via the communication interface 5004. The display system 15 and the speaker 535 can be integrated into a single unit with other components of the system 13 in an electronic device such as a television, for example. In various embodiments, the display interface 532 includes a display driver, such as a timing controller (T Con) chip, for example.
[0125] Alternatively, the display system 15 and the speaker 535 can be separated from one or more of the other components. In various embodiments where the display system 15 and the speaker 535 are external components, an output signal can be provided via a dedicated output connection that includes, for example, an HDMI port, a USB port, or a COMP output.
[0126] Figure 5B shows a block diagram of an example of system 11 in which various aspects and embodiments are implemented. System 11 is very similar to system 13. System 11 can be embodied as a device that includes various components described hereinafter and is configured to execute one or more of the aspects and embodiments described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, cameras, and servers. The elements of system 11 can be embodied in a single integrated circuit (IC), multiple ICs, and / or separate components, either alone or in combination. For example, in at least one embodiment, system 11 includes one processing module 500 that implements an encoding module. In various embodiments, system 11 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input ports and / or output ports. In various embodiments, system 11 is configured to implement one or more of the aspects described herein.
[0127] Inputs to processing module 500 can be provided via various input modules as shown in block 531 already described with respect to Figure 5D.
[0128] The various elements of system 11 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected using internal buses known in the art, including suitable connection arrangements such as an inter-integrated circuit (I2C) bus, wiring, and printed circuit boards, and can transmit data therebetween. For example, in system 11, processing module 500 is interconnected to the other elements of system 11 by bus 5005.
[0129] The communication interface 5004 of processing module 500 enables system 11 to communicate on communication channel 12.
[0130] In various embodiments, data is streamed to, or otherwise provided to, system 11 using a Wi-Fi network, such as a wireless network like IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of such embodiments are received via communication channel 12 and communication interface 5004 that are adapted for Wi-Fi communication. Typically, the communication channel 12 of such embodiments is connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, the RF connection of input block 531 is used to provide streaming data to system 11.
[0131] As indicated above, various embodiments provide data in a non-streaming fashion. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0132] The data provided to system 11 can be provided in different formats. In various embodiments, these data are encoded and conform to known video compression formats such as AV1, VP9, VVC, HEVC, AVC. In various embodiments, these data are raw data provided, for example, by picture and / or audio acquisition modules connected to or included in system 11. In that case, processing module 500 is responsible for encoding these data.
[0133] System 11 can provide output signals to various output devices that can store and / or decode output signals such as system 13.
[0134] Various implementations involve decoding. When used in this application, "decoding" may include, for example, all or part of the process executed on the received encoded video stream to generate a final output suitable for display. In various embodiments, such a process may include one or more of the processes typically executed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and prediction. In various embodiments, such a process may also or alternatively include the processes executed by the decoders of the various implementations described in this application, for example, processes for applying NN-based coding modes such as NN-based post-processing or in-loop filtering.
[0135] Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a broader decoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.
[0136] Various implementations involve encoding. Similar to the above considerations regarding "decoding", when used in this application, "encoding" may include, for example, all or part of the process executed on the input video sequence to generate an encoded video stream. In various embodiments, such a process may include one or more of the processes typically executed by an encoder, such as partitioning, prediction, transformation, quantization, and entropy encoding. In various embodiments, such a process may also or alternatively include the processes executed by the encoders of the various implementations described in this application, for example, processes for signaling to the decoder possible block size variations for which the inference process of NN-based post-processing or in-loop filtering can operate.
[0137] Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to a broader encoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.
[0138] Note that the syntax element names used in this specification are for illustrative purposes. Therefore, they do not exclude the use of other syntax element names. For example, in the following, the following syntax element names are used: nn_subblock_dx, nn_subblock_dy, nn_subblock_dxb, nn_subblock_dyb. These names can be replaced by nn_subblock_dL, nn_subblock_dT, nn_subblock_dxR, nn_subblock_dyB, respectively.
[0139] When a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.
[0140] Various embodiments refer to rate distortion optimization. In particular, during the encoding process, the balance or trade-off between rate and distortion is typically considered. Rate distortion optimization is usually formulated to minimize a rate distortion function that is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, these approaches can be based on an extensive test of all encoding options including all considered modes or encoding parameter values, involving their encoding cost and a complete evaluation of the associated distortion of the reconstructed signal after encoding and decoding. Also, to reduce the encoding complexity, faster approaches can be used, in particular, using the calculation of approximate distortion based on the predicted or prediction residual signal instead of the reconstructed signal. A mixture of these two approaches can also be used, such as by using approximate distortion for only some of the considered encoding options and complete distortion for other encoding options. In other approaches, only a subset of the considered encoding options is evaluated. More generally, many approaches employ any of various techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.
[0141] The implementations and aspects described in this specification can be implemented, for example, as a method or process, an apparatus, a software program, a data stream, or a signal. Even if considered only in the context of a single form of implementation (for example, considered only as a method), the implementation of the features considered can also be implemented in other forms (for example, an apparatus or a program). For example, an apparatus can be implemented in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor, which refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes, for example, communication devices such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information among end users.
[0142] References to "one embodiment", "an embodiment", "one implementation", "an implementation", and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment", "in an embodiment", "in one implementation", "in an implementation", and other variations that appear at various places throughout this application do not necessarily all refer to the same embodiment.
[0143] In addition, this application may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, obtaining information from memory, or obtaining information from, for example, another device, module, or user.
[0144] Furthermore, this application may refer to "accessing" various information. Accessing information can include, for example, one or more of receiving information, obtaining information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0145] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information can include, for example, one or more of accessing information or obtaining information (e.g., from memory). Furthermore, "receiving" typically involves, in some way, during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information.
[0146] The use of any of " / ", "and / or", "at least one of", "one or more of", for example, in the case of "A / B", "A and / or B", "at least one of A and B", "one or more of A and B", is intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", "one or more of A, B, and C", such phrases are intended to encompass the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art of this and related arts, this can be extended for as many listed items as desired.
[0147] Also, as used herein, the term "signaling" means, among other things, indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals the use of some encoding tools. In this way, in embodiments, the same parameters can be used on both the encoder side and the decoder side. Thus, for example, the encoder can send the decoder certain parameters (explicit signaling) so that the decoder can use the same specific parameters. On the other hand, if the decoder already has other parameters along with that specific parameter, signaling can be used that does not perform the transmission (implicit signaling) simply to enable the decoder to know and select that specific parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the word "signal", but the word "signal" can also be used as a noun herein.
[0148] As will be apparent to those skilled in the art, the implementation can result in various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for implementing a method or data generated by one of the described implementations. For example, the signals can be formatted to carry the encoded video stream and SEI messages of the described embodiments. Such signals can be formatted, for example, as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as baseband signals. Formatting can include, for example, encoding the encoded video stream and modulating a carrier wave with the encoded video stream. The information carried by the signals can be analog information or digital information, for example. As is known, the signals can be transmitted over various different wired or wireless links. The signals can be stored on a processor-readable medium.
[0149] In the following, various embodiments are proposed that signal to a decoder various patch dimensions in which the inference process of an NN-based process can operate.
[0150] FIG. 14A schematically illustrates the application of an embodiment during an encoding process.
[0151] The embodiment of FIG. 14A is executed, for example, by the processing module 500 of the system 11.
[0152] In step 1401, the processing module 500 obtains the original video data to be encoded.
[0153] In step 1402, the processing module 500 encodes the original video data into an encoded video stream, for example using the encoding process of FIG. 3.
[0154] In step 1403, the processing module 500 signals information called NN information, which represents an acceptable margin around a patch for the inference process of the NN-based image processing tool, in the form of metadata associated with the encoded video stream.
[0155] In step 1404, the processing module 500 provides the receiving device with an encoded video stream having the information signaled in the form of metadata.
[0156] In an embodiment, the encoding process of step 1402 includes at least one NN-based image processing tool, such as at least one NN-based in-loop filtering process.
[0157] In an embodiment, the signaled NN information in step 1403 provides information regarding each NN-based image processing tool executed during the encoding process of step 1402, or information regarding at least one NN-based post-processing process to be performed by the receiving device on the decoded video data.
[0158] In an embodiment, the signaled NN information is signaled in an SEI message, and in the SEI message, the NN information is represented in the form of syntax elements described in relation to the following tables TAB2, TAB3, and TAB4.
[0159] FIG. 14B schematically represents the application of an embodiment during the decoding process.
[0160] The embodiment of FIG. 14B is executed, for example, by the processing module 500 of the system 13.
[0161] In step 1411, the processing module 500 acquires a video stream.
[0162] In step 1412, the processing module 500 obtains signaled NN information representing an acceptable margin around a patch for at least one inference process of the NN-based image processing tool in the form of metadata associated with the video stream.
[0163] In step 1413, the processing module 500 decodes the video stream by applying, for example, the decoding process of FIG. 4. Each NN-based image processing tool uses information representing an acceptable margin around a patch of the inference process of the NN-based image processing tool during the decoding process (e.g., during the loop filter process) or during post-processing of the decoded video data, and is applied by the processing module 500.
[0164] In step 1414, the processing module 500 provides the decoded video data to a display module such as the display system 15.
[0165] In the following, several types of additional NN information including receptive field information and offsets are proposed. A syntax for SEI messages adapted to transport this type of information is also proposed. However, as will be described below, the transport of NN information is not limited to SEI messages.
[0166] Receptive field information As can be seen above, the NNs used in NN-based image processing tools are generally defined (trained) using a given block (or patch) size for the input and output of the inference process. When changing the input or output block (or patch) size of the inference process with respect to the original input or output block (or patch) size considered when defining the NN, it is necessary to know which regions of the input to the inference process should be considered in order to calculate an output having the same sample values as the corresponding sample values of the output provided by the inference process when applied to blocks (or patches) of a given size. Below, the term receptive field is used to represent an additional margin value used to define a margin around a block (or patch) such that all values within the block (or patch) in the output depend only on the input block (or patch) with this margin added. Generally, the receptive field value depends on the NN network. In that case, it is called the NN receptive field. However, below, a broader definition is used and the receptive field is a value large enough to define a margin around each sample of a block (or patch) equal to or including the NN receptive field. Below, the terms block or patch are used interchangeably and a block is a particular case of a patch.
[0167] Depending on whether sub-blocks or super-blocks are considered, several cases arise.
[0168] In a first embodiment corresponding to the case shown in FIG. 7A, the original inference process of the NN-based image processing tool considered a block B having a given size of w×h samples and a margin overlapSize of d samples, but this inference process is applied to sub-blocks of block B to which separate inference processes are applied. w and h may be different such that the block and sub-blocks are rectangular, or may be equal such that the block and sub-blocks are square.
[0169] In the first embodiment, when the margin overlapSize is smaller than the receptive field of NN (i.e., the NN receptive field), an asymmetric margin around the sub-block is used.
[0170] FIG. 8A schematically shows an example of region expansion for a sub-block based inference process.
[0171] In FIG. 8A, the original block B is divided into “4” square sub-blocks B1 to B4 of equal size. FIG. 8A shows two upper sub-blocks B1 and B2. For sub-block B1, · The upper margin and the left margin are d, which is the same margin as the original block B. · The lower margin and the right margin are defined by a single value r, where r corresponds to the receptive field of NN.
[0172] For sub-block B2, the left margin and the lower margin use the receptive field size r, and the right margin and the upper margin use the margin value d.
[0173] Note that block B may correspond to the entire picture.
[0174] Regarding the syntax, in the first embodiment, it is proposed to add a syntax element nn_receptive_field to the existing syntax of the SEI message described in Table TAB1 to encode the value representing the receptive field value r. In that case, the margin around a patch (such as block B1, B2, B3, or B4) is defined by the margin value d and the receptive field value r given by the syntax element nn_receptive_filed.
[0175] In a variant of the syntax adapted to the first embodiment, the syntax element nn_receptive_field is replaced by the syntax element nn_receptive_field_plus1. Therefore, the value of the receptive field is nn_receptive_field_plus1 - 1. If the value of the receptive field is "-1" (i.e., nn_receptive_field_plus1 = 0), it indicates that the NN does not have a receptive field.
[0176] In a variant of the first embodiment, depending on the architecture of the NN, the receptive field may have different ranges in the horizontal and vertical directions. For example, if the NN includes a convolutional layer having a horizontal convolutional stride that is twice as large as the vertical convolutional stride, this convolutional layer may make the receptive field of the NN different horizontally and vertically. This case is considered in FIG. 8B.
[0177] FIG. 8B schematically represents an example of area expansion for a sub-block-based inference process when the receptive field of the NN features has different ranges horizontally and vertically.
[0178] For sub-block B1, · The upper margin and the left margin are d, which is the same margin as the original block B. · The lower margin is r h and is defined as such, and the right margin is r w ,r h and r w are defined as such, defining the receptive field vertically and horizontally, respectively.
[0179] For sub-block B2, the left margin is equal to r w and the lower margin is equal to r h , and the right margin and the upper margin use the margin d.
[0180] Regarding the syntax related to the variant of the first embodiment, the value r of the syntax element nn_receptive_field_vertically for encoding the value hand the value r of the syntax element nn_receptive_field_horizontally w is proposed to be added to the existing syntax of TAB1.
[0181] In a second embodiment corresponding to the case shown in FIG. 7B, the original inference process of the NN-based image processing tool considered a block B having a given size of w×h samples and a margin overlapSize of d samples, but this inference process is applied to the superblock of block B. The superblock B can have any size including the size corresponding to the entire frame containing block B. In the case of the second embodiment, in order to recover the same result as the processing based on block B, the margin needs to be greater than the receptive field of the NN.
[0182] FIG. 9 schematically represents the receptive field and boundary of block B.
[0183] In FIG. 9, a block with a margin of d samples is shown together with the receptive field of r samples. When d is smaller than r, the samples within the margin between d and r are assumed to be equal to a certain value. Usually, 0 padding is used, that is, the samples outside the block are assumed to be 0. Other padding can be used (such as mirroring). In that case (when d is smaller than r), the block cannot be processed with a size larger than the size defined by the original NN without changing the sample values.
[0184] In a first variation of the second embodiment, the ability to process blocks larger than the block size considered in the definition of the NN network used in the NN-based image processing tool is determined by comparing the receptive_field value r with the overlapSize value d. If d < r, the block cannot be processed with a larger size. Otherwise, if d ≧ r, the block can be processed with a larger size.
[0185] In a modification of the second embodiment, the ability to process blocks larger than the block size considered during the definition of the NN network used in the NN-based image processing tool is explicitly signaled by adding the flag nn_possible_superblock_processing_flag to the syntax of Table TAB1.
[0186] Offset information Problems arise when the NN used by the NN-based image processing tool involves some tensor size variations along the NN inference process. FIG. 10 schematically represents an example of an NN with tensor size variations. · Assume an input tensor in the form of a 3D (three-dimensional) tensor corresponding to a square block B100 of size w×w and depth D as the input to the NN inference process. · First, in step 101, a 2D convolution using a kernel of size 3×3 with a stride of "1" is applied to the input tensor. The stride is a parameter of the NN that modifies the amount of movement of the convolution kernel over the image or video. · Then, in step 102, max pooling with a stride of "2" is applied to the output tensor of step 101. Max pooling is a pooling operation that calculates the maximum value of a sub-part of the tensor and uses it to create a downsampled (pooled) tensor. This is typically used after a convolutional layer. It adds a small amount of translation invariance, i.e., changing the picture by a small amount will have little significant impact on the values of most of the pooled outputs. After this max pooling layer, the downsampled tensor is w / 2×w / 2×D, where / 2 is integer division and D is the depth of the downsampled tensor. Note that max pooling can be replaced by any downscaling layer (convolution with stride, other pooling with stride, etc.). · In step 103, a 2D convolution using a kernel of size 3×3 with a stride of "1" is applied to the downsampled tensor. · In step 104, for example, upscaling such as pixel shuffling (exchanging width / height and tensor depth), transposed convolution with a stride of "2", or interpolation is applied. After upscaling, the size of the upscaled tensor becomes w×w×D’, where D’ is the depth of the tensor. · In step 105, a 2D convolution using a 3×3 kernel with a stride of "1" is applied to the upscaled tensor to obtain an output tensor. Next, the resulting block B’ is extracted from the output tensor.
[0187] Some output tensor characteristics may be different from the input tensor characteristics if the NN inference process involves some internal tensor size variations, which is typically the case for the NN inference process in FIG. 10.
[0188] These characteristics include an output tensor size that may be different from the input tensor size, and / or the sample positions within the output tensor may be different from the corresponding sample positions within the input tensor from the offset value. These differences can be caused by NN parameters such as the stride policy in the NN inference process, the final use of padding, etc.
[0189] For example, consider the case shown in FIG. 10. · The original block size of block B is 16×16 and the margin is 2, that is, the input size is 20×20. · After the max pooling step 102, the subsampled tensor size is 10×10×D. · After the upscaling step 104, the upscaled tensor size returns to 20×20.
[0190] For this particular NN, the receptive field is "5".
[0191] Considering the case of the first embodiment having sub - blocks B1 to B4, the input size is as follows. · Sub - block size = 16 / 2×16 / 2 = 8×8 For the upper - left sub - block B1, · The margin is "2". · The bottom margin and the right margin are "5" (equal to the receptive field). · (2 + 8 + 5)×(2 + 8 + 5)=15×15 gives the input tensor size.
[0192] For sub - block B1, the output sub - block B1’’ obtained by applying the process of FIG. 10 to the input sub - block B1 is the same as the sub - block B1’ extracted from the block B’ obtained by applying the process of FIG. 10 to the full - block B.
[0193] However, for the upper - right sub - block B2, · The upper - right boundary, the margin is 2. · The bottom margin and the left margin are 5 (equal to the receptive field). · (5 + 8 + 2)×(2 + 8 + 5)=15×15 gives the total input tensor size.
[0194] However, due to the offset introduced during the process of FIG. 10, the output sub - block B2’’ obtained by applying the process of FIG. 10 to the sub - block B2 is different from the sub - block B2’ extracted from the block B’.
[0195] FIG. 13 schematically shows the influence of the offset introduced in the NN inference process with some tensor size variations.
[0196] In this case, the size of block B is 8×8, the value of margin d is "2", and the value of receptive field r is "5". As shown in FIG. 7A, block B is divided into four sub-blocks B1 to B4 with a size of 4×4. Therefore, the margin around block B is 2 samples. The left margin and the top margin of sub-block B1 are 2 samples. The right margin and the bottom margin of sub-block B1 are 5 samples. The right margin and the top margin of sub-block B2 are 2 samples. The left margin and the bottom margin of sub-block B2 are 5 samples.
[0197] Assuming that block B, and sub - blocks B1 and B2 are designed such that this process processes 8×8 sized blocks, they are input into the NN inference process of FIG. 10. For block B, the input tensor size is (2 + 2+8)×(2 + 2+8)=12×12. For blocks B1 and B2, the input tensor size is (2 + 4+5)×(2 + 4+5)=11×11. As can be seen in the description of FIG. 10, the NN inference process includes a down - scaling step 102. Assume that the down - scaling process used in the case of FIG. 13 takes one sample out of two samples of the input tensor starting from the top - left sample. In FIG. 13, the sample marked with "B" represents the sample retained by the down - sampling process of step 102 when applied to block B. The sample marked with "B1" represents the sample retained by the down - sampling process of step 102 when applied to sub - block B1. The sample marked with "B2" represents the sample retained by the down - sampling process of step 102 when applied to sub - block B2. As can be seen, when the NN inference process of FIG. 10 is applied to block B and sub - block B1, the same sample is retained by the down - sampling process. However, when the NN inference process of FIG. 10 is applied to sub - block B2, the sample retained by the down - sampling process is different from the sample retained by the down - sampling process when the NN inference process of FIG. 10 is applied to block B. This explains why the same result cannot be obtained when the NN inference process of FIG. 10 is applied to sub - block B2 and when it is applied to block B.
[0198] In this example, the following parameters are needed to systematically restore the same sub - block when applying the inference process of FIG. 10 at the full - block level or sub - block level. · For the upper - right boundary, assume the margin is "2". · Assume the bottom margin is "5". · Assume that the left margin is 5 + 1. · Give the full input tensor size of (6 + 8 + 2)×(2 + 8 + 5)=16×15.
[0199] As can be seen from this example, more parameters need to be signaled to better consider these offsets. In the third embodiment, additional parameters specify the input size of the sub - block and the offset within the output tensor.
[0200] Examples of the additional parameters according to the first variation of the third embodiment are shown in relation to FIG. 11 and provided in Table TAB2 below.
[0201] FIG. 11 schematically represents an example of the layout for the sub - block inference process of the bottom - right sub - block.
[0202] In FIG. 11, an example is shown in which the original block B with a margin (overlapSize) of d is processed using "4" sub - blocks. In FIG. 11, the parameters of the bottom - right sub - block B4 are shown. dT, dB, dL, dR are the top margin, bottom margin, left margin, and right margin, respectively. These margins are usually obtained by adding / subtracting offset values ox and oy to the horizontal and vertical margin / receptive field values of the original block.
[0203] By using offsets, it becomes possible to encode the margins with respect to the default parameters (depending on the original margin and receptive field). By representing the receptive field of the NN as r and the original margin as d, the following equations are used to calculate the margins of the "4" sub - blocks B1 - B4 of block B (shown in FIG. 7A). · Block B1: ○ dT = d+oy1 ○ dL = d+ox1 ○ dB = r ○ dR = r · Block B2: ○ dT = d+oy2 ○ dL = r + ox2 ○ dB = r ○ dR = d · Block B3: ○ dT = r + oy3 ○ dL = d + ox3 ○ dB = d ○ dR = r · Block B4: ○ dT = r + oy4 ○ dL = r + ox4 ○ dB = d ○ dR = d
[0204] Typically, the parameters are as follows. · ox1 = oy1 = 0 · ox2 = 1, oy2 = 0 · ox3 = 0, oy3 = 1 · ox4 = oy4 = 1
[0205] In the second modification of the third embodiment, the left offset dx1 (and the upper offset dy1 respectively) and the right offset dx1b (and the lower offset dy1b respectively) are different. · Block B1: ○ dT = d + oy1 ○ dL = d + ox1 ○ dB = r + oy1b ○ dR = r + ox1b · Block B2: ○ dT = d + oy2 ○ dL = r + ox2 ○ dB = r + oy2b ○ dR = d + ox2b · Block B3: ○ dT = r + oy3 ○ dL = d + ox3 ○ dB = d + oy3b ○ dR = r + ox3b · Block B4: ○ dT = r + oy4 ○ dL = r + ox4 ○ dB = d + oy4b ○ dR = d + ox4b
[0206] The second variant of the third embodiment is associated with the additional syntax provided in Table TAB3.
[0207] In the variant, the margins are defined absolutely instead of being defined relative to the margin, i.e., for each sub-block Bi, the margins dT, dB, dL, and dR are signaled.
[0208] Note that the input size of the sub-block is derived from the size of the original block.
[0209] In the third variant of the third embodiment, additional parameters are signaled for further flexibility. FIG. 12 schematically represents an embodiment having offsets wx and wy on the output. wx and wy represent the vertical and horizontal offsets of the block within the output tensor for extracting the output block B', respectively. In other words, wx and wy represent the position of the output block B' in the output tensor of the NN inference process. The above offsets are typically applied when only operations using samples within the input tensor are used and the output size of a layer within the NN, such as a convolutional layer, is reduced compared to the input tensor.
[0210] [Table 4]
[0211] Table TAB2 represents the first embodiment of the additional syntax added to the syntax of Table TAB1. The following semantics are used. ·nn_receptive_field: Specify the size that is half of the receptive field size of the network. This represents the number of samples around the input, so that the same useful output can be obtained by transforming the input. Note that the border or overlap parameters are usually below the receptive field. The nn_receptive_field value is in the range from "0" to nn_patch_size / 2. Assuming nn_patch_size is the input patch size and border_size is the border around the input patch, it is assumed that the output patch size is equal to nn_patch_size and is obtained by cropping the central part of the output tensor. ·nn_possible_superblock_processing: When set to "1", it enables the decoder to perform processing for a larger block size using a multiple of nn_patch_size and add the border_size. For example, the input tensor has the following width and height. ○ W = nn_patch_size * 2 + border_size ○ H = nn_patch_size * 2 + border_size ·nn_possible_halfblock_processing: When set to "1", it enables the decoder to perform processing for a smaller (half) block size. For example, the input tensor has a width and height (without margin). ○ W = nn_patch_size / 2 ○ H = nn_patch_size / 2 Assume that the sub-block Bi (Bi is from "1" to "4") is as shown in Figure 7A. · In the first modification of the third embodiment, nn_half_block_dx[i] and nn_half_block_dy[i] are used to derive the boundary of each sub-block Bi by setting oxi = nn_half_block_dx[i] and oyi = nn_half_block_dy[i]. r is set to nn_receptive_field, and d is set to border_size. The output sample result is obtained by cropping the output using the margin defined as the input. By default, the output sample is placed at the position of the input sample, removing the added margin.
[0212] Table TAB3 represents the second embodiment of the additional syntax added to the syntax of Table TAB1.
[0213]
Table 5
[0214] The following semantics are used. · In the second modification of the third embodiment, nn_half_block_dx[i], nn_half_block_dy[i], nn_half_block_dxb[i], and nn_half_block_dyb[i] are used to derive the boundary of each sub-block Bi by setting oxi = nn_half_block_dx[i], oyi = nn_half_block_dy[i], oxib = nn_half_block_dxb[i], and oyib = nn_half_block_dyb[i]. r is set to nn_receptive_field, and d is set to border_size.
[0215] Table TAB4 represents a third embodiment of the additional syntax added to the syntax of Table TAB1. The additional syntax of Table TAB4 indicates that it is easy to extend to a sub-block size of 1 / 4 or smaller, and the margins around each sub-block are signaled using the same syntax as shown in the following table for a 1 / 4 size.
[0216]
Table 6
[0217] The additional syntax of Tables TAB2, TAB3, and TAB4 can be applied to the post-processing NN inference process or to any NN inference process within the prediction loop of an encoder or decoder, such as a loop filter.
[0218] Note that so far in this disclosure, block B has been noted to be divided into 4 or "16" sub-blocks. The following Table TAB5 provides a flexible syntax added to the syntax of Table TAB1 that allows the block to be divided into 4, 16, 64, 256, etc. sub-blocks.
[0219]
Table 7
[0220] In Table TAB5, nn_nb_subdivision_minus1 + 1 is the number of possible recursive divisions of the patch. Each offset for each level is read in the corresponding variable nn_sublock dzz[level][i], where dzz is either dx, dy, dxb, or dyb. The value of nn_nb_subdivision_minus1 + 1 is in the range of "1" to floor(log2(nn_patch_size)). If it does not exist, the value is assumed to be equal to 0.
[0221] When sub - blocks are possible, for a specific level L where L is in the range of "1" to nn_nb_subdivision_minus1 + 1, each sub - block of size (nn_patch_size >> L) is used as the input to the NN with the following margins according to the process already described. · dT[i]=defaultMarginT+nn_subblock_dy[L][i]; · dL[i]=defaultMarginL+nn_subblock_dx[L][i]; · dB[i]=defaultMarginB+nn_subblock_dyb[L][i]; · dR[i]=defaultMarginR+nn_subblock_dxb[L][i].
[0222] Just in case, dT, dL, dB, and dR are the top margin, left margin, bottom margin, and right margin respectively. · defaultMarginT is equal to d(original block border_size) for sub - blocks above the original block, and equal to r(receptive field) otherwise. · defaultMarginL is equal to d(original block border_size) for sub - blocks to the left of the original block, and equal to r(receptive field) otherwise. · defaultMarginR is equal to d(original block border_size) for sub - blocks to the right of the original block, and equal to r(receptive field) otherwise. · defaultMarginB is equal to d(original block border_size) for sub - blocks below the original block, and equal to r(receptive field) otherwise.
[0223] Table TAB6 provides the same flexible syntax when the syntax element nn_receptive_field is replaced by the syntax element nn_receptive_field_plus1.
[0224]
Table 8
[0225] In addition, so far, the shape of the sub-block depends on the shape of block B. When block B is square, the sub-block is square. When block B is rectangular, the sub-block is rectangular. The syntax proposed in the present disclosure is also adapted when the shape of the sub-block is independent of the shape of block B. For example, · The square block B can be divided into rectangular sub-blocks. · The rectangular block B can be divided into square sub-blocks. · Block B can also be divided into any combination of square and rectangular sub-blocks of various sizes.
[0226] In the present disclosure, for example, various NN information such as syntax that can be transmitted or stored has been described. This NN information can be packaged or arranged in various ways including common ways in video standards, such as putting the NN information into an SEI message such as the SEI message described in the present disclosure. Other ways are also available, including common ways in system-level or application-level standards such as putting the information below. · SDP (Session Description Protocol), for example, a format for describing a multimedia communication session for the purpose of session announcement and session invitation, as described in an RFC and used in conjunction with RTP (Real-time Transport Protocol) transmission. · The DASH MPD (Media Presentation Description) descriptor, for example, is used in DASH and transmitted via HTTP. The descriptor is associated with a representation or a set of representations to provide additional characteristics to the content representation. · An RTP header extension, such as that used during RTP streaming. · An ISO Base Media File Format, such as that used in OMAF, which in some specifications uses boxes that are object - oriented building blocks defined by a unique type identifier and length, also known as atoms. · An HLS (HTTP live Streaming) manifest transmitted via HTTP. The manifest is associated with a version or set of versions of the content, for example, to provide characteristics of the version or set of versions.
[0227] The above describes several embodiments. The features of these embodiments can be provided alone or in any combination. Further, an embodiment can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types. · A bitstream or signal includes one or more of the described syntax elements or a variant thereof. · Create and / or transmit and / or receive and / or decode a bitstream or signal that includes one or more of the described syntax elements or a variant thereof. · A television, set - top box, mobile phone, tablet, or other electronic device that executes at least one of the described embodiments. · A television, set - top box, mobile phone, tablet, or other electronic device that executes at least one of the described embodiments and displays the resulting picture (e.g., using a monitor, screen, or other type of display). · A television, set - top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal including an encoded video stream and executes at least one of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives (e.g., using an antenna) a signal including an encoded video stream and executes at least one of the described embodiments. A server, camera, mobile phone, tablet, or other electronic device that wirelessly receives (e.g., using an antenna) a signal including an encoded video stream and executes at least one of the described embodiments. A server, camera, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to transmit a signal including an encoded video stream and executes at least one of the described embodiments.
Claims
**Claim 1** Obtaining a video stream (1411); Obtaining metadata associated with the video stream representing a first margin around a patch for an inference process of a neural network-based image processing tool (1412); Applying the neural network-based image processing tool using the metadata to decode the video stream (1413). A method comprising the steps of: **Claim 2** The method according to claim 1, wherein the first margin depends on a receptive field that depends on a neural network used in the neural network-based image processing tool. **Claim 3** The method according to claim 2, wherein the metadata includes at least one syntax element representing the receptive field that depends on the neural network used in the neural network-based image processing tool. **Claim 4** The method according to claim 3, wherein the at least one syntax element includes a first syntax element that vertically defines the receptive field and a second syntax element that horizontally defines the receptive field. **Claim 5** The method according to any one of claims 2 to 4, wherein the ability of the inference process to process a patch larger than a patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network with a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool. **Claim 6** The method according to any one of claims 2 to 4, wherein the ability of the inference process to process a patch larger than a patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified in the metadata by a syntax element. **Claim 7** The metadata includes at least one syntax element representing at least one offset added to a value representing the receptive field that depends on the neural network, or added to a second margin around a patch considered in the definition of the neural network used in the neural network-based image processing tool. The offset is used in response to the current patch being processed by the inference process and has a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch. The method according to any one of claims 2 to 6.
8. The metadata includes at least one syntax element representing the position of the output patch of the inference process of the neural network-based image processing tool in the output tensor generated by the inference process. The method according to any one of claims 1 to 7.
9. Obtaining a video stream (1402); Signaling information representing a first margin around a patch for an inference process of a neural network-based image processing tool in the form of metadata associated with the video stream (1403). A method comprising:
10. The first margin depends on a receptive field that depends on the neural network used in the neural network-based image processing tool. The method according to claim 9.
11. The metadata includes at least one syntax element representing the receptive field that depends on the neural network used on the neural network-based image processing tool. The method according to claim 10.
12. The at least one syntax element includes a first syntax element that vertically defines the receptive field and a second syntax element that horizontally defines the receptive field. The method according to claim 11.
13. The ability of the inference process to process patches larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network with a second margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool, the method according to any one of claims 10 to 12.
14. The ability of the inference process to process patches larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified in the metadata by a syntax element, the method according to any one of claims 10 to 12.
15. The metadata includes at least one syntax element representing at least one offset added to a value representing the receptive field that depends on the neural network or added to a margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool, the offset being used in response to the current patch being processed by the inference process and having a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch, the method according to any one of claims 9 to 14.
16. The metadata includes at least one syntax element representing the position of the output patch of the inference process of the neural network-based image processing tool in the output tensor generated by the inference process, the method according to any one of claims 9 to 15.
17. The video stream is obtained by applying a video compression process to the original video, the video compression process including the neural network-based image processing tool within a prediction loop of the video compression process or the neural network-based image processing tool being a post-processing tool, the method according to any one of claims 9 to 16.
18. A signal including metadata associated with a video stream, representing a margin around a patch for an inference process of a neural network-based image processing tool.
19. A computer program including program code instructions for implementing the method according to any one of Claims 1 to 17.
20. A non-transitory information storage medium storing program code instructions for implementing the method according to any one of Claims 1 to 17.
21. A device comprising an electronic circuit, wherein the electronic circuit is configured to acquire a video stream (1411), acquire metadata associated with the video stream, representing a first margin around a patch for an inference process of a neural network-based image processing tool (1412), and apply the neural network-based image processing tool using the metadata to decode the video stream (1413).
22. The device according to Claim 21, wherein the first margin depends on a receptive field that depends on a neural network used in the neural network-based image processing tool.
23. The device according to Claim 22, wherein the metadata includes at least one syntax element representing the receptive field that depends on the neural network used in the neural network-based image processing tool.
24. The device according to Claim 23, wherein the at least one syntax element includes a first syntax element defining the receptive field vertically and a second syntax element defining the receptive field horizontally.
25. The ability of the inference process to process a patch larger than a patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network and a second margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool. The device according to any one of Claims 22 to 24.
26. The device according to any one of claims 22 to 24, wherein the ability of the inference process to process patches larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified in the metadata by a syntactic element.
27. The device according to any one of claims 21 to 26, wherein the metadata includes at least one syntactic element representing at least one offset added to a value representing the receptive field that depends on the neural network or added to a second margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool, the offset being used in response to the current patch to be processed by the inference process and having a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch.
28. The device according to any one of claims 21 to 27, wherein the metadata includes at least one syntactic element representing the position of the output patch of the inference process of the neural network-based image processing tool in the output tensor generated by the inference process.
29. A device comprising an electronic circuit, wherein the electronic circuit is configured to acquire a video stream (1402) and signal information representing a first margin around a patch for an inference process of a neural network-based image processing tool in the form of metadata associated with the video stream (1403).
30. The device according to claim 29, wherein the first margin depends on a receptive field that depends on the neural network used in the neural network-based image processing tool.
31. The device according to claim 30, wherein the metadata includes at least one syntactic element representing the receptive field that depends on the neural network used on the neural network-based image processing tool.
32. The device according to claim 31, wherein the at least one syntax element includes a first syntax element that vertically defines the receptive field and a second syntax element that horizontally defines the receptive field.
33. The ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is determined by comparing at least one value representing the receptive field that depends on the neural network and a second margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool. The device according to any one of claims 30 to 32.
34. The ability of the inference process to process a patch larger than the patch size considered in the definition of the neural network used in the neural network-based image processing tool is specified in the metadata by a syntax element. The device according to any one of claims 30 to 32.
35. The metadata includes at least one syntax element representing at least one offset added to a value representing the receptive field that depends on the neural network or added to a second margin around the patch considered in the definition of the neural network used in the neural network-based image processing tool. The offset is used in response to the current patch being processed by the inference process and has a size smaller than the patch size considered in the definition of the neural network used in the neural network-based image processing tool based on the position of the current patch. The device according to any one of claims 29 to 34.
36. The metadata includes at least one syntax element representing the position of the output patch of the inference process of the neural network-based image processing tool in the output tensor generated by the inference process. The device according to any one of claims 29 to 35.
37. The device according to any one of claims 29 to 36, wherein the video stream is obtained by applying a video compression process to an original video, and the video compression process includes the neural network-based image processing tool within a prediction loop of the video compression process or the neural network-based image processing tool is a post-processing tool.