Position-dependent filtering

By introducing position-related weights into the neural network, the trade-off between the complexity and compression efficiency of the video coding filter is optimized, solving the problem of high complexity in existing technologies and achieving more efficient video coding and reduced hardware costs.

CN120917754APending Publication Date: 2025-11-07TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024459.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-12
Filing Date
2024-03-28
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing neural network-based video coding filters are difficult to optimize in terms of the trade-off between complexity and compression efficiency, resulting in high hardware implementation costs and short battery life.

Method used

By introducing position-related weights into the neural network, the network can process the input differently depending on the position of the sample within the image, reducing the reuse of weights, lowering computational complexity, and maintaining or improving compression efficiency.

Benefits of technology

It achieves a reduction in bit rate and an increase in coding efficiency, while extending battery life and reducing hardware costs without increasing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917754A_ABST
    Figure CN120917754A_ABST
Patent Text Reader

Abstract

A method includes identifying a first group within a set of inputs, where the first group includes a first plurality of values. The method includes selecting a first weight tensor, wherein the first weight tensor is selected based at least in part on a location of a first value. The method includes generating a first weighted sum by applying the first weight tensor to the first plurality of values. The method includes identifying a second group within the set of inputs, wherein the second group includes a second plurality of values. The method includes selecting a second weight tensor, where the second weight tensor is selected based at least in part on a location of a second value, where the second weight tensor is different from the first weight tensor. The method includes generating a second weighted sum by applying the second weight tensor to the second plurality of values. The method includes identifying a third group within the set of inputs, wherein the third group includes a third plurality of values. The method includes selecting a third weight tensor, wherein the third weight tensor is selected based at least in part on a location of a third value. The method includes generating a third weighted sum by applying the third weight tensor to the third plurality of values. Wherein the first weight tensor is the same as the third weight tensor. The method further includes performing one or more of generating a first output value by processing the first weighted sum with a non-linear activation function, generating a second output value by processing the second weighted sum with a non-linear activation function, and generating a third output value by processing the third weighted sum with a non-linear activation function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to methods and devices for performing filtering and in particular for filtering for video encoding and decoding. Aspects of the present disclosure relate to, for example, video compression, video coding, image compression, image coding, neural network based video coding, neural network based in-loop filtering, in-loop filtering, and post filtering. However, embodiments can also be applicable to non-video technologies, including audio and text. BACKGROUND

[0002] Video is the dominant form of data traffic in today's networks and is expected to increase its share continuously. One way to reduce the data traffic from video is compression. In compression, a source video is encoded into a bitstream, which can then be stored and transmitted to end users. Using a decoder, end users can extract the video data and display it on a screen.

[0003] However, since the encoder does not know to which kind of device the encoded bitstream is to be sent, the encoder has to compress the video into a standardized format. Then, all devices that support the chosen standard can successfully decode the video. The compression can be lossless, i.e., the decoded video will be identical to the source video provided to the encoder, or it can be lossy, where a certain degree of content degradation is accepted. Using lossy compression can allow for significantly lower bitrates, i.e., the compression ratio can be much higher. This is because perfectly reproducing image noise would make lossless compression quite expensive.

[0004] A video sequence contains a sequence of pictures. A commonly used color space in video sequences is YCbCr, where Y is the luma component and Cb and Cr are the chroma components. Sometimes, the Cb and Cr components are referred to as U and V. Other color spaces are also used, such as ICtCp (aka IPT) (where I is the luma component and Ct and Cp are the chroma components), constant luminance YCbCr (where Y is the luma component and Cb and Cr are the chroma components), RGB (where R, G, and B correspond to the blue, green, and blue components, respectively), YCoCg (where Y is the luma component and Co and Cg are the chroma components), and so on.

[0005] The order in which pictures are placed in a video sequence is referred to as the "display order". Each picture is assigned a picture order count (POC) value to indicate its display order. In this disclosure, the terms "image", "picture", or "frame" can be used interchangeably.

[0006] Video compression is used to compress a video sequence into a sequence of coded pictures. In many existing video codecs, pictures are divided into blocks of different sizes. A block is a two-dimensional array of samples. The blocks are used as a basis for encoding. A video decoder then decodes the coded pictures into pictures containing sample values.

[0007] Video standards are typically developed by international organizations, as these organizations represent different companies and research institutions with different areas of expertise and interests. The most widely applied video compression standard at present is H.264 / AVC (Advanced Video Coding), which was jointly developed by ITU-T and ISO. The first version of H.264 / AVC was finalized in 2003, with several updates in the following years. The successor of H.264 / AVC is called H.265 / HEVC (High Efficiency Video Coding), which was also developed by ITU-T (International Telecommunication Union - Telecommunication) and International Organization for Standardization (ISO), and finalized in 2013. MPEG and ITU-T created a joint video exploration team (JVET) for a successor to HEVC. The name of this video codec is Versatile Video Coding (VVC), and version 1 of the VVC specification has been published as Rec. ITU-T H.266 | ISO / IEC (International Electrotechnical Commission) 23090-3, “Versatile Video Coding”, 2020.

[0008] The VVC video coding standard is a block-based video codec and utilizes both temporal and spatial prediction. Spatial prediction is achieved using intra (I) prediction within the current picture. Temporal prediction is achieved using single-directional (P) or bi-directional inter (B) prediction at block level from previously decoded reference pictures. In the encoder, the difference between the original sample data and the predicted sample data, referred to as the residual, is transformed to the frequency domain, quantized, and then entropy coded before being transmitted together with the necessary prediction parameters, such as prediction mode and motion vectors, which can also be entropy coded. The decoder performs entropy decoding, inverse quantization, and inverse transformation to obtain the residual, which is then added to the intra or inter prediction to reconstruct the picture.

[0009] The VVC video coding standard uses a block structure called quad-tree plus binary-tree plus ternary-tree block structure (QTBT+TT), where each picture is first divided into squares called coding tree units (CTU). All CTUs have the same size, and dividing the picture into CTUs is done without any syntax controlling it. Each CTU is further divided into coding units (CU) that can have square or rectangular shape. The CTU is first divided by a quad-tree structure, and then it can be further divided vertically or horizontally in a binary structure with equally sized divisions to form coding units (CU). Thus, the blocks can have square or rectangular shape. The depth of the quad-tree and binary-tree can be set by the encoder in the bitstream. The ternary-tree (TT) part adds the possibility to divide a CU into three divisions instead of two equally sized divisions. This increases the possibility of using a block structure that better follows the content structure of the picture, e.g., roughly following important edges in the picture.

[0010] Intra coded blocks are I-blocks. Single predicted blocks are P-blocks, and bi-predicted blocks are B-blocks. For some blocks, the encoder decides that the residual does not need to be coded, possibly because the prediction is close enough to the original block. Then, the encoder signals to the decoder that the transform coding of that block should be bypassed (i.e., skipped). Such blocks are referred to as skipped blocks.

[0011] At the 20th JVET meeting, it was decided to set up exploratory experiments (EE) for neural network based (NN-based) video coding. The exploratory experiments continued at subsequent JVET meetings 21-29 with additional tests, including NN-based in-loop filtering, NN-based post filtering, NN-based super-resolution, and NN-based intra prediction.

[0012] Regarding in-loop filtering, VVC contains three in-loop filters that are not currently based on neural networks: a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). The deblocking filter is used to remove blocking artifacts by smoothing discontinuities across block boundaries in horizontal and vertical directions. The deblocking filter uses a block boundary strength (BS) parameter to determine the filtering strength. The BS can have values 0, 1, and 2, where a larger value indicates stronger filtering. The output of the deblocking filter is further processed by the SAO, and then the output of the SAO is processed by the ALF. The output of the ALF filter can then be put into the decoded picture buffer (DPB), which is used for prediction of pictures that are coded (or decoded) subsequently. Since the deblocking filter, the SAO filter, and the ALF can affect pictures in the DPB that are used for prediction, they are classified as in-loop filters, also referred to as in-loop filters. This means that the changes made by the in-loop filters affect not only the current picture but also future pictures. Decoders can further filter pictures in the DPB, but do not store the filtered output in the DPB, but only send it to the display / decoding file. In contrast to in-loop filters, such filters do not affect future predictions, and thus are classified as post-processing filters, also referred to as post filters.

[0013] Contributions JVET-X0066 and JVET-Y0143 are two consecutive contributions describing neural network based in-loop filtering. Both contributions use the same NN model for filtering. The NN based in-loop filter is placed before SAO and ALF, and the samples before the deblocking filter are used as input to the filter. One purpose of using the NN based filter is to improve the quality of the reconstructed samples. Here, it is helpful that the neural network model is non-linear. While deblocking, SAO and ALF all contain non-linear elements such as conditions, they are thus not strictly linear, but they are all based on linear filters. In contrast, a NN model large enough in principle can learn any non-linear mapping, and thus can represent a much wider range of functions compared to deblocking, SAO and ALF. In JVET-X0066 and JVET-Y0143, there are four neural network models (i.e., four neural network based in-loop filters). In an improved version of this work presented in contribution JVET-AB0052, only two models are used: one for luma samples and one for chroma samples. The use of NN filtering can be controlled (turned on or off) on a block (CTU) level or picture level. The encoder can further determine the filtering strength for the blocks / pictures it turns on.

[0014] The two NN models are convolutional neural networks. The terms "neural network", "neural network model" and "NN model" can be used interchangeably. Using the model for luma samples as an example, the model from JVET-AB0052 has five inputs, the reconstructed sample of luma before deblocking (called "rec"), the predicted sample of luma ("pred"), the BS information of luma (BS), the quantization parameter (qp), and the information about whether the particular sample is intra predicted, uni-predicted or bi-predicted (called "IPB"). The five inputs are first passed through a convolutional layer and a parametric rectified linear unit (PReLU) layer each, and then they are concatenated and fused together to generate the signal y as Figure 4A The figure shows the "head" of the neural network, including the inputs, the fusion part and the transition part.

[0015] Each input is connected to a convolutional layer with kernel size 3x3, denoted as "conv3x3" in the figure, with multiple output channels. As an example, in JVET-AB0052, 96 channels are used for each input, while in JVET-AC0126, more channels are used for the "rec" input and fewer channels are used for the "QP" input, since the "rec" input typically carries more information about the signal than the "QP" input. The fusion section contains a "conv1x1" layer, which is a convolutional layer with kernel size 1x1. In JVET-AB0052 and JVET-AC0126, this layer has 5*96 = 480 input channels and 96 output channels. PReLU constitutes an activation layer. "Unsqueeze, expand" is a dimension operation that expands qp (which is just a number in the experimental software NNVC) to the same size as the other inputs (e.g., 128x128 or 256x256 sample arrays). It should be noted that in many applications, it can be the case that different samples within a CTU will be associated with different QPs. In this case, qp would effectively have the same resolution as the other inputs (rec, pred, part, and bs), and the "un-squeeze, expand" layer would not be needed. Finally, "concat" denotes concatenation, and "2 " denotes down-sampling by a factor of 2.

[0016] After the head, there are N sequential residual blocks of the same structure. In JVET-AB0052 and JVET-AC0126, N = 8. The first residual block inputs y and outputs z0. The i-th residual block outputs zi, i = 0,.., 7. In Figure 4B The structure of the residual block is depicted in FIG. 2. After the eight residual blocks, the signal z7 is processed by a convolutional layer, PReLU, another convolutional layer, pixel shuffle, and a final scaling to generate the output, as shown in Figure 4C The chroma model differs from the luma model. For example, the chroma model takes the reconstructed luma samples ("rec") as input. However, embodiments of the present disclosure apply to both luma and chroma processing.

[0017] There is still a need for improved filtering. For example, the NN-based in-loop filters proposed in JVET-X0066, JVET-AB0053, JVET-AB0052, and JVET-AC0126 significantly improve the codec’s compression efficiency (i.e., they significantly reduce the bit rate without reducing the objective quality as measured by MSE-based PSNR). The increase in compression efficiency, often simply referred to as “gain”, is typically measured as Bjontegaard-delta rate (BDR) against an anchor. As an example, a BDR of -1% means that the same PSNR distortion can be achieved with 1% less bits. As reported in JVET-Y0143, the BDR gain for the luma component (Y) is -9.80% for random access (RA) configuration, and -7.39% for all intra (AI) configuration. The complexity of a NN model used for compression is typically measured by MACs / pixel (multiply-accumulate operations per pixel). The high gain of a NN model is directly related to the high complexity of the NN model. The luma intra model described in JVET-Y0143 has a complexity of 430 kMACs / pixel, i.e., 430000 multiply-accumulate operations per pixel. There are other measures of complexity, e.g., the total model size in terms of stored parameters. As another example, in some frameworks, e.g., in Keras, one can have convolutional layers where the weights are different in each output location. In Keras, such a layer is called LocallyConnected2D. This means that the number of parameters increases very quickly with the size of the input, making only very small input image sizes feasible. SUMMARY

[0018] In neural network based video coding, methods to improve the complexity-performance trade-off are highly appreciated. The structure of existing neural network models can be improved. For example, one can keep the compression efficiency performance the same (or even improve the compression efficiency performance) while reducing the complexity. The high complexity of current NN models is a significant challenge for practical hardware implementation. Therefore, it is highly desirable to reduce the complexity while keeping the NN compression efficiency performance the same. For example, a lower number of KMACs / sample means that less power is consumed in a wireless device, increasing the battery life. It also means that a smaller chip surface area needs to be used for video decoding, reducing the manufacturing cost. One can typically trade off a reduced bit rate at constant complexity for a reduced complexity at constant bit rate, and vice versa. Therefore, it is also highly desirable to improve the coding efficiency while keeping the complexity the same.

[0019] The current state-of-the-art neural network based loop filter is built on a convolutional neural network. These networks have the property that the same weights are used for all convolutional positions. However, the output signal from a video codec (“rec”) is not translation invariant. As an example, the smallest transform is 4x4 samples large, and the top-left sample of a block that goes through such a transform must therefore have x and y coordinates that are divisible by 4. The result of this is that samples close to the edge of a 4x4 block are more likely to be affected by block artifacts than samples in the center of the block. Therefore, according to embodiments, methods and apparatuses are provided that better adapt to this position dependent property of the signal by allowing position dependent weights in the network in an efficient way. Another aspect of the embodiments is to make the network aware of the position within the image so that it can treat different positions differently. In certain aspects, the network can have the potential to exploit known and existing signal properties to improve performance.

[0020] According to a first aspect of the disclosure, a method is provided. The method comprises identifying a first group within a set of inputs, wherein the first group comprises a first plurality of values. The method comprises selecting a first weight tensor, wherein the first weight tensor is selected based at least in part on positions of the first values. The method comprises generating a first weighted sum by applying the first weight tensor to the first plurality of values. The method comprises identifying a second group within the set of inputs, wherein the second group comprises a second plurality of values. The method comprises selecting a second weight tensor, wherein the second weight tensor is selected based at least in part on positions of the second values, wherein the second weight tensor is different from the first weight tensor. The method comprises generating a second weighted sum by applying the second weight tensor to the second plurality of values. The method comprises identifying a third group within the set of inputs, wherein the third group comprises a third plurality of values. The method comprises selecting a third weight tensor, wherein the third weight tensor is selected based at least in part on positions of the third values. The method comprises generating a third weighted sum by applying the third weight tensor to the third plurality of values, wherein the first weight tensor and the third weight tensor are the same. The method further comprises performing one or more of: generating a first output value by processing the first weighted sum with a non-linear activation function, generating a second output value by processing the second weighted sum with a non-linear activation function, and generating a third output value by processing the third weighted sum with a non-linear activation function.

[0021] According to a second aspect of the disclosure, a device configured to perform the method according to the first aspect is provided.

[0022] According to a third aspect of the disclosure, a computer program comprising instructions, which when executed by processing circuitry, causes the processing circuitry to perform the method according to the first aspect is provided.

[0023] According to a fourth aspect of the present disclosure, there is provided a carrier containing the computer program according to the third aspect, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.

[0024] Certain embodiments can provide one or more of the following technical advantages. For example, one advantage of the proposed solution is an improved performance vs. complexity trade-off in the codec. The BD rate can be reduced by more than half a percentage point without increasing the complexity in kMACs / sample. This means that, according to some embodiments, the bit rate is reduced by 0.5% without changing the picture quality in terms of PSNR, and without increasing the complexity in terms of kMACs / sample. Another aspect is that the proposed solution only requires at most a quarter of the computation and a quarter of the number of parameters compared to other techniques. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate various embodiments.

[0026] Figure 1A Figures 1A, 1B, 1C each illustrate a system according to some embodiments.

[0027] Figure 2 Figures 2A, 2B, 2C each illustrate a schematic block diagram of an encoder according to some embodiments.

[0028] Figure 3 Figures 3A, 3B, 3C each illustrate a schematic block diagram of a decoder according to some embodiments.

[0029] Figure 4A Figure 4 illustrates a schematic block diagram of an example head using five inputs. Figure 4B Figure 5 illustrates a schematic block diagram of an example body with repeated residual blocks. Figure 4C Figure 6 illustrates a schematic block diagram of an example tail with final processing including pixel shuffle.

[0030] Figure 5A Figures 7A, 7B, 7C each illustrate a schematic block diagram of a filter according to some embodiments. 5B Figure 8 illustrates an aspect of filtering.

[0031] Figure 6 Figure 9 illustrates a division of a picture into blocks according to an embodiment.

[0032] Figure 7 Figures 10A, 10B, 10C each illustrate a schematic block diagram of a weight tensor according to some embodiments. Figure 8 Figures 11A, 11B, 11C each illustrate a schematic block diagram of a weight tensor according to some embodiments. Figure 9A Figures 12A, 12B, 12C each illustrate a schematic block diagram of a weight tensor according to some embodiments. Figure 9B Figures 13A, 13B, 13C each illustrate a schematic block diagram of a weight tensor according to some embodiments.

[0033] Figure 10 Figures 14A, 14B, 14C each illustrate a schematic block diagram of a weight tensor according to some embodiments. 11 Figures 15A, 15B, 15C each illustrate a schematic block diagram of a weight tensor according to some embodiments.

[0034] Figure 12、 Figure 13 、 Figure 14A and Figure 14B shows the use of a weight tensor according to an embodiment.

[0035] Figure 15 and 16 shows an example of a shared weight matrix within a repeater matrix according to an embodiment.

[0036] Figure 17 and 18 shows the use of a weight tensor according to an embodiment.

[0037] Figures 19-25 shows the use of pixel shuffle and unshuffle operations according to an embodiment.

[0038] Figure 26 and 27 shows a repeater matrix pattern according to some embodiments.

[0039] Figures 28-30 shows the use of input for matrix selection according to some embodiments.

[0040] Figures 31A-31C and Figures 32A-32C shows a process according to some embodiments.

[0041] Figure 33 shows a schematic block diagram of a device according to an embodiment. DETAILED DESCRIPTION

[0042] One or more embodiments described herein can take advantage of the fact that a video codec can not treat all sample locations equally. In certain aspects, this is done by having the neural network know the location of the sample within the picture, for example by having different weights for different locations. Contrary to how non-shared convolutional layers such as Keras LocallyConnected2D work, according to an embodiment, weights are periodically reused. In this way, the neural network can differentiate between different sample locations in an efficient manner and provide many advantages.

[0043] Figure 1A A system 100 according to some embodiments is shown. The system 100 comprises a first entity 102, a second entity 104, and a network 110. The first entity 102 is configured to transmit a video stream (a.k.a. “video bitstream”, “bitstream”, “encoded video”) 106 to the second entity 104.

[0044] The first entity 102 can be any computing device (e.g., a network node, such as a server) capable of encoding video using encoder 112 and transmitting the encoded video to the second entity 104 via network 110. The second entity 104 can be any computing device (e.g., a network node) capable of receiving the encoded video and decoding it using decoder 114. Each of the first entity 102 and the second entity 104 can be a single physical entity or a combination of multiple physical entities. Multiple physical entities can be located in the same location or can be distributed in the cloud.

[0045] In some embodiments, such as Figure 1B As shown, the first entity 102 is a video streaming server 132, and the second entity 104 is a user equipment (UE) 134. The UE 134 can be any of a desktop computer, laptop computer, tablet computer, mobile phone, or any other computing device. The video streaming server 132 is capable of transmitting video bitstream 136 (e.g., YouTube) to the video streaming client 134. TM (Video stream). After receiving the video bitstream 136, the UE 134 can decode the received video bitstream 136 to generate and display the video for the video stream.

[0046] In other embodiments, such as Figure 1C As shown, the first entity 102 and the second entity 104 are the first UE 152 and the second UE 154, respectively. For example, the first UE 152 can be a provider of a video conferencing session or a caller of a video chat, while the second UE 154 can be a responder of a video conferencing session or a responder of a video chat. Figure 1C In the illustrated embodiment, the first UE 152 is capable of transmitting data to the second UE 154 for video conferencing (e.g., Zoom™, Skype). TM MS Teams TM (etc.) or video chat (e.g., FaceTime) TM The video bitstream 156 is received. After receiving the video bitstream 156, the UE 154 can decode the received video bitstream 156 to generate and display video for video conferencing sessions or video chats.

[0047] Figure 2A schematic block diagram of an encoder 112 is shown in accordance with some embodiments. The encoder 112 is configured to encode blocks of sample values (hereinafter referred to as "blocks") in video frames of a source video 202. In the encoder 112, a current block (e.g., a block contained in a video frame of the source video 202) is predicted by performing motion estimation from blocks already provided in the same frame or a previous frame by a motion estimator 250. In the case of inter-frame prediction, the result of the motion estimation is a motion or displacement vector associated with a reference block. A motion compensator 250 utilizes the motion vector to output an inter-frame prediction of the block.

[0048] An intra-frame predictor 249 computes an intra-frame prediction of the current block. The outputs from the motion estimator / compensator 250 and the intra-frame predictor 249 are input to a selector 251 which selects either the intra-frame prediction or the inter-frame prediction for the current block. The output from the selector 251 is input to an error calculator in the form of a summer 241 which also receives the sample values of the current block. The summer 241 calculates and outputs a residual error as the difference in sample values between the block and its prediction. The error is transformed in a transformer 242, e.g., by a discrete cosine transform, and quantized by a quantizer 243, followed by encoding in an encoder 244, e.g., by an entropy encoder. In inter-frame coding, the estimated motion vector is brought to the encoder 244 for generating an encoded representation of the current block.

[0049] The residual error for the current block, which is transformed and quantized, is also provided to an inverse quantizer 245 and an inverse transformer 246 to retrieve the original residual error. This error is added by a summer 247 to the block prediction output from the motion compensator 250 or the intra-frame predictor 249 to create a reconstructed sample block 280, which can be used for prediction and encoding of the next block. According to embodiments, the reconstructed sample block 280 is processed by a NN filter 230 in order to perform filtering. The filtering can be, for example, against blocking artifacts. The output of the NN filter 230, i.e., the output data 290, is then temporarily stored in a frame buffer 248, where it can be used by the intra-frame predictor 249 and the motion estimator / compensator 250. In some embodiments, the encoder 112 can include a SAO unit 270 and / or an ALF 272. The SAO unit 270 and the ALF 272 can be configured to receive the output data 290 from the NN filter 230, perform additional filtering on the output data 290, and provide the filtered output data to the buffer 248.

[0050] Even in Figure 2In the embodiment shown in FIG. 3, the NN filter 230 is arranged between the SAO unit 270 and the adder 247. In other embodiments, the NN filter 230 can replace the SAO unit 270 and / or the ALF 272. In some embodiments, the NN filter 230 can be after the SAO unit 270 and before the ALF 272, or after both. Alternatively, in other embodiments, the NN filter 230 can be arranged between the buffer 248 and the motion compensator 250. Furthermore, in some embodiments, a deblocking filter (not shown) can be arranged between the NN filter 230 and the adder 247, such that the reconstructed sample block 280 is deblocked before being provided to the NN filter 230. In embodiments, the NN output can be combined with the deblocking output. For example, where the output of the NN is: outNN = f(rec, other inputs to the NN), the output of the deblocking filter is: outDBLK = deblock(rec), then the combination combNN is: combNN = c*outNN + (1-c)*outDBLK where c is a number between 0 and 1 that can be signaled from the encoder to the decoder, and where * is scalar multiplication. This can ensure that the combNN output has been filtered through the NN filter (e.g., if c = 1) or through the deblocking filter (c = 0) or a weighted between the two filter outputs (e.g., if c = 0.5, we consider combNN to be the average of outNN and outDBLK). In certain aspects, this combNN is then fed to Figure 2 the SAO unit 270 of FIG. 3. The foregoing discussion regarding NN filter 230 filter placement and functionality can also apply to the NN filter 330 as shown in FIG. 3. Alternatively, the NN filter 230 can alternatively be placed as a post-filter, i.e., it will not affect the samples in the display buffer 248 used for prediction. Figure 2

[0051] Figure 3 is a schematic block diagram of a decoder 114 according to some embodiments. The decoder 114 includes a decoder 361, e.g., an entropy decoder, to decode the encoded representation of a block to obtain a set of quantized and transformed residual errors. These residual errors are inverse quantized in an inverse quantizer 362 and inverse transformed by an inverse transformer 363 to obtain a set of residual errors. These residual errors are added in an adder 364 to sample values of a reference block. The reference block is determined by a motion estimator / compensator 367 or an intra predictor 366, depending on whether inter or intra prediction is performed.

[0052] ​The selector 368 is thus interconnected with the adder 364 and the motion estimator / compensator 367 as well as the intra predictor 366. According to embodiments, the resulting decoded block 380 output from the adder 364 is input to the NN filter unit 330 to filter any blocking artifacts. The filtered block 390 is the output from the NN filter 330 and is furthermore preferably temporarily provided to the frame buffer 365 and can be used as a reference block for subsequent blocks to be decoded. The frame buffer (e.g. decoded picture buffer (DPB)) 365 is thus connected to the motion estimator / compensator 367 to make the stored sample blocks available to the motion estimator / compensator 367. The output from the adder 364 is preferably also input to the intra predictor 366 to be used as an unfiltered reference block. In some embodiments, the decoder 114 can comprise a SAO unit 380 and / or an ALF 382. The SAO unit 380 and the ALF 382 can be configured to receive the output data 390 from the NN filter 330, perform additional filtering on the output data 390, and provide the filtered output data to the buffer 365.

[0053] Even in embodiments as shown in Fig. 1 and Figure 3 , the NN filter 330 is arranged between the SAO unit 380 and the adder 364, in other embodiments, the NN filter 330 can replace the SAO unit 380 and / or the ALF 382. Alternatively, in other embodiments, the NN filter 330 can be arranged between the buffer 365 and the motion compensator 367. Furthermore, in some embodiments, a deblocking filter (not shown) can be arranged between the NN filter 330 and the adder 364, such that the reconstructed sample block 380 is deblocked before being provided to the NN filter 330. Alternatively, in some embodiments, the NN filter 330 can be used as a post filter. It would then be placed between the frame buffer and the display / output and the samples processed by the NN filter would not be used for prediction of the motion estimation compensation 367.

[0054] According to embodiments, one or more of the arrangements shown in Fig. 1 and Figure 2 . Figure 4A , 4B and / or 4C can be used in the NN filter 230, 330 shown in Fig. 1 and Figure 2 . Figure 2 and Figure 3 and the system 100.

[0055] As described with respect to Fig. 1 and Figure 2 . With reference to Figures 4A-4CFor example, the NN solution can be based on a convolutional neural network (CNN), such as a neural network built using building blocks named conv3x3, convlxl, conv3x3, 2↓, etc. The function of such a block can be shown as in Figure 5A 00 An 8x8 block of input samples (501), containing samples s 01 (503), etc. will be processed by a 3x3 weight matrix (505). The terms weight matrix and weight tensor can be used interchangeably. The term sample here can for example refer to an input to a neural network or an output signal from a previous neural network layer.

[0056] Starting from sample s 11 (506), the 3x3 neighborhood (504) centered on sample s 11 (506) is multiplied element-wise with the weight matrix W (505). The result is the weighted sum t 11 (507): .

[0057] A bias term is added to the weighted sum t 11 and the result is fed through a non-linear activation function, for example a parametric rectified linear unit PReLU: The output value u11 is then placed in the output 8x8 block (not shown) in a position corresponding to the position of s 11 in the input block. Sometimes there is more than one output channel, in which case the process is repeated again with another weight matrix W and another PReLU function. Likewise, sometimes there is more than one input channel, in which case the weight matrix W becomes a three-dimensional tensor. However, to simplify the representation, Figure 5A only a CNN with one channel input and one channel output is shown.

[0058] Figure 5B The processing of the next element s 12 (516) is shown. Again, the 3x3 neighborhood centered on element s 12 ​(506) The 3x3 neighborhood (514) around the center is used with the weight matrix W (505) to create the weighted sum (517). Here, the same weight matrix is used regardless of which sample is the center sample, which means that the weight matrices W (505) and W (515) are identical. This reuse of weights can be an important aspect of a convolutional neural network, as it means that the number of parameters in the network can be kept low compared to a fully connected neural network. It also means that a single weight (such as w10) will be trained on more training examples (i.e. more sample locations) compared to the case where the network is fully connected. This makes training faster, and easier to converge. In many practical applications, it also makes sense to reuse weights for different locations in the picture. For example, in the original image data, there can be no unique features that cause samples with e.g. odd x-coordinate to behave differently compared to samples with even x-coordinate.

[0059] For pictures that are output as a video decoder, this is typically no longer true. Many aspects of the VVC codec (e.g. transform size, partitioning) mean that the decoded output picture is no longer independent of the location within the picture. As an example, samples close to a 4x4 boundary are more likely to exhibit block artifacts than samples located in the center of a 4x4 block. However, using a fully connected layer or a non-shared convolutional layer (like Keras LocallyConnected2D) would mean that the number of parameters would grow extremely fast as a function of the picture size, which is not feasible. According to embodiments, a solution is provided in which weights are repeatedly reused.

[0060] Figure 6 Aspects of one or more embodiments are illustrated. In this example, instead of assuming that the picture is divided into 4x4 blocks, it is assumed that an 8x8 picture (601) in Figure 6 is divided into 2x2 blocks. For example, samples a 00 , b 01 , c 10 and d 11 together form a 2x2 block (602). Likewise, samples a 02 , b 03 , c 12 and d 13 together form a 2x2 block (603). Note that in each 2x2 block, the top-left sample is always denoted a, the top-right sample is always denoted b, the bottom-left sample is always denoted c, and the bottom-right sample is always denoted d. This division into blocks with repeated sample notation can be repeated over the entire picture 601.

[0061] Figure 7 is shown how filtering can be performed according to an embodiment of the invention. When processing the 2x2 block (603), the filter (604) is applied to the samples d 11When monitoring the position (1,1) in the representation (e.g. in a CTU or other value / sample block 701), a 3x3 neighborhood (704) is used with a weighted weight matrix Δ (705) to produce a weighted sum t according to .

[0062] Similarly, Figure 8 The processing of the sample in position (1,2) is shown, which is represented by the next sample c 12 (806) in the input group 801. Again, a 3x3 neighborhood (804) is used with a weight matrix Γ (805) to produce a weighted sum t according to 12 (807) However, and according to some embodiments, the weight matrix Γ (805) is different from the weight matrix Δ (705). The weight matrix Δ (705) is reserved only for the case when the intermediate sample is a d sample (i.e. a sample having a lower right position within the 2x2 block). Likewise, the weight matrix Γ (805) is reserved for the case when the intermediate sample is a c sample (i.e. a sample having a lower left position within its 2x2 block). Furthermore, but not shown, the weight matrix A is used for the case when the intermediate sample is an A sample (i.e. a sample having an upper left position within its 2x2 block), and the weight matrix b is used for the case when the intermediate sample is a b sample (i.e. a sample having an upper right position within its 2x2 block). It should be noted that in all these cases, not only different weight tensors can be used, but also different bias terms. Thus, if the intermediate sample is an a sample, a bias value bias α may be used if the intermediate sample is a b sample, a bias value bias β may be used for c samples and d samples, respectively, and bias γ and bias bias δ may be used. It should also be noted that if the convolution is 4x4 instead of 3x3, there is no exact center sample. However, in this case, the switching of the weight tensor can be done with reference to another sample of the 4x4 neighborhood, such as the upper left sample. In the following, we will mainly describe how the weight tensors are handled, but it is part of the invention that the bias values can be handled in a similar way.

[0063] Reference is now made to Figure 9A , where the center sample of the 3x3 neighborhood (904) is the sample c 54of (906). Since it is a c sample, i.e. has its lower-left position within its 2x2 block in this example, its 3x3 neighborhood (904) will be multiplied by the matrix (905) to form the weighted sum (907). Note that this matrix 905 is the same as matrix (805) because it deals with the same kind of sample (lower-left sample, c sample in this example). This is in contrast to how non-shared convolutional layers like Keras LocallyConnected2D work - in that case, these matrices would be different. It should be noted that the specific letters and corresponding positions used in the example are for illustration purposes only.

[0064] According to some embodiments, the coefficients in the matrix Γ will always multiply the same kind of sample. As an example, the coefficient γ 00 in the sum (907) will multiply a sample of type "lower-right" b 43 Similarly, in Figure 8 the sum (807) the coefficient γ 00 will multiply a sample of type "lower-right" b 01 b 01 is also a sample of "lower-right" type. In some aspects, this makes it possible for the weights in different matrices to be specialized for a specific position within the 2x2 block. As one example, if the lower-left sample is always darker than the other samples, the invention can compensate for this by increasing the weight in the matrix Γ. This is not possible with traditional convolutional layers, nor with LocallyConnected2D layers, because different matrices are used everywhere.

[0065] Now referring to Figure 9B and according to some embodiments, in cases where the two neighborhoods of block 911 have overlapping positions, the matrices can be reused. In Figure 9B the example, the region 914 centered on sample c14 (916) overlaps with the region 804 shown in Figure 8 which is also demarcated by the dashed line in Figure 9B . As shown, they use the same matrix Γ because they are both centered on a c sample.

[0066] For the purpose of simplifying the illustration, the example is illustrated using the case where there is one input channel and one output channel, but it should be noted that the embodiments described herein can be generalized to any number of input channels and output channels. For example, in Figure 10In the illustrated test implementation, 96 input channels and 96 output channels are used with the position-dependent convolutional layer. In addition, 2x2 has been used for simplicity, but pictures can use different sizes. For example, VVC video decoding pictures are more likely to exhibit regularity on 4x4 sample grids. However, due to downsampling in the transition part of the head, 4x4 regularity in the input data corresponds to 2x2 regularity in the backbone part of the neural network.

[0067] According to some embodiments, Figure 4B The first conv3x3 in the backbone illustrated in the above can be implemented in Pytorch in a position-dependent manner using the following code: Here, tl stands for "top left", tr stands for "top right", bl stands for "bottom left", and br stands for "bottom right". The convl_tl, convl_tr, convl_bl, and convl_br can be implemented using regular CNN layers with a stride of 2: While the first conv3x3 in the backbone is used as an example here, embodiments are applicable to other parts of the NN (e.g., the head or the tail) and other convolutional layers.

[0068] Figure 10 Performance of the position-dependent NN of one or more embodiments is shown compared to a regular implementation without position-dependent convolution. Figure 10 The curve with circles in the above shows the performance of the NN in JVET-AB0053 with traditional convolution. The solid line shows the performance with L1 loss function and learning rate le-4, and the dashed line shows the performance at the end of training when the Mean Squared Error (MSE) loss function is used with learning rate le-5 to fine-tune the training. In the graph, lower BDR difference is better, and the last circle indicates a 7.61% bitrate reduction. Figure 10 The curve with crosses in the above shows the performance of the NN according to embodiments. Again, the solid line shows the performance with L1 loss function and learning rate le-4, and the dashed line shows the performance at the end of training when the MSE loss function is used with learning rate le-5 to fine-tune the training. As can be seen from the graph, the rightmost cross of -8.12% is significantly lower (better) than the rightmost circle. The difference is -8.12% - (-7.61%) = -0.51%. However, the same number of Multiply Accumulate (MAC) operations are performed in both cases, which means that the complexity in terms of kMACs / sample remains the same.

[0069] In the above example, four matrices instead of one are needed, and also four bias terms instead of one, so the number of parameters is roughly quadrupled. In some cases, this increase in parameter count can be a reasonable tradeoff for a 0.5% BD rate gain. However, one can wish to reduce the number of parameters. Thus, in another embodiment, instead of using position-dependent convolutions in all conv3x3 layers in the trunk, one uses them only in a few of the eight residual blocks. As an example, Figure 11 Performance is shown for using position-dependent convolutions in only the last residual block of the eight residual blocks. Other subsets can be chosen. Here, in addition to the previous curves, the curve marked with triangles shows the performance of the approach where position-dependent filtering occurs only in the last residual block of the trunk. It can be seen that this curve is in between the two other curves, which means that some gain can be preserved while the number of parameters is significantly reduced.

[0070] In the above example, weights in the trunk are reused after 2x2 samples. For example, referring to Figure 7 When the filter is centered at position d 11 (706), the same matrix Δ (705) is used as when centered two samples to the right of position d 13 (705) or two samples down of position d 31 (705). It is also possible to reuse weights less, for example every 4x4 samples, as shown in Figure 12 and with the set of inputs 1201. Here, the 8x8 block of samples (1201) is divided into four 4x4 blocks, indicated by the dashed dividers (1202). These four blocks are examples of relay blocks, in that the weight tensor is repeated (reused) between them. As an example, the sample f 11 (1206) is one step to the right and one step down in its 4x4 block (1203), and the sample f 55 (1208) is in the same relative position (one step to the right and one step down) in its 4x4 relay block (1211). Thus, when processing t 11 (1207) (where the convolution filter is centered at f 11 (1206)) and when processing t55 (1210) (where the convolution filter is centered at f 55 (1208)), the same weight tensor Φ (1205) will be used. Using 4x4 relay blocks can increase the number of required parameters, and thus memory consumption, compared to the case of using 2x2 relay blocks. In contrast, as shown in Figure 13 and block 1301, when filtering is centered at another position g 12 (1306) and another position within its 4x4 relay block (1303) to produce t 12(1307) At this point, a different weight tensor H (1305) is used, and this weight tensor is alternatively shared with other positions having similar positions, such as g 56 (1308) t 56 (1310).

[0071] According to some embodiments, the weight tensors at different positions within a repeater block (or sub-block) can be grouped together. Figure 14A An example is shown in which samples 1401 are grouped according to how far they are from a block boundary. Samples in the corner of a repeater block (such as a77 (1411)) that touch both dashed dividers (1420) are labeled with “a”, and all of these samples use the same weight matrix A when a filter is centered on them. Samples at the edge of a repeater block but not at a corner, thus touching exactly one dashed divider, such as sample b 13 (1406) use a different weight tensor B (1405). As an example, when a convolutional filter is centered on sample b 13 (1406) and when a convolutional filter is centered on sample b 45 (1408) both use B because both are edge samples. Finally, samples in the middle of a repeater block that do not touch any dashed dividers are denoted as “c”, and use a third weight tensor G for filtering. By sharing filters in this way, memory consumption can be reduced while still preserving a relevant position dependence, in this case distance to an edge. Other relevant position dependencies can be used. In some embodiments, various patterns of reusing weight tensors can be repeated. Referring to Figure 14B as an example, an 8x8 repeater block 1451 is used. When centered on an A value, such as a 13 (1452), weight tensor A (1453) is used. This can be repeated for most (or almost all) positions. However, there can be limited exceptions in a given embodiment or application. For example, here there is only one exception— the value centered on b 56 (1454), in which case weight tensor B (1455) is used. That is, one tensor can be used for most (but not all) positions.

[0072] Figure 15 and Figure 16 Two other examples of shared weight matrices within a repeater matrix are shown in Figure 15 In Figure 16In this case, the size of the repeater block is 4x4 and only two different weight tensors are used, one for A samples and one for b samples. Yet another example is a repeater matrix in a checkerboard pattern as shown in Figure 26 Switching the matrix on a per sample basis can put a burden on some implementations. Therefore, in another embodiment, a repeater matrix can be used that keeps the same weights in a 2x2 area as shown in Figure 27 In some embodiments, three-dimensional repeater blocks can be used.

[0073] In another embodiment, the weight tensors can not only have different coefficients but also different shapes for different positions within the repeater block. Figure 17 An example of this is shown in Fig. 17, where an input set 1701. Here, when calculating t 12 (1707), the filter region (1703) centered at sample g12 (1706) has a shape of 3x3 and a 3x3 weight tensor (1705) is used. However, when calculating t64 (1710), the filter region (1709) centered at sample i64 (1708) has a shape of 1x3 and a 1x3 weight tensor (1711) is used. By having smaller filter regions at some positions, the complexity in terms of KMAC can be reduced and also the number of parameters (number of tensor weights).

[0074] In some embodiments, the signal to be compressed is not two-dimensional but one-dimensional. Some examples are sound signals, radio signals, text, etc. Also in these cases, position dependent filtering can be used to enhance the CNN. Figure 18 An example is shown in Fig. 18. Here, an 8-sample frame (1801) is divided into four repeater blocks (1802), each having 2 samples. When the convolution filter is centered at b1 (1804), a weight tensor B (1805) is used, where the same weight tensor is used for all samples labeled B (b1, b3, b5 and b7). In contrast, if the convolution filter is centered at a sample labeled “A” such as a6 (1807), another weight tensor A (1808) is used and the same weight tensor is used for (a0, a2 and a4) as well.

[0075] Although different numbers are used to label the input sets (e.g. blocks or frames) 801, 901,... in the various figures, the used figure 1801 can be the same in different figures according to embodiments. That is, although shown with different numbers, different figures can show different filtering steps (or sub-steps) applied to the same input. As an example, Fig. 9 shows the same input set 901 as Fig. 8. Figure 8The selection of neighborhoods (904) is different from the neighborhoods (804) shown in the middle, but they can be selected in the same way as in the monitoring 801, 901.

[0076] Pixel shuffle and pixel unshuffle can be built-in functions in certain modern neural network training frameworks, such as PyTorch (torch.nn.PixelShuffle and torch.nn.PixelUnshuffle). Figure 19 An example is shown of how pixel unshuffle and pixel shuffle can be implemented. In this example, a single-channel sample block of size 4x4 (1901) undergoes pixel unshuffle (1902). This creates a four-channel output, where all samples with even x and y coordinates end up in the first channel (channel 0, 1903), i.e. all samples named “a” in the original sample block 1901 end up in channel 0 (1903). Likewise, all samples named “b” end up in channel 1 (1904), all samples named “c” end up in channel 2 (1905), and all samples named “d” end up in channel 3 (1906). According to some embodiments, pixel shuffle and pixel unshuffle can be used to perform filtering.

[0077] In the one-dimensional case, pixel shuffle works as shown in Figure 20 An example is provided using the one-dimensional case; however, the same principles apply to the two-dimensional case as well. Here, an input block of one channel (2001) that undergoes pixel unshuffle results in an output tensor (2002) containing two channels: channel 0 (2003) will hold samples from block 2001 with even x coordinates (“a” samples), and channel 1 (2004) will hold samples with odd x coordinates (“b” samples). According to certain embodiments, a position-dependent NN filter can be utilized to process the 1x8 sample block 2001 from Figure 20 This is shown in the example of Figure 21 Here, when the convolution is centered on an even sample (named “a”), a weight tensor is used to process the 1x8 sample block, and when the convolution is centered on an odd sample (named “b”), another weight tensor is used. When the convolution is outside the block, e.g. when centered on sample a0, the convolution is padded with zeros in this example. This results in: Adding a bias term and applying PReLU, the result is: .

[0078] According to embodiments, and further with reference to Figure 20, pixel unshuffling can be applied to the 1 -channel 1 x8 input tensor (2001 ), which generates a two-channel output tensor (2002) with each channel (2003, 2004) of size 1 x4. Next, a non-position dependent convolutional layer is used that takes 2 input channels and outputs 2 channels. For output channel 0, the kernel tensor is used, where the matrix for input channel 0 is , and where the matrix for input channel 1 is . Figure 22 It is shown how the convolution is performed for the first position of output channel 0. When centered at the first sample position (2201 ), the padded zeros and the first two samples from input channel 0 will be multiplied with the matrix . Likewise, for input channel 1, will be multiplied with , and the two dot products will be added together to form the first element c0 in output channel 0. Thus, the result will be: where denotes the regular multiplication of two scalars. This equals: which equals . Figure 23 The next position of output channel 0 is shown, where: In the same way, and .

[0079] Similarly, Figure 24 the case for output channel 1 is shown. Here, the weight tensor is used, where the matrix for input channel 0 is and where the matrix for input channel 1 is . This gives: In the same way, one obtains , and . After adding a bias term and applying PReLU, pixel unshuffling is used to return from two channels to one channel as shown in Figure 25 . The result (2501 ) is: which (with applied bias and PReLU) equals: .

[0080] In this regard, and in accordance with some embodiments, filters are applied by using pixel de-shuffling, regular CNN layers, and pixel shuffling. In this implementation with pixel shuffling, instead of using a single dot product to produce each value, as when using position-dependent filtering according to other embodiments, two dot products are used, one per input channel, where In this example, because half of the weights are zero, they do not contribute to the final result. Thus, in the one-dimensional case, the number of KMACs per sample is doubled. In the 2D case, this translates to an increase of 4 kMACs per sample. As an example, if a single-channel input of size 8x8 and a single-channel output of the same size is used, along with a position-dependent convolution with a 3x3 kernel, and the repeat block size is 2x2, then the position-dependent convolution will simply be 3x3=9 MACs per sample. To implement the filter using the pixel shuffling / de-shuffling method as described above, the 8x8 input is first de-shuffled into four channels of 4x4. Then a conv2 with a kernel of 3x3 and 4 channels is needed to obtain four output channels. The final pixel shuffling returns the 8x8 input. This is 3x3x4x8x8=2304 MACs per sample or 2304 / (8x8)=36 MACs. The number of parameters can also be four times that of other embodiments and 16 times that of regular (non-position-dependent) CNNs.

[0081] Aspects of one or more embodiments described herein can be combined. For example, according to embodiments, a network using pixel shuffling can be combined with a network using position-dependent filtering. This is reflected in Figure 10 Convolution with stride=2 is used to reduce the resolution, and finally pixel shuffling is used to go back to the original resolution, but since the backbone is in between, position-dependent convolution is used on top of the pixel shuffling. However, these results can be further improved by completely removing the pixel shuffling and replacing it with position-dependent filtering only.

[0082] In some embodiments, not all channels of a convolution are position-dependent. As an example, if there are four output channels, the first two can use position-dependent convolution, while the last two can use ordinary convolution.

[0083] ​In another embodiment, the choice of which filter to use is not predetermined by location, but rather the network selects the weight tensors. For example, the network may select one of four potential weight tensors based on a signal. In this embodiment, the signal may come from the neural network itself or may be an input to the network. As an example, in one embodiment, the image gradient is computed outside the neural network and forwarded as input to the network. If the gradient exceeds a certain value in both the x and y directions, a first weight tensor is used. If the gradient exceeds a threshold only in the x direction, a second weight tensor is used. If the gradient exceeds a threshold only in the y direction, a third weight tensor is used, and if the gradient does not exceed a threshold in any direction, a fourth weight tensor is used. While gradients are used in this example, other properties, including other image-related properties, can be the basis for selecting a particular weight tensor.

[0084] In some embodiments, the neural network is given input that enables it to distinguish different relative positions. For example, the input may identify positions within a repeater block. Figure 28 An example is shown below. In this example, four additional inputs are concatenated with the input of each convolutional layer. These inputs are "1" if the sample location is at a specific position within the repeater block. As an example, the signal "is_top_left" is "1" if the sample coordinates are even in both the x and y dimensions, as shown below. Figure 29 As shown. Similarly, if the sample position is odd in both the x and y directions, the signal "is_bottom_right" is 1, otherwise it is zero. This allows the convolution to distinguish between sample positions. In an alternative embodiment, (1-is_top_left), (1-is_top_right), (1-is_bottom_right), and (1-is_bottom_right) are alternatively fed as inputs into the convolution. As illustrated, including such inputs means the neural network can learn a mapping similar to using a specific set of weights when "is_top_left" is 1 and another set of weights when "is_bottom_right" is 1. While four inputs are used as an example, different numbers (e.g., based on the repeater block size) can be used depending on the embodiment. Figure 30 Another example of how to input location information into the network, including x and y coordinates, is provided. In this case, only two inputs, "give_y_coord" and "give_x_coord", are needed instead of four inputs, "is_top_left", "is_top_right", "is_bottom_left", and "is_bottom_right".

[0085] In some embodiments, if such input is fed to a network, it can not be strictly required to have different weight matrices at different positions. Instead, the network can learn this behavior even if regular convolutional layers are used. This can be illustrated, for example, using the one-dimensional correlation filter example described in conjunction Figure 21 with the exception that RELU (Rectified Linear Unit) has been used instead of PRELU (Parametric Rectified Linear Unit). In this example, three channels are used as input: channel 0 is the sample input [a0 b1 a2 b3 a4 b5 a6 b7] (as in 2001 of Figure 21 with the exception that RELU (Rectified Linear Unit) has been used instead of PRELU (Parametric Rectified Linear Unit). In this example, three channels are used as input: channel 0 is the sample input [a0 b1 a2 b3 a4 b5 a6 b7] (as in 2001 of This means that the output after ReLU will be: If is set large enough, the expression containing in the ReLU function will become negative and the ReLU output will be zero. Thus, s0 = s2 = s4 = s6 = 0. However, s1 will be the same as the value u1 discussed in conjunction Figure 21 with the exception that RELU (Rectified Linear Unit) has been used instead of PRELU (Parametric Rectified Linear Unit). In this example, three channels are used as input: channel 0 is the sample input [a0 b1 a2 b3 a4 b5 a6 b7] (as in 2001 of This means that, if is large enough, the output will be equal to: The subsequent layer can add these two channels together with equal weights and obtain a value similar to the output u discussed in conjunction Figure 21 with the exception that RELU (Rectified Linear Unit) has been used instead of PRELU (Parametric Rectified Linear Unit). In this example, three channels are used as input: channel 0 is the sample input [a0 b1 a2 b3 a4 b5 a6 b7] (as in 2001 of Thus, at least in the case where RELU is used instead of PRELU, a filter with regular convolution can be provided in accordance with some embodiments.

[0086] Referring now to Figure 31A , a process 3100 is provided in accordance with some embodiments. This process can be performed, for example, in the device 3300, including in a decoder or encoder. In some embodiments, the process 3100 is applied in a convolutional layer of a neural network filter (e.g., within a head, stem / body, or tail of a NN feature). This can include, for example, as a combination Figure 2 ,Figure 3 and Figure 3 Part of the process described. 4A-4C. Process 3100 can begin at step s3102, where a first group within a set of inputs is identified, where the first group includes a plurality of values. The set of inputs can be divided into relay blocks. In some embodiments, the set of inputs is a block, frame, or group of samples or other values. The inputs can correspond to (e.g., from) one or more samples of a video picture. However, in some embodiments, process 3100 is not limited to video processing. In step s3104, a first weight tensor is selected based on a position of a first value of the group. For example, the weight tensor can be selected based on the value’s position in the relay block (e.g., whether it is near an edge, corner, center, etc.) or based on the value’s absolute position (e.g., row-column position in the set of inputs). In some embodiments, the weight tensor can be selected at least in part according to a received position signal. In step s3106, a first weighted sum is generated by applying the position-dependent weight tensor to the first plurality of values. In step s3108 (which can be optional in some embodiments), an output value is generated from the weighted sum. This can include, for example, processing the first weighted sum with a non-linear activation function such as a parametric rectified linear unit.

[0087] Referring now to FIG. 1 and Figure 2 Referring now to FIG. 1 and Figure 31B and 31C Additional processing can be performed on the set of inputs according to processes 3120 and 3130. For example, additional groups can be identified, additional weight tensors selected, and additional weighted sums generated. These weighted sums can be used to generate additional outputs.

[0088] Process 3120 includes identifying (s3122) a second group within the set of inputs, where the second group includes a plurality of values; selecting (s3124) a second weight tensor based on a position; generating (s3126) a second weighted sum by applying the position-dependent weight tensor to the second plurality of values; and optionally, generating (s3128) a second output value from the weighted sum. According to embodiments, the first weight tensor is different from the second weight tensor. For example, process 3120 can be performed with one or more steps of process 3100.

[0089] Process 3130 includes identifying (s3132) a third group within the set of inputs, where the third group includes a plurality of values; selecting (s3134) a third weight tensor based on a position; generating (s3136) a third weighted sum by applying the position-dependent weight tensor to the third plurality of values; and generating (s3138) a third output value from the weighted sum. Process 3130 can be performed, for example, with one or more steps of process 3100, and process 3130 can also be performed with one or more steps of process 3120. According to some embodiments, the first weight tensor and the third weight tensor are the same when the position of the first value and the position of the third value are the same (e.g., they are in the same relative position in the relay block, even if they are at different absolute positions in the set of inputs). Additionally, in embodiments, the first group and the third group can partially overlap or not overlap.

[0090] Reference is now made to Figure 32A Process 3200 is provided according to some embodiments. The process can be performed, for example, in device 3300, including in a decoder or encoder. In some embodiments, process 3200 is applied in a convolutional layer of a neural network filter (e.g., within a head, stem / body, or tail of a NN feature). This can include, for example, as part of the processing described in connection with other figures. Process 3200 can start at step s3202, in which an input block of values is obtained. In some embodiments, the input is a block, frame, or group of samples or other values. The values can correspond to (e.g., from) one or more samples of a video picture. However, in some embodiments, process 3200 is not limited to video processing. In step s3204, a pixel de-shuffling operation is performed on the input to generate an output tensor having a first channel and a second channel. In step s3206, a first weight tensor is applied to the first channel, and in step s3208, a second weight tensor is applied to the second channel. In step s3210, a de-shuffling operation is performed on the first and second channels to generate an output (e.g., an output block of values).

[0091] Reference is now made to Figure 32BAccording to some embodiments, process 3250 is provided. The process can be performed, for example, in device 3300, including in a decoder or encoder. In some embodiments, process 3250 is applied in a convolutional layer of a neural network filter (e.g., within a head, stem / body, or tail of NN features). This can include, for example, as part of the processes described in connection with other figures. Process 3200 can begin at step s3252, where a set comprising a plurality of values is identified. In some embodiments, the set is a block, frame, or group of samples or other values. The values can correspond to (e.g., from) one or more samples of a video picture. However, in some embodiments, process 3250 is not limited to video processing. At step s3254, a weight tensor is selected based on a signal. At step s3256, a weighted sum is generated by applying the signal-dependent weight tensor to the plurality of values. At step s3258, an output is generated from the weighted sum.

[0092] Referring now to Figure 32C According to some embodiments, process 3270 is provided. The process can be performed, for example, in device 3300, including in a decoder or encoder. In some embodiments, process 3250 is applied in a convolutional layer of a neural network filter (e.g., within a head, stem / body, or tail of NN features). This can include, for example, as part of the processes described in connection with other figures. Process 3200 can include obtaining (s3272) a set of inputs comprising a plurality of values, and processing (s3274) the set of inputs using a first weight tensor and a second weight tensor. According to embodiments, values corresponding to a first position are processed with the first weight tensor, values corresponding to a second position are processed with the first weight tensor, and values corresponding to a third position are processed with the second weight tensor. In certain aspects, process 3270 can have two features. In a first aspect, at least one position within a repeater block (or other similar grouping) uses a different weight matrix than another position. In a second aspect, the repeater block is repeated at least once.

[0093] Figure 33 is a block diagram of a device 3300 for implementing an encoder 112, a decoder 114, or a component included in an encoder 112 or a decoder 114 (e.g., a NN filter 280 or 330) according to some embodiments. When device 3300 implements a decoder, device 3300 can be referred to as a “decoding device 3300,” and when device 3300 implements an encoder, device 3300 can be referred to as an “encoding device 3300.” As Figure 33As shown in FIG. 33, the apparatus 3300 can include processing circuitry (PC) 3302, which can comprise one or more processors (P) 3355 (e.g., a general- purpose microprocessor, and / or one or more other processors, such as an application specific integrated circuit (ASIC), field programmable gate array (FPGA), and / or the like), which processors can be collocated in a single housing or in a single data center, or can be geographically distributed (i.e., the apparatus 3300 can be a distributed computing apparatus); at least one network interface 3348, including a transmitter (Tx) 3345 and a receiver (Rx) 3347, for enabling the apparatus 3300 to transmit data to, and to receive data from, other nodes connected to a network 110 (e.g., an Internet Protocol (IP) network), the network interface 3348 being connected (directly or indirectly) to the network (e.g., the network interface 3348 can be wirelessly connected to the network 110, in which case the network interface 3348 is connected to an antenna arrangement); and a storage unit (a.k.a., "data storage system") 3308, which can comprise one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments in which the PC 3302 comprises a programmable processor, a computer program product (CPP) 3341 can be provided. The CPP 3341 comprises a computer readable medium (CRM) 3342 storing a computer program (CP) 3343 comprising computer readable instructions (CRI) 3344. The CRM 3342 can be a non-transitory computer readable medium, such as magnetic

[0094] Summary of embodiments A1. A method comprising: identifying a first group within a set of inputs, wherein the first group comprises a first plurality of values; selecting a first weight tensor, wherein the first weight tensor is selected based at least in part on a position of a first value; and generating a first weighted sum by applying the first weight tensor to the first plurality of values.

[0095] A2. The method of A1, further comprising: identifying a second group within the set of inputs, wherein the second group comprises a second plurality of values; selecting a second weight tensor, wherein the second weight tensor is selected based at least in part on the locations of the second values; and generating a second weighted sum by applying the second weight tensor to the second plurality of values.

[0096] A3. The method of A2, wherein the first weight tensor and the second weight tensor are different.

[0097] A4. The method of A2 or A3, wherein the first group and the second group partially overlap in the set of inputs.

[0098] A5. The method of any of Al-A4, further comprising: identifying a third group within the set of inputs, wherein the third group comprises a third plurality of values; selecting a third weight tensor, wherein the third weight tensor is selected based at least in part on the locations of the third values; and generating a third weighted sum by applying the third weight tensor to the third plurality of values.

[0099] A6. The method of A5, wherein the first weight tensor and the third weight tensor are the same, and wherein the locations of the first values and the locations of the third values are the same (e.g., in the same locations in the repeater block).

[0100] A7. The method of A5 or A6, wherein the first weight tensor and the third weight tensor are the same, and wherein the locations of the first values in the set of inputs and the locations of the third values in the set of inputs are different.

[0101] A8. The method of any of A5-A7, wherein the first group and the third group do not overlap in the set of inputs.

[0102] A9. The method of any of A5-A7, wherein the first group and the third group partially overlap in the set of inputs.

[0103] A10. The method of any of Al-A9, further comprising performing one or more of: generating a first output value (e.g., by processing the first weighted sum with a non-linear activation function, such as a parametric rectified linear unit); and / or generating a second output value (e.g., by processing the second weighted sum with a non-linear activation function, such as a parametric rectified linear unit); and / or A third output value is generated (e.g., by processing the third weighted sum with a non-linear activation function, such as a parametric rectified linear unit).

[0104] Al 1. The method of any one of Al-AlO, wherein the set of inputs is partitioned into n x n (e.g., wherein n = 2, 3, 4, or 8) or n x m (e.g., wherein n≠m ) relay block.

[0105] A12. The method of Al l, wherein the first weight tensor is selected based on the first value having one of the following positions in the relay block of the set of inputs: (i) a center position in the relay block; (ii) a lower-left position in the relay block; (iii) a lower-right position in the relay block; (iv) an upper-middle position in the relay block; (v) a lower-middle position in the relay block; (vi) a left-middle position in the relay block; (vii) a right-middle position in the relay block; (viii) a left-most position in the relay block; (ix) a right-most position in the relay block; (x) an upper-most position in the relay block; (xi) a lower-most position in the relay block; (xii) an interior position in the relay block; (xiii) an upper-left position in the relay block; or (xiv) an upper-right position in the relay block.

[0106] A13. The method of any one of Al-Al2, wherein the first weight tensor is selected based on a distance of the first value to a boundary of the relay block (e.g., whether the first value is located at an edge of the relay block, at a corner of the relay block, or away from an edge of the relay block).

[0107] A14. The method of any one of Al-Al3, wherein the first value is a center-located value within the first group.

[0108] A15. The method of any of A1-A14, wherein the second value is a center- positioned value within the second group, and the third value is a center-positioned value within the third group.

[0109] A16. The method of any of A1-A15, wherein the method is applied in a convolutional layer of a neural network filter.

[0110] A17. The method of any of A1-A16, wherein the method is applied to a head, stem / body, and / or tail of a neural network feature.

[0111] A18. The method of any of A1-A17, wherein the method is applied only to a subset of convolutional layers of a given neural network filter or neural network portion (e.g., head, stem / body, or tail), or only to a subset of channels.

[0112] A19. The method of any of A1-A18, wherein the set of inputs are values corresponding to (e.g., derived from) a video picture.

[0113] A20. The method of any of A1-A18, wherein the set of inputs are values corresponding to (e.g., derived from) a one-dimensional source signal (e.g., a sound signal or a text signal).

[0114] A21. The method of any of A1-A20, wherein the method is performed as (or as part of) a filtering step in an encoding or decoding process.

[0115] A22. The method of any of A1-A21, wherein one or more weight values within the weight tensor are defined based on a position of a value to which the weight tensor is applied.

[0116] A23. The method of any of A1-A21, wherein positions within the set of inputs are defined in a checkerboard arrangement.

[0117] A24. The method of any of A1-A23, further comprising: performing a shuffle and / or unshuffle operation on one or more values (e.g., values of the groups or output values).

[0118] A25. The method of any of A1-24, wherein the weight tensor is selected based at least in part on an input signal indicative of a position of a value.

[0119] A26. The method of A25, wherein the input signal is concatenated with an input to a convolutional layer in which the method is performed.

[0120] A27. The method of any of A1-A26, wherein the first weight tensor is selected based at least in part on a position of the first value within the set of inputs.

[0121] B1. A method comprising: obtaining an input block of values; performing a pixel de-shuffling operation on the input block of values to generate an output tensor, wherein the output tensor comprises at least a first channel and a second channel; applying a first weight tensor to the first channel; applying a second weight tensor to the second channel; and performing a pixel shuffling operation on the first channel and the second channel to obtain an output block of values, wherein the first weight tensor and the second weight tensor are different.

[0122] B2. The method of A1, wherein each of the first weight tensor and the second weight tensor comprises one or more zero padding values.

[0123] C1. A method comprising: identifying a group within a set of inputs, wherein the group comprises a plurality of values; selecting a weight tensor, wherein the weight tensor is selected based at least in part on a signal; and generating a weighted sum by applying the weight tensor to the plurality of values.

[0124] C2. The method of C1, further comprising: receiving the signal (e.g., as input to the network).

[0125] C3. The method of C1 or C2, wherein the signal comprises at least one value.

[0126] C4. The method of C3, wherein selecting the weight tensor comprises: comparing the value to a threshold.

[0127] C5. The method of C3 or C4, wherein: the value exceeds the threshold in both an x-direction and a y-direction, and a first weight tensor is selected; the value exceeds the threshold only in the x-direction, and a second weight tensor is selected; the value exceeds the threshold only in the y-direction, and a third weight tensor is selected; and / or the value does not exceed the threshold in any direction, and a fourth weight tensor is selected.

[0128] C6. The method of any of C1-C5, further comprising: The output is generated (e.g., by processing the weighted sum with a non-linear activation function, such as a parametric rectified linear unit).

[0129] C7. The method of any of CI -C6, wherein the signal indicates an image property (e.g., a gradient).

[0130] D1. A method comprising: obtaining a set of inputs comprising a plurality of values; and processing the set of inputs using a first weight tensor and a second weight tensor, wherein: (i) values corresponding to a first position are processed with the first weight tensor, (ii) values corresponding to a second position are processed with the first weight tensor, and (iii) values corresponding to a third position are processed with the second weight tensor.

[0131] D2. The method of Dl, wherein the positions are positions within a relay block of the set of inputs.

[0132] D3. The method of D2, wherein the first position and the third position are positions within the same relay block.

[0133] D4. The method of D2, wherein the first position and the third position are positions within different relay blocks.

[0134] D5. The method of any of D2-D4, wherein the first position and the second position are equivalent positions within two different (i.e., non-repeating) relay blocks of the set of inputs.

[0135] D6. The method of any of D2-D4, wherein the first position and the second position are within the same relay block.

[0136] D7. The method of Dl, wherein the positions are positions within the set of inputs.

[0137] E1. A device configured to perform the method of any of embodiments Al-A27, Bl- B2, CI-C7, and Dl-D7.

[0138] E2. The device of El, wherein the device is an encoder or a decoder.

[0139] E3. The device of El or E2, comprising: a memory; and processing circuitry coupled to the memory, wherein the device is configured to perform the method of any of embodiments A1-A27, B1-B2, C1-C7, and D1-D7.

[0140] F1. A computer program comprising instructions which, when executed by processing circuitry, cause the processing circuitry to perform the method of any of embodiments A1-A27, B1-B2, C1-C7, and D1-D7.

[0141] F2. A carrier containing the computer program of F1, wherein the carrier is one of an electronic signal, optical signal, radio signal, and computer readable storage medium.

[0142] G1. A device configured to: identify a first group within a set of inputs, wherein the first group comprises a first plurality of values; select a first weight tensor, wherein the first weight tensor is selected based at least in part on a location of a first value; and generate a first weighted sum by applying the first weight tensor to the first plurality of values.

[0143] G2. The device of embodiment G1, wherein the device is further configured to perform the method of any of embodiments A2-A27.

[0144] G3. The device of G1 or G2, wherein the device is configured to generate an encoded or decoded video.

[0145] H1. A device configured to: obtain an input block of values; perform a pixel de-shuffling operation on the input block of values to generate an output tensor, wherein the output tensor comprises at least a first channel and a second channel; apply a first weight tensor to the first channel; apply a second weight tensor to the second channel; and perform a pixel shuffling operation on the first channel and the second channel to obtain an output block of values, wherein the first weight tensor and the second weight tensor are different.

[0146] H2. The device of embodiment H1, wherein the device is further configured to perform the method of embodiment B2.

[0147] H3. The device of H1 or H2, wherein the device is configured to generate an encoded or decoded video.

[0148] I1. A device configured to: identifying a group within a set of inputs, wherein the group comprises a plurality of values; selecting a weight tensor, wherein the weight tensor is selected based at least in part on a signal; and generating a weighted sum by applying the weight tensor to the plurality of values.

[0149] I2. The device of embodiment II, wherein the device is further configured to perform the method of any of embodiments C2-C7.

[0150] I3. The device of I I or I2, wherein the device is configured to generate an encoded or decoded video.

[0151] J1. A device configured to: obtain a set of inputs comprising a plurality of values; and process the set of inputs with a first weight tensor and a second weight tensor, wherein: (i) values corresponding to a first location are processed with the first weight tensor, (ii) values corresponding to a second location are processed with the first weight tensor, and (iii) values corresponding to a third location are processed with a second weight tensor.

[0152] J2. The device of embodiment Jl, wherein the device is further configured to perform the method of any of embodiments D2-D7.

[0153] J3. The device of Jl or J2, wherein the device is configured to generate an encoded or decoded video.

[0154] While various embodiments are described herein, it should be understood that they are presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above described exemplary embodiments. Furthermore, to the extent that any of the above described elements are presented in a particular combination, it should be understood that this combination is merely exemplary and that other combinations are also possible.

[0155] In addition, while the processes described above and illustrated in the drawings are shown as a series of steps, it will be appreciated that some of the steps could be performed in an order other than the order shown, and that some of the steps could be performed in parallel or with no time in between.

Claims

1. A method comprising: identifying a first group within a set of inputs, wherein the first group comprises a first plurality of values; selecting a first weight tensor, wherein the first weight tensor is selected based at least in part on a location of a first value; generating a first weighted sum by applying the first weight tensor to the first plurality of values; identifying a second group within the set of inputs, wherein the second group comprises a second plurality of values; selecting a second weight tensor, wherein the second weight tensor is selected based at least in part on a location of a second value, wherein the second weight tensor is different from the first weight tensor; generating a second weighted sum by applying the second weight tensor to the second plurality of values; identifying a third group within the set of inputs, wherein the third group comprises a third plurality of values; selecting a third weight tensor, wherein the third weight tensor is selected based at least in part on a location of a third value; and generating a third weighted sum by applying the third weight tensor to the third plurality of values; wherein the first weight tensor and the third weight tensor are the same, further comprising performing one or more of: generating a first output value by processing the first weighted sum with a non-linear activation function; generating a second output value by processing the second weighted sum with a non-linear activation function; generating a third output value by processing the third weighted sum with a non-linear activation function.

2. The method of claim 1, wherein, The set of inputs is divided into n x n or n x m relay blocks.

3. The method of claim 2, wherein, The first weight tensor is selected based on the first value having one of the following locations in a relay block of the set of inputs: top-left location, top-right location, bottom-left location, and bottom-right location.

4. The method of any one of claims 1-3, wherein, The location of the first value in the relay block and the location of the third value in the relay block are the same.

5. The method of any one of claims 1-4, wherein, The location of the first value in the relay block and the location of the second value in the relay block are different.

6. The method of any one of claims 1-5, wherein, The first weight tensor is selected based on a distance of the first value to a relay block boundary.

7. The method of any one of claims 1-6, wherein, The first value is a centrally located value within the first group.

8. The method of any one of claims 1-7, wherein, The second value is a centrally located value within the second group and the third value is a centrally located value within the third group.

9. The method of any one of claims 1-8, wherein, The method is applied in a convolutional layer of a neural network filter.

10. The method of any one of claims 1-9, wherein, The method is applied to a head, stem / body, and / or tail of a neural network feature.

11. The method of any one of claims 1-10, wherein, The method is applied only to a subset of convolutional layers of a given neural network filter or neural network portion, or only to a subset of channels.

12. The method of any one of claims 1-11, wherein, The set of inputs are values corresponding to a video picture.

13. The method of any one of claims 1-12, wherein, The set of inputs are values corresponding to a one-dimensional source signal.

14. The method of any one of claims 1-13, wherein, The method is performed as a filtering step in an encoding or decoding process or as part of a filtering step.

15. The method of any one of claims 1-14, wherein, One or more weight values within a weight tensor are defined based on the location of the values to which the weight tensor is applied.

16. The method of any one of claims 1-14, wherein, Locations within the set of inputs are defined in a checkerboard arrangement.

17. The method of any of claims 1-16, further comprising: performing a shuffle and / or unshuffle operation on one or more values.

18. The method of any one of claims 1-17, wherein, selecting a weight tensor based at least in part on an input signal indicative of the location of a value.

19. The method of claim 18, wherein, The input signal is concatenated with the input of a convolutional layer performing the method.

20. The method of any one of claims 1-19, wherein, The first weight tensor is selected based at least in part on the position of the first value within the set of inputs.

21. A device configured to perform the method according to any one of claims 1-20.

22. The apparatus of claim 21, wherein, The device is an encoder or a decoder.

23. A computer program comprising instructions which, when executed by processing circuitry, cause the processing circuitry to perform the method according to any one of claims 1-20.

24. A carrier containing the computer program of claim 23, wherein the carrier is one of an electronic signal, an optical signal, a radio frequency signal or a computer readable storage medium. The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium. The carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.