Spatial Frequency Conversion-Based Image Correction Using Inter-Channel Correlation Information

By applying spatial frequency transformations and separate neural networks to primary and secondary image channels, the method enhances image quality and processing efficiency in image and video coding, addressing optimization challenges in existing technologies.

JP7717985B2Active Publication Date: 2025-08-04HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024540624
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-28
Filing Date
2022-07-18
Publication Date
2025-08-04
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

Existing image and video coding technologies, including hybrid codecs and neural network-based approaches, can benefit from improved encoding and decoding methods that optimize transformation, quantization, and entropy coding, particularly in handling feature maps across different devices or cloud environments.

Method used

A method involving spatial frequency transformation of primary and secondary image channels, processed by separate neural networks, followed by inverse transformations to enhance image correction, allowing independent optimization of neural networks and reducing processing time and load.

Benefits of technology

This approach improves image quality by reducing distortion and processing time, facilitating flexible channel selection and adaptation, and enabling efficient encoding and decoding of images and videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717985000037
    Figure 0007717985000037
  • Figure 0007717985000038
    Figure 0007717985000038
  • Figure 0007717985000039
    Figure 0007717985000039
Patent Text Reader

Abstract

The present disclosure relates to image modification, such as image correction, where the processing is based at least in part on a neural network. In particular, the image modification includes multi-channel processing where a primary channel is processed separately and a secondary channel is processed based on the processed primary channel. The primary channel is processed based on a first spatial frequency transformation to obtain a transformed primary channel, and the secondary channel is processed based on a second spatial frequency transformation to obtain a transformed secondary channel. The transformed primary channel is processed by a first neural network to obtain a modified transformed primary channel, and the transformed secondary channel is processed by a second neural network based on the transformed primary channel to obtain a modified transformed secondary channel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of databaseized encoding and decoding on neural network architectures. In particular, some embodiments relate to methods and apparatuses for such encoding and decoding of images and / or videos from bitstreams, particularly for image correction, using multiple processing layers.

Background Art

[0002] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, signals are typically encoded block by block by predicting the block and further encoding only the difference between the original block and its prediction. In particular, such encoding may include transformation, quantization, and bitstream generation and usually includes some form of entropy coding. Typically, the three components of a hybrid coding method, namely, transformation, quantization, and entropy coding, are optimized separately. Recent video compression standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use the transformed representation to code the residual signal after prediction.

[0003] In recent years, neural network architectures have been applied to image and / or video coding. Generally, these neural network (NN)-based approaches can be applied to image and video coding in various different ways. For example, some end-to-end optimized image or video coding frameworks have been discussed. Further, deep learning has been used to determine or optimize some parts of an end-to-end coding framework, such as the selection or compression of prediction parameters. Moreover, some neural network-based approaches have also been discussed for use in hybrid image and video coding frameworks, for example, for implementation as a trained deep learning model for intra or inter prediction in image or video coding.

[0004] The end-to-end optimized image or video coding applications discussed above have in common that they generate some feature map data to be transmitted between an encoder and a decoder.

[0005] A neural network is a machine learning model that uses one or more layers of non-linear units that can predict an output based on the received input. Some neural networks include one or more hidden layers in addition to an output layer. As an output of each hidden layer, a corresponding feature map can be provided. Such corresponding feature maps of each hidden layer can be used as an input to subsequent layers in the network, i.e., subsequent hidden layers or the output layer. Each layer of the network generates an output from the received input according to the current values of its respective parameter set. In a neural network split between different devices, e.g., between an encoder and a decoder, or between a device and the cloud, the feature maps at the output of the split location (e.g., the first device) are compressed and transmitted to the remaining layers of the neural network (e.g., to the second device).

[0006] Further improvements in encoding and decoding using a trained network architecture may be desirable. SUMMARY OF THE INVENTION

[0007] The present invention relates to a method and apparatus for modifying, e.g., correcting, an image or video. MEANS FOR SOLVING THE PROBLEM

[0008] The above and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description, and the drawings.

[0009] Particular embodiments are outlined in the appended independent claims, and other embodiments are defined in the dependent claims.

[0010] In particular, embodiments of the present invention provide an approach for modifying an image based on a neural network system that processes multiple image channels. The primary channel is processed individually. One or more secondary channels are processed taking into account information from the primary channel. Before processing by the neural network system, the first channel and the secondary channels undergo a spatial frequency transformation. Before processing by the neural network system, which image channel is the primary channel can be selected.

[0011] According to a first aspect, the present disclosure relates to a method for modifying an image region represented by two or more image channels, the method comprising processing a primary channel among the two or more image channels based on a first spatial frequency transformation to obtain a transformed primary channel, and processing a secondary channel among the two or more image channels different from the primary channel based on a second spatial frequency transformation to obtain a transformed secondary channel.

[0012] Two or more image channels may include color channels and / or feature channels. The color channels and the feature channels reflect image characteristics. Each type of channel may provide information not present in other channels, and thus cooperative processing may improve those channels with respect to the primary channel. For example, two or more image channels may be YUV channels, and the primary channel may be the Y channel. The image region may be a patch of a predetermined size corresponding to a part of an image or parts of a plurality of images, or the image region may be an image or a plurality of images. The image may be a still image or a frame of a video sequence.

[0013] Furthermore, the method according to the first aspect includes processing the transformed primary channel by a first neural network to obtain a modified transformed primary channel, and processing the transformed secondary channel based on the transformed primary channel (used as auxiliary information) by a second neural network to obtain a modified transformed secondary channel. The first network and the second network may be different from each other and may operate independently of each other.

[0014] The method according to the first aspect further includes processing the modified transformed primary channel based on a first inverse spatial frequency transform to obtain a modified primary channel, and processing the modified transformed secondary channel based on a second inverse spatial frequency transform to obtain a modified secondary channel. A modified image region is obtained based on the modified primary channel and the modified secondary channel.

[0015] The first spatial frequency conversion and the second spatial frequency conversion provide information in the spatial frequency domain, without being limited thereto. It should be noted that the conversion by the spatial frequency conversion can be regarded as a kind of "preconditioning" in which the signal is converted into a more redundant format before being further processed. A more redundant signal is easier to process for a neural network.

[0016] To improve the quality of the modified image region on the primary channel, information regarding the primary channel is used as auxiliary information for processing the secondary channel. The auxiliary information is provided in the spatial frequency conversion region and is conveniently added to the spatial frequency conversion secondary channel before the input to the second neural network. The converted primary channel can be processed independently by the first neural network from the second neural network. Thus, the coefficients of one of the neural networks can be changed / optimized without affecting the output of the other network. Thereby, the overall conditioning / optimization of the neural network can be done quite rapidly. Further, when the first neural network and the second neural network are implemented as convolutional networks, different kernels may be used for them. In particular, according to the method of the first aspect, in order to obtain the modified channel, it is not necessary to process all channels by exactly the same network at a certain stage. Thereby, the overall processing time consumption can be reduced compared to the art.

[0017] According to one implementation form, each of the first neural network and the second neural network is a convolutional neural network (CNN) or includes a CNN. CNN has been proven to be superior to other networks, such as multiple perceptrons, in many image processing applications and is known for relatively robust and fast processing. Each of the convolutional neural networks may include at least one residual network component that enables residual learning with reduced memory requirements. One or more of the convolutional neural networks may use a scaling layer represented by one or more scaling values. Accordingly, the scaling layer may be adapted to signal one or more scaling values.

[0018] In one possible implementation form, one or both of the first spatial frequency conversion and the second spatial frequency conversion are selected from the group consisting of energy compression conversions including wavelet transform, discrete Fourier transform, fast Fourier transform, and discrete cosine transform. The spatial frequency conversion can be selected according to the actual application that may require different conversions to achieve the desired quality of the modified image region. The first spatial frequency conversion and the second spatial frequency conversion may be the same (one of wavelet transform, discrete Fourier transform, fast Fourier transform, energy compression conversion, and discrete cosine transform).

[0019] Depending on the actual application, and with regard to processing speed and memory requirements, the selection of wavelet transform may be appropriate in some cases. In this case, one or both of the first spatial frequency conversion and the second spatial frequency conversion may be a wavelet transform selected from the group consisting of discrete wavelet transform (DWT) and stationary wavelet transform. If DWT is to be used, Haar (for simplicity) or Daubechies (for accuracy) wavelets may be selected.

[0020] In another possible implementation of the method of the first aspect, the primary channel is selected (not fixedly predetermined in advance) from two or more image channels. By processing the image in terms of patches or multiple images, regions of the image or video sequence can be processed differently, and in particular, the selection of the primary channel can be appropriately changed. Since the content within the image and / or within the video sequence can change, it may be advantageous for the image correction / correction to adapt the primary channel accordingly.

[0021] According to another implementation, the secondary channel can also be selected from two or more image channels. By providing this additional selection option, the flexibility of the processing is improved.

[0022] When at least one secondary channel is selected, exactly one primary channel can be selected. If a flag in the encoded stream indicates that the other channels among the two or more channels should not be processed, it may not be necessary to label the selected channel as the primary channel or the secondary channel.

[0023] According to another implementation, the primary channel and the secondary channel can be selected from two or more image channels based on the output of a classifier operating based on another neural network. Using a classifier makes it possible to train or design such a classifier in order to appropriately select the image channel that will be the primary channel so that the quality of the image correction (such as image correction) can be improved.

[0024] In principle, the method according to the first aspect is suitable for processing primary and secondary channels of the same size and different sizes. When the primary channel and the secondary channel are of the same size, according to one implementation form, the processing of the converted secondary channel based on the converted primary channel includes the step of concatenating a second three-dimensional tensor representing the converted secondary channel with a first three-dimensional tensor representing the converted primary channel. The concatenation is performed along the first non-spatial dimension of the tensor. The spatial dimensions of the tensor are the height dimension and the width dimension of the image region. The first non-spatial dimension results from the spatial frequency conversion. For example, when a discrete wavelet transform is used for the spatial frequency conversion, the first non-spatial dimension of the tensor is given by the spatial low-frequency subband LL and the spatial high-frequency subbands HL (vertical features), LH (horizontal features), and HH (diagonal features).

[0025] Such concatenation for using auxiliary information in the image correction process is relatively fast and can be performed in a memory-efficient manner.

[0026] The size of the primary channel can be made larger than the size of the secondary channel (thus, having excellent resolution). In this case, according to one implementation form, the conversion primary channel is processed based on at least one additional first spatial frequency conversion in order to obtain an auxiliary conversion primary channel having the same size as the conversion secondary channel in the height direction and width direction of the image area (the first spatial frequency conversion and the additional first spatial frequency conversion form a cascade spatial frequency conversion). In this case, the processing of the conversion secondary channel is based on the auxiliary conversion primary channel. On the other hand, when the size of the secondary channel is larger than the size of the primary channel, according to one implementation form, the conversion secondary channel is processed based on at least one additional second spatial frequency conversion in order to obtain an auxiliary conversion secondary channel having the same size as the conversion primary channel in the height direction and width direction of the image area (the second spatial frequency conversion and the additional second spatial frequency conversion form a cascade spatial frequency conversion). In this case, the processing of the conversion secondary channel includes the processing of the auxiliary conversion secondary channel based on the conversion primary channel.

[0027] In either case, the cascade conversion enables the processing of channels of different sizes without significantly increasing the processor load and processing time. Further, the concatenation for using the auxiliary information in the image correction process can also be used for the processing of channels of different sizes by the cascade conversion. Thus, the processing of the conversion secondary channel based on the conversion primary channel by the method of the first aspect according to one implementation form includes the step of concatenating a second three-dimensional tensor representing the conversion secondary channel with a first three-dimensional tensor representing the auxiliary conversion primary channel when the size of the primary channel is larger than the size of the secondary channel, and, on the other hand, the step of concatenating a second three-dimensional tensor representing the auxiliary conversion secondary channel with a first three-dimensional tensor representing the conversion primary channel when the size of the secondary channel is larger than the size of the primary channel.

[0028] In many applications, the primary channel has a larger size (if any), but this is not always the case. For example, in the case of a combination of a low-resolution grayscale camera and a high-resolution noisy color camera, the low-resolution channel of the low-resolution grayscale camera can be selected as the primary channel having a smaller size than the secondary channel provided by the high-resolution noisy color camera accordingly.

[0029] It should be noted that generally, the overall processing can be facilitated by restricting the image region to the shape of a square region in the height and width dimensions of the image region. According to one embodiment, a step of dividing the image into image regions including the image region thereof, and a step of padding the image regions resulting from the division, which are not square in the height and width dimensions of the image region, so that they become square in the height and width dimensions of the image region are performed. Alternatively, a step of dividing the image into image regions including the image region thereof, and a step of padding the image so that the image is divided only into image regions that are all square in the height and width dimensions of the image region including the image region, if the image cannot be divided only into image regions that are square in the height and width dimensions of the image region including the image region are performed.

[0030] Furthermore, it should be noted that before receiving each spatial frequency conversion, the primary channel and the secondary channel may receive a pixel shift as described in the following detailed description. The pixel shift can further increase the processing efficiency. In this case, the corrected channel after the inverse obtained by the inverse spatial frequency conversion receives a pixel unshift.

[0031] In some exemplary implementations, the method further includes a step of selecting a minimum size of an image region based on the number of hidden layers of a neural network, where the minimum size is at least 2 * ((kernel_size - 1) / 2 * n_layers) + 1, kernel_size is the size of the kernel of the neural network which is a convolutional neural network, and n_layers is the number of layers of the neural network.

[0032] Such a lower bound on the patch size selection enables the full utilization of the information of the processed image without adding redundancy, such as by padding, depending on the design of the neural network.

[0033] According to one embodiment, (which can be combined with any of the foregoing or the following embodiments and examples), the method further includes a step of rearranging the pixels of each of at least two image channels of an image region into a plurality S of sub-regions, where each of the sub-regions of the image channels among the at least two image channels contains a subset of the samples of the image channel, and for all the image channels, the horizontal dimension of the sub-regions is the same and equal to an integer multiple mh of the greatest common divisor of the horizontal dimension of the image, and for all the image channels, the vertical dimension of the sub-regions is the same and equal to an integer multiple mv of the greatest common divisor of the vertical dimension of the image.

[0034] Such rearrangement enables the neural network to be used to process images with different dimensions / resolutions of their image channels.

[0035] In particular, the S sub-regions of the image region are relatively prime to each other with S = mh * mv, have a horizontal dimension dimh and a vertical dimension dimv, and the sub-regions contain samples of the image region at positions {kh * mh + offh, kv * mv + offv}, where kh ∈ [0, dimh - 1] and kv ∈ [0, dimv - 1], and each combination of offh and offv specifies a respective sub-region with offk ∈ [1, mh] and offv ∈ [1, mv].

[0036] By determining the patch size as described above, even when the resolution and / or dimensions (vertical and / or horizontal) of the channels are different from each other, it is possible to utilize the image and effectively adapt the patch size to the dimensions of the image for each channel.

[0037] As already described, the first neural network and the second neural network can be operated independently of each other. According to one implementation form, the weights (and activation functions) of one of the first neural network and the second neural network are determined and used independently of the weights of the other of the first neural network and the second neural network. An individual adaptation of one network does not affect the configuration of the other network.

[0038] According to a second aspect, there is provided a method for encoding an image or video sequence or an image, including the steps of obtaining an original image region, encoding the obtained image region into a bitstream, and applying a correction of the image region obtained by reconstructing the encoded image region as described above.

[0039] Using image correction in the coding of an image or video enables an improvement in the quality of the decoded image. This can be the quality in terms of reduced distortion. However, in some applications, there may be some special effects that can be desired, and the correction may lead to those improvements (not necessarily reducing the distortion with respect to the original picture).

[0040] For example, the encoding may include the step of including an indication of the selected primary channel in the bitstream. This enables a potentially better reconstruction on the decoder side, that is, better with respect to the distortion with respect to the original (undistorted) image.

[0041] According to one exemplary implementation, the method further includes the step of including in the bitstream an adaptation of one or more weights of at least one of the weights of the first neural network and the weights of the second neural network.

[0042] According to one exemplary implementation, the method includes the steps of obtaining a plurality of image regions, applying the method of modifying the obtained image regions individually to the image regions of the plurality of obtained image regions, and including in the bitstream of each of the plurality of image regions at least one of an indication that the method for modifying the obtained image regions should not be applied to the image regions, an adaptation of one or more weights of at least one of the first neural network and the second neural network, or an indication of a selected primary channel of the region. The region-based processing facilitates adaptation to images or video content.

[0043] When applying the method for modifying the obtained image regions, the selection of the primary channel and the secondary channel can be performed based on the reconstructed image regions without referring to the obtained image regions input to the encoding step. This avoids additional overhead (rate requirements). According to a third aspect, there is provided a method for decoding an image or video sequence or image from a bitstream, including the step of reconstructing the reconstruction from the bitstream and the step of applying the method for modifying the image regions as described above.

[0044] The application of image or video modification on the decoder side can improve the image quality of the decoded image.

[0045] A method for decoding an image or video sequence in some embodiments includes an instruction indicating that a method for modifying an acquired image region should not be applied to the image region, an instruction of a selected primary channel of the region, parsing a bitstream to obtain at least one of at least one adaptation of weights of at least one of a first neural network and a second neural network, reconstructing an image region from the bitstream, and modifying the reconstructed image region using the indicated primary channel as the selected primary channel when the instruction indicates the selected primary channel.

[0046] Reconstruction based on side information can provide better performance in terms of quality, as described above for the corresponding encoding method. The modification can be applied as an in-loop filter or as a post-processing filter in the encoder and / or decoder.

[0047] According to an exemplary implementation, the method further includes modifying the weights of each neural network accordingly when an adaptation of weights of at least one of the first neural network and the second neural network exists in the bitstream.

[0048] According to a fourth aspect, there is provided a computer program product including program code stored on a non-transitory medium, which, when executed on one or more processors, performs a method according to any one of the above aspects and implementations.

[0049] According to a fifth aspect, there is provided an apparatus for modifying an image region represented by two or more image channels, the apparatus including a circuit configured to perform steps by a method according to any one of the above-described aspects and implementation forms. The apparatus provides technical means for implementing the operations in the method defined according to the first aspect. This function may be implemented by hardware or by hardware that executes corresponding software.

[0050] According to a sixth aspect, there is provided an apparatus for modifying an image region represented by two or more image channels, the apparatus including: a first spatial frequency conversion unit configured to process a primary channel among the two or more image channels to obtain a converted primary channel; and a second spatial frequency conversion unit configured to process a secondary channel among the two or more image channels different from the primary channel to obtain a converted secondary channel. Further, the apparatus includes: a first neural network configured to process the converted primary channel to obtain a modified converted primary channel; and a second neural network configured to process the converted secondary channel based on the converted primary channel to obtain a modified converted secondary channel. Further, the apparatus includes: a first inverse spatial frequency conversion unit configured to process the modified converted primary channel to obtain a modified primary channel; a second inverse spatial frequency conversion unit configured to process the modified converted secondary channel to obtain a modified secondary channel; and a combining unit configured to obtain a modified image region based on the modified primary channel and the modified secondary channel.

[0051] Further features and implementation forms of the method according to the first aspect of the present disclosure correspond to respective possible features and implementation forms of the apparatus according to the sixth aspect of the present disclosure. The advantages of the apparatus according to the sixth aspect can be made the same as the corresponding implementation form advantages of the method according to the first aspect.

[0052] According to a seventh aspect, there is provided an encoder for encoding an image or a video sequence or an image, the encoder including: an input module for obtaining an original image region; a compression module for encoding the obtained image region into a bit stream; a reconstruction module for reconstructing the encoded image region; and one of the above-described apparatuses for modifying the reconstructed image region according to the fifth and sixth aspects.

[0053] According to an eighth aspect, there is provided a decoder for decoding an image or a video sequence or an image from a bit stream, the decoder including: a reconstruction module for reconstructing an image region from the bit stream; and an apparatus for modifying the reconstructed image region according to the fifth and sixth aspects.

[0054] According to a ninth aspect, the present disclosure relates to a video stream decoding apparatus including a processor and a memory. The memory stores instructions for causing the processor to perform the method according to the first aspect and its implementation forms.

[0055] According to a tenth aspect, the present disclosure relates to a video stream encoding apparatus including a processor and a memory. The memory stores instructions for causing the processor to perform the method according to the first aspect and its implementation forms.

[0056] According to an eleventh aspect, there is proposed a computer-readable storage medium storing instructions which, when executed, cause one or more processors to encode video data. The instructions cause the one or more processors to perform the method according to the first aspect or the second aspect, or any possible implementation form of the first aspect and its implementation forms.

[0057] The above-described apparatus may be embodied on an integrated chip.

[0058] Any of the above-described embodiments and exemplary implementations may be combined with each other if considered appropriate.

[0059] In the following, embodiments of the present invention will be described in more detail with reference to the accompanying drawings and figures.

Brief Description of the Drawings

[0060]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

DETAILED DESCRIPTION OF THE INVENTION

[0061] Similar reference numerals and names in different drawings may indicate similar elements.

[0062] In the following description, reference is made to the accompanying drawings, which form a part of the present disclosure and, by way of illustration, show specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and may include structural or logical changes not shown in the figures. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the present disclosure is defined by the appended claims.

[0063] For example, it is understood that the disclosure related to the described method may also apply to the corresponding device or system configured to perform the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units for performing the one or more described method steps, such as functional units (e.g., one unit for performing one or more steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, if a specific device is described based on one or more units, such as functional units, the corresponding method may include one step for performing the functions of the one or more units (e.g., one step for performing the functions of one or more units, or multiple steps each performing one or more of the functions of the multiple units), even if such one or more steps are not explicitly described or illustrated in the drawings. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.

[0064] The following provides an overview of some of the technical terms used and the frameworks in which embodiments of the present disclosure may be used.

[0065] Artificial neural network An artificial neural network (ANN) or connectionist system is a computing system inspired vaguely by the biological neural circuits that make up animal brains. Such systems generally "learn" to perform tasks by considering examples, without being programmed with task-specific rules. For example, in image recognition, an ANN can learn to identify images containing cats by analyzing exemplary images manually labeled as "cat" or "not a cat" and using the results to identify cats in other images. The ANN can do this without any prior knowledge of cats, such as that cats have fur, tails, whiskers, and cat-like faces. Instead, the ANN automatically generates discriminative features from the examples it processes.

[0066] An ANN is based on a collection of connected units or nodes called artificial neurons that roughly model neurons in a biological brain. Each connection can transmit a signal to other neurons, similar to a synapse in a biological brain. An artificial neuron that receives a signal can then process the signal and signal neurons connected to that neuron.

[0067] In an ANN implementation, the "signal" on a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. This connection is called an edge. Neurons and edges typically each have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal on a connection. A neuron can have a threshold such that a signal is transmitted only if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers can perform different transformations on their inputs. Signals proceed from the first layer (input layer) to the last layer (output layer), sometimes traversing multiple layers.

[0068] The original goal of the ANN approach was to solve problems in the same way as the human brain. Over time, attention shifted to performing specific tasks, leading to a deviation from biology. ANNs have been used for a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, board games and video games, medical diagnosis, and even in activities that were previously considered to be exclusive to humans, such as drawing pictures.

[0069] The name "Convolutional Neural Network" (CNN) indicates that this network uses a mathematical operation called convolution. Convolution is a special type of linear operation. A convolutional network is a neural network that uses convolution in at least one of its layers instead of the general matrix multiplication.

[0070] FIG. 1 schematically illustrates a general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input layer, an output layer, and a plurality of hidden layers. The input layer is a layer to which an input (such as a portion 11 of an input image as shown in FIG. 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved with multiplication or other dot products. The output of a layer is one or more feature maps (illustrated by empty solid rectangles), which may also be called channels. There may be resampling (such as subsampling) involved in some or all of the operations of a layer. As a result, the feature maps can become smaller, as illustrated in FIG. 1. Note that convolution with a stride can also reduce the size (resampling) of the input feature map. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, normalization layers, etc., and these inputs and outputs are masked by the activation function and the final convolution, so they are called hidden layers. The layer is colloquially called a convolutional layer, but this is just by convention. Mathematically, it is technically a sliding dot product or cross-correlation. This is important for the indices within the matrix in terms of how the weights are determined at a particular index point.

[0071] When programming a CNN for processing images, as shown in FIG. 1, the input is a tensor having dimensions (number of images)×(image width)×(image height)×(image depth). It should be understood that the image depth can be composed of the channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map having dimensions (number of images)×(feature map width)×(feature map height)×(feature map channels). The convolutional layers within a neural network should have the following attributes. A convolutional kernel (hyperparameter) defined by width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) must be equal to the number of channels (depth) of the input feature map.

[0072] In the past, traditional multi-layer perceptron (MLP) models have been used for image recognition. However, due to the full connectivity between nodes, MLP models have suffered from high dimensionality and have not been able to handle high-resolution images well. A 1000×1000 pixel image with RGB color channels has 3 million weights, which is too high to be efficiently and conveniently processed in correspondence with full connectivity. Also, such a network architecture does not take into account the spatial structure of the data and treats input pixels that are far apart from each other in the same way as pixels that are close to each other. This ignores the locality of reference in image data, both computationally and semantically. Therefore, for purposes such as image recognition where spatially local input patterns occupy a large number, the full connectivity of neurons is wasted.

[0073] Convolutional neural networks are a biologically inspired variation of multi-layer perceptrons specifically designed to emulate the behavior of the visual cortex. These models reduce the problems posed by the MLP architecture by exploiting the strong spatially local correlations present in natural images. The convolutional layer is the core building block of the CNN. The parameters of the layer consist of a set of learnable filters (the above-mentioned kernels), which have small receptive fields but extend across the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume to compute the dot product between the entries of the filter and the input, generating a two-dimensional activation map for that filter. As a result, the network learns filters that activate when detecting some specific type of feature at a certain spatial position within the input.

[0074] Stacking the activation maps of all filters along the depth dimension forms the full output volume of the convolutional layer. Thus, all entries within the output volume can also be interpreted as the output of neurons that look at small regions within the input and share neurons and parameters within the same activation map. A feature map, or activation map, is the output activation of a given filter. Feature maps and activations have the same meaning. This is why in some papers it is called an activation map because it is a mapping corresponding to the activation of different parts of an image, and also a feature map because it is a mapping of where certain types of features are found within an image. High activation means that a particular feature has been found.

[0075] Another important concept in CNNs is pooling, which is a form of non-linear downsampling. There are several non-linear functions for implementing pooling, among which max pooling is the most common. Max pooling divides the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.

[0076] Intuitively, the exact location of a feature is not as important as its approximate location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. The pooling layer serves to gradually reduce the spatial size of the representation, reduce the number of parameters, memory footprint, and computational load within the network, and thus also control overfitting. In CNN architectures, it is common to periodically insert a pooling layer between successive convolutional layers. The pooling operation provides another form of translation invariance.

[0077] The pooling layer operates independently for each depth slice of the input and spatially resizes it. The most common form is a pooling layer with a 2×2 filter, which is applied with a stride of 2 for each depth slice of the input along both width and height, discarding 75% of the activations. In this case, all max operations are over four numbers. The depth dimension remains invariant. In addition to max pooling, the pooling unit can also use other functions such as average pooling or L2 norm pooling. Average pooling has been commonly used, but recently it has become less preferred compared to max pooling, which often works better in practice. For aggressive reduction of the size of the representation, there is a recent trend to use smaller filters or even discard the pooling layer completely. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the fast R-CNN architecture.

[0078] The above-mentioned ReLU is an abbreviation for rectified linear unit and applies a non-saturating activation function. ReLU effectively removes negative values from the activation map by setting negative values to 0. ReLU increases the non-linear characteristics of the decision function and the network as a whole without affecting the receptive field of the convolutional layer. Other functions, such as the saturated hyperbolic tangent and sigmoid functions, are also used to increase non-linearity. ReLU is often preferred over other functions because it can train neural networks several times faster without significantly degrading generalization accuracy.

[0079] After several convolutional layers and max-pooling layers, high-level inference in the neural network is performed through fully connected layers. The neurons in the fully connected layers are connected to all the activations of the previous layer, as seen in a normal (non-convolutional) artificial neural network. Thus, those activations can be computed as an affine transformation, followed by a bias offset (vector addition of a learned or fixed bias term) after matrix multiplication.

[0080] The "loss layer" (including the calculation of the loss function) specifies how training penalizes the deviation between the predicted (output) label and the true label, and is usually the final layer of the neural network. Various loss functions suitable for different tasks can be used. The softmax loss is used to predict a single class out of K mutually exclusive classes. The sigmoid cross-entropy loss is used to predict K independent probability values in [0,1]. The Euclidean loss is used for regression to real-valued labels.

[0081] In summary, FIG. 1 shows the data flow in a typical convolutional neural network. First, the input image is passed through a convolutional layer and abstracted into a feature map that includes multiple channels corresponding to some of the filters within a set of learnable filters in this layer. Next, the feature map is subsampled, for example, using a pooling layer, which reduces the dimensions of each channel within the feature map. Next, the data reaches another convolutional layer that can have a different number of output channels. As described above, the number of input channels and output channels are hyperparameters of the layer. To establish the connectivity of the network, these parameters need to be synchronized between two connected layers such that the number of input channels of the current layer equals the number of output channels of the previous layer. For the first layer that processes input data, such as an image, the number of input channels is usually equal to the number of channels of the data representation. For example, for an RGB or YUV representation of an image or video, there are three channels, or for a grayscale image or video representation, there is one channel. The channels obtained by one or more convolutional layers (and optionally one or more resampling layers) can be passed to an output layer. Such an output layer can be convolutional or resampling in some implementations. In an exemplary and non-limiting implementation, the output layer is a fully connected layer.

[0082] Autoencoders and Unsupervised Learning An autoencoder is a type of artificial neural network used to learn efficient data coding in an unsupervised manner. A schematic diagram thereof is shown in FIG. 2. An autoencoder includes an encoder side 210 where an input x is input to an input layer of an encoder sub-network 220, and a decoder side 250 where an output x' is output from a decoder sub-network 260. The purpose of the autoencoder is to learn a representation (encoding) 230 of a dataset x, usually for dimensionality reduction, by training the networks 220, 260 to ignore signal "noise". Along with the reduction (encoder) side sub-network 220, a reconstruction (decoder) side sub-network 260 is learned, and the autoencoder attempts to generate from the reduced encoding 230 its original input x, and thus a representation x' as close as possible to its name. In the simplest case, when one hidden layer is given, the encoder stage of the autoencoder takes the input x and maps it to h, h = σ(Wx + b) where

[0083] This image h is usually called the code 230, latent variable, or latent representation. Here, σ is an element-wise activation function such as the sigmoid function or the rectified linear unit. W is a weight matrix and b is a bias vector. The weights and biases are usually initialized randomly and then updated iteratively during training via backpropagation. Then, the decoder stage of the autoencoder maps h to a reconstruction x' of the same shape as x: x' = σ'(W'h' + b') wherein the σ', W' and b' of the decoder may be independent of the corresponding σ, W and b of the encoder.

[0084] Variational autoencoder models make strong assumptions about the distribution of latent variables. They use a variational approach to latent representation learning, resulting in an additional loss component and a specific estimator for a training algorithm called the stochastic gradient variational Bayes (SGVB) estimator. The data is a directed graphical model p θ(x|h), and the encoder uses the posterior distribution p θ Approximation q to (h|x) φ Assume we are learning (h|x), where φ and θ represent the parameters of the encoder (recognition model) and decoder (generative model), respectively. The probability distribution of the latent vectors in a VAE typically matches the probability distribution of the training data much more closely than a standard autoencoder. The objective function of a VAE has the following form: L(φ,θ,x)=D KL (q φ (h|x)||p θ (h))-E qφ(h│x) (log p θ (x|h))

[0085] In the formula, D KL denotes the Kullback-Leibler divergence. The prior distribution for the latent variables is usually a centrally isotropic multivariate Gaussian distribution p θ It is set so that (h) = N(0,I). In general, the shapes of the variational and likelihood distributions are chosen so that they are factored Gaussian distributions: q φ (h|x)=N(ρ(x),ω 2 (x)I) p φ (x|h)=N(μ(h),σ 2 (h)I) where ρ(x) and ω 2 (x) is the encoder output, μ(h) and σ 2 (h) is the decoder output.

[0086] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, have enabled researchers to apply neural network-based techniques to image and video compression tasks. For example, end-to-end optimized image compression using networks based on variational autoencoders has been proposed.

[0087] Accordingly, data compression is regarded as a fundamental and well-studied problem in engineering and is generally formulated with the aim of designing a code for a given discrete data ensemble with minimum entropy. This solution relies heavily on knowledge of the probabilistic structure of the data, and thus the problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as vectors of image pixel intensities) must be quantized into a finite set of discrete values, which introduces error.

[0088] In this context, two competing costs, the entropy (rate) of the discretized representation and the error (distortion) resulting from quantization, known as the irreversible compression problem, must be traded off. Different compression applications, such as data storage and transmission over channels with limited capacity, require different rate-distortion tradeoffs.

[0089] Simultaneous optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional spaces is intractable. For this reason, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-valued representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This approach is called transform coding because of the central role of the transformation.

[0090] For example, JPEG uses the discrete cosine transform for blocks of pixels, while JPEG2000 uses multi-scale orthogonal wavelet decomposition. Typically, the three components of the transform coding method, the transform, quantization, and entropy coding, are optimized separately (often by manual parameter adjustment). The latest video compression standards such as HEVC, VVC, and EVC also use a transformed representation to code the residual signal after prediction. Several transforms are used for that purpose, such as the discrete cosine transform and discrete sine transform (DCT, DST), as well as the low-frequency non-separable manually optimized transform (LFNST).

[0091] Variational image compression The variational autoencoder (VAE) framework can be regarded as a non-linear transform coding model. The transformation process can mainly be divided into four parts. This is illustrated in FIG. 3A showing the VAE framework.

[0092] The transformation process can mainly be divided into four parts, and FIG. 3A illustrates the VAE framework. In FIG. 3A, the encoder 101 maps the input image x (represented by y) to a latent representation via the function y = f(x). This latent representation may hereinafter be referred to as part of the "latent space" or a point within the "latent space". The function f() is a transformation function that transforms the input signal x into a more compressible representation y. The quantizer 102 uses Q representing a quantization function to transform the latent representation y into a quantized latent representation having a (discrete) value by

Number

Number

Number

Number

Number

[0093] The latent space can be understood as a representation of compressed data where similar data points are close to each other within the latent space. The latent space is useful for learning data features and finding a simpler representation of the data for analysis. Quantized latent representation [Number] , and side information of the hyper-prior distribution 3 [Number] is included (quantized) in the bitstream 2 using arithmetic coding (AE). Further, the quantized latent representation is reconstructed into an image [Number] A decoder 104 for converting to is provided. The signal [Number] is an estimate of the input image x. x is preferably [Number] as close as possible to, in other words, it is desirable that the reconstruction quality is as high as possible. However, [Number] The higher the similarity between and x, the greater the amount of side information that needs to be transmitted. The side information includes the bitstream 1 and the bitstream 2 shown in FIG. 3A, which are generated by the encoder and transmitted to the decoder. Usually, the greater the amount of side information, the higher the reconstruction quality. However, a large amount of side information means a low compression rate. Therefore, one objective of the system described in FIG. 3A is to balance the reconstruction quality and the amount of side information transmitted in the bitstream.

[0094] In FIG. 3A, component AE105 is an arithmetic encoding module, and the quantized latent representation

Number

Number

Number

[0095] Arithmetic decoding (AD) 106 is a process of reversing the binarization process, and the binary number is converted into a sample value. Arithmetic decoding is provided by arithmetic decoding module 106.

[0096] Note that the present disclosure is not limited to this specific framework. Furthermore, the present disclosure is not limited to image or video compression, and can also be applied to object detection, image generation, and recognition systems.

[0097] In FIG. 3A, there are two sub-networks connected to each other. A sub-network in this context is a logical division between parts of the entire network. For example, in FIG. 3A, modules 101, 102, 104, 105, and 106 are called the "encoder / decoder" sub-network. The "encoder / decoder" sub-network is responsible for encoding (generating) and decoding (parsing) the first bitstream "bitstream 1". The second network in FIG. 3A includes modules 103, 108, 109, 110, and 107 and is called the "hyper encoder / decoder" sub-network. The second sub-network is responsible for generating the second bitstream "bitstream 2". The purposes of the two sub-networks are different.

[0098] The first sub-network is responsible for the following: · Conversion 101 of the input image x to its latent representation y (which is easier to compress x), · Quantizing the latent representation y to a quantized latent representation

Number

Number

Number

[0099] The purpose of the second sub-network is to obtain the statistical characteristics (e.g., the mean value, variance, and correlation between samples of bitstream 1) of the samples of "bitstream 1" so that the compression of bitstream 1 by the first sub-network becomes more efficient. The second sub-network generates a second bitstream "bitstream 2" that includes the information (e.g., the mean value, variance, and correlation between samples of bitstream 1).

[0100] The second network converts the quantization latent representation

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

Number

[0101] Figure 3A illustrates an example of a VAE (Variational Autoencoder), the details of which may vary in different implementations. For example, in a particular implementation, there may be additional components to more efficiently obtain the statistical characteristics of the samples of bitstream 1. In one such implementation, there may be a context model aimed at extracting the cross-correlation information of bitstream 1. The statistical information provided by the second subnetwork may be used by the AE (arithmetic encoder) 105 and AD (arithmetic decoder) 106 components.

[0102] Figure 3A shows the encoder and decoder in a single figure. As will be apparent to those skilled in the art, the encoder and decoder may be incorporated into different devices, and this is very often the case.

[0103] Figure 3B shows the encoder, and Figure 3C shows the decoder component of the VAE framework separately. As an input, the encoder receives a picture according to some embodiments. The input picture may include one or more channels such as color channels or other types of channels, for example, depth channels, motion information channels, etc. The output of the encoder (as shown in Figure 3B) is bitstream 1 and bitstream 2. Bitstream 1 is the output of the first subnetwork of the encoder, and bitstream 2 is the output of the second subnetwork of the encoder.

[0104] Similarly, in Figure 3C, two bitstreams, bitstream 1 and bitstream 2, are received as inputs, and the reconstructed (decoded) image

Number

[0105] Specifically, as seen in Figure 3B, the encoder includes an encoder 121 that converts the input x into a signal y that is then provided to the quantizer 322. The quantizer 122 provides information to the arithmetic encoding module 125 and the hyper-encoder 123. The hyper-encoder 123 provides the bitstream 2 described above to the hyper-decoder 147, and the hyper-decoder 147 provides this information to the arithmetic encoding module 105 (125).

[0106] The output of the arithmetic encoding module is bitstream 1. Bitstream 1 and bitstream 2 are the outputs of the signal encoding, and these outputs are then provided (sent) to the decoding process. Unit 101 (121) is called the "encoder", but it is also possible to call the complete sub-network described in Figure 3B the "encoder". An encoder generally means a unit (module) that converts an input into an encoded (e.g., compressed) output. As can be seen from Figure 3B, unit 121 actually can be regarded as the core of the entire sub-network in order to perform the conversion of input x to y which is a compressed version of x. The compression in encoder 121 can be achieved, for example, by applying a neural network, or generally any processing network having one or more layers. In such a network, the compression can be performed by a cascade process including downsampling that reduces the size and / or number of channels of the input. Thus, an encoder may sometimes be called, for example, a neural network (NN)-based encoder.

[0107] The remaining parts in the figure (quantization unit, hyper-encoder, hyper-decoder, arithmetic encoder / decoder) are all parts responsible for improving the efficiency of the encoding process or for the conversion of the compressed output y into a series of bits (bitstream). Quantization can be provided to further compress the output of the NN encoder 121 by irreversible compression. AE125, in combination with the hyper-encoder 123 and hyper-decoder 127 used to constitute AE125, can perform binarization that can further compress the quantized signal by reversible compression. Therefore, it is also possible to call the entire sub-network in Figure 3B the "encoder".

[0108] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of the signal before converting it into binary digits (bits). For example, in the VAE framework, the encoder, which is a non-linear transformation, maps the input image x to y, where y has a smaller width and height than x. Since y has a smaller width and height, and thus a smaller size, the dimensionality (size) of the signal is reduced, and thus it is easier to compress the signal y. Note that in general, the encoder does not necessarily need to reduce the size of both (or generally all) dimensions. Rather, some exemplary implementations may provide an encoder that reduces the size of only one dimension (or generally, a subset of dimensions).

[0109] In "Density Modeling of Images Using a Generalized Normalization Transformation," presented at the 4th Int. Conf. for Learning Representations, 2016, by J. Balle, L. Valero Laparra, and E. P. Simoncelli (2015), In: arXiv e-prints (hereinafter referred to as "Balle"), the authors proposed a framework for the end-to-end optimization of an image compression model based on non-linear transformations. The authors optimize for the mean squared error (MSE), but use a more flexible transformation constructed from a cascade of linear convolutions and non-linearity. Specifically, the authors use a generalized divisive normalization (GDN) coupled non-linearity, which has been proven to be effective in the Gaussian distribution of image density, inspired by models of neurons in the biological visual system. Following this cascade transformation, uniform scalar quantization is performed (i.e., each element is rounded to the nearest integer), effectively implementing vector quantization in a parametric form over the original image space. The compressed image is reconstructed from these quantized values using an approximate parametric non-linear inverse transformation.

[0110] Such an example of a VAE framework is shown in Figure 4, which utilizes six downsampling layers marked 401-406. The network architecture includes a hyperprior model. a ,g s ) shows the image autoencoder architecture, and the right side (h a ,h s ) corresponds to an autoencoder that implements a hyperprior. The factorized prior model is then used to compute the analytic and synthetic transformation g a and g s The encoder uses the same architecture as the arithmetic encoder and decoder. Q stands for quantization, and AE and AD stand for arithmetic encoder and decoder, respectively. The encoder applies g a , resulting in responses y (latent representations) with spatially varying standard deviations. a contains multiple convolutional layers with subsampling and generalized decomposition normalization (GDN) as the activation function.

[0111] The response is h a , which summarizes the distribution of standard deviations in z. z is then quantized, compressed, and transmitted as side information. The encoder then generates the quantized vector

number

number

number

number

[0112] The layer including downsampling is indicated by a downward arrow in the layer description. The layer description "Conv N,k1,2↓" means that the layer is a convolutional layer, has N channels, and the size of the convolutional kernel is k1×k1. For example, k1 may be equal to 5 and k2 may be equal to 3. As described above, 2↓ means that downsampling by a factor of 2 is performed in this layer. Downsampling by a factor of 2 results in one of the dimensions of the input signal being reduced by half in the output. In Figure 4, 2↓ indicates that both the width and height of the input image are reduced by a factor of 2. Since there are six downsampling layers, if the width and height of the input image 414 (also represented by x) are given by w and h, the output signal z^413 has a width and height equal to w / 64 and h / 64, respectively. The modules represented by AE and AD are an arithmetic encoder and an arithmetic decoder, respectively, and are described with reference to FIGS. 3A to 3C. The arithmetic encoder and decoder are specific implementations of entropy coding. AE and AD can be replaced by other means of entropy coding. In information theory, entropy encoding is a reversible data compression method used to convert the values of symbols into a binary representation and is a recoverable process. Also, "Q" in the figure corresponds to the quantization operation also described above in relation to Figure 4 and is further explained in the "Quantization" section above. Also, the quantization operation and the corresponding quantization unit as part of component 413 or 415 do not necessarily exist and / or can be replaced by another unit.

[0113] Figure 4 also shows a decoder including upsampling layers 407 to 412. Although implemented as convolutional layers, an additional layer 420 that does not provide upsampling to the received input is provided between upsampling layers 411 and 410 in the order of processing of the input. A corresponding convolutional layer 430 is also shown for the decoder. Such layers can be provided in the NN to perform operations on the input that do not change the size of the input but change certain characteristics. However, it is not necessary for such layers to be provided.

[0114] Looking at the processing order of bitstream 2 through the decoder, the upsampling layers follow in reverse order, i.e., from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide upsampling at an upsampling ratio of 2, indicated by ↑. Of course, not all upsampling layers necessarily have the same upsampling ratio, and other upsampling ratios such as 3, 4, 8, etc. may be used. Layers 407 to 412 are implemented as convolutional layers (conv). Specifically, since these may be intended to provide an operation opposite to that of the encoder to the input, the upsampling layer may apply a transposed convolution operation to the received input such that its size is increased by a factor corresponding to the upsampling ratio. However, the present disclosure is not generally limited to transposed convolution, and upsampling may be performed in any other way, such as by bilinear interpolation between two neighboring samples or by nearest neighbor sample copy, etc.

[0115] In the first subnetwork, after some convolutional layers (401 to 403), generalized divisive normalization (GDN) follows on the encoder side and inverse GDN (IGDN) follows on the decoder side. In the second subnetwork, the activation function applied is ReLu. It should be noted that the present disclosure is not limited to such an implementation form, and generally, other activation functions may be used instead of GDN or ReLu.

[0116] Cloud Solutions for Machine Tasks Machine-oriented video coding (VCM) is another direction of computer science that is currently widespread. The main idea behind this approach is to transmit a coded representation of image or video information targeted for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to conventional image and video coding targeted at human perception, the quality characteristic is not the reconstruction quality but the performance of computer vision tasks, such as object detection accuracy. This is illustrated in Figure 5.

[0117] Video coding for machines, also known as collaborative intelligence, is a relatively new paradigm for the efficient placement of deep neural networks across a mobile cloud infrastructure. By partitioning the network between the mobile side 510 and the cloud side 590 (e.g., a cloud server), it is possible to distribute the computational workload such that the overall energy and / or latency of the system is minimized. In general, collaborative intelligence is a paradigm in which the processing of neural networks is distributed among two or more different computing nodes, e.g., between devices, but generally between any functionally defined nodes. Here, the term "node" does not refer to the neural network nodes described above. Rather, the (computing) nodes here refer to separate devices / modules (physically or at least logically) that implement a part of the neural network. Such devices may be different servers, different end-user devices, a mix of servers and / or user devices, and / or the cloud and / or processors, etc. In other words, the computing nodes may be considered as nodes that belong to the same neural network and communicate with each other to transfer coding data within / for the neural network. For example, one or more layers may be executed on a first device (such as a device on the mobile side 510) to enable complex computations, and one or more layers may be executed on another device (such as a cloud server on the cloud side 590). However, the distribution may also be finer, and a single layer may be executed on multiple devices. In the present disclosure, the term "plurality" refers to two or more. In one existing solution, a part of the neural network function is executed on a device (such as a user device or an edge device) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a set of processing or computing systems located outside the device operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training.In this case, data flows in both directions, from the cloud to the mobile during backpropagation in training, during the forward pass in training, and from the mobile to the cloud (illustrated in FIG. 5) during inference.

[0118] Some research has presented semantic image compression by encoding deep features and then reconstructing the input images from them. Compression based on uniform quantization was shown, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, instead of sending compressed natural image data to the cloud and performing object detection using the reconstructed images, it may be more efficient to send the output of the hidden layer (deep feature map) 550 from the mobile unit 510 to the cloud 590. Thus, it may be advantageous to compress the data (features) generated by the mobile side 510, which may include a quantization layer 520 for this purpose. Correspondingly, the cloud side 590 may include an inverse quantization layer 560. Efficient compression of feature maps is beneficial for image and video compression and reconstruction for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a common approach for compressing deep features (i.e., feature maps).

[0119] Today, video content contributes to over 80% of Internet traffic, and this percentage is expected to further increase. Therefore, it is important to build an efficient video compression system to generate higher-quality frames with a given bandwidth budget. Furthermore, most video-related computer vision tasks, such as video object detection and video object tracking, are sensitive to the quality of compressed video, and efficient video compression can potentially benefit other computer vision tasks. On the other hand, video compression techniques are also useful for action recognition and model compression. However, in the past few decades, video compression algorithms have relied on manually crafted modules, such as block-based motion estimation and discrete cosine transform (DCT), to reduce the redundancy of video sequences as described above. Each module is well designed, but the overall compression system is not optimized end-to-end. It is desirable to further improve video compression performance by optimizing the entire compression system together.

[0120] End-to-End Image or Video Compression DNN-based image compression methods can utilize large-scale end-to-end training and advanced non-linear transformations that are not used in conventional approaches. However, it is not straightforward to directly apply these techniques to build an end-to-end learning system for video compression. First, how to generate and compress motion information adjusted for video compression remains an open problem. Video compression methods rely heavily on motion information to reduce temporal redundancy in video sequences.

[0121] A direct solution is to represent the motion information using learning-based optical flow. However, current learning-based optical flow approaches aim to generate the most accurate flow field possible. Precise optical flow is often not optimal for specific video tasks. Furthermore, the data volume of optical flow increases significantly compared to motion information in conventional compression systems, and directly applying existing compression approaches to compress optical flow values will significantly increase the number of bits required to store the motion information. Second, it is unclear how to build a DNN-based video compression system by minimizing rate-distortion-based objectives for both residual information and motion information. Rate-distortion optimization (RDO) aims to achieve higher quality (i.e., less distortion) of the reconstructed frames when a given number of bits (or bitrate) for compression is provided. RDO is important for video compression performance. To utilize the ability of end-to-end training for learning-based compression systems, an RDO strategy is required to optimize the entire system.

[0122] In "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pages 11006 - 11015, the authors proposed an end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual coding.

[0123] Such an encoder is shown in Figure 6. In particular, Figure 6 shows the overall structure of an end-to-end trainable video compression framework. To compress the motion information, the optical flow v tConvert to the corresponding representation m that is more suitable for better compression t The CNN was specified to convert to. Specifically, an autoencoder-type network is used to compress the optical flow.

[0124] Generally, video compression may degrade the perceptual quality of images, and an image correction filter may generally be used to improve the output quality of the compressed video.

[0125] One type of image correction filter improves the quality of multi-channel images by exploiting the similarity between channels. The performance of multi-channel image correction algorithms varies depending on several parameters of the input multi-channel image (e.g., the number of channels, the quality of the channels), and also varies across the image data of each channel.

[0126] There are various image correction algorithms. Only a very small number of them utilize inter-channel correlation information for image correction. In the present disclosure, focus is placed on multi-channel image correction filters that use neural networks such as convolutional neural networks. In a neural network-based correction filter, the network is trained with two sets of images, one representing the original (target, desired) quality and the other representing the expected range and type of distortion. Such a network can be trained to improve images degraded by, for example, sensor noise, or by video compression, or by other types of distortion. Usually, different (individual separate) training is required for each type of distortion. A more general network (e.g., one that processes a wider range and type of distortion) has a lower average performance. Here, performance refers to the quality of reconstruction that can be measured by, for example, an objective criterion such as PSNR, or by some metrics that also take into account human vision.

[0127] In some embodiments of the present application, a deep convolutional neural network (CNN) is trained to reduce compression artifacts and improve the visual quality of an image while maintaining a high compression rate. In particular, according to one embodiment, a method for modifying an input image region is provided. Here, modifying refers to any modification, typically a modification obtained by filtering or other image correction approaches. The type of modification may depend on the specific application.

[0128] One network that yields good results for a wide range of distortions without the need to be trained for each specific case is known from Cui, Kai and Steinbach, Eckehard (2018), "Decoder Side Image Quality Enhancement exploiting Inter-channel Correlation in a 3-stage CNN", Submission to CLIC 2018, IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2018. In it, a three-stage convolutional neural network (CNN)-based approach is proposed that can utilize inter-channel correlation to improve image quality on the decoder side.

[0129] Figure 7 illustrates such a three-channel CNN framework. A CNN is a neural network that uses convolutions instead of general matrix multiplications in at least one of its layers. The convolutional layers convolve the input and pass the result to the next layer. They have several beneficial features particularly for image / video processing. The applied stages are described, including the CNN of Figure 7, known configurations, and possible configurations and alternatives that may facilitate the application of the CNN in some embodiments of the present disclosure.

[0130] The input image is stored in the RGB (Red, Green, Blue) format (color space). The input image may be a still image or an image that is a frame of a video sequence (video).

[0131] The circled numbers in Figure 7 represent the stages of processing. In stage 1, patches are selected from the input image. In a specific example, the patches have a predetermined size such as a size of 240×240 samples (pixels). The patch size may be fixed or may be predetermined. For example, the patch size may be a parameter that can be configured by the user or application, or may be transmitted in a bitstream and set accordingly. The selection of the patch size may be made according to the image size and / or the amount of detail in the image.

[0132] Here, the term "patch" refers to a part of the image that is processed by filtering, and the processed part is then pasted back at the position of the patch. The patch may be regular, such as rectangular or square. However, the present disclosure is not limited to this, and the patch may have any shape, such as a shape according to the shape of the object to be detected / recognized that is to be filtered. In some embodiments, the entire image is filtered (corrected) patch by patch. In other embodiments, only the selected patches (for example, corresponding to objects) may be filtered, and the remaining part of the image may not be filtered or may be filtered by another approach. Filtering means any kind of correction.

[0133] The selection may be the result of sequential processing or parallel processing in which all patches are filtered. In such a case, the selection may be made in a predetermined order such as from left to right, top to bottom. However, the selection may also be made by the user or application, and only a part of the image may be considered. The patches may be continuous or may be dispersed.

[0134] When the entire image is divided into patch regions, padding may be applied if an integer multiple of the patch dimension (vertical or horizontal) does not match the image size. Padding may include mirroring of the available image portions across the axis formed by the image boundary (horizontal or vertical) to achieve a size that fits an integer number of patches. More specifically, the padding is performed such that the vertical dimension (number of samples after padding) is an integer multiple of the vertical patch dimension. Further, the horizontal dimension (number of samples after padding) is an integer multiple of the horizontal patch dimension.

[0135] Image correction can then be performed by sequentially or in parallel selecting and processing each of the patches. The patches may not overlap as proposed above. However, the patches may also overlap, which can improve quality and reduce possible boundary effects between separately processed patches.

[0136] In stage 2, the pixels of the patch are rearranged to make processing easier. The rearrangement may include so-called pixel shifting. Pixel shifting rearranges the pixels within each channel, whereby a channel having dimension N×N×1 is transformed into a 3D array having dimension N / 2×N / 2×4 (where the symbol "×" represents "times", i.e., multiplication). This is done by subsampling a channel in which a single value is taken from each non-overlapping block of 2×2 values. The first layer of the 3D array is created from the top-left values, the second layer is created from the top-right, the third layer is created from the bottom-left, and the fourth layer is created from the bottom-right. At the end of the pixel shift operation, a 3-channel RGB patch having size N×N pixels becomes a stack (3D array) having dimension N / 2×N / 2×4. This pixel shift is done mainly for computational reasons, i.e., because it is easier for a processor (e.g., a graphical processing unit, GPU) to process a narrow and deep stack than a wide and shallow stack. Note that the present disclosure is not limited to subsampling by 4. In general, more or less subsampling may be done, resulting in a corresponding stack depth of more or fewer subsampled images of the result.

[0137] In stage 3 of FIG. 7, the green channel is processed in single-channel mode. In this example of FIG. 7, it is assumed that the green channel is the primary channel and the remaining channels are secondary channels. It is assumed that the distortion of each channel is strongly correlated with the color of the channel. It is assumed that the green channel has the least distortion. Since the RGB image is captured using an RGGB sensor pattern (Bayer pattern) and captures more green samples than red and blue samples, this can be a reasonable assumption on average for RGB images affected by sensor noise or compression. However, according to the present disclosure, as discussed below, the green channel is not necessarily the best, and it may be advantageous to perform channel selection. Further, channels other than color channels may be used and may provide higher-quality information.

[0138] In FIG. 7, the primary (i.e., green) channel is processed only in stage 3. The red channel is processed together with the improved green channel (stage 4). The blue channel is processed together with the improved green channel (stage 5). The improved red, green, and blue channels are then stacked (combined) together (stage 6) and processed with or without side information.

[0139] Overall, the framework has four NN stages: stage 3, stage 4, stage 5, and stage 7 of FIG. 7. Stages 3 and 7 take as input a single 3D array of values and output a 3D array having the same size as the input. Stages 4 and 5 take as input two 3D arrays, one primary and one auxiliary. The output is of the same size as the primary output and is intended to be a processed (e.g., corrected) version of the primary input, and the auxiliary input is used only to assist in the processing and is not output.

[0140] In stage 4, the final hidden layer is processed by N convolutional kernels of 3×3×64 and outputs a stack of size Z. In stage 5, since the network is trained to approximate the difference between the input and the desired output, the input is added to the processed output.

[0141] In stage 5 of FIG. 7, the blue channel is processed with emphasis on the corrected green channel. This processing is similar to that described with reference to the red channel and FIG. 7.

[0142] Note that in this disclosure, the terms "channel" or "image channel" do not necessarily refer to color channels. Other (tensor) channels such as depth channels or other feature channels may be corrected using the embodiments described herein.

[0143] In stage 6 of FIG. 7, all processed channels (red, green, blue) are combined into a Z = 12 (4 subsampled images per color channel) image of size 120×120. Here, combining means stacking together as a common input for the next stage 7. In stage 7, the combined channels are processed together by the network (S).

[0144] In stage 8, the pixels are rearranged and returned to form the processed patch. This may include removing padding by cropping the mirrored portions. Note that the above padding by mirroring is only one of the possible ways to perform padding. The present disclosure is not limited to such specific examples.

[0145] In stage 9, the processed patch is inserted back into the original image. In other words, the original image is updated with the corrected patch. During the training procedure, a set of the original image and the distorted image are used as input, and all convolutional kernels of all four networks are selected. The purpose of the network is to obtain the distorted image and generate a close match to the original image.

[0146] The CNN-based multi-channel correction filter as described above operates in a strict and non-adaptive manner, i.e., the processing parameters are set up during the design (or training) of the filter and are applied in the same way regardless of the content of the image passing through the filter. However, the optimal selection of the primary channel can vary from image to image or even for different parts of the same image. For the quality of image correction, it is advantageous if the primary (first) channel is the channel with the highest quality, which means the lowest distortion. This is because the primary channel is also involved in the correction of the remaining channels. The inventors recognized that better performance can be achieved by carefully selecting the primary channel, such as for each patch or for each image, which was confirmed by experiments. Second, the correction performance (of all correction filters) varies depending on the input image quality. The correction filter functions optimally for a certain range of distortion intensities and does not bring much improvement for very high or very low distortion levels. This is because high-quality inputs can hardly be improved further, and low-quality inputs are too distorted to be reliably improved. As a result, for some inputs, it is beneficial to skip the correction process completely. Third, the above-described (referring to FIG. 7) image correction is designed for RGB inputs where the resolution of all three channels is the same. However, in some video standard formats (e.g., YUV 4:2:0 or YUV 4:2:2, etc.), different channels can have different resolutions and numbers of pixels. It should be noted that the image correction described herein is not necessarily applied during encoding or decoding. The image correction may be used for preprocessing. For example, the image correction may be used to correct images or videos in raw formats such as Bayer pattern-based formats. In other words, the correction of the image may be applied to any image or video.

[0147] The above problem is addressed in International Publication No. WO 2021 / 249684 (A1), which provides an image correction filter to which a content analysis module is added to adapt to multi-channel image formats having different input formats and different contents, any number of channels (e.g., n channels as illustrated in FIG. 8), and different numbers of pixels within each channel, and to adjust the filter to process the input channels in a specific order or to skip the processing entirely. In particular, the primary channel can be selected, for example, based on content analysis. The configuration shown in FIG. 8 enables processing of n channels. After selection 810 of a patch of the image to be processed, application 820 of pixel shift to the selected patch, and selection of the primary channel, the primary channel is processed by network 1, and the corrected corrected primary channel thus obtained is used as auxiliary information when processing the secondary channels. The output m’1 of network 1 is concatenated 830, 840 with the 3D array from m n to m 2c from m nc and the resulting concatenated 3D array m n from m to m is processed by networks 2 to N, respectively. The outputs from networks 1 to N are concatenated 850, m’1 + m’2 + … + m’ n and processed by network M. The output of network M is pixel unshifted 860 to obtain a corrected multi-channel image region.

[0148] Unlike the related art, the present disclosure provides spatial frequency conversion-based image correction (e.g., correction) using inter-channel correlation information. Spatial frequency conversion provides information regarding spatial frequency, but is not necessarily limited to spatial frequency only (e.g., wavelet transform also provides information regarding position).

[0149] Suitable spatial frequency conversions that can be used include energy compression conversions including wavelet transforms (discrete wavelet transform, DWT, or stationary wavelet transform), discrete Fourier transform, fast Fourier transform, and discrete cosine transform.

[0150] Certain embodiments are illustrated in FIG. 9. This and other embodiments can be adapted to different input formats and different contents, and represent an adjustable correction filter capable of processing multi-channel image formats having any number of channels and different numbers of pixels within each channel. The channel representing the image area to be processed can be any feature channel considered appropriate, such as a color channel, and can have any size (depth), for example different sizes.

[0151] For the provided (distorted) input image, a patch (or image area) to be processed is selected 910. It should be noted that the selection of the patch does not necessarily mean that there is any intelligence in selecting the patch. Rather, the selection may be made sequentially (e.g., in a loop), or, for example, when parallel processing is applied, may be made for two or more patches (or even all patches) at once. The selection step simply corresponds to determining which patch is to be processed. For an image of size (height dimension and width dimension having) k...2 p ×l...2 p in the case of the size of the image, the image is of size 2 p ×2 pIt is divided into k×l square patches each having. If the image cannot be divided into k×l square patches, the image is padded before division. Alternatively, the image is divided into patches, and if the patches resulting from the division are not square in the width and height dimensions, the patches are padded to obtain square patches. Depending on the type of image modification, there may be several criteria that are advantageously observed when selecting the patch size (the size of the image area to be modified). In particular, there may be some minimum size given by the type of processing applied during the image modification (see WO 2021 / 249684(A1) for details).

[0152] The input image may be a still image or part of a video sequence. The dimension 2 to be processed p ×2 p Patches having are divided into color channels as (distorted) channels, for example, RGB or YUV (luminance and chrominance) channels (components). One of the channels is selected as the primary channel (for example, luma in the case of the YUV channel). The other channels are secondary channels that are processed based on the information regarding the primary channel. The channels are of size 2 p ×2 pIt has, and the integer p does not have to be the same for each channel, that is, different sample resolutions can be provided for different channels. In the illustrated embodiment of FIG. 9, p is the same for all channels, for example, to process a YUV444 image (patch). When the channel is a YUV channel, the Y channel can be selected as the primary channel. The selection of the primary channel can be made according to the teachings provided in International Publication No. 2021 / 249684 (A1). In particular, the primary channel may be selected based on an analysis, for example, a content analysis of the patch or the selected patch. For example, a classifier based on a neural network or a convolutional neural network may be used for the selection of the primary channel (and the secondary channel). The classifier may also be implemented by several algorithms such as an algorithm for determining the level of detail, and / or the intensity and / or direction of edges, the distribution of gradients, the movement characteristics (in the case of video) or other image features. Based on the comparison of such features of each image channel, the primary channel can then be selected. For example, as the primary channel, an image channel containing most edges or the sharpest edges (corresponding to most details) may be selected.

[0153] The selection is made on the encoder side and can also be signaled to the decoder side, can be predetermined, or can be made independently on the encoder side and the decoder side using an image analysis method performed on the patch. In the latter two cases, the selection of the primary channel is not signaled. Further, note that the selection of the primary channel for a patch can be different for each or part of the patch into which the image is divided. Alternatively, a fixed predetermined primary channel may be used.

[0154] All channels 920, 930, 940 are processed by a two-dimensional discrete wavelet transform (DWT). Below this, DWT is used as an example for using spatial frequency transformation. Other examples as described above may be used. Any type of wavelet considered appropriate can be used for DWT. For example, the Haar wavelet or the Daubechies wavelet may be considered appropriate choices. The application of DWT 920 to the pixel values of the pixels of the primary channel results in a DWT transform of the primary channel of size 2 p-1 ×2 p-1 ×4. The third dimension (of size 4) of the three-dimensional array output by the DWT is given by the spatial low-frequency subband LL and the spatial high-frequency subbands HL (vertical features), LH (horizontal features), and HH (diagonal features). All subbands can be arranged in a single matrix (layer) such that the output of the DWT transform remains a single layer, or can be represented by individual layers. Correspondingly, the application of the DWT 930, 940 to the pixel values of the pixels of the secondary channel results in a DWT transform of the secondary channel of size 2 p-1 ×2 p-1 ×4.

[0155] The DWT transform primary channel is processed independently by Network 1, which is the first network, from the DWT transform secondary channels. Further, the DWT transform primary channel is concatenated 950, 960 with each of the DWT transform secondary channels (along the third dimension). Thus, the DWT transform primary channel is used as auxiliary information for the processing of the (DWT transform) secondary channels. The arrays resulting from the concatenation process are input into Networks 2 to N, which are networks that operate independently of each other. Network 1 is designed to have inputs and outputs of the same size (2 p-1 ×2 p-1 ×4). Networks 2 to N are designed to have inputs with more (tensor) layers than the output (due to the concatenation process of 2 p-1 ×2 p-1×8 input and 2 p-1 ×2 p-1 ×4 output).

[0156] The output of Network 1 undergoes an inverse DWT 970, and the outputs from Network 2 to Network N also undergo an inverse DWT 980, 990. After the application of the inverse DWT 970, 980, 990, a corrected primary channel of size 2 p ×2 p and a corrected secondary channel of size 2 p ×2 p are obtained and combined to achieve a corrected, e.g., image-corrected (substantially distortion-free) patch of size 2 p ×2 p as described above. The next patch may be selected and processed. All processed patches can be assembled to construct a reconstructed image, and any padding is removed (if necessary).

[0157] The above filter includes N DWT transforms, N neural networks, and N inverse DWT transforms. For example, it is trained to receive a distorted image (e.g., by compression) and output an approximation of the original image. Using the processed (e.g., restored) primary channel to assist in the processing (e.g., restoration) of the secondary channel enables the filter to utilize the correlation between channels and achieve good reconstruction with a relatively small number of processing steps. The use of DWT transforms helps decorrelate the input data (into high spatial frequency components and low spatial frequency components), which helps the network reach good performance with fewer parameters and makes training easier compared to the prior art. During the training process, each network learns to take a distorted input (e.g., by compression) and output the best approximation of the original (undistorted) image. This is done by training the network with pairs of distorted (source) patches and original (target) patches. One common approach is to train with multiple sets of distorted images, thus obtaining multiple sets of neural network coefficients, where each set is optimized for a specific type of content (e.g., computer game, video conference, etc.) or a specific level of distortion (e.g., high compression, medium compression, etc.). During compression, the encoder can analyze the content and select the best set of trained coefficients. This selection can vary patch-by-patch and channel-by-channel. For example, Network 1 may use the set of coefficients for "highly compressed computer screen content", Network 2 may use the set for "medium compressed video conference screen content", and the processing by Network N may be turned off and operate in "pass-through" mode. Also, any of the networks (from Network 1 to Network N) can be independently turned off and the processing can be completely skipped.

[0158] Furthermore, the use of the DWT results in a rearrangement of the array elements input to the DWT, i.e., four layers along a quarter of the spatial dimension and the third dimension. In the subsequent neural network, the DWT transform is applied to the first two dimensions, which enables the neural network processing unit to efficiently process the data in parallel.

[0159] In fact, compared with the prior art, the overall processing can be accelerated by more than five times depending on the actual application.

[0160] Furthermore, in the prior art, all channels are combined and then processed together (see network M in FIG. 8). A set of networks is trained, and each network is optimized for different content. It may happen that the optimal processing setup for a given image includes turning off the processing of the primary channel and different setups for each of the secondary channels. In the prior art, it was necessary to run all the networks multiple times and assemble the outputs (e.g., take the original for channel 1 and version 4 for channel 2, etc.). According to the present disclosure, channel processing can be independently turned on / off, and the selection / training of the weight coefficients can advantageously be performed independently for each of the networks / channels. Optimal parameter selection usually requires different processing for each of the channels, e.g., the three YUV channels. In the prior art, it was necessary to run the filter three times (for the three channels), use the optimal parameters of one of the components each time, and then combine the desired parts to form the output. According to the present disclosure, the filter only needs to be run once regardless of the filter selection. One coefficient of the neural network can be changed without affecting the output of other neural networks.

[0161] As already described, channels of different sizes corresponding to different channel resolutions can also be handled. FIG. 10 shows an embodiment in which the size of the primary channel is four times the size of each of the secondary channels, having integers k, l, and p where k...2 p ×l...2 p An image of size is divided into k×l square patches each having size 2 p ×2 p If the image cannot be divided into k×l square patches, the image is padded before division. A patch of size 2 p ×2 p is selected for processing and divided into color channels, for example, (distorted) channels such as RGB or YUV (luminance and chrominance) channels. One of the channels is selected or predetermined as the primary channel. The other channels are secondary channels that are processed based on information regarding the primary channel. The primary channel has size 2 p ×2 p and each of the secondary channels has size 2 p-1 ×2 p-1 For example, the processed pass of the image has a YUV420 format.

[0162] All channels are processed 1020, 1030, 1040 using a two-dimensional DWT. Any type of wavelet considered appropriate can be used for the DWT. For example, a Haar wavelet or a Daubechies wavelet may be considered an appropriate choice. Application 1020 of the DWT to the pixel values of the pixels of the primary channel results in a DWT-transformed primary channel of size 2 p-1 ×2 p-1 ×4 having a third dimension given by the spatial low-frequency subbands LL and spatial high-frequency subbands HL, LH, HH. Correspondingly, application 1030, 1040 of the DWT to the pixel values of the pixels of the secondary channels results in a DWT-transformed secondary channel of size 2 p-2 ×2 p-2 ×4. A size of 2 p-1×2 p-1 The DWT-transformed primary channel of ×4 is processed by network 1 designed to have the same size of input and output. 2 p-1 ×2 p-1 The output of network 1 of size ×4 is subjected to inverse DWT 1080 to obtain the corrected / correction primary channel of size 2 p-1 ×2 p-1 ×1080.

[0163] In order to use the DWT-transformed primary channel as auxiliary information for the processing of the (DWT-transformed) secondary channel, the first two dimensions must be the same. Therefore, the DWT-transformed primary channel of size 2 p-1 ×2 p-1 ×4 is subjected to further (cascaded) DWT 1050 to obtain the auxiliary DWT-transformed primary channel of size 2 p-2 ×2 p-2 ×16. The three-dimensional array thus obtained is concatenated with the array of the DWT-transformed secondary channel 1060, 1070, and the concatenated array thus obtained is supplied to network 2 to network N respectively from the network, the network, the network. The network, the network, the network 2 to network N are designed to have a larger input in the third dimension and an output having a size exactly 4 in the third dimension (by DWT / inverse DWT), and the output of the network, the network, the network 2 to network N is of size 2 p-2 ×2 p-2 Each inverse DWTS 1090 and 1100 is subjected to obtain the corrected / correction secondary channel of size ×2. Based on the combination of the corrected primary channel and the secondary channel, a corrected (substantially distortion-free) patch of size 2 p ×2 p ×2. As described above, the following patches may be selected and processed. All processed patches can be assembled to construct the reconstructed image, and any padding is removed (if necessary).

[0164] The configuration illustrated in FIG. 10 provides the same advantages as the configuration described with reference to FIG. 9. In the embodiments described with reference to FIGS. 9 and 10, DWT is used, but alternatively, any other spatial frequency transform deemed appropriate may be used. In particular, a transform that does not result in downsampling by a factor of 2 in the height and width dimensions of the image region may be selected.

[0165] The neural network used in the embodiments described with reference to FIGS. 9 and 10 may be a convolutional neural network. It should be further noted that the embodiments described with reference to FIGS. 9 and 10 may be modified to perform the above-described pixel shift operation and pixel anti-shift operation, respectively, before the DWT and after the inverse DWT (or any other spatial frequency transform used).

[0166] FIG. 11 illustrates a specific example of a convolutional neural network that may be used in the embodiments described with reference to FIGS. 9 and 10. Except for the number of input tensor layers, the topologies of all the networks illustrated in FIGS. 9 and 10 may be the same or similar to each other.

[0167] For the processing of the DWT transform primary channel, the number of input tensor layers is 4. When processing the DWT transform secondary channel, the input to the neural network involved is larger to accommodate both the DWT transform secondary channel itself and the DWT transform primary channel concatenated in the third dimension. Further, if one of the channels (e.g., the primary channel) is larger and a cascaded DWT is used to equalize the first two dimensions, each of the two concatenated channels may have a third dimension larger than 4. For example, if one channel is twice as large as the other channel, the number of input tensor layers of the secondary channel is 4…(2 n times larger), the number of input tensor layers of the secondary channel is 4…(2 nIt is (+1). When processing YUV420 content, the size of the primary channel is 4 times the size of the secondary channel (2 times in each direction), N = 2, and the input tensor layer is 20. The number of output tensor layers is 4 (as required by the subsequent IDWT block) and is controlled by the size of the output convolutional block ("Convolution 2" in FIG. 11). When the number of input tensor layers and output tensor layers are different from each other, the "Select m" block shown in FIG. 11 extracts only the first m tensor layers from p (> m) input tensor layers (thus omitting the information derived from the auxiliary input) and sends them to the summation block at the output.

[0168] Several cascaded residual blocks (ResBlocks) are arranged between the input convolutional block Convolution 1 and the output convolutional block Convolution 2 and can be configured for residual learning during the training phase and neural network inference. According to one embodiment, 8 ResBlocks each having 48 features are used, although other combinations are possible. Rectified linear unit (ReLU) layers are used, for example, to implement activation functions. Batch normalization layers and scaling layers well-known in the context of CNNs may be used to facilitate the training process. According to some embodiments, such a scaling layer can apply a linear transformation to the preceding layer, for example, by multiplying each element of the preceding layer by a scaling factor. This can be achieved by a single scaling factor or by a layer having multiple identical weights. According to some embodiments, the scaling may be selectively applied, for example, using a scaling layer, to some elements of the preceding layer, and only the selected elements have weights that achieve the desired scaling value.

[0169] According to some embodiments, such a scaling layer can be represented by one or more scaling values used as a multiplier for the preceding block elements. Different arrangements of the scaling layer in the convolutional neural network layer configuration are contemplated. According to one configuration, the scaling layer can be arranged after the output convolutional block (“Convolution 2” in FIG. 11) as shown in FIG. 11 above. According to another configuration, the scaling layer can also be arranged before the output of each ResBlock, as shown in FIG. 11 below. The scaling layer can be adapted to control the ratio of the processed to the unprocessed input to be summed before output. According to some embodiments, one or more scaling values of the scaling layer can have default values obtained during the training process. According to some embodiments, the encoder can also use different values for a given image, image component, or image region. In such cases, the different values can be provided by signaling one or more scale values from the encoder to the decoder. In some embodiments, a configuration where a single scaling element is used for all elements of the preceding layer may be preferred. In these embodiments, signaling a single scale value per ResBlock allows for great flexibility in controlling the output of the neural network.

[0170] In particular, as illustrated in FIG. 12, a method for modifying an image region represented by two or more image channels is provided. Here, the term "modifying" refers to any modification such as image filtering or image correction. In principle, the image can be a patch of a predetermined size corresponding to a part of the image or parts of a plurality of images, or can be an image or a plurality of images. The two or more channels may be color channels, or other channels such as depth channels or multispectral image channels, or any other feature channels. One of the channels is a primary channel, and another one of the channels is a secondary channel. There may be a plurality of secondary channels. The method may include the step of selecting one of the two or more image channels as the primary channel and selecting at least another one of the two or more image channels as the secondary channel. It should be noted that the primary channel can (according to some embodiments) also be regarded as the leading channel. The secondary channel can (according to some embodiments) also be regarded as the response channel. For example, two or more secondary channels can be selected.

[0171] A method 1200 for modifying an image region represented by two or more image channels illustrated in FIG. 12 includes a step S1210 of processing a primary channel among the two or more image channels based on a first spatial frequency conversion to obtain a conversion primary channel. Similarly, a secondary channel among the two or more image channels different from the primary channel is processed based on a second spatial frequency conversion in step S1220 to obtain a conversion secondary channel. Needless to say, a plurality of secondary channels can be processed by a plurality of second spatial frequency conversions (and a plurality of second neural networks, see below). The first spatial frequency conversion and the second spatial frequency conversion can be selected from the group consisting of energy compression conversions including wavelet conversion (DWT or stationary wavelet conversion), discrete Fourier conversion, fast Fourier conversion, and discrete cosine conversion. According to an example, the first spatial frequency conversion and the second spatial frequency conversion can be the same type of spatial conversion among the group.

[0172] Furthermore, the method 1200 includes a step S1230 of processing the conversion primary channel by a first neural network to obtain a modified conversion primary channel, and a step S1240 of processing the conversion secondary channel based on the conversion primary channel (used as auxiliary information) by a second neural network to obtain a modified conversion secondary channel. The first neural network and the second neural network can be different from each other. When the primary channel and the secondary channel have different sizes (according to different resolutions of the image regions in different channels), cascade spatial frequency conversion can be used for the larger one of the channels to adjust the height dimension and the width dimension of the converted channels with respect to each other (see the description with reference to FIG. 10 above).

[0173] The output of the first neural network, i.e., the corrected conversion primary channel, is processed based on a first inverse spatial frequency conversion (corresponding to the first spatial frequency conversion) S1250 to obtain the corrected primary channel. Similarly, the corrected conversion secondary channel is processed based on a second inverse spatial frequency conversion S1260 to obtain the corrected secondary channel. Subsequently, the corrected image region can be obtained based on the corrected primary channel and the corrected secondary channel S1270. This procedure can be repeated for another selected image region. After all the image regions of the image to be processed have been processed by the method 1200 illustrated in FIG. 12, a corrected version of the entire image can be obtained.

[0174] Although the operations are shown in the drawings in a particular order, it should not be understood that such operations are required to be performed in the particular order or sequence shown, or that all illustrated operations be performed, to achieve the desired result. In certain circumstances, multitasking or parallel processing may be advantageous. Further, the separation of the various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and the described program components and systems may generally be integrated together into a single software product or packaged into multiple software products.

[0175] Particular embodiments of the subject matter have been described. There are other embodiments within the scope of the appended claims. For example, the acts recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order or sequence shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.

[0176] The method 1200 illustrated in FIG. 12 can be implemented in an apparatus 1300 configured to modify the image region illustrated in FIG. 13, and this apparatus 1300 can be configured to execute the steps of the method 1200 illustrated in FIG. 12. The apparatus 1300 can be constituted by an encoder (e.g., encoder 20 shown in FIGS. 14 and 15) or a decoder (e.g., decoder 20 shown in FIGS. 14 and 15), or can be included by the video coding device 8000 shown in FIG. 17 or the apparatus 9000 shown in FIG. 18.

[0177] The apparatus 1300 for modifying an image region represented by two or more image channels illustrated in FIG. 13 includes a first spatial frequency conversion unit 1310 configured to process a primary channel among the two or more image channels to obtain a conversion primary channel, and a second spatial frequency conversion unit 1320 configured to process a secondary channel among the two or more image channels different from the primary channel to obtain a conversion secondary channel.

[0178] Furthermore, the apparatus 1300 includes a first neural network (NN) 1330 configured to process the conversion primary channel to obtain a modified conversion primary channel, and a second neural network (NN) 1340 configured to process the conversion secondary channel based on the conversion primary channel to obtain a modified conversion secondary channel. The apparatus 1300 also includes a first inverse spatial frequency conversion unit 1350 configured to process the modified conversion primary channel to obtain a modified primary channel. Similarly, the apparatus 1300 also includes a second inverse spatial frequency conversion unit 1360 configured to process the modified conversion secondary channel to obtain a modified secondary channel.

[0179] Furthermore, the apparatus 1300 includes a combining unit 1370 configured to obtain a modified image region based on the modified primary channel and the modified secondary channel.

[0180] Some exemplary implementations in hardware and software A corresponding system that can place the above-described processing, in particular, by an encoder-decoder processing chain is illustrated in FIG. 14. FIG. 14 is a schematic block diagram illustrating an exemplary coding system that can utilize the technology of the present application, for example, a video, image, audio, and / or other coding system (or short coding system). The video encoder 20 (or short encoder 20) and the video decoder 30 (or short decoder 30) of the video coding system 10 represent examples of devices that can be configured to perform the techniques according to various examples described in the present application. For example, video coding and decoding may use a neural network that can apply the above-described bitstream parsing and / or bitstream generation to transmit feature maps between distributed (two or more) computing nodes.

[0181] As shown in FIG. 14, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, to decode encoded picture data 13.

[0182] The source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.

[0183] The picture source 16 may include, or be, any kind of picture capture device, such as a camera for capturing real-world pictures, and / or any kind of picture generation device, such as a computer graphics processor for generating computer animation pictures, or real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any kind of other device for acquiring and / or providing those or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may be any kind of memory or storage for storing any of the aforementioned pictures.

[0184] Distinguished from the processing performed by the pre-processor 18 and the pre-processing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.

[0185] The pre-processor 18 is configured to receive the (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain pre-processed picture 19 or pre-processed picture data 19. The pre-processing performed by the pre-processor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It will be understood that the pre-processing unit 18 may be an optional component. Note that the pre-processing may also use a neural network that uses presence indicator signaling (as shown in any of FIGS. 1 to 7).

[0186] The video encoder 20 is configured to receive the pre-processed picture data 19 and provide encoded picture data 21.

[0187] The communication interface 22 of the source device 12 receives the encoded picture data 21 and may be configured to transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0188] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally, i.e., optionally, include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0189] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof) directly, for example, from the source device 12 or from any other source, such as a storage device, e.g., an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.

[0190] The communication interfaces 22 and 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any type of network, such as a wired or wireless network or any combination thereof, or any type of private and public network, or any combination of any type thereof.

[0191] The communication interface 22 may be configured to, for example, package the encoded picture data 21 into a suitable format, such as a packet, and / or process the encoded picture data using any type of transmission encoding or processing for transmission via the communication link or communication network.

[0192] The communication interface 28 that forms the corresponding part of the communication interface 22 may be configured to receive, for example, the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or packaging to obtain the encoded picture data 21.

[0193] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow for the communication channel 13 in FIG. 14 that points from the source device 12 to the destination device 14. For example, it may be configured to send and receive messages, for example, set up a connection, check and exchange any other information regarding the communication link and / or data transmission, for example, the encoded picture data transmission. The decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture 31.

[0194] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain the post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31, for example, for display by the display device 34.

[0195] The display device 34 of the destination device 14 is configured to receive, for example, post - processed picture data 33 for displaying pictures to a user or a viewer. The display device 34 may be any kind of display for representing a reconstructed picture, for example, an integrated or external display or monitor, or may include the same. The display may include, for example, a liquid crystal display (LCD), an organic light - emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other kind of display.

[0196] FIG. 14 shows the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both or both functions, the source device 12 or corresponding functions and the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions and the destination device 14 or corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software or any combination thereof.

[0197] As will be apparent to those skilled in the art based on the description, the functions or the presence and (exact) partitioning of different units within the source device 12 and / or destination device 14 as shown in FIG. 14 may vary depending on the actual device and application.

[0198] The encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30) or both the encoder 20 and decoder 30 can be implemented via processing circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. The encoder 20 can be implemented via processing circuitry 46 to embody various modules including a neural network or a portion thereof. The decoder 30 can be implemented via processing circuitry 46 that embodies any of the coding systems or subsystems described herein. The processing circuitry can be configured to perform various operations as described hereinafter. If the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 and the video decoder 30 may be integrated as part of an encoder / decoder (codec) combined in a single device, for example, as shown in FIG. 15.

[0199] Source device 12 and destination device 14 may include any of a wide range of devices, such as any type of handheld or stationary device, e.g., a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video gaming console, a video streaming device (such as a content service server or a content delivery server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use an operating system, or may use any type of operating system. In some cases, source device 12 and destination device 14 may be equipped with wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.

[0200] In some cases, the video coding system 10 illustrated in FIG. 14 is merely an example, and the technology of the present application may be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory or streamed over a network. A video encoding device may encode data and store it in memory and / or a video decoding device may retrieve and decode data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data in memory and / or retrieve and decode data from memory.

[0201] FIG. 16 is a schematic diagram of a video coding device 8000 according to an embodiment of the present disclosure. The video coding device 8000 is suitable for implementing the embodiments of the disclosure described herein. In one embodiment, the video coding device 8000 can be a decoder such as the video decoder 30 of FIG. 14, or an encoder such as the video encoder 20 of FIG. 14.

[0202] The video coding device 8000 includes a receiving port 8010 (or input port 8010) and a receiver unit (Rx) 8020 for receiving data, a processor, logic unit, or central processing unit (CPU) 8030 for processing data, a transmitter unit (Tx) 8040 and a transmitting port 8050 (or output port 8050) for transmitting data, and a memory 8060 for storing data. The video coding device 8000 may also include optical - electrical (OE) components and electro - optical (EO) components coupled to the receiving port 8010, the receiver unit 8020, the transmitter unit 8040, and the transmitting port 8050 for transmitting or receiving optical or electrical signals.

[0203] Processor 8030 is implemented by hardware and software. Processor 8030 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. Processor 8030 communicates with a reception port 8010, a receiver unit 8020, a transmitter unit 8040, a transmission port 8050, and a memory 8060. Processor 8030 includes a neural network-based codec 8070. The neural network-based codec 8070 implements the embodiments of the disclosure described above. For example, the neural network-based codec 8070 implements, processes, prepares, or provides various coding operations. Therefore, including the neural network-based codec 8070 provides a substantial improvement to the functionality of the video coding device 8000 and results in the conversion of the video coding device 8000 to different states. Alternatively, the neural network-based codec 8070 is stored in the memory 8060 and implemented as instructions executed by the processor 8030.

[0204] The memory 8060 may include one or more disks, tape drives, and solid state drives, and be used as an overflow data storage device that stores such programs when selected for execution and stores instructions and data read during program execution. The memory 8060 may be, for example, volatile and / or non-volatile, and may be a read-only memory (ROM), a random access memory (RAM), a ternary content addressable memory (TCAM), and / or a static random access memory (SRAM).

[0205] FIG. 17 is a simplified block diagram of an apparatus that can be used as either or both of the source device 12 and the destination device 14 from FIG. 14 according to an exemplary embodiment.

[0206] The processor 9002 within the apparatus 9000 may be a central processing unit. Alternatively, the processor 9002 may be any other type of device, or multiple devices, that can manipulate or process information that exists or will be developed in the future. The disclosed implementations can be carried out using a single processor as illustrated, e.g., processor 9002, but advantages in terms of speed and efficiency can be achieved using multiple processors.

[0207] In one implementation, the memory 9004 within the apparatus 9000 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 9004. The memory 9004 can include code and data 9006 that are accessed by the processor 9002 using the bus 9012. The memory 9004 can further include an operating system 9008 and an application program 9010, and the application program 9010 includes at least one program that enables the processor 9002 to perform the methods described herein. For example, the application program 9010 can include applications 1 through N, which further include video coding applications that perform the methods described herein.

[0208] The apparatus 9000 can also include one or more output devices, such as a display 9018. In one example, the display 9018 can be a touch-sensitive display that combines a touch-sensitive element operable to sense touch input with a display. The display 9018 can be coupled to the processor 9002 via the bus 9012.

[0209] Although it is shown here as a single bus, the bus 9012 of the device 9000 can be composed of a plurality of buses. Further, the secondary storage can be directly connected to other components of the device 9000 or can be accessed via a network, and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Therefore, the device 9000 can be implemented in a wide variety of configurations.

Explanation of Signs

[0210] 11 Portion of the input image 210 Encoder side 220 Encoder sub-network 230 Representation (encoding) of the dataset x 250 Decoder side 260 Decoder sub-network 101 Encoder 102 Quantizer 103 Hyper-encoder 104 Decoder 105 Arithmetic encoder 106 Arithmetic decoder 107 Hyper-decoder 108 Quantizer 109 Arithmetic encoder 110 Arithmetic decoder 121 Encoder 122 Quantizer 123 Hyper-encoder 125 Arithmetic encoding module 127 Hyper-decoder 144 Decoder 147 Hyper-decoder 401 Downsampling layer 402 Downsampling layer 403 Downsampling layer 404 Downsampling layer 405 Downsampling layer 406 Downsampling layer 407 Upsampling Layer 408 Upsampling Layer 409 Upsampling Layer 410 Upsampling Layer 411 Upsampling Layer 412 Upsampling Layer 413 Component 414 Input Image 415 Component 420 Further Layer 430 Corresponding Convolutional Layer 510 Mobile Side 520 Quantization Layer 550 Hidden Layer (Deep Feature Map) 560 Inverse Quantization Layer 590 Cloud Side 810 Select Patch 820 Pixel Shift 830 Concatenate 840 Concatenate 850 Concatenate 860 Pixel Anti-Shift 910 Select Patch 920 Discrete Wavelet Transform (DWT) 930 DWT 940 DWT 950 Concatenate 960 Concatenate 970 Inverse DWT 980 Inverse DWT 990 Inverse DWT 1010 Select Patch 1020 DWT 1030 DWT 1040 DWT 1050 DWT 1060 Concatenate 1070 Concatenate 1080 Inverse DWT 1090 Inverse DWT 1100 Inverse DWT 1200 Method for Modifying Image Regions Represented by Two or More Image Channels 1300 An apparatus for modifying an image region represented by two or more image channels 1310 First spatial frequency conversion unit 1320 Second spatial frequency conversion unit 1330 First neural network 1340 Second neural network 1350 First inverse spatial frequency conversion unit 1360 Second inverse spatial frequency conversion unit 1370 Combining unit 10 Video coding system 12 Source device 13 Communication channel 14 Destination device 16 Picture source 17 Picture data 18 Preprocessor 19 Preprocessed picture data 20 Encoder / Video encoder 21 Encoded picture data 22 Communication interface 28 Communication interface 30 Decoder / Video decoder 31 Decoded picture data 32 Postprocessor 33 Postprocessed picture data 34 Display device 40 Video coding system 41 (One or more) imaging devices 42 Antenna 43 (One or more) processors 44 (One or more) memory stores 45 Display device 46 Processing circuit 8000 Video coding device 8010 Receiving port 8020 Receiver unit (Rx) 8030 Processor 8040 Transmitter Unit (Tx) 8050 Transmission Port 8060 Memory 8070 Neural Network-based Codec 9000 Device 9002 (One or more) Processors 9004 Memory 9006 Data 9008 Operating System 9010 Application Program 9012 Bus 9018 Display

Claims

1. A method for modifying an image region represented by two or more image channels, the method comprising: processing a primary channel among the two or more image channels based on a first spatial frequency transform to obtain a transformed primary channel (S1210); processing a secondary channel among the two or more image channels different from the primary channel based on a second spatial frequency transform to obtain a transformed secondary channel (S1220); processing the transformed primary channel by a first neural network to obtain a modified transformed primary channel (S1230); processing the transformed secondary channel based on the transformed primary channel by a second neural network to obtain a modified transformed secondary channel (S1240); processing the modified transformed primary channel based on a first inverse spatial frequency transform to obtain a modified primary channel (S1250); processing the modified transformed secondary channel based on a second inverse spatial frequency transform to obtain a modified secondary channel (S1260); and obtaining a modified image region based on the modified primary channel and the modified secondary channel (S1270). A method comprising the above steps.

2. The method according to claim 1, wherein one or both of the first spatial frequency transform and the second spatial frequency transform are selected from the group consisting of energy compression transforms including wavelet transform, discrete Fourier transform, fast Fourier transform, and discrete cosine transform.

3. The method according to claim 2, wherein both the first spatial frequency transform and the second spatial frequency transform are one of wavelet transform, discrete Fourier transform, fast Fourier transform, energy compression transform, and discrete cosine transform.

4. The method according to claim 2, wherein one or both of the first spatial frequency transform and the second spatial frequency transform are a wavelet transform selected from the group consisting of discrete wavelet transform and stationary wavelet transform.

5. The method according to claim 1, further comprising the step of selecting the primary channel from the two or more image channels.

6. The method according to claim 5, further comprising the step of selecting the secondary channel from the two or more image channels.

7. The method according to claim 6, wherein the primary channel and the secondary channel are selected from the two or more image channels based on an output of a classifier operating based on another neural network.

8. The processing (S1240) of the transformed secondary channel based on the transformed primary channel includes concatenating (950, 960) a second three-dimensional tensor representing the transformed secondary channel with a first three-dimensional tensor representing the transformed primary channel, according to the method of claim 1.

9. The method according to claim 1, wherein the size of the primary channel is different from the size of the secondary channel.

10. a) When the size of the primary channel is larger than the size of the secondary channel, processing the transformed primary channel (S1230) based on at least one additional first spatial frequency transformation to obtain an auxiliary transformed primary channel of the same size as the transformed secondary channel in the height and width directions of the image region, the processing (S1240) of the transformed secondary channel is based on the auxiliary transformed primary channel, b) When the size of the secondary channel is larger than the size of the primary channel, processing the transformed secondary channel (S1240) based on at least one additional second spatial frequency transformation to obtain an auxiliary transformed secondary channel of the same size as the transformed primary channel in the height and width directions of the image region, the processing (S1240) of the transformed secondary channel includes processing the auxiliary transformed secondary channel based on the transformed primary channel, The method according to claim 9.

11. The processing (S1240) of the transformed secondary channel based on the transformed primary channel is a) When the size of the primary channel is larger than the size of the secondary channel, the step of concatenating (1060, 1070) a second three-dimensional tensor representing the transformed secondary channel with a first three-dimensional tensor representing the auxiliary transformed primary channel, b) When the size of the secondary channel is larger than the size of the primary channel, concatenating a second three-dimensional tensor representing the auxiliary conversion secondary channel with a first three-dimensional tensor representing the conversion primary channel; The method according to claim 10, comprising: **Claim 12** The method according to claim 1, wherein the image region is a square region in the height dimension and the width dimension of the image region. **Claim 13** The method according to claim 1, further comprising dividing an image into an image region including the image region, and padding an image region resulting from the division that is not square in the height dimension and the width dimension of the image region so that the image region becomes square in the height dimension and the width dimension of the image region. **Claim 14** dividing an image into an image region including the image region; padding the image so that, when the image cannot be divided only into image regions that are square in the height dimension and the width dimension of the image region including the image region, the image is divided only into image regions that are all square in the height dimension and the width dimension of the image region including the image region; The method according to claim 1, further comprising: **Claim 15** The method according to claim 1, wherein the first neural network and the second neural network operate independently of each other. **Claim 16** The method according to claim 15, wherein the weights of one of the first neural network and the second neural network are determined and used independently of the weights of the other of the first neural network and the second neural network. **Claim 17** The method according to claim 1, wherein each of the first neural network and the second neural network is a convolutional neural network or includes a convolutional neural network. **Claim 18** The method according to claim 17, wherein each of the convolutional neural networks includes at least one residual network component. **Claim 19** The method according to claim 17, wherein one or more of the convolutional neural networks use a scaling layer represented by one or more scaling values. **Claim 20** The method according to claim 19, wherein the scaling layer is adapted to signal the one or more scaling values. **Claim 21** The method according to claim 1, wherein the two or more image channels include color channels and / or feature channels. **Claim 22** The image region is a patch of a predetermined size corresponding to a part of an image or parts of a plurality of images, or an image or a plurality of images and is one of them. The method according to claim 1. **Claim 23** obtaining an original image region; encoding the obtained image region into a bitstream; applying the method according to any one of claims 1 to 22 to correct the image region obtained by reconstructing the encoded image region and including a method for encoding an image or a video sequence of images. **Claim 24** Obtaining an original image region; encoding the obtained image region into a bitstream; applying the method according to any one of claims 5 to 7 to correct the image region obtained by reconstructing the encoded image region including the step of including an indication of the selected primary channel in the bitstream and including a method for encoding an image or a video sequence of images. **Claim 25** The method for encoding an image or a video sequence of images according to claim 23, including the step of including in the bitstream an adaptation of one or more weights of at least one of the first neural network and the second neural network. **Claim 26** Obtaining an original image region; encoding the obtained image region into a bitstream; applying the method according to any one of claims 5 to 7 to correct the image region obtained by reconstructing the encoded image region including the step of including in the bitstream an adaptation of one or more weights of at least one of the first neural network and the second neural network; obtaining a plurality of image regions; applying the method for correcting the obtained image region individually to the image regions of the plurality of obtained image regions To each of the bitstreams of the plurality of image regions, an indication indicating that the method for modifying the acquired image region should not be applied to the image region, an indication of the selected primary channel of the image region, adapting at least one of the one or more weights of at least one of the first neural network and the second neural network including at least one of the steps of A method for encoding an image or a video sequence of images, comprising: **Claim 27** A method for decoding an image or a video sequence of images from a bitstream, comprising: reconstructing an image region from the bitstream, applying the method according to claim 1 to modify the image region, and A method comprising: **Claim 28** A method for decoding an image or a video sequence of images from a bitstream, comprising: an indication indicating that the method for modifying the acquired image region should not be applied to the image region, an indication of the primary channel of the image region, adapting at least one of the one or more weights of at least one of the first neural network and the second neural network parsing the bitstream to obtain at least one of the steps of reconstructing an image region from the bitstream, when the indication indicates a selected primary channel, modifying the reconstructed image region according to the method according to any one of claims 1 to 22 using the indicated primary channel as the selected primary channel, and A method comprising: **Claim 29** When the adaptation of the weights of at least one of the first neural network and the second neural network is present in the bitstream, modifying the weights of the first and second neural networks accordingly A method for decoding an image or a video sequence of images from a bitstream according to claim 28, further comprising: **Claim 30** The method for decoding an image or a video sequence of images from a bitstream according to claim 27, wherein the modification of the image region is applied by an in-loop filter or a post-processing filter. **Claim 31**: A computer program which, when executed on one or more processors, performs the method according to any one of claims 1 to 22. **Claim 32** An apparatus for modifying an image region represented by two or more image channels, the apparatus including a circuit configured to perform the steps according to the method of any one of claims 1 to 22. **Claim 33** An apparatus (1300) for modifying an image region represented by two or more image channels, a first spatial frequency conversion unit (1310) configured to process a primary channel among the two or more image channels to obtain a conversion primary channel; a second spatial frequency conversion unit (1320) configured to process a secondary channel among the two or more image channels different from the primary channel to obtain a conversion secondary channel; a first neural network (1330) configured to process the conversion primary channel to obtain a modified conversion primary channel; a second neural network (1340) configured to process the conversion secondary channel based on the conversion primary channel to obtain a modified conversion secondary channel; a first inverse spatial frequency conversion unit (1350) configured to process the modified conversion primary channel to obtain a modified primary channel; a second inverse spatial frequency conversion unit (1360) configured to process the modified conversion secondary channel to obtain a modified secondary channel; and a combining unit (1370) configured to obtain a modified image region based on the modified primary channel and the modified secondary channel The apparatus (1300) includes. **Claim 34** An encoder for encoding an image or a video sequence or an image, the encoder comprising: an input module for obtaining an original image region; a compression module for encoding the obtained image region into a bitstream; a reconstruction module for reconstructing the encoded image region An apparatus for modifying an image region represented by two or more image channels, comprising a circuit configured to perform the steps of the method according to any one of claims 1 to 22, or an apparatus (1300) for modifying an image region represented by two or more image channels according to claim 33, and An encoder comprising the same. **Claim 35** A decoder for decoding an image or a video sequence or an image from a bitstream, the decoder comprising: A reconstruction module for reconstructing an image region from the bitstream; and An apparatus for modifying an image region represented by two or more image channels, comprising a circuit configured to perform the steps of the method according to any one of claims 1 to 22, or an apparatus (1300) for modifying an image region represented by two or more image channels according to claim 33, and A decoder comprising the same.

Citation Information

Patent Citations

  • Adaptive image enhancement using inter-channel correlation information

    WO2021249684A1