Visual data processing method, device and medium
By including indicators in the bitstream and sharing processing parameters of multiple components of visual data, the problem of redundancy in parameter transmission during processing of multiple components in the prior art is solved, and the encoding and decoding efficiency and compression performance are improved.
Patent Information
- Application Number
- CN202380074019.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-10-19
AI Technical Summary
Existing neural network-based image/video codec techniques still have challenges in improving codec quality and efficiency, especially when processing multiple components, the parameters need to be transmitted separately, resulting in increased bit rate and reduced compression performance.
A method is proposed to avoid transmission of parameters separately by including a first indication in the bitstream indicating whether a set of values for a set of parameters for a neural network-based model is common to the processing of multiple components of the visual data.
This method improves encoding and decoding efficiency, reduces the edge information required to process multiple components of visual data, and thus improves compression performance.
Smart Images

Figure CN120077655A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to visual data processing technologies, and more particularly, to neural network-based visual data encoding and decoding. Background Art
[0002] The past decade has witnessed the rapid development of deep learning in various fields, especially in computer vision and image processing. Neural networks were initially invented with the interdisciplinary research of neuroscience and mathematics. It has demonstrated powerful capabilities in the context of non-linear transformation and classification. Neural network-based image / video compression technologies have made significant progress in the past five years. It is reported that the latest neural network-based image compression algorithms achieve rate-distortion (R-D) performance comparable to that of Versatile Video Coding (VVC). As the performance of neural image compression continues to improve, neural network-based video compression has become an actively developed research area. However, there is generally a desire to further improve the encoding and decoding quality and efficiency of neural network-based image / video encoding and decoding. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method includes: performing conversion between visual data and a bitstream of the visual data by using a neural network (NN)-based model, where the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing multiple components of the visual data.
[0005] According to the method of the first aspect of the present disclosure, the indication is included in the bitstream and is used to indicate whether a set of values for a set of parameters of the NN-based model is common for processing multiple components of the visual data. By means of this indication, it is possible to avoid separately transmitting the set of values for multiple components of the visual data through a signal. Therefore, the proposed method can advantageously improve the encoding and decoding efficiency.
[0006] In a second aspect, a device for visual data processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for visual data processing of visual data. The method includes: performing a conversion between visual data and the bitstream using a neural network (NN)-based model, where the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing multiple components of the visual data.
[0009] In a fifth aspect, a method for storing a bitstream of visual data is proposed. The method includes: performing a conversion between visual data and the bitstream using a neural network (NN)-based model, where the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing multiple components of the visual data; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram showing an exemplary visual data codec system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 A typical transform coding scheme is shown;
[0014] Figure 3 An image from the Kodak dataset and different representations of the image are shown;
[0015] Figure 4 The network architecture of an autoencoder implementing a hyperprior model is shown;
[0016] Figure 5 A block diagram of a combined model is shown;
[0017] Figure 6 The encoding process of the combined model is shown;
[0018] Figure 7 The decoding process of the combined model is shown;
[0019] Figure 8 illustrates an example of a decoding process according to some embodiments of the present disclosure;
[0020] Figure 9 illustrates a flowchart of a method for visual data processing according to an embodiment of the present disclosure; and
[0021] Figure 10 illustrates a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.
[0022] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0023] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.
[0024] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0025] As used herein, the terms "one embodiment", "embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment includes that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0026] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0027] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example environment
[0028] Figure 1 is a block diagram showing an example visual data encoding and decoding system 100 that can utilize the techniques of the present disclosure. As shown, the visual data encoding and decoding system 100 can include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 can be configured to generate encoded visual data, and the destination device 120 can be configured to decode the encoded visual data generated by the source device 110. The source device 110 can include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.
[0029] The visual data source 112 can include sources such as visual data acquisition devices. Examples of visual data acquisition devices include, but are not limited to, an interface for receiving visual data from a visual data provider, a computer graphics system for generating visual data, and / or combinations thereof.
[0030] The visual data can include one or more pictures or one or more images of a video. The visual data encoder 114 encodes the visual data from the visual data source 112 to generate a bitstream. The bitstream can include a sequence of bits forming an encoded representation of the visual data. The bitstream can include encoded pictures and associated visual data. An encoded picture is an encoded representation of a picture. The associated visual data can include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded visual data can be directly transmitted to the destination device 120 via the I / O interface 116 through the network 130A. The encoded visual data can also be stored on a storage medium / server 130B for access by the destination device 120.
[0031] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded visual data from a source device 110 or a storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0032] The visual data encoder 114 and the visual data decoder 124 may operate according to a visual data encoding / decoding standard (such as a video encoding / decoding standard or a still picture encoding / decoding standard, as well as other existing and / or future standards).
[0033] Some exemplary embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific visual data codecs, the disclosed techniques are also applicable to other coding techniques. In addition, although some embodiments describe the encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. In addition, the term visual data processing includes visual data encoding or compression, visual data decoding or decompression, and visual data transcoding, in which visual data is represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview A neural network-based image and video compression method includes separate processing of color components of an image, where control parameters used for processing one component are also used for another component. 2. Introduction The past decade has witnessed the rapid development of deep learning in various fields, especially in computer vision and image processing. From the great success of deep learning techniques in the field of computer vision, many researchers have shifted their attention from traditional image / video compression techniques to neural image / video compression techniques. Neural networks were originally invented along with the interdisciplinary research of neuroscience and mathematics. It has demonstrated powerful capabilities in the context of non-linear transformation and classification. Neural network-based image / video compression techniques have made significant progress in the past five years. It is reported that the latest neural network-based image compression algorithm achieves comparable R-D performance to Versatile Video Coding (VVC), which is the latest video coding standard developed by the Joint Video Exploration Team (JVET) of experts with MPEG and VCEG. As the performance of neural image compression continues to improve, neural network-based video compression has become an actively developed research area. However, due to the inherent difficulties of the problem, neural network-based video coding and decoding is still in its infancy. 2.1. Image / Video Compression Image / video compression generally refers to the computational technique of compressing an image / video into binary codes for easy storage and transmission. The binary codes may or may not support lossless reconstruction of the original image / video, known as lossless compression and lossy compression. Most efforts are devoted to lossy compression because lossless reconstruction is not required in most scenarios. Generally, the performance of an image / video compression algorithm is evaluated from two aspects, i.e., compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes, and the fewer the number of binary codes, the better; the reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, and the higher the reconstruction quality, the better. Image / video compression techniques can be divided into two branches, classical video coding and decoding methods and neural network-based video compression methods. Classical video coding and decoding schemes adopt transformation-based solutions, where researchers utilize the statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully manually designing entropy codes to model the dependencies in the quantization domain. Neural network-based video compression has two forms: neural network-based coding and decoding tools and neural network-based video compression based on end-to-end neural networks. The former is embedded in existing classical video codecs as coding and decoding tools and is only used as part of the framework, while the latter is a separate framework developed based on neural networks without relying on classical video codecs. In the past three decades, a series of classical video coding and decoding standards have been developed to adapt to the increasing visual content. The International Organization for Standardization ISO / IEC has two expert groups, namely the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), and ITU-T also has its own Video Coding Experts Group (VCEG) for standardizing image / video coding and decoding technologies. Influential video coding and decoding standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. After H.265 / HEVC, the Joint Video Exploration Team (JVET) formed by MPEG and VCEG has been working on a new video coding and decoding standard, Versatile Video Coding (VVC). The first version of VVC was released in July 2020. Compared with HEVC, VVC is reported to have an average 50% bitrate reduction at the same visual quality. Image / video compression based on neural networks is not a new invention as there are many researchers working on neural network-based image coding and decoding. However, the network architectures are relatively shallow and the performance is not satisfactory. Benefiting from the support of large amounts of data and powerful computing resources, neural network-based methods are better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, the technology is still far from mature and there are many challenges to be solved. 2.2. Neural Networks A neural network (also known as an artificial neural network (ANN)) is a computational model used in machine learning techniques, which usually consists of multiple processing layers, and each layer consists of multiple simple but non-linear basic computational units. One benefit of such a deep network is considered to be the ability to process data with multi-level abstractions and transform the data into different kinds of representations. Note that these representations are not manually designed; instead, the deep network including the processing layers learns from massive data using general machine learning processes. Deep learning eliminates the need for manual representations and is therefore considered useful, especially for processing native unstructured data such as acoustic and visual signals, while processing such data has been a long-term difficulty in the field of artificial intelligence. 2.3. Neural Networks for Image Compression Existing neural networks for image compression methods can be classified into two categories, namely pixel probability modeling and autoencoders. The former belongs to the prediction coding and decoding strategy, while the latter is a transform-based solution. Sometimes, these two methods are combined together in the literature. 2.3.1. Pixel Probability Modeling According to Shannon information theory, the best method for lossless coding and decoding can achieve the minimum coding rate log 2p(x), where p(x) is the probability of symbol x. Many lossless encoding and decoding methods have been developed in the literature, and among these lossless encoding and decoding methods, arithmetic encoding and decoding is considered to be the best. Given the probability distribution p(x), arithmetic encoding and decoding ensures that the coding rate is as close as possible to its theoretical limit — log 2 p(x). Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to dimensionality reduction. Following the predictive coding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is the image. p(x) = p(x 1 )p(x 2 |x 1 )...p(x i |x 1 ,..., x i-1 )...p(x m×n |x 1 ,..., x m×n-1 ) (1) where m and n are the height and width of the image respectively. The previous observations are also called the context of the current pixel. When the image is large, it may be difficult to estimate the conditional probability, so a simplified method is to limit the scope of its context. p(x) = p(x 1 )p(x 2 |x 1 )...p(x i |x i-k ,..., x i-1 )...p(x m×n |x m×n-k ,..., x m×n-1 ) (2) where k is a predefined constant that controls the context scope. It should be noted that the condition can also consider the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on the previously encoded and decoded pixels (including R / G / B samples), the current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample, and when encoding and decoding the current B sample, the previously encoded and decoded pixels and the current R sample and the current G sample can also be considered. Neural networks were initially introduced for computer vision tasks and have been proven to be effective in regression and classification problems. Therefore, it has been proposed to use neural networks to estimate the probability p(x 1 ,n 2 ,…,x i-1 given its context x i)。In the existing design, pixel probabilities are proposed for binary images (i.e., x i ∈{-1, +1}). The Neural Autoregressive Distribution Estimator (NADF) is designed for pixel probability modeling, which is a feed-forward network with a single hidden layer. Similar work is presented in the existing design, where the feed-forward network also has connections that skip the hidden layer and the parameters are also shared. Experiments are performed on the binarized MNIST dataset. In the existing design, NADF is extended to the real-valued model RNADE, where the probability p(x i |x 1 ,…,x i-1 ) is derived through a mixture of Gaussians. Their feed-forward network also has a single hidden layer, but the hidden layer is rescaled to avoid saturation and uses the Rectified Linear Unit (ReLU) instead of the sigmoid. In the existing design, NADF and RNADE are improved by reordering the pixels and using deeper neural networks. Most of the above methods directly model the probability distribution in the pixel domain. Some researchers also attempt to model the probability distribution as a condition of either an explicit representation or a latent representation. That is, it can be estimated: where h is the additional condition, and p(x) = p(h)p(x|h) means that the modeling is split into an unconditional one and a conditional one. The additional condition can be image label information or a high-level representation. 2.3.2. Autoencoders Autoencoders originate from the existing design. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part converts the high-dimensional input signal into a low-dimensional representation, usually with a reduced spatial size but a larger number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders achieve automatic learning of the representation and eliminate the need for handcrafted features, which is also considered one of the most important advantages of neural networks. Figure 2 A typical transform coding-decoding scheme is shown. The original image x is transformed by the analysis network g a to achieve the latent representation y. The latent representation y is quantized and compressed into bits. The number of bits R is used to measure the coding rate. Then the quantized latent representation is inverse-transformed by the synthesis network g s to obtain the reconstructed image The distortion is calculated in the perceptual space by transforming x and using the function g p . Applying an autoencoder network to lossy compression is intuitive. It only requires encoding the learned latent representation from a trained neural network. However, adapting the autoencoder to image compression is not straightforward because the original autoencoder is not optimized for compression, and thus it is not efficient to directly use the trained autoencoder. Additionally, there are other major challenges: First, the low-dimensional representation should be quantized before being encoded, but quantization is not differentiable, which is required in backpropagation when training a neural network. Second, the objectives in the compression scenario are different because both distortion and bitrate need to be considered. Estimating the bitrate is challenging. Third, practical image codec schemes need to support variable bitrate, scalability, encoding speed / decoding speed, and interoperability. In response to these challenges, many researchers have actively contributed to this field. A prototype autoencoder for image compression is in Figure 2 which can be regarded as a transform coding strategy. The original image x is transformed using the analysis network y = g a (x), where y is the latent representation to be quantized and coded. The synthesis network will inverse-transform the quantized latent representation to obtain the reconstructed image The framework is trained using a distortion loss function, i.e., where D is the distortion between x and , R is the bitrate calculated or estimated from the quantized representation , and λ is the Lagrange multiplier. It should be noted that D can be calculated in the pixel domain or the perceptual domain. All existing research works follow this prototype, and the differences may only be in the network structure or the loss function. 2.3.3. Hyperprior Model In the transform coding method for image compression, the encoder subnetwork (Section 2.3.2) uses a parametric analysis transform to transform the image vector x into the latent representation y, which is then quantized to form Since are discrete values, they can be losslessly compressed using entropy coding techniques (e.g., arithmetic coding) and transmitted as a bit sequence. As is obvious from the left-middle and right-middle images of Figure 3 , there is significant spatial dependence among the elements of . It is worth noting that their scales (right-middle image) seem to be spatially coupled. In existing designs, another set of random variables is introduced to capture the spatial dependence and further reduce redundancy. In this case, the image compression network is shown in Figure 4 . In Figure 4 , the left hand side of the model is the encoder ga and decoder g s (explained in Section 2.3.2). The right - hand side is for obtaining an additional encoder h that exploits hyper - prior information a and decoder h that exploits hyper - prior information s network. In this architecture, the encoder feeds the input image x to g a to produce a response y with spatially varying standard deviation. The response y is fed into h a to aggregate the standard - deviation distribution in z. Then z is quantized compressed and transmitted as side information. Then, the encoder uses the quantized vector to estimate σ, the spatial distribution of the standard deviation, and uses it to compress and transmit the quantized image representation The decoder first recovers from the compressed signal and then it uses h s to obtain σ, which provides it with the correct probability estimate to also successfully recover and then feeds it to g s to obtain the reconstructed image. When an encoder that exploits hyper - prior information and a decoder that exploits hyper - prior information are added to the image - compression network, the spatial redundancy of the quantized latent values is reduced. The right - most image in Figure 3 corresponds to the quantized latent values when using an encoder / decoder that exploits hyper - prior information. Compared with the middle - right image, the spatial redundancy is significantly reduced because the samples of the quantized latent values are less correlated. Figure 3 shows an image from the Kodak dataset and different representations of the image. Figure 3 The left - most image in Figure 3 shows an image from the Kodak dataset. Figure 3 The middle - left image in Figure 3 shows a visualization of the latent representation y of the image. Figure 4 shows the network architecture of an auto - encoder that implements the hyper - prior model. The left - hand side shows the image auto - encoder network and the right - hand side corresponds to the hyper - prior sub - network. The analysis and synthesis transforms are denoted as g a and g a . Q represents quantization, AE and AD represent the arithmetic encoder and arithmetic decoder respectively. The hyper - prior model consists of two sub - networks, an encoder that exploits hyper - prior information (denoted as ha represented) and a decoder that utilizes hyperprior information (denoted by h s The hyperprior model generates quantized latent values of the hyperprior information which includes information about the probability distribution of the samples of the quantized latent values and is included in the bitstream and transmitted to the receiver (decoder) together with 2.3.4. Context Model Although the hyperprior model improves the modeling of the probability distribution of the quantized latent values additional improvements can be obtained by utilizing an autoregressive model that predicts the quantized latent values from the causal context (context model) of the latent values. The term autoregressive means that the output of the process is later used as its input. For example, the context model subnetwork generates a sample of the latent value, which is later used as input to obtain the next sample. In existing designs, a joint architecture is utilized, in which both a hyperprior model subnetwork (an encoder that utilizes hyperprior information and a decoder that utilizes hyperprior information) and a context model subnetwork are used. The hyperprior and context models are combined to learn the probability model of the quantized latent values and then used for entropy coding and decoding. As shown, the outputs of the context subnetwork and the decoder subnetwork that utilizes hyperprior information are referred to as the subnetwork combination of the entropy parameters, which generates the mean μ and scale (or variance) σ parameters for a Gaussian probability model. The Gaussian probability model is then used to encode the samples of the quantized latent values into the bitstream by means of an arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized latent values from the bitstream by means of an arithmetic decoder (AD) module Figure 5 Figure 5 A block diagram of the combined model is shown. The combined model jointly optimizes the autoregressive component, which, together with the hyperprior and the underlying autoencoder, estimates the probability distribution of the latent values from the causal context (context model) of the latent values. The real-valued latent representation is quantized (Q) to create the quantized latent values and the quantized latent values of the hyperprior information which are compressed into the bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted regions correspond to the components performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream. Typically, the latent samples are modeled as a Gaussian distribution or a Gaussian mixture model (not limited to this). In existing designs, according to Figure 5 , the context model and the hyperprior are jointly used to estimate the probability distribution of the latent samples. Since a Gaussian distribution can be defined by the mean and variance (also known as sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ). 2.3.5. Gain Variational Autoencoder (G-VAE) Typically, neural network-based image / video compression methods require training multiple models to adapt to different bitrates. The Gain Variational Autoencoder (G-VAE) is a variational autoencoder with a pair of gain units. The G-VAE is designed to achieve continuous variable bitrate adaptation using a single model. The G-VAE includes a pair of gain units, which are typically inserted between the output of the encoder and the input of the decoder. The output of the encoder is defined as the latent representation y ∈ R c*h*w , where c, h, w represent the number of channels, the height, and the width of the latent representation. Each channel of the latent representation is denoted as y (i) ∈ R h*w , where i = 0, 1, …, c - 1. A pair of gain units includes a gain matrix M ∈ R c*n and an inverse gain matrix, where n is the number of gain vectors. The gain vectors can be denoted as m s = {α s(0) , α s(1) , …, α s(c-1)}, α s(i) ∈ R, where s represents the index of the gain vector in the gain matrix. The motivation for the gain matrix is similar to the quantization table in JPEG that controls the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector. where ⊙ is the element-wise multiplication, i.e., and α s(i) is the i-th gain value in the gain vector m s . The inverse gain matrix used on the decoder side can be denoted as M' ∈ R c*n , which consists of n inverse gain vectors, i.e., M' = {δ s(0) , δ s(1) , …, δ s(c-1)}, δ s(i) ∈ R. The inverse gain process is denoted as: where is the quantized latent representation after decoding, and y' s is the quantized latent representation of the inverse gain to be fed into the synthesis network. To achieve continuous variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors {m t , m′ t} and {m r , m′ r}, the interpolated gain vectors can be obtained via the following equations. m v = [(m r ) l ·(m t ) 1-l m′ v = [(m′ r ) l ·(m′ t ) 1-l where l ∈ R is the interpolation coefficient that controls the corresponding bit rate of the generated pair of gain vectors. Since l is a real number, any bit rate between two given pairs of gain vectors can be achieved. 2.3.6. Encoding Process Using the Joint Autoregressive Hyperprior Model Figure 5 Corresponds to the state of the technical compression method proposed in the existing design. In this section and the next section, the encoding process and the decoding process will be described respectively. Figure 6 Depicts the encoding process. The input image is first processed using the encoder sub-network. The encoder transforms the input image into a transformed representation called the latent value, denoted by y. Then y is input into the quantizer block denoted by Q to obtain the quantized latent value Then is converted into a bitstream (bits1) using the arithmetic coding module (denoted as AE). The arithmetic coding block sequentially converts each sample point of into the bitstream (bits1). The module using the encoder of the hyperprior information, the context, the decoder using the hyperprior information, and the entropy parameter sub-network are used to estimate the probability distribution of the sample points of the quantized latent value The latent value y is input into the encoder using the hyperprior information, and the encoder using the hyperprior information outputs the hyperprior information latent value (denoted by z). Then the hyperprior information latent value is quantized and a second bitstream (bits2) is generated using the arithmetic coding (AE) module. The factored entropy module generates a probability distribution that is used to encode the quantized hyperprior information latent value into the bitstream. The quantized hyperprior information latent value includes information about the probability distribution of the quantized latent value . The entropy parameter sub-network generates for encoding the quantized latent value Probability distribution estimation. The information generated by the entropy parameter generally includes the mean μ and the scale (or variance) σ parameters of the Gaussian probability distribution, which are used together to obtain the Gaussian probability distribution. The Gaussian distribution of the random variable x is defined as where the parameter μ is the mean or expectation of the distribution (as well as its median and mode), and the parameter σ is its standard deviation (or variance or scale). To define the Gaussian distribution, the mean and variance need to be determined. In the existing design, the entropy parameter module is used to estimate the mean and variance values. The sub-network uses the decoder of the hyper-prior information to generate a part of the information used by the entropy parameter sub-network, and the other part of the information is generated by the autoregressive module called the context module. The context module uses the samples that have been encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the samples of the quantized latent values. The quantized latent values are usually a matrix composed of many samples. The samples can be indicated using indices such as or , depending on the dimension of the matrix . The samples are encoded one by one by AE, usually using the raster scan order. In the raster scan order, the rows of the matrix are processed from top to bottom, and the samples in one row are processed from left to right. In this case (where the raster scan order is used by AE to encode the samples into the bitstream), the context module uses the samples encoded before the raster scan order to generate information related to the samples . The information generated by the context module and the decoder using the hyper-prior information is combined by the entropy parameter module to generate the probability distribution for encoding the quantized latent values into the bitstream (bits1). Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder. Note that other names can be used for the above modules. In the above description, Figure 6 all the elements in 2.3.7. Decoding Process Using the Joint Autoregressive Hyper-Prior Model Figure 7 The decoding process is depicted separately. In the decoding process, the decoder first receives the first bitstream (bits1) and the second bitstream (bits2) generated by the corresponding encoder. bits2 is first decoded by the arithmetic decoding (AD) module by using the probability distribution generated by the factored entropy sub-network. The factored entropy module usually generates the probability distribution using a predetermined template, for example, using predetermined mean and variance values in the case of the Gaussian distribution. The output of the arithmetic decoding process of bits2 is is the quantized latent value of the hyperprior information. The AD process reverts to the AE process applied in the encoder. The processes of AE and AD are lossless, meaning that the quantized latent value of the hyperprior information generated by the encoder can be reconstructed at the decoder without any changes. After obtaining it, it is processed by a decoder that utilizes hyperprior information, and the output of the decoder that utilizes hyperprior information is fed into an entropy parameter module. The three sub-networks, context, decoder that utilizes hyperprior information, and entropy parameter employed in the decoder are the same as those in the encoder. Thus, the exact same probability distribution can be obtained in the decoder (as in the encoder), which is necessary for reconstructing the quantized latent value without any loss is necessary. As a result, the same version of the quantized latent value obtained in the encoder can be obtained in the decoder. After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter sub-network, the arithmetic decoding module decodes the samples of the quantized latent value one by one from the bitstream bits1. From a practical perspective, the autoregressive model (context model) is inherently serial and thus cannot be accelerated using techniques such as parallelization. Finally, the fully reconstructed quantized latent value is input into the synthesis transform (represented as the decoder in Figure 7 ) module to obtain the reconstructed image. In the above description, Figure 7 all the elements in 2.4. Neural Networks for Video Compression Similar to conventional video coding techniques, neural image compression serves as the basis for intra-frame compression in neural network-based video compression. Thus, the development of neural network-based video compression techniques lags behind that of neural network-based image compression, but more effort is needed to address the challenges arising from its complexity. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared with image compression, video compression requires effective methods to remove inter-picture redundancy. Subsequently, inter-picture prediction is a key step in these works. Motion estimation and compensation are widely adopted but have only recently been implemented by trained neural networks. The research on neural network-based video compression can be classified into two categories according to the target scenarios: random access and low latency. In the case of random access, decoding needs to be possible from any point in the sequence. Usually, the whole sequence is divided into multiple individual segments, and each segment can be decoded independently. In the case of low latency, the aim is to reduce the decoding time, so usually only the frames that are temporally previous can be used as reference frames to decode subsequent frames. 2.5. Basic knowledge Almost all natural images / videos are in digital format. A grayscale digital image can be represented by , where is a set of values for pixels, m is the image height and n is the image width. For example, is a common setting, and in this case So pixels can be represented by 8-bit integers. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while the compressed bits are necessarily fewer. Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by , using three separate channels to store red, green, and blue information. Similar to 8-bit grayscale images, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Neural network-based video compression schemes are mainly developed in the RGB color space, while traditional codecs usually use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels, namely Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. Since the human visual system is less sensitive to chrominance components, Cb and Cr are usually downsampled to achieve pre-compression. A color video sequence consists of multiple color images (called frames) to record the scenes at different timestamps. For example, in the RGB color space, a color video can be represented by X = {x 0 , x 1 , …, x t , …, x T-1}, where T is the number of frames in the video sequence, If m = 1080, n = 1920, and the video has 50 frames per second (fps), then the data bitrate of this uncompressed video is 1920×1080×8×3×50 = 2,488,320,000 bits per second (bps), approximately 2.32 Gbps, which requires a large amount of storage and therefore necessarily needs to be compressed before transmission over the Internet. Typically, for natural images, lossless methods can achieve a compression ratio of about 1.5 to 3, which is clearly lower than the requirements. Therefore, lossy compression is developed to achieve further compression ratios, but at the cost of introducing distortion. The distortion can be measured by calculating the mean squared difference between the original image and the reconstructed image (i.e., the mean squared error (MSE)). For grayscale images, the MSE can be calculated using the following equation. Therefore, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR): where is the maximum value in, e.g., 255 for an 8-bit grayscale image. There are other quality assessment metrics, such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM). To compare different lossless compression schemes, it is sufficient to compare the compression ratios for a given resulting bitrate or vice versa. However, to compare different lossy compression methods, both the bitrate and the reconstruction quality must be considered. For example, it is a common method to calculate the relative bitrates for several different quality levels and then average these bitrates; the average relative bitrate is called the Bjontegaard's delta-rate (BD-rate). There are other important aspects for evaluating image / video codec schemes, including encoding / decoding complexity, scalability, robustness, etc. 2.5.1. Separate Processing of the Luminance Component and Chrominance Component of an Image According to one embodiment, separate sub-networks can be used to decode the luminance component and chrominance component of an image. Figure 8 An example of the decoding process according to some embodiments of the present disclosure is shown. In Figure 8 the luminance component of the image is processed by sub-networks such as "Synthesis", "Predictive Fusion", "Mask Convolution", "Hyper Decoder using Hyperprior Information", "Hyper Scale Decoder using Hyperprior Information", etc. While the chrominance component is processed by sub-networks: "Synthesis UV", "Predictive Fusion UV", "Mask Convolution UV", "Hyper Decoder UV using Hyperprior Information", "Hyper Scale Decoder UV using Hyperprior Information", etc. The advantage of the above separate processing is that by applying separate processing, the computational complexity of image processing is reduced. Usually in neural network-based image and video decoding, the computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is equal to 192, the computational complexity will be proportional to 192 x 192. On the other hand, if the feature maps are partitioned into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128 x 128 + 64 x 64, which corresponds to a 45% reduction in complexity. Usually, separate processing of the luminance component and chrominance component of an image does not result in an excessive reduction in performance because the correlation between the luminance component and chrominance component is usually very small. Figure 8 The processing (decoding process) in can be explained as follows: Figure 8 1. First, a factored entropy model is used to decode the quantized latent values for luminance and chrominance, i.e., in 2. The probability parameters (such as variance) generated by the second network are used to generate the quantized residual latent values by performing an arithmetic decoding process. 3. The quantized residual latent values are de-gained using an inverse gain unit (iGain), as shown in orange in Figure 8 . For the luminance component and chrominance component, the outputs of the inverse gain unit are denoted as and 4. For the luminance component, the following steps are performed in a loop until all elements of are obtained: a. The first sub-network is used to estimate the mean parameter of using the samples of the already obtained quantized latent values . b. The quantized residual latent values and the mean are used to obtain the next element of . 5. After all samples of are obtained, an inverse synthesis transform can be applied to obtain the reconstructed image. 6. For the chrominance component, steps 4 and 5 are the same, but using a separate set of networks. 7. The decoded luminance component is used as additional information to obtain the chrominance component. Specifically, an inter-channel correlation information (ICCI) filter sub-network is used for chrominance component recovery. The luminance is fed into the ICCI sub-network as additional information to assist in chrominance component decoding. 8. After the luminance component and chrominance component are reconstructed, an adaptive color transform (ACT) is performed. The module named ICCI is a neural network-based post-processing module. This disclosure is not limited to the UCCI sub-network, and any other neural network-based post-processing module can also be used. In Figure 8 (the decoding process), an example implementation of some embodiments of this disclosure is depicted. The framework includes two branches for the luminance component and the chrominance component respectively. In each branch, the first sub-network includes a context, a prediction, and an optional decoder module that utilizes hyperprior information. The second network includes a variance decoder module that utilizes hyperprior information. The quantized hyperprior information latent value is and the arithmetic decoding process generates a quantized residual latent value, which is further sent to the iGain unit to obtain a gain-quantized residual latent value and After the residual latent value is obtained, a recursive prediction operation is performed to obtain the latent value and The following steps describe how to obtain the samples of the latent value And the chrominance component is processed in the same way but with a different network. 1. The autoregressive context module is used to generate the first input of the prediction module using the samples where the (m, n) pair is the index of the already obtained samples of the latent value. 2. Optionally, the second input of the prediction module is obtained by using a decoder that utilizes hyperprior information and the quantized hyperprior information latent value and 3. Using the first input and the second input, the prediction module generates the mean mean[:,i,j]. 4. The mean mean[:,i,j] and the quantized residual latent value are added together to obtain the latent value 5. Steps 1 to 4 are repeated for the next sample. Whether to apply and / or how to apply at least one method disclosed in this document can be signaled from the encoder to the decoder in the bitstream, for example. Alternatively, whether to apply and / or how to apply at least one method disclosed in the document can be determined by the decoder based on codec information such as size, color format, etc. Alternatively or additionally, (in Figure 8Modules named MS1, MS2, or MS3+O in (China) can be included in the processing flow. The module can perform operations on the input of the module by multiplying the input by a scalar or adding an additive component to the input to obtain an output. The scalar or additive component used by the module can be indicated in the bitstream. The module named RD or the module named AD in Figure 8 can be an entropy decoding module. The entropy decoding module can be a range decoder, an arithmetic decoder, etc. The solution described herein is not limited to Figure 8 the specific combination of units illustrated in. Some modules can be missing, and some modules can be shifted in the processing order. Additional modules can also be included. For example: 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined by another module, which can be neural network-based. 2. One or more of the modules named MS1, MS2, or MS3+O can be removed. The core of the solution is not affected by the removal of one or more of the scaling and adding modules. In Figure 8 other operations performed during the processing of the luminance component and the chrominance component are also indicated using a star symbol. These processes are denoted as MS1, MS2, MS3+O. These processes can be, but are not limited to, adaptive quantization, latent sample scaling, and latent sample compensation operations. For example, in the adaptive quantization process, it can correspond to scaling the samples using a multiplier before the prediction process, where the multiplier is predefined or the value of the multiplier is indicated in the bitstream. The latent scaling process can correspond to the process of scaling the samples using a multiplier after the prediction process, where the value of the multiplier is predefined or indicated in the bitstream. The compensation operation can correspond to adding an additive element to the samples, again where the value of the additive element can be indicated, inferred, or predefined in the bitstream. Another operation can be a slicing operation, where the samples are first sliced (grouped) into overlapping or non-overlapping regions, and each region is processed independently. For example, the samples corresponding to the luminance component can be divided into slices with a slice height of 20 samples, while the chrominance component can be divided into slices with a slice height of 10 samples for processing. Another operation can be the application of wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the amount of samples that can be processed in parallel can be indicated by a control parameter. The control parameter can be indicated in the bitstream, inferred, or can be predefined. In the case of separate luminance and chrominance processing, the number of samples that can be processed in parallel can be different, so different indicators can be signaled in the bitstream to control the operations of the luminance and chrominance processing respectively. 3. Problem 3.1 Core Problem In Section 2.5.1, different processes were illustrated in the case of separate processing of the luminance component and the chrominance component. Since the luminance component and the chrominance component are processed separately, different sets of control parameters need to be signaled in the bitstream to control different processing stages. For example, for the adaptive quantization process of the luminance component, a set of indicators will need to be signaled, and for the adaptive quantization process of the chrominance component, a second set of parameters will need to be signaled. Similarly, for the scaling of samples, the compensation of samples, the slicing of samples, etc., two sets of control parameters will be required. One set is used to control the behavior of the processing steps for the luminance component, and the second set is used to control the behavior of the processing steps for the chrominance component. The necessity of signaling two sets of control parameters separately for the luminance component and the chrominance component leads to an increase in the bitrate and thus a reduction in the compression performance. 4. Detailed Solutions The following detailed solutions should be considered as examples to explain the general concepts. These solutions should not be interpreted in a narrow way. In addition, these solutions can be combined in any way. The proposed solution is related to the separate processing of the components of an image using a neural network. A mechanism for sharing control parameters is disclosed, where the processing of the luminance component and the chrominance component can share some or all of a set of control parameters. 4.1 Core of the Solution The goal of the proposed solution is to provide a mechanism for signaling the control parameters of codec tools that can be used to process at least two components of an image. In one example, one component can be the luminance component, and the second component can be the chrominance component. 4.2 Details of the Solution In Figure 8 , a compression network is depicted, where the luminance component and the chrominance component of the image are reconstructed separately. In addition, example operations for processing residual latent samples, such as MS2, are depicted. The process of MS2 can be, for example, an adaptive quantization process, where the input of the MS2 process (quantized residual samples) is multiplied by a multiplier signaled in the bitstream. The adaptive quantization process generally includes: · At the encoder, a scaling process of the residual samples before quantization. · Quantization of the scaled residual samples at the encoder. · Reception of the quantized residual samples at the decoder. · An inverse scaling operation at the decoder. The value of the scalar (used in inverse scaling) needs to be included in the bitstream so that the decoder can successfully perform the inverse scaling operation. In this case, the control parameter of the adaptive quantization process is a scalar value that controls the amplitude of the scaling operation. According to the proposed solution: A neural network-based image or video decoding method, wherein the image or video at least includes a plurality of components, comprising the following steps: - Obtain an indicator from the bitstream, - Obtain a set of control parameters, - If the value of the indicator is equal to a predefined value, use the set of control parameters in the processing of both components of the image or video. - If the value of the indicator is equal to a second predefined value, use the set of control parameters in the processing of only one of the plurality of components. Combine at least 2 processed components to reconstruct the image or video. According to an example implementation of the proposed solution, an indicator is included in the bitstream. The indicator controls the use of control parameters associated with codec tools (such as the adaptive quantization illustrated above). If the indicator assumes a predefined value "A", for example, "A" can be equal to 1, the same set of control parameters is used in the processing of both components of the image. The plurality of components can be a luminance component and a chrominance component. In another example, the plurality of components can be a red component and / or a green component and / or a blue component. On the other hand, if the value of the indicator is equal to a second predefined value "B", for example, "B" can be equal to 0, the set of control parameters is applied in the processing of only one of the plurality of components. In this case, the other component can be processed without using the process utilizing the set of control parameters, or a second set of control parameters is used. According to an example implementation of the proposed solution, 3 indicators can be included in the bitstream, having the following functions: 1. The first indicator indicates whether the codec tool is used in the processing of one component of the image. 2. The second indicator indicates whether the codec tool is used in the processing of other components of the image. 3. If both indicators indicate that the codec tool is used in two components, a third indicator is included in the bitstream to indicate whether the control parameters used to process the two components are the same. If the control parameters are indicated to be the same, only a single set of control parameters is included in the bitstream to process the two components. The codec tools controlled by the control parameter set can be (but are not limited to) adaptive quantization, sample scaling, skip mode, slicing, sample compensation, wavefront parallel processing, slicing, etc. For example, the skip mode corresponds to processing samples in a way that determines whether a sample is included in the bitstream based on a threshold. The control parameter set corresponding to the skip mode can include the value of the threshold. Slicing corresponds to grouping samples into at least two groups and processing them independently or in parallel. The control parameter set corresponding to slicing can include the slice size or the number of slices or the slice segmentation mode. According to another exemplary embodiment of the proposed solution, an indicator can be included in the bitstream to indicate one of the following options: 1. For the processing of all components, the codec tool is turned off. 2. For the processing of only one component, the codec tool is turned on. For this codec tool, only one set of control parameters is included in the bitstream. 3. For the processing of two components, the codec tool is turned on. For this codec tool, only one set of control parameters is included in the bitstream and is applied to the processing of both components. 4. For the processing of two components, the codec tool is turned on. For this codec tool, only one set of control parameters is included in the bitstream. The first component and / or the second component adopt the control parameters transmitted through the signal. The second component and / or the first component adopt a set of derived control parameters. The derived operations can be but are not limited to rescaling, resampling, etc. 5. For the processing of two components, the codec tool is turned on. For this codec tool, two sets of control parameters are included in the bitstream, and one set of control parameters is used in the processing of one component, and the second set of control parameters is used in the processing of the second component. According to another exemplary embodiment of the proposed solution, an indicator can be included in the bitstream to indicate whether a set of control parameters of one component is the same as those of other components. As an example: 1. The first set of control parameters is included in the bitstream and is used for the processing of the first component. 2. An indicator is included in the bitstream to indicate whether the same set of control parameters as the first component is used for the processing of the second component. If so, the same set of control parameters is used to process the second component. Otherwise, the second set of control parameters is included in the bitstream and is used for the processing of the second component. Alternatively, the second set of control parameters can be derived based on the first set of control parameters, and the second set of control parameters is used for the processing of the second component. According to another exemplary embodiment of the proposed solution, a first indicator may be included in the bitstream to indicate the number of control parameter sets, which may be represented as N. After the first indicator, N second indicators may be included in the bitstream corresponding to each control parameter set. Each second indicator may indicate: · Whether the corresponding control parameter set is only applied to the first component, · Whether the corresponding control parameter set is only applied to the second component, · Whether the corresponding control parameter set is applied to both components. In the above example, N control parameter sets are included in the bitstream. The control parameter sets may control the same codec tool. If N control parameter sets are included in the bitstream to control the codec tool, this may correspond to the repeated application of the same codec tool N times, with different sets used each time. The inclusion of N control parameter sets in the bitstream may correspond to different codec tools. One parameter set may correspond to adaptive quantization, and other parameter sets may correspond to wavefront parallel processing, etc. In the above example, an indicator is included in the bitstream corresponding to at least one of the control parameter sets, and the indicator controls whether the control parameter set is applied to the first component, the second component, or both. According to another exemplary embodiment of the proposed solution, control parameter sets are included in the bitstream to control the codec tool. Corresponding to the control parameter sets, an indicator is included in the bitstream to indicate: · Whether the corresponding control parameter set is only applied to the first component, · Whether the corresponding control parameter set is only applied to the second component, · Whether the corresponding control parameter set is applied to both components. 4.3. Benefits of the Solution According to the proposed solution, the side information required to process multiple components of an image is reduced. 5. Examples 1. Decoder Example: An image or video decoding method including a neural network, the method comprising the following steps: - Obtaining control parameter sets from the bitstream, - Obtaining an indicator from the bitstream that indicates whether the control parameter set is used during the processing of the first component, or during the processing of the second component, or during the processing of both the first and second components. - Processing one component using a first neural subnetwork and processing a second component using a second neural subnetwork, The reconstructed image is obtained by combining the output of the first neural sub-network and the output of the second neural sub-network.
[0034] More details of embodiments of the present disclosure related to neural network-based visual data encoding and decoding will be described below. As used herein, the term "visual data" may refer to video, images, pictures in a video, or any other visual data suitable for being encoded and decoded.
[0035] As described above, in the prior art design, the values of the control parameter sets are transmitted separately for the luminance component and the chrominance component by signals. For example, even if the scaling factor for the sample scaling process is the same for processing the luminance component and the chrominance component. The value of this scaling factor needs to be transmitted by signal twice, that is, once for processing the luminance component and once for processing the chrominance component. This results in an increase in the bit rate and thus a decrease in the compression performance and encoding / decoding efficiency.
[0036] To solve the above problems and some problems not mentioned, a visual data processing solution as described below is disclosed. Embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow way. In addition, these embodiments can be applied individually or in combination in any way.
[0037] Figure 9 A flowchart of a method 900 for visual data processing according to some embodiments of the present disclosure is shown. As Figure 9 shown, at 902, a transformation is performed between visual data and a bitstream of the visual data using a neural network (NN)-based model. In some embodiments, the transformation may include encoding the visual data into the bitstream. Additionally or alternatively, the transformation may include decoding the visual data from the bitstream. For example, Figure 8 the decoding model shown can be used to decode visual data from the bitstream.
[0038] In some embodiments, the bitstream includes a first indication for indicating whether a set of values for a set of parameters for the NN-based model is common for processing multiple components of the visual data. By way of example and not limitation, the first indication may be a flag, a syntax element, etc. For example, if a set of values for a set of parameters is common for processing multiple components, each of the multiple components can be processed by using the set of values for the set of parameters. In other words, a set of values for a set of parameters is the same for processing multiple components.
[0039] In some embodiments, the plurality of components may include a luminance component, a chrominance component, a red component, a green component, a blue component, a Y component, a U component, a V component, a chrominance blue (Cb) component, and / or a chrominance red (Cr) component. It should be understood that the plurality of components may include any other suitable components, such as an alpha component for transparency. The scope of the present disclosure is not limited in this regard.
[0040] In some embodiments, the plurality of components of the visual data may be at least partially processed separately. For example, the plurality of components may include a luminance component and a chrominance component of the visual data. Referring to Figure 8 , the luminance component and the chrominance component may be processed separately using a module composed of the same sequence of the same neural network layers, where the same neural network layers have differences in the dimensions on the input tensors and the number of tensor channels. A first indication may be used to indicate whether the value of at least one parameter for the NN-based model is common for the processing of the luminance component and the chrominance component. If the value of at least one parameter is common for the processing of the luminance component and the chrominance component, the value of at least one parameter may be signaled only once, rather than being signaled separately for the luminance component and the chrominance component. Thereby, the side information required for processing the plurality of components of the visual data is reduced.
[0041] In view of the above, an indication is included in the bitstream, and the indication is used to indicate whether a set of values of a set of parameters for the NN-based model is common for the processing of the plurality of components of the visual data. By means of this indication, separately signaling the set of values for the plurality of components of the visual data can be avoided. Therefore, the proposed method can advantageously improve the coding and decoding efficiency.
[0042] In some embodiments, a set of parameters may be used for a first coding and decoding tool of the NN-based model. As used herein, the term "coding and decoding tool" may refer to any suitable sub-process for processing visual data, and the coding and decoding tool may be implemented as a tool, a unit, a module, etc. For example, the first coding and decoding tool may include adaptive quantization, sample scaling, skip mode, slicing, sample compensation, gain unit, inverse gain unit, filter, inter-channel correlation information (ICCI) filter, gain process, inverse gain process, post-processing module, decoder module using hyperprior information, variance decoder module using hyperprior information, and / or wavefront parallel processing. It should be understood that the possible implementations of the first coding and decoding tool described here are merely illustrative and should not be construed as limiting the present disclosure in any way.
[0043] In some embodiments, a set of parameters may include a single parameter. In this case, the set of values may include a single value of the single parameter. Alternatively, the set of parameters may include multiple parameters. In this case, the set of parameters may also include multiple values, and each value in the multiple values corresponds to one of the multiple parameters. For example, the number of values in the set of values may be equal to the number of parameters in the set of parameters.
[0044] In some embodiments, a set of parameters may include any suitable parameters for a first codec tool. By way of example and not limitation, depending on the first codec tool, the set of parameters may include a scalar, a multiplier, a vector, a scaling factor, a threshold, a slice size, a number of slices, a slice partitioning mode, an index, a model, an offset, an additive coefficient, a subtractive coefficient, and / or a number of samples for parallel processing. For example, if the first codec tool is a skip mode, the set of parameters may include a threshold for determining whether a sample is included in the bitstream. In another example, if the first codec tool is a slicing process, the set of parameters may include a slice size, a number of slices, and / or a slice partitioning mode. In another example, the offset may be implemented as a displacement term. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0045] In some embodiments, a set of values of a set of parameters may be indicated in the bitstream. For example, on the decoder side, the set of values may be obtained by decoding the bitstream. Alternatively, a set of values of a set of parameters may be predefined. In this case, the set of values does not need to be signaled, and thus the bitrate for encoding and decoding visual data may be advantageously further reduced.
[0046] In some embodiments, a first indicated value being equal to a first value may be used to indicate that a set of values of a set of parameters is common for processing multiple components. Additionally, a first indicated value being equal to a second value may be used to indicate that a set of values of a set of parameters is not common for processing multiple components. The second value is different from the first value. In one example, the first value may be 1, and the second value may be 0. In another example, the first value may be 0, and the second value may be 1. It should be understood that the specific values recited herein are intended to be exemplary and not to limit the scope of the present disclosure.
[0047] In some embodiments, the bitstream may further include a second indication for indicating whether a first codec tool is applied to a first component (e.g., a chrominance component, etc.) among a plurality of components. Additionally, the bitstream may further include a third indication for indicating whether the first codec tool is applied to a second component (e.g., a luminance component, etc.) among the plurality of components. The second component is different from the first component. If both indicators are used to indicate that the codec tool is used in two components, a first indication may be included in the bitstream to indicate whether a set of values for a set of parameters for an NN-based model is common for the processing of the plurality of components.
[0048] In some embodiments, the first indication may further indicate at least one of the following: whether the first codec tool is applied to the plurality of components, at least one component to which the first codec tool is applied among the plurality of components, or at least one component to which a set of values of a set of parameters among the plurality of components is applied.
[0049] By way of example and not limitation, in the case where the plurality of components includes two components of visual data. The first indication may be used to indicate one of the following options:
[0050] First option: The first codec tool is not applied to these two components.
[0051] Second option: The first codec tool is only applied to one of these two components. For the first codec tool, only a set of values of a set of parameters is included in the bitstream.
[0052] Third option: The first codec tool is applied to these two components. For the first codec tool, only a set of values of a set of parameters is included in the bitstream and is used for the processing of both components.
[0053] Fourth option: The first codec tool is applied to these two components. For the first codec tool, only a set of values of a set of parameters is included in the bitstream. One of the two components is processed by using the set of values transmitted through the signal. The other of the two components is processed by using another set of values determined based on the set of values transmitted through the signal.
[0054] Fifth option: The first codec tool is applied to these two components. For the first codec tool, two sets of values of a set of parameters are included in the bitstream. One of the two components is processed by using the first set of the two sets of values. The other of the two components is processed by using the second set of the two sets of values.
[0055] It should be understood that the above options are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.
[0056] In some embodiments, the bitstream may further include a fourth indication indicating the number of multiple sets of parameters for at least one codec tool for the NN-based model. The multiple sets of parameters may include a set of parameters for a first codec tool, and the at least one codec tool may include the first codec tool. In this case, for each set of values of each set of parameters among the multiple sets of parameters, an indication may be included in the bitstream to indicate whether the set of values is common for processing multiple components of the visual data.
[0057] In some embodiments, the bitstream may further include a fifth indication for indicating the number of multiple sets of values of a set of parameters for the first codec tool. In this case, for each value among the multiple sets of values, an indication may be included in the bitstream to indicate whether the set of values is common for processing multiple components of the visual data.
[0058] In some embodiments, if a set of values of a set of parameters is not common for processing multiple components, the first indication may further be used to indicate one or more components to which the set of values of the set of parameters among the multiple components is applied.
[0059] In some embodiments, a set of values of a set of parameters may be common for processing multiple components. In this case, each of the multiple components may be processed by using the first codec tool having the set of values of the set of parameters.
[0060] In some embodiments, a set of values of a set of parameters may not be common for processing multiple components, and the set of values may be applied to a first component among the multiple components. In this case, the second component different from the first component among the multiple components may not be processed by using the first codec tool. Alternatively, the second component may be processed by using the first codec tool having another set of values of the set of parameters. In one example, the another set of values may be determined based on the set of values, for example, by rescaling the set of values, resampling the set of values, etc. In another example, the another set of values may be indicated in the bitstream. In yet another example, the another set of values may be predefined.
[0061] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for visual data processing for visual data. In the method, a conversion is performed between visual data and the bitstream by using a neural network (NN)-based model. The bitstream includes a first indication for indicating whether a set of values of a set of parameters for the NN-based model is common for processing multiple components of the visual data.
[0062] According to still other embodiments of the present disclosure, a method for storing a bitstream of visual data is provided. According to the method, a transformation is performed between visual data and a bitstream using a neural network (NN)-based model. The bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for the processing of multiple components of the visual data. Further, the bitstream is stored in a non-transitory computer-readable recording medium.
[0063] Embodiments of the present disclosure may be described according to the following items, and the features of the following items may be combined in any reasonable manner.
[0064] Item 1. A method for visual data processing, comprising: performing a transformation between visual data and a bitstream of the visual data using a neural network (NN)-based model, wherein the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for the processing of multiple components of the visual data.
[0065] Item 2. The method according to Item 1, wherein the set of parameters is used for a first codec tool of the NN-based model.
[0066] Item 3. The method according to any one of Items 1 to 2, wherein the multiple components include at least one of the following: a luminance component, a chrominance component, a red component, a green component, a blue component, a Y component, a U component, a V component, a chrominance blue (Cb) component, or a chrominance red (Cr) component.
[0067] Item 4. The method according to any one of Items 2 to 3, wherein the first codec tool includes at least one of the following: adaptive quantization, sample scaling, skip mode, slicing, sample compensation, gain unit, anti-gain unit, filter, inter-channel correlation information (ICCI) filter, gain process, anti-gain process, post-processing module, decoder module using hyperprior information, variance decoder module using hyperprior information, or wavefront parallel processing.
[0068] Item 5. The method according to any one of Items 1 to 4, wherein the set of parameters includes at least one of the following: scalar, multiplier, vector, scaling factor, threshold, slice size, number of slices, slice segmentation mode, index, model, offset, additive coefficient, or number of samples for parallel processing.
[0069] Item 6. The method according to any one of Items 1 to 5, wherein the set of values of the set of parameters is indicated or predefined in the bitstream.
[0070] Item 7. The method according to any one of Items 1 to 6, wherein the value of the first indication is equal to a first value for indicating that a set of values of a set of parameters is common for the processing of multiple components, or the value of the first indication is equal to a second value for indicating that a set of values of a set of parameters is not common for the processing of multiple components, and the second value is different from the first value.
[0071] Item 8. The method according to Item 7, wherein the first value is 1 and the second value is 0, or wherein the first value is 0 and the second value is 1.
[0072] Item 9. The method according to any one of Items 2 to 8, wherein the bitstream further comprises: a second indication for indicating whether a first codec tool is applied to a first component among multiple components, and a third indication for indicating whether the first codec tool is applied to a second component among multiple components, the second component being different from the first component.
[0073] Item 10. The method according to any one of Items 2 to 9, wherein the first indication is further used to indicate at least one of the following: whether a first codec tool is applied to multiple components, at least one component among multiple components to which the first codec tool is applied, or at least one component among multiple components to which a set of values of a set of parameters is applied.
[0074] Item 11. The method according to any one of Items 2 to 10, wherein the bitstream further comprises a fourth indication for indicating the number of multiple sets of parameters for at least one codec tool for an NN-based model, the multiple sets of parameters including a set of parameters for the first codec tool, and the at least one codec tool includes the first codec tool, or wherein the bitstream further comprises a fifth indication for indicating the number of multiple sets of values of a set of parameters for the first codec tool.
[0075] Item 12. The method according to any one of Items 1 to 11, wherein if a set of values of a set of parameters is not common for the processing of multiple components, the first indication is further used to indicate the component among multiple components to which a set of values of a set of parameters is applied.
[0076] Item 13. The method according to any one of Items 2 to 12, wherein a set of values of a set of parameters is common for the processing of multiple components, and each component among multiple components is processed by using a first codec tool having a set of values of a set of parameters.
[0077] Item 14. The method according to any one of Items 2 to 12, wherein a set of values of a set of parameters is not common for the processing of multiple components, and a set of values of a set of parameters is applied to a first component among multiple components, and a second component different from the first component among multiple components is not processed by using the first codec tool.
[0078] Item 15. The method according to any one of Items 2 to 12, wherein a set of values of a set of parameters is not common for the processing of a plurality of components, and a set of values of a set of parameters is applied to a first component among the plurality of components, and a second component different from the first component among the plurality of components is processed by using a first codec tool having another set of values different from the set of values of the set of parameters.
[0079] Item 16. The method according to Item 15, wherein the another set of values is determined based on the set of values, or the another set of values is indicated in the bitstream, or the another set of values is predefined.
[0080] Item 17. The method according to any one of Items 1 to 16, wherein a plurality of components of the visual data are at least partially processed separately.
[0081] Item 18. The method according to any one of Items 1 to 17, wherein a set of parameters includes a single parameter, or wherein a set of parameters includes a plurality of parameters.
[0082] Item 19. The method according to any one of Items 1 to 18, wherein the visual data includes video, pictures or images of video.
[0083] Item 20. The method according to any one of Items 1 to 19, wherein the conversion includes encoding the visual data into a bitstream.
[0084] Item 21. The method according to any one of Items 1 to 19, wherein the conversion includes decoding the visual data from the bitstream.
[0085] Item 22. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Items 1 to 21.
[0086] Item 23. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 to 21.
[0087] Item 24. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for visual data processing for visual data, wherein the method includes: performing a conversion between the visual data and the bitstream by using a model based on a neural network (NN), wherein the bitstream includes a first indication for indicating whether a set of values of a set of parameters for the model based on the NN is common for the processing of a plurality of components of the visual data.
[0088] Article 25. A method for storing a bitstream of visual data, comprising: performing a conversion between visual data and a bitstream using a neural network (NN)-based model, wherein the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing multiple components of the visual data; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0089] Figure 10 FIG. shows a block diagram of a computing device 1000 in which various embodiments of the present disclosure may be implemented. The computing device 1000 may be implemented as a source device 110 (or a visual data encoder 114) or a destination device 120 (or a visual data decoder 124) or may be included in the source device 110 (or a visual data encoder 114) or the destination device 120 (or a visual data decoder 124).
[0090] It should be understood that Figure 10 the computing device 1000 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0091] As Figure 10 shown, the computing device 1000 includes a general-purpose computing device 1000. The computing device 1000 may include at least one or more processors or processing units 1010, a memory 1020, a storage unit 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060.
[0092] In some embodiments, the computing device 1000 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 1000 may support any type of interface to the user (such as a "wearable" circuitry, etc.).
[0093] The processing unit 1010 can be a physical processor or a virtual processor, and can implement various processes based on the programs stored in the memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1000. The processing unit 1010 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0094] The computing device 1000 generally includes various computer storage media. Such media can be any media accessible by the computing device 1000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1020 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 1030 can be any removable or non-removable media, and can include machine-readable media, such as memory, flash drives, magnetic disks, or other media that can be used to store information and / or visual data and can be accessed in the computing device 1000.
[0095] The computing device 1000 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 10 it, a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In this case, each drive can be connected to a bus (not shown) via one or more visual data media interfaces.
[0096] The communication unit 1040 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 1000 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Therefore, the computing device 1000 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0097] The input device 1050 can be one or more of various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and so on. The output device 1060 can be one or more of various output devices, such as a display, a speaker, a printer, and so on. With the aid of the communication unit 1040, the computing device 1000 can also communicate with one or more external devices (not shown), such as a storage device and a display device, the computing device 1000 can also communicate with one or more devices that enable a user to interact with the computing device 1000, or if needed, the computing device 1000 can also communicate with any device (such as a network card, a modem, etc.) that enables the computing device 1000 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0098] In some embodiments, some or all components of the computing device 1000 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, visual data access, and storage services, which do not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses a suitable protocol to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding visual data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at locations in a remote visual data center. The cloud computing infrastructure can provide services through a shared visual data center, although to the user, they appear as a single access point. Thus, a cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0099] In an embodiment of the present disclosure, the computing device 1000 can be used to implement visual data encoding / decoding. The memory 1020 may include one or more visual data encoding / decoding modules 1025 having one or more program instructions. These modules are accessible and executable by the processing unit 1010 to perform the functions of various embodiments described herein.
[0100] In an example embodiment of performing visual data encoding, the input device 1050 may receive visual data as input 1070 to be encoded. The visual data may be processed, for example, by the visual data codec module 1025 to generate an encoded bitstream. The encoded bitstream may be provided as output 1080 via the output device 1060.
[0101] In an example embodiment of performing visual data decoding, the input device 1050 may receive the encoded bitstream as input 1070. The encoded bitstream may be processed, for example, by the visual data codec module 1025 to generate decoded visual data. The decoded visual data may be provided as output 1080 via the output device 1060.
[0102] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for visual data processing, comprising: performing a conversion between visual data and a bitstream of the visual data by using a neural network (NN)-based model, wherein the bitstream includes a first indication for indicating whether a set of values for a set of parameters for the NN-based model is common for processing of multiple components of the visual data.
2. The method according to claim 1, wherein the set of parameters is used for a first codec tool of the NN-based model.
3. The method according to any one of claims 1 to 2, wherein the multiple components include at least one of the following: a luminance component, a chrominance component, a red component, a green component, a blue component, a Y component, a U component, a V component, a chrominance blue (Cb) component, or a chrominance red (Cr) component.
4. The method according to any one of claims 2 to 3, wherein the first codec tool includes at least one of the following: adaptive quantization, sample scaling, skip mode, slicing, sample compensation, gain unit, inverse gain unit, filter, inter-channel correlation information (ICCI) filter, gain process, inverse gain process, post-processing module, decoder module using hyperprior information, variance decoder module using hyperprior information, or wavefront parallel processing.
5. The method according to any one of claims 1 to 4, wherein the set of parameters includes at least one of the following: scalar, multiplier, vector, scaling factor, threshold, slice size, number of slices, slice segmentation mode, index, model, offset, additive coefficient, or number of samples for parallel processing.
6. The method according to any one of claims 1 to 5, wherein the set of values of the set of parameters is indicated or predefined in the bitstream.
7. The method according to any one of claims 1 to 6, wherein a value of the first indication equals a first value for indicating that the set of values of the set of parameters is common for processing of the multiple components, or the value of the first indication equals a second value for indicating that the set of values of the set of parameters is not common for processing of the multiple components, the second value being different from the first value.
8. The method according to claim 7, wherein the first value is 1 and the second value is 0, or wherein the first value is 0 and the second value is 1.
9. The method according to any one of claims 2 to 8, wherein the bitstream further comprises: a second indication for indicating whether the first codec tool is applied to a first component among the multiple components, and a third indication for indicating whether the first codec tool is applied to a second component among the multiple components, the second component being different from the first component.
10. The method according to any one of claims 2 to 9, wherein the first indication is further used to indicate at least one of the following: whether the first codec tool is applied to the multiple components, at least one component among the multiple components to which the first codec tool is applied, or At least one component to which the set of values of the set of parameters among the plurality of components is applied.
11. The method according to any one of claims 2 to 10, wherein the bitstream further comprises a fourth indication for indicating the number of sets of parameters of at least one codec tool for the NN-based model, the sets of parameters comprising the set of parameters for the first codec tool, and the at least one codec tool comprises the first codec tool, or wherein the bitstream further comprises a fifth indication for indicating the number of sets of values of the set of parameters for the first codec tool.
12. The method according to any one of claims 1 to 11, wherein if the set of values of the set of parameters is not common for the processing of the plurality of components, the first indication is further used to indicate the component to which the set of values of the set of parameters among the plurality of components is applied.
13. The method according to any one of claims 2 to 12, wherein the set of values of the set of parameters is common for the processing of the plurality of components, and each of the plurality of components is processed by using the first codec tool having the set of values of the set of parameters.
14. The method according to any one of claims 2 to 12, wherein the set of values of the set of parameters is not common for the processing of the plurality of components, and the set of values of the set of parameters is applied to a first component among the plurality of components, and a second component different from the first component among the plurality of components is not processed by using the first codec tool.
15. The method according to any one of claims 2 to 12, wherein the set of values of the set of parameters is not common for the processing of the plurality of components, and the set of values of the set of parameters is applied to a first component among the plurality of components, and a second component different from the first component among the plurality of components is processed by using the first codec tool having another set of values different from the set of values.
16. The method according to claim 15, wherein the another set of values is determined based on the set of values, or the another set of values is indicated in the bitstream, or the another set of values is predefined.
17. The method according to any one of claims 1 to 16, wherein the plurality of components of the visual data are at least partially processed separately.
18. The method according to any one of claims 1 to 17, wherein the set of parameters comprises a single parameter, or wherein the set of parameters comprises a plurality of parameters.
19. The method according to any one of claims 1 to 18, wherein the visual data comprises video, pictures or images of the video.
20. The method according to any one of claims 1 to 19, wherein the conversion comprises encoding the visual data into the bitstream.
21. The method according to any one of claims 1 to 19, wherein the conversion comprises decoding the visual data from the bitstream.
22. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 21.
23. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 21.
24. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for visual data processing of visual data, wherein the method comprises: performing a conversion between the visual data and the bitstream using a neural network (NN)-based model, wherein the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing of multiple components of the visual data.
25. A method for storing a bitstream of visual data, comprising: performing a conversion between the visual data and the bitstream using a neural network (NN)-based model, wherein the bitstream includes a first indication for indicating whether a set of values for a set of parameters of the NN-based model is common for processing of multiple components of the visual data; and storing the bitstream in a non-transitory computer-readable recording medium.
Citation Information
Patent Citations
Image encoding device, image decoding device, image encoding method, and image decoding method
CN101222645A
Method and apparatus of neural network for video coding
CN111133756A
Using neural network filtering in video coding
CN114390288A
Neural network-based post-filter for video coding and decoding
CN115209143A