Visual data processing method, device and medium
By determining the target weight based on the relevant information of visual data in the image/video encoding and decoding system, the problem of poor module weight determination in the prior art is solved, and the encoding and decoding efficiency and effect are improved.
Patent Information
- Application Number
- CN202380074376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively determine the module weight in image/video encoding and decoding, resulting in poor encoding and decoding efficiency and effect.
By determining the target weights for the target module in the codec system based on the information associated with the visual data, the codec system is implemented using at least one neural network.
Improve the performance of the module and improve the encoding and decoding effect and encoding and decoding efficiency of the encoding and decoding system.
Smart Images

Figure CN120112958A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to visual data processing techniques, and more particularly to weight determination for modules in a coding system. Background Art
[0002] Image / video compression is an important technology to reduce the cost of image / video transmission and storage in a lossless or lossy manner. Image / video compression technology can be divided into two branches: classical video codec methods and neural network-based video compression methods. Classical video codec schemes adopt transform-based solutions, in which researchers exploit statistical dependencies in latent variables (e.g., wavelet coefficients) by carefully manually designing entropy codes that model dependencies in the quantization domain. Neural network-based video compression has two styles: neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded into existing classical video codecs as codec tools and is only used as part of the framework; while the latter is a separate framework developed based on neural networks and does not rely on classical video codecs. It is generally expected to further improve the codec efficiency of image / video coding. Summary of the invention
[0003] Embodiments of the present disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method includes: for conversion between visual data and a bit stream of visual data, based on information associated with the visual data, determining a target weight for use by a target module in a codec system, the codec system being implemented using at least one neural network; and performing conversion based on the target weight by using the codec system. The method according to the first aspect of the present disclosure determines the weights for a module in a codec system based on information associated with video data (e.g., visual data content). In this way, the performance of the module can be improved. Therefore, the coding and decoding effect and coding and decoding efficiency of the codec system can be improved.
[0005] In a second aspect, a device for visual data processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, and the instructions enable a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for visual data processing. The method includes: determining a target weight for use by a target module in a codec system based on information associated with the visual data, the codec system being implemented using at least one neural network; and generating a bitstream based on the target weight by using the codec system.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a target weight for use by a target module in a codec system based on information associated with visual data, the codec system being implemented using at least one neural network; generating a bitstream based on the target weight by using the codec system; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example visual data encoding and decoding system according to some embodiments of the present disclosure is shown;
[0012] Figure 2 A diagram showing a typical transform coding scheme;
[0013] Figure 3 An image from the Kodak dataset and different representations of the image are shown;
[0014] Figure 4 The network architecture of the autoencoder implementing the hyper-prior model is shown;
[0015] Figure 5 shows a block diagram of a combined model that jointly optimizes autoregressive components that estimate probability distributions of latent values from their causal context (context model) as well as hyper-priors and underlying autoencoders;
[0016] Figure 6 The encoding process is shown;
[0017] Figure 7 The decoding process is shown;
[0018] Figure 8 The decoding process with independent subnetwork weight selection for the synthetic network is shown;
[0019] Fig. 9 A flowchart of a method for visual data processing according to an embodiment of the present disclosure is shown;
[0020] Fig.10 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.
[0021] Same or similar reference numbers generally refer to same or similar elements throughout the drawings. DETAILED DESCRIPTION
[0022] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0023] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0024] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.
[0025] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0026] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment
[0027] Figure 1 1 is a block diagram illustrating an example visual data encoding and decoding system 100 that can utilize the techniques of the present disclosure. As shown, the visual data encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a data encoding device or a visual data encoding device, and the destination device 120 may also be referred to as a data decoding device or a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a data source 112, a data encoder 114, and an input / output (I / O) interface 116.
[0028] The data source 112 may include a source such as a data acquisition device. Examples of data acquisition devices include, but are not limited to, an interface for receiving data from a data content provider, a computer graphics system for generating data, and / or a combination thereof.
[0029] The data may include one or more pictures or one or more images of a video. The data encoder 114 encodes the data from the data source 112 to generate a bit stream. The bit stream may include a bit sequence that forms a coded representation of the data. The bit stream may include coded pictures and associated data. The coded pictures are coded representations of pictures. The associated data may include sequence parameter sets, picture parameter sets, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0030] The destination device 120 may include an I / O interface 126, a data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded data from the source device 110 or the storage medium / server 130B. The data decoder 124 may decode the encoded data. The display device 122 may display the decoded data to the user. The display device 122 may be integrated with the destination device 120, or may be outside the destination device 120, which is configured to be connected to an external display device interface.
[0031] The data encoder 114 and the data decoder 124 may operate according to a data codec standard, such as a video codec standard or a still picture codec standard and other existing and / or future standards.
[0032] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the section titles used in this document are for ease of understanding, and the embodiments disclosed in the section are not limited to that section. In addition, although some embodiments are described with reference to multifunctional video codecs or other specific data codecs, the disclosed technology is also applicable to other coding and decoding technologies. In addition, although some embodiments describe the coding and decoding steps in detail, it should be understood that the corresponding decoding steps of decoding will be implemented by the decoder. In addition, the term data processing includes data encoding or compression, data decoding or decompression, and data transcoding, in which data is represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview A neural network-based image and video compression method includes an autoregressive subnet and an entropy encoding and decoding engine. In the present disclosure, an independent subnet weight selection for neural network-based image and video compression is proposed. During the decoding process, a synthetic network or a predicted network can be selected from a set of pre-trained network weights instead of using a specific subnet. 2. Introduction The past decade has witnessed the rapid development of deep learning in various fields, especially in computer vision and image processing. From the great success of deep learning techniques in the field of computer vision, many researchers have shifted their attention from traditional image / video compression techniques to neural image / video compression techniques. Neural networks were originally invented as a result of interdisciplinary research in neuroscience and mathematics. It has demonstrated powerful capabilities in the context of nonlinear transformation and classification. Image / video compression techniques based on neural networks have made significant progress in the past five years. It is reported that the latest neural network-based image compression algorithms achieve RD performance comparable to that of Versatile Video Codec (VVC), which is the latest video codec standard developed by the Joint Video Experts Group (JVET) with experts from the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). With the continuous improvement of the performance of neural image compression, neural network-based video compression has become an actively developed research area. However, due to the inherent difficulty of the problem, neural network-based video codec is still in its infancy. 2.1 Image / Video Compression Image / video compression generally refers to a computing technique that compresses an image / video into a binary code for storage and transmission. Binary codes may or may not support lossless reconstruction of the original image / video, referred to as lossless compression and lossy compression. Most efforts are devoted to lossy compression because lossless reconstruction is not required in most scenarios. Typically, the performance of an image / video compression algorithm is evaluated from two aspects, compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes, and the fewer the number of binary codes, the better; the reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, and the higher the reconstruction quality, the better. Image / video compression techniques can be divided into two branches: classical video codec methods and neural network-based video compression methods. Classical video codec schemes adopt transform-based solutions, in which researchers exploit statistical dependencies in latent variables (e.g., discrete cosine transform (DCT) or wavelet coefficients) by carefully manually designing entropy codes that model the dependencies in the quantization domain. Neural network-based video compression has two flavors: neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded into existing classical video codecs as codec tools and is only used as part of the framework; while the latter is a separate framework developed based on neural networks and does not rely on classical video codecs. Over the past three decades, a series of classic video codec standards have been developed to accommodate the increased visual content. The international standards organization ISO / IEC has two expert groups, the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), and ITU-T also has its own Video Codec Experts Group (VCEG), which is used for the standardization of image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. After H.265 / HEVC, the Joint Video Experts Group (JVET) formed by MPEG and VCEG has been working on a new video codec standard, Versatile Video Codec (VVC). The first version of VVC was released in July 2020. Compared with HEVC, VVC is reported to have an average bit rate reduction of 50% at the same visual quality. Neural network-based image / video compression is not a new invention, as there are many researchers working on neural network-based image encoding and decoding. However, the network architecture is relatively shallow and the performance is not satisfactory. Benefiting from the support of large amounts of data and powerful computing resources, neural network-based methods are better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, proving its feasibility. However, the technology is still far from mature and there are still many challenges to be solved. 2.2 Neural Networks A neural network (also called an artificial neural network (ANN)) is a computational model used in machine learning techniques that is typically composed of multiple processing layers, and each layer is composed of multiple simple but nonlinear basic computational units. One benefit of such deep networks is believed to be the ability to process data with multiple levels of abstraction and convert the data into different kinds of representations. Note that these representations are not manually designed; instead, the deep network comprising the processing layers learns from massive amounts of data using a general machine learning process. Deep learning eliminates the necessity for manual representations and is therefore believed to be useful, particularly for processing native unstructured data such as acoustic and visual signals, while processing such data has been a long-standing difficulty in the field of artificial intelligence. 2.3 Neural Networks for Image Compression Existing neural networks for image compression methods can be divided into two categories, namely pixel probability modeling and autoencoders. The former belongs to the predictive encoding and decoding strategy, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature. 2.3.1 Pixel Probabilistic Modeling According to Shannon information theory, the optimal method for lossless coding and decoding can achieve the minimum coding rate log 2p(x), where p(x) is the probability of symbol x. Many lossless codecs have been developed in the literature, and among these, arithmetic coding is considered the best. Given a probability distribution p(x), arithmetic coding ensures that the coding rate is as close as possible to its theoretical limit without taking into account round-off errors—log 2 p(x). Hence, the remaining question is how to determine the probability, which is very challenging for natural images / videos due to the reduced dimensionality. Following the predictive coding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is the image. p(x)=p(x 1 )p(x 2 |x 1 )...p(x i |x 1 , …, x i-1 )…p(x m×n |x 1 ,…,x m×n-1 ) (1) Where m and n are the height and width of the image, respectively. The previous observation is also called the context of the current pixel. When the image is large, it may be difficult to estimate the conditional probability, so a simplified approach is to limit the scope of its context. p(x)=p(x 1 )p(x 2 |x 1 )...p(x i |x i-k , ..., x i-1 )...p(x m×n |x m×n-k , …, x m×n-1 ) (2) Where k is a predefined constant that controls the scope of the context. It should be noted that the condition may also consider the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on the previously encoded and decoded pixel (including R / G / B samples), the current G sample can be encoded and decoded based on the previously encoded and decoded pixel and the current R sample, and when encoding and decoding the current B sample, the previously encoded and decoded pixel and the current R sample and the current G sample may also be considered. Neural networks were initially introduced for computer vision tasks and have been shown to be effective in regression and classification problems. Therefore, it has been proposed to use neural networks to estimate a given its context x. 1 ,x 2 ,…,x i-1 The probability p(x i). For a binary image (i.e., x i ∈{-1,+1}) pixel probabilities are proposed. A neural autoregressive distribution estimator (NADF) is designed for pixel probability modeling, where is a feed-forward network with a single hidden layer. Similar work is presented, where the feed-forward network also has connections that skip the hidden layer and parameters are also shared. NADF is extended to a real-valued model RNADE, where the probability p(x i |x 1 ,…,x i-1 ) is derived using a mixture of Gaussians. Their feed-forward networks also have a single hidden layer, but the hidden layer has rescaling to avoid saturation and uses rectified linear units (ReLU) instead of sigmoids. NADF and RNADE are improved by reorganizing the order of pixels and using deeper neural networks. Designing advanced neural networks plays an important role in improving pixel probability modeling. A multidimensional long short-term memory (LSTM) is proposed, which works with a mixture of conditional Gaussian scale mixtures for probabilistic modeling. LSTM is a special recurrent neural network (RNN) and has been shown to be good at modeling sequential data. The spatial variant of LSTM is used for images. Several different neural networks have been studied, including RNN and CNN, namely PixelRNN and PixelCNN. In PixelRNN, two variants of LSTM are proposed, called row LSTM and diagonal BiLSTM, the latter of which is specifically designed for images. PixelRNN incorporates residual connections to help train deep neural networks with up to 12 layers. In PixelCNN, masked convolutions are used to adapt to the shape of the context. Compared with previous work, PixelRNN and Pix-elCNN are more specialized for natural images: they treat pixels as discrete values (e.g., 0, 1, ..., 255) and predict a multinomial distribution of discrete values; they process color images in the RGB color space; they work well on the large-scale image dataset ImageNet. Gated PixelCNN is proposed to improve PixelCNN and achieve comparable performance to PixelRNN, but with much lower complexity. PixelCNN++ is proposed, which has the following improvements on PixelCNN: using discretized logistic mixture likelihood instead of 256-way multinomial distribution; using downsampling to obtain structure at multiple resolutions; introducing additional shortcut connections to speed up training; using dropout for regularization; RGB is combined for one pixel. PixelSNAIL is proposed, where causal convolution is combined with self-attention. Most of the above methods directly model the probability distribution in the pixel domain. Some researchers have also tried to model the probability distribution as a conditional probability distribution based on explicit representation or latent representation. That is, one can estimate: Where h is an additional condition, and p(x)=p(h)p(x|h) means that the modeling is divided into an unconditional one and a conditional one. The additional condition can be image label information or a high-level representation. 2.3.2 Autoencoder Autoencoders are proposed. The method for autoencoders is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part converts the high-dimensional input signal into a low-dimensional representation, usually with reduced spatial size but with a larger number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders enable automatic learning of representations and eliminate the need for hand-crafted features, which is also considered one of the most important advantages of neural networks. Figure 2 A diagram of a typical transform coding scheme 200 is shown. The original image x is analyzed by the analysis network g a The latent representation y is quantized and compressed into bits. The number of bits R is used to measure the coding rate. The quantized latent representation By synthetic network g s Inverse transform to obtain the reconstructed image The distortion is obtained by transforming x and Using the function g p is calculated in perceptual space. It is intuitive to apply autoencoder networks to lossy compression. The learned latent representation can be encoded from a trained neural network. However, adapting the autoencoder to image compression is not trivial because the original autoencoder is not optimized for compression and is therefore not efficient by directly using the trained autoencoder. In addition, there are other major challenges: First, the low-dimensional representation should be quantized before being encoded, but quantization is not differentiable, which is required in backpropagation when training neural networks. Second, the goals in the compression scenario are different because both distortion and bitrate need to be considered. Estimating the bitrate is challenging. Third, practical image coding and decoding schemes need to support variable bitrate, scalability, encoding speed / decoding speed, and interoperability. In response to these challenges, many researchers have actively contributed to this field. Prototype autoencoder for image compression in Figure 2 It can be regarded as a transform coding strategy. The original image x is analyzed using the network y = g a (x) is transformed, where y is the latent representation to be quantized and encoded. The synthesis network transforms the quantized latent representation Perform inverse transformation to obtain the reconstructed image The framework is trained using the distortion loss function, i.e. where D is x and The distortion between them, R is expressed from the quantized is the bit rate calculated or estimated, and λ is the Lagrange multiplier. It should be noted that D can be calculated in the pixel domain or the perceptual domain. All existing research works follow this prototype, and the difference may only be the network structure or the loss function. In terms of network structure, RNN and CNN are the most widely used architectures. In the RNN-related category, a general framework for variable rate image compression using RNN is proposed. Binary quantization is used to generate codes and the bitrate is not considered during training. The framework actually provides scalable codec capabilities, where RNNs with convolutional and deconvolutional layers are reported to perform well. An improved version is then proposed that utilizes a neural network upgrade encoder similar to PixelRNN to compress binary codes. It is reported that the performance is better than JPEG on the Kodak image dataset using the MS-SSIM evaluation metric. The RNN-based solution is further improved by introducing hidden state activation. In addition, an SSIM weighted loss function is designed and a spatially adaptive bitrate mechanism is enabled. They achieve better results than BPG on the Kodak image dataset using MS-SSIM as the evaluation metric. Spatially adaptive bitrate is supported by training a stop code tolerant RNN. A general framework for rate-distortion optimized image compression is proposed. Multivariate quantization is used to generate integer codes and the bitrate during training is considered, i.e., the loss is a joint rate-distortion cost, which can be MSE or other. They add random uniform noise during training to motivate quantization and use the differential entropy of the noise code as a proxy for the bitrate. They use generalized divisive normalization (GDN) as the network structure, which includes a linear mapping followed by a nonlinear parameter normalization. The effectiveness of GDN for image coding and decoding is verified. An improved version is then proposed, in which they use 3 convolutional layers, each followed by a downsampling layer and a GDN layer as a forward transform. Therefore, they use 3 layers of inverse GDN, each followed by an upsampling layer and a convolutional layer to motivate the inverse transform. In addition, an arithmetic coding method is designed to compress integer codes. It is reported that the performance is better than JPEG and JPEG 2000 on the Kodak dataset in terms of MSE. In addition, they improve the method by designing the variance hyper-prior into the autoencoder. They utilize the subnet h a Transform the latent representation y into z = h a (y), and z will be quantized and transmitted as side information. Therefore, we try to use the quantized side information Decoded to quantized The standard deviation of the subnet h s To achieve the inverse transformation, this will The Z.Z. Gaussian mixture model is used to further remove redundancy in the residual. The reported performance is comparable to that of VVC on the Kodak image set using PSNR as the evaluation metric. 2.3.3 Super Prior Model Figure 3 An example latent representation of an image is shown, including an image 300 from the Kodak dataset, a visualization of the latent value 310 representation y of the image 300, the standard deviation σ 320 of the latent value 310, and the time delay y 330 after the introduction of the hyper-prior network. The hyper-prior network includes an encoder and a decoder that utilize hyper-prior information. In the transform codec method for image compression, the encoder subnet (Section 2.3.2) uses a parameter analysis transform The image vector x is transformed into a latent representation y, which is then quantized to form because is a discrete value, so it can be losslessly compressed using entropy coding techniques (such as arithmetic coding) and transmitted as a bit sequence. As from Figure 3 The potential value 310 and standard deviation σ320 are obvious. There is significant spatial dependence between the elements of . Notably, their scales (center-right image) appear to be spatially coupled. Introducing another set of random variables to capture spatial dependencies and further reduce redundancy. In this case, the image compression network Figure 4 Shown in. Figure 4 4 is a schematic diagram showing an example network architecture of an autoencoder implementing a super-prior model. The upper side shows the image autoencoder network, while the lower side corresponds to the super-prior subnet. The analysis transform and the synthesis transform are denoted as g a and g a Q represents quantization, AE and AD represent arithmetic encoder and arithmetic decoder respectively. The super prior model consists of two subnetworks, an encoder using super prior information (using h a denoted by h) and a decoder using hyper-prior information (denoted by h s denoted by ). The hyper-prior model generates a quantized hyper-prior information potential value It includes information about the quantified potential value The probability distribution information of the sample points. is included in the bitstream and is related to are transmitted together to the receiver (decoder). exist Figure 4The left hand side of the model is the encoder g a and decoder g s (Explained in Section 2.3.2). The right-hand side is used to obtain The additional encoder h that utilizes the hyper-prior information a and a decoder using hyper-prior information h s In this architecture, the encoder subjects the input image x to g a , resulting in a response y with a spatially varying standard deviation. The response y is fed to h a In the example above, the standard deviation distribution in z is summarized. Then z is quantized The encoder then uses the quantized vector to estimate the spatial distribution of the standard deviation σ and use it to compress and transmit the quantized image representation The decoder first recovers the Then it uses h s to obtain σ, which provides the correct probability estimate for the successful recovery Then to g s feed to obtain a reconstructed image. When an encoder using super-prior information and a decoder using super-prior information are added to an image compression network, the quantized potential value The space redundancy is reduced. Figure 3 The rightmost image in (i.e., potential value y 330) corresponds to the quantized potential value when using an encoder with super-prior information / decoder with super-prior information. Compared to the center right image (i.e., standard deviation σ 320), the spatial redundancy is significantly reduced because the samples of the quantized potential value are less correlated. 2.3.4 Context Model Although the hyper-prior model improves the quantified potential , but additional improvement can be obtained by utilizing an autoregressive model that predicts the quantized potential value from the causal context of the potential value (context model). The term autoregressive means that the output of a process is later used as its input. For example, the context model subnet generates a sample of potential values that is later used as input to get the next sample. Figure 55 is a schematic diagram illustrating an example combined model configured to jointly optimize a context model along with a hyper-prior and an autoencoder. The combined model jointly optimizes an autoregressive component that, along with the hyper-prior and the underlying autoencoder, estimates a probability distribution of a potential value from its causal context (context model). The real-valued potential representation is quantized (Q) to create a quantized potential value and the quantified super-prior information potential value They are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD).The dashed area corresponds to the components executed by a receiver (eg, a decoder) to recover the image from the compressed bitstream. The joint architecture utilizes both the super-prior model subnet (encoder utilizing super-prior information and decoder utilizing super-prior information) and the context model subnet. The super-prior and context models are combined to learn the quantized latent values The probability model of is then used for entropy coding and decoding. Figure 5 As shown, the output of the context subnetwork and the decoder subnetwork using hyper-prior information is called the subnetwork combination of entropy parameters, which generates the mean μ and scale (or variance) σ parameters for the Gaussian probability model. The Gaussian probability model is then used to encode the quantized potential value samples into the bitstream with the help of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized potential value from the bitstream through the arithmetic decoder (AD) module Typically, the potential sample points are modeled as a Gaussian distribution or a Gaussian mixture model (but not limited to this). Figure 5 In , the context model and the hyper-prior are jointly used to estimate the probability distribution of the underlying samples. Since a Gaussian distribution can be defined by a mean and a variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ). 2.3.5 Gain Variational Autoencoder (G-VAE) Typically, neural network-based image / video compression methods require training multiple models to adapt to different bitrates. Gain Variational Autoencoder (G-VAE) is a variational autoencoder with a pair of gain units. G-VAE is designed to achieve continuous variable bitrate adaptation using a single model. G-VAE includes a pair of gain units, which are usually inserted into the output of the encoder and the input of the decoder. The output of the encoder is defined as the latent representation y∈R c*h*w , where c, h, w represent the number of channels, the height and width of the latent representation. Each channel of the latent representation is represented as y (i) ∈R h*w , where i = 0, 1, ..., c-1. A pair of gain units includes a gain matrix M∈R c*nand the inverse gain matrix, where n is the number of gain vectors. The gain vector can be represented as m s ={α s(0) ,α s(1) ,…,α s(c-1)}, α s(i) ∈R, where s represents the index of the gain vector in the gain matrix. The motivation of the gain matrix is similar to the quantization table in JPEG, which controls the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector. where ⊙ is the channel-by-channel multiplication, i.e. And α s(i) is the gain vector m s The inverse gain matrix used at the decoder side can be expressed as M′∈R c*n , which consists of n inverse gain vectors, namely, M′={δ s(0) ,δ s(1) ,…,δ s(c-1)}, δ s(i) ∈R. The inverse gain process is expressed as: in is the decoded quantized latent representation, and y′ s is the quantized potential representation of the inverse gain that will be fed into the synthesis network. To achieve continuously variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors {m t ,m′ t} and {m r ,m′ r}, and the interpolated gain vector can be obtained via the following equation. m v =[(m r ) l ·(m t ) 1-l ] m′ v =[(m′ r ) l ·(m′ t ) 1-l ] Where l∈R is the interpolation coefficient that controls the corresponding bit rate of the generated gain vector pair. Since l is a real number, any bit rate between two given gain vector pairs can be achieved. 2.3.6 Encoding process using joint autoregressive hyper-prior model Figure 5 This corresponds to a compression method using a joint autoregressive hyper-prior model. In this section and the next, the encoding process and the decoding process will be described respectively. Figure 6 The encoding process 600 is depicted. The input image is first processed using the encoder subnet. The encoder transforms the input image into a transformed representation called a potential value, represented by y. Then y is input to a quantizer block represented by Q to obtain a quantized potential value Then The arithmetic coding block (denoted as AE) is converted into a bit stream (bits1). The arithmetic coding block converts the Each sample point is converted into the bit stream (bits1). The encoder using hyper-prior information, context, decoder using hyper-prior information, and entropy parameter subnetworks are used to estimate the quantized potential value The potential value y is input to the encoder using super-prior information, and the encoder using super-prior information outputs the super-prior information potential value (represented by z). The super-prior information potential value is then quantized And a second bitstream (bits2) is generated using an arithmetic coding (AE) module. The decomposition entropy module generates a probability distribution that is used to encode the quantized super-prior information potential value into the bitstream. The quantized super-prior information potential value includes information about the quantized potential value information about the probability distribution of . Entropy parameter subnetwork generation is used to encode quantized latent values The information generated by the entropy parameters usually includes the mean μ and scale (or variance) σ parameters which are used together to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as Where the parameter μ is the mean or expectation of the distribution (as well as its median and mode), and the parameter σ is its standard deviation (or variance or scale). In order to define a Gaussian distribution, the mean and variance need to be determined. The Entropy Parameters module is used to estimate the mean and variance values. The decoder of the subnet using the hyper-prior information generates part of the information used by the entropy parameter subnet, the other part of which is generated by an autoregressive module called the context module. The context module generates information about the probability distribution of the samples of the quantized potential value using the samples that have been encoded by the arithmetic coding (AE) module. The quantized potential value It is usually a matrix consisting of many sample points. The sample points can be represented by or The index to be indicated depends on the matrix The dimension of . The samples are encoded one by one by the AE, usually using raster scan order. In raster scan order, the rows of the matrix are processed from top to bottom, where the samples in a row are processed from left to right. In this scenario (where raster scan order is used by the AE to encode the samples into the bitstream), the context module uses the samples encoded before in raster scan order to generate the same samples as The information generated by the context module and the decoder using the hyper-prior information is combined by the entropy parameter module to generate a parameter for converting the quantized potential value Probability distribution encoded into the bitstream (bits1). Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to a decoder. Note that other names can be used for the above modules. In the above description, Figure 6 All elements in are collectively referred to as encoders. The analysis transformation that transforms the input image into a latent representation is also called an encoder (or autoencoder). 2.3.7 Decoding process using the joint autoregressive hyper-prior model Figure 7 The decoding process is depicted separately corresponding to the encoding process 600 . In the decoding process 700, the decoder first receives a first bit stream (bits1) and a second bit stream (bits2) generated by a corresponding encoder. Bits2 is first decoded by an arithmetic decoding (AD) module using a probability distribution generated by a factorized entropy subnet. The factorized entropy module typically generates the probability distribution using a predetermined template, such as a predetermined mean and variance value in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is is the quantized super-prior information potential value. The AD process is restored to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized super-prior information potential value generated by the encoder is can be reconstructed at the decoder without any changes. In obtaining after, is processed by a decoder utilizing super-prior information, the output of which is fed to the entropy parameter module. The three subnetworks, context, decoder utilizing super-prior information, and entropy parameters used in the decoder are the same as those in the encoder. Therefore, exactly the same probability distribution can be obtained in the decoder (as in the encoder), which is very important for reconstructing the quantized potential values without any loss. As a result, the quantized potential value obtained in the encoder The same version of can be obtained in the decoder. After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of the quantized potential values one by one from the bitstream bits1. From a practical point of view, the autoregressive model (context model) is inherently serial and therefore cannot be accelerated using techniques such as parallelization. Finally, the fully reconstructed quantized potential value Input to the composite transform (in Figure 7 The decoder module is used to obtain the reconstructed image. In the above description, Figure 7 All elements in are collectively referred to as a decoder. The synthetic transform that converts the quantized potential values into the reconstructed image is also called a decoder (or auto-decoder). 2.4 Neural Networks for Video Compression Similar to conventional video codec technology, neural image compression serves as the basis for intra-frame compression in neural network-based video compression, so the development of neural network-based video compression technology is later than that of neural network-based image compression, but more efforts are needed to solve the challenges caused by its complexity. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared with image compression, video compression requires effective methods to remove inter-picture redundancy. Subsequently, inter-picture prediction is a key step in these works. Motion estimation and compensation are widely adopted, but have only recently been implemented by trained neural networks. The research on neural network-based video compression can be divided into two categories according to the target scenario: random access and low latency. In the case of random access, decoding can be started from any point in the sequence, usually the entire sequence is divided into multiple separate segments, and each segment can be decoded independently. In the case of low latency, the goal is to reduce the decoding time, so that usually only the temporally previous frame can be used as a reference frame to decode the subsequent frame. 2.4.1 Low Latency A video compression scheme with a trained neural network is proposed. The video sequence frames are divided into blocks and each block will choose a mode from two available modes (intra-frame codec or inter-frame codec). If intra-frame codec is selected, there is an associated autoencoder for compressing the block. If inter-frame codec is selected, motion estimation and compensation are performed using traditional methods and the trained neural network will be used for residual compression. The output of the autoencoder is directly quantized and encoded by the Huffman method. Another neural network based video codec scheme with PixelMotionCNN is proposed. Frames are compressed in temporal order and each frame is divided into blocks which are compressed in raster scan order. Each frame will first be extrapolated with the previous two reconstructed frames. When a block is to be compressed, the extrapolated frame is fed into PixelMotionCNN along with the context of the current block to derive the latent representation. The residual is then compressed by a variable rate image scheme. The performance of this scheme is comparable to H.264. A truly end-to-end neural network-based video compression framework is proposed, where all modules are implemented using neural networks. The scheme accepts the current frame and the previously reconstructed frame as input, and the optical flow will be derived using a pre-trained neural network as motion information. The motion information will be warped with the reference frame, and then the neural network will generate a motion compensated frame. Two separate neural autoencoders are used to compress the residual and motion information. The entire framework is trained using a single rate-distortion loss function. It achieves better performance than H.264. A video compression scheme based on advanced neural networks is proposed. This scheme inherits and extends the traditional video coding scheme of neural networks, and has the following main features: 1) only one autoencoder is used to compress motion information and residuals; 2) multiple frames and multiple optical flows are used for motion compensation; 3) online states are learned and propagated through subsequent frames over time. The scheme achieves better performance than the HEVC reference software in MS-SSIM. An extended end-to-end neural network-based video compression framework is proposed. In this solution, multiple frames are used as references. Therefore, it is possible to provide a more accurate prediction of the current frame by using multiple reference frames and the associated motion information. In addition, motion field prediction is deployed to remove motion redundancy along the temporal channel. A post-processing network is also introduced in this work to remove reconstruction artifacts from the previous process. The performance is significantly better than H.265 in terms of PSNR and MS-SSIM. Scale-space flow was proposed to replace the commonly used optical flow by adding a scaling parameter based on the frame. It is reported to achieve better performance than H.264. A multi-resolution representation for optical flow is proposed. Specifically, the motion estimation network generates multiple optical flows with different resolutions and lets the network learn which one to choose under a loss function. The performance is slightly improved compared to H.265. 2.4.2 Random Access A neural network-based video compression scheme with frame interpolation is proposed. Key frames are first compressed using a neural image compressor, and the remaining frames are compressed in a hierarchical order. They perform motion compensation in the perceptual domain, i.e., feature maps are derived at multiple spatial scales of the original frame, and the motion is used to warp the feature maps that will be used in the image compressor. The method is reported to be comparable to H.264. An interpolation-based video compression method is proposed, in which the interpolation model combines motion information compression and image synthesis, and the same autoencoder is used for both the image and the residual. A neural network-based video compression method based on a variational autoencoder with a deterministic encoder is proposed. Specifically, the model consists of an autoencoder and an autoregressive prior. Unlike previous methods, this method accepts group of pictures (GOP) as input and incorporates a 3D autoregressive prior by considering temporal correlations while encoding and decoding latent representations. It provides performance comparable to H.265. 2.5 Preliminary Knowledge Almost all natural images / videos are in digital format. Grayscale digital images can be obtained by Indicates that is a set of values for pixels, m is the image height and n is the image width. For example, is a common setting, and in this case Therefore a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed bits are necessarily less. Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by Representation, using three separate channels to store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24bpp. Digital images / videos can be represented in different color spaces. Neural network-based video compression schemes are mainly developed in the RGB color space, while traditional codecs typically use the YUV color space to represent video sequences. In the YUV color space, the image is decomposed into three channels, namely Y, Cb, and Cr, where Y is the brightness component and Cb / Cr are the chrominance components. Since the human visual system is less sensitive to the chrominance components, Cb and Cr are usually downsampled for pre-compression. A color video sequence consists of multiple color images (called frames) to record scenes at different time stamps. For example, in the RGB color space, a color video can be represented by X = {x 0 ,x 1 ,…,x t ,…,x T-1}, where T is the number of frames in the video sequence, If m=1080, n=1920, And the video has 50 frames per second (fps), then the data bit rate of the uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits per second (bps), about 2.32 Gbps, which requires a lot of storage, so it must be compressed before transmission over the Internet. Typically, for natural images, lossless methods can achieve a compression ratio of about 1.5 to 3, which is clearly lower than required. Therefore, lossy compression is developed to achieve further compression ratios, but at the expense of inducing distortion. Distortion can be measured by calculating the average squared difference between the original image and the reconstructed image, i.e., the mean square error (MSE). For grayscale images, MSE can be calculated using the following equation. Therefore, the quality of the reconstructed image compared to the original image can be measured by the Peak Signal-to-Noise Ratio (PSNR): in for The maximum value in , for example 255 for an 8-bit grayscale image. There are other quality assessment metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM). In order to compare different lossless compression schemes, it is sufficient to compare the compression ratio at a given resulting bit rate or vice versa. However, in order to compare different lossy compression methods, both the bit rate and the reconstruction quality must be considered. For example, to calculate the relative bit rates of several different quality levels and then average these bit rates is a common approach; the average relative bit rate is called Bjontegaard's delta-rate (BD rate). There are other important aspects of evaluating image / video codec schemes, including encoding / decoding complexity, scalability, robustness, etc. 3. Question 3.1 Core Issues like Figure 7 As shown, the synthesis network is responsible for the last step of the decoding process, that is, synthesizing the reconstructed image from the decoded feature map. The entire decoder usually includes multiple synthesis networks to cover the applicable range of bit rates. The current method has the following problems. Problem 1. Due to the varying content of image / video content, using a single synthesis network for that particular rate point may not be optimal. For example, the characteristics of screen content are different from natural content images. Using a synthesis network to process screen content images is likely to result in performance degradation. Problem 2. In some scenarios, the synthesis network is heavy in terms of the number of parameters and computational complexity. Therefore, more synthesis networks mean more burden on storage space and computational consumption. 4. Detailed solution To solve problem 1, a set of N pre-trained sub-network weights can be trained based on a classification method. For example, the training set is classified into two categories based on whether it is screen content; the training set is classified into multiple categories based on object categories (such as people, landscapes, buildings, etc.). During decoding, a synthetic network is selected from the group based on the current content. To solve problem 2, we can use fewer synthetic networks and use the entropy encoding and decoding part (such as Figure 8 The prediction fusion part in ) is used to compensate for the potential performance loss. The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. 4.1 Objectives of the Implementation Example An objective of embodiments of the present disclosure is to improve codec efficiency and reduce model size for neural network-based image / video compression. 4.2 Core of the Implementation Example The synthesis can be selected from a set of multiple pre-trained sub-network weights. The total number of synthesis networks can be reduced, and more prediction options in the entropy codec part can be used to implement variable rate codecs. 4.2.1 Decoding process Figure 8 A decoding process 800 is shown with independent subnetwork weight selection 810 for a synthetic network. According to the present disclosure, the decoding operation is performed as follows: 1. First, the decomposition entropy model is used to decode the quantized latent values, i.e., Fig. 9 In 2. The probability parameters (eg, variance) generated by the second network are used to generate quantized residual potential values by performing an arithmetic decoding process. 3. Such as Figure 8 As shown, the quantized residual potential value is inversely gained using an inverse gain unit (iGain). The output of the inverse gain unit is expressed as 4. Perform the following steps in a loop until you get All elements of: a. The first subnet is used to use the obtained quantized potential values The sample points are used to estimate The mean parameter of . b. The quantized residual potential value and mean are used to obtain The next element of . 5. Select a synthetic network from the set of pre-trained sub-network weights. The index of the synthetic network can be encoded in the bitstream. 6. After the synthesis network is determined and After all the sample points of are obtained, the reconstruction can be derived. In the above Figure 8 An exemplary implementation of the embodiment is depicted in (Decoding Process). The framework shows how to decode the luma component and the chroma components can be decoded using the same structure. The first subnet includes context, prediction and optionally a decoder module using super prior information. The second network includes a variance decoder module using super prior information. The quantized super prior information potential is The arithmetic decoding process generates a quantized residual potential value, which is further fed into the iGain unit to obtain the gained quantized residual potential value After obtaining the residual latent value, a recursive prediction operation is performed to obtain the latent value and The following steps describe how to obtain the potential value Sample points. 1. The autoregressive context module is used to use the sample points The first input to the prediction module is generated where the (m,n) pair is the index of the sample point where the potential value has been obtained. 2. Optionally, the second input to the prediction module is obtained by using a decoder using super-prior information and a quantized super-prior information potential value Obtained. 3. Using the first input and the second input, the prediction module generates the mean mean[:,i,j]. 4. The mean mean[:,i,j] and the quantized residual potential value Add together to get the potential value 5. Repeat steps 1-4 for the next sample point. Whether and / or how to apply at least one method disclosed in this document may be signaled from an encoder to a decoder, for example in a bitstream. Alternatively, whether and / or how to apply at least one method disclosed in the document may be determined by a decoder based on codec information such as size, color format, etc. Alternatively or additionally, the designation MS1, MS2 or MS3+O (in Figure 8A module (in) may be included in a process flow. The module may perform an operation on its input by multiplying the input with a scalar or adding an additive component to the input to obtain an output. The scalar or additive component used by the module may be indicated in the bitstream. exist Figure 8 The module named RD or the module named AD in the example may be an entropy decoding module. It may be a range decoder or an arithmetic decoder, etc. The invention described herein is not limited to Figure 8 Some modules may be missing and some modules may be shifted in the processing order. In addition, additional modules may be included, such as post-processing filters, using the reconstructed luma component to help chroma reconstruction in the synthesis network, etc. 4.2.2 Reducing the number of synthetic transformations Typically multiple synthesis transforms involve different bit rate ranges to support continuously variable rates. However, it is memory consuming and computationally heavy. In terms of memory efficiency and computational efficiency, it is always beneficial to reduce the number of synthesis networks used in the decoding process. To achieve this goal, the number of synthesis networks is reduced, for example from 4 to 3 in some cases. In order to compensate for possible performance degradation and support variable rate coding, there are more options for the prediction / entropy coding part. 4.3 Example Embodiments The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. 1. In one example, it is possible to include N 1 The composite transformation weight is selected from the set of subnetwork weights. Figure 8 As shown, it can be obtained from N 2 The prediction fusion network weight is selected from the set of subnetwork weights. 2. In one example, one synthesis can be trained for screen content and another for natural content, and a single entropy codec network can be used for both. 3. In one example, there may be a reduced number of synthesis networks. To support continuous variable rate coding, the prediction fusion part may be enhanced by changing the network architecture (such as the number of convolutional layers, resampling layer type, activation layer type). 5. Other embodiments 1.Decoder: a) Independent subnetwork weight selection scheme for neural network based image or video compression. b) A continuous rate codec scheme using less synthetic networks, where the entropy coding part is enhanced To compensate for the resulting performance degradation.
[0033] Fig. 9 A flow chart of a method 900 for visual data processing according to an embodiment of the present disclosure is shown. The method 900 is implemented for conversion between visual data and a bit stream of the visual data.
[0034] At block 910, target weights are determined for use by a target module in a codec system based on information associated with the visual data. The codec system is implemented using at least one neural network (NN). For example, the temporal data may be an image or a video. The codec system may be an end-to-end image or video compression system. For example, the codec system may perform Figure 6 The encoding process 600 and / or Figure 8 The decoding process 800 in .
[0035] At block 920, the conversion is performed based on the target weights using a codec system. In some embodiments, the conversion includes decoding the visual data from a bitstream. Alternatively or additionally, in some embodiments, the conversion includes encoding the visual data into a bitstream.
[0036] Method 900 enables weights for modules in a codec system to be determined based on information associated with visual data (such as visual data content). In this way, the performance of the modules can be improved. Therefore, the codec effect and codec efficiency of the codec system can be improved.
[0037] In some embodiments, the information associated with the visual data includes at least one of the following: the content of the visual data, or the category of the object of the visual data. For example, the category of the object includes at least one of a person, a landscape, or a building. The content of the visual data may include one of a screen content or a natural content.
[0038] In some embodiments, the target weight is selected from a plurality of candidate weights, which are determined by training at least one synthesis module based on a training data set associated with at least one of a plurality of contents of the visual data or a plurality of categories of objects of the visual data, and the at least one synthesis module is used to determine the reconstruction of the visual data. In some example embodiments, the candidate weights are synthesis transformation weights. The target synthesis transformation weight can be selected from a set of subnetwork weights based on classification training. For example, the training data set can be classified into two categories based on whether it is screen content. For another example, the training data set can be classified into multiple categories based on object categories. During the decoding process, a target module such as a synthesis network can be selected from a pre-trained weight set based on the current content of the visual data. In some example embodiments, one synthesis can be trained for screen content and another synthesis can be trained for natural content. A single entropy encoding and decoding network is used for both synthesis networks.
[0039] In some embodiments, performing the conversion includes: determining an index of a target module from a bitstream; determining the target module from a plurality of candidate modules based on the index, the plurality of candidate modules being trained based on a plurality of candidate weights; and performing the conversion by determining a reconstruction of the visual data using the target module based on the target weights and a representation of the visual data. For example, the target synthesis network may be selected from a set of pre-trained subnetwork weights. The index of the target synthesis network may be encoded in the bitstream, e.g. Figure 8 Index 1, 2, 3, 4, ... or N as shown.
[0040] In some embodiments, the target weights include a set of weight values, and determining the reconstruction of the visual data includes determining the reconstruction of the visual data using a target module based on the set of weight values and a plurality of samples of the representation of the visual data.
[0041] In some embodiments, method 900 further includes determining an index of the target module based on information associated with the visual data.
[0042] In some embodiments, the first number of multiple candidate modules is less than the second number of candidate modules in another codec system that encodes and decodes visual data without determining the target weight. In other words, fewer synthesis networks are used. For example, the number of synthesis networks is reduced, such as from 4 to 3 in some cases. More options for prediction in the entropy codec part can be used to implement variable rate codecs. In this way, codec efficiency can be improved. In addition, the model size for neural network-based image or video compression can be reduced.
[0043] In some embodiments, performing the conversion includes: determining at least one sample of a representation of the visual data using a prediction module in the codec system; and determining a reconstruction of the visual data using the target module based on a target weight and the at least one sample.
[0044] In some embodiments, determining the at least one sample point includes: determining a prediction weight from a plurality of candidate prediction weights based on information associated with the visual data; and determining the at least one sample point by using a prediction module based on the prediction weight.
[0045] In some embodiments, method 900 further includes: updating at least one of the first architecture of the prediction module in the codec system or the second architecture of the entropy codec module in the codec system by modifying at least one of the following: the number of convolutional layers in at least one architecture, the type of resampling layers in at least one architecture, or the type of activation layers in at least one architecture. For example, there may be a reduced number of synthetic networks. In order to support continuous variable rate coding, the prediction fusion part can be enhanced by changing the network architecture, such as the number of convolutional layers, the type of resampling layers, the type of activation layers, etc. In this way, a continuous rate coding scheme using fewer synthetic networks is implemented, in which the entropy codec part is enhanced to compensate for the induced performance degradation.
[0046] In some embodiments, the codec system includes a decomposition entropy module implemented using at least one neural network, a variance codec module using super-prior information, a context module and a target module, and wherein performing the conversion includes: determining a first representation of the visual data based on the bitstream by using the decomposition entropy module; determining a first probability parameter of the visual data based on the first representation by using the variance codec module using super-prior information; determining a residual representation of the visual data based on the first probability parameter and the bitstream; determining a second representation of the visual data based on the first representation and the residual representation by using the context module; and determining reconstruction of the visual data based on the second representation by using the target module based on the target weight.
[0047] In some embodiments, the residual representation is also determined based on the gain module.
[0048] In some embodiments, the residual representation comprises a quantized residual representation.
[0049] In some embodiments, the context module includes an autoregressive context module and a prediction module, and determining the second representation includes: determining a first intermediate representation based on a first sample point of the second representation by using the autoregressive context module; determining a second probability parameter of the visual data based on at least the first intermediate representation by using the prediction module; and determining a second sample point of the second representation based on the second probability parameter and the residual representation.
[0050] In some embodiments, the context module also includes a codec module utilizing super-prior information, and determining the second probability parameter includes: determining a second intermediate representation based on the first representation by using the codec module utilizing super-prior information; and determining the second probability parameter based on the first intermediate representation and the second intermediate representation by using a prediction module.
[0051] In some embodiments, the second probability parameter is further determined by the prediction module based on a prediction weight selected from a plurality of prediction weights.
[0052] In some embodiments, the visual data includes a luminance component and a chrominance component. The luminance component and the chrominance component may use the same structure (e.g. Figure 8 The structure shown in ) is used for encoding and decoding.
[0053] In some embodiments, the coding system further comprises a scaling module, the scaling module being configured to scale an input of the scaling module based on a scaling factor. In some embodiments, the scaling factor is included in the bitstream.
[0054] In some embodiments, the codec system further comprises an addition module, the addition module being configured to add the addition factor to an input of the addition module. In some embodiments, the addition factor is included in the bitstream.
[0055] In some embodiments, the coding and decoding system further includes at least one of the following: an entropy coding and decoding module, a range coding and decoding module, or an arithmetic coding and decoding module.
[0056] In some embodiments, the target module includes a synthesis module for determining a reconstruction of the visual data.
[0057] In some embodiments, further information about the application method is included in the bitstream. In some embodiments, the further information indicates at least one of: whether to apply the method, or how to apply the method.
[0058] In some embodiments, the method 900 further includes: determining another information based on the encoding and decoding information of the visual data.
[0059] In some embodiments, the codec information includes at least one of the following: the size of the visual data, or the color format of the visual data.
[0060] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a device for visual data processing. In the method, based on information associated with the visual data, a target weight used by a target module in a codec system is determined. The codec system is implemented using at least one neural network. A bitstream is generated based on the target weight by using a coding system.
[0061] According to another embodiment of the present disclosure, a method for storing a bitstream of visual data is provided. In the method, a target weight for use by a target module in a codec system is determined based on information associated with the visual data. The codec system is implemented using at least one neural network. A bitstream is generated based on the target weight using a coding system. The bitstream is stored in a non-transitory computer-readable recording medium.
[0062] Implementations of the present disclosure may be described according to the following items, and features of the following items may be combined in any reasonable manner.
[0063] Item 1. A method for visual data processing, comprising: determining, for conversion between visual data and a bit stream of the visual data, a target weight for use by a target module in a codec system based on information associated with the visual data, the codec system being implemented using at least one neural network; and performing the conversion based on the target weight by using the codec system.
[0064] Item 2. A method according to Item 1, wherein the information associated with the visual data includes at least one of the following: content of the visual data, or a category of an object of the visual data.
[0065] Item 3. A method according to Item 2, wherein the category of the object includes at least one of a person, a landscape, or a building.
[0066] Item 4. A method according to Item 2, wherein the content of the visual data includes one of screen content or natural content.
[0067] Item 5. A method according to any one of Items 2-4, wherein the target weight is selected from a plurality of candidate weights, and the plurality of candidate weights are determined by training at least one synthesis module based on a training data set associated with a plurality of contents of the visual data or at least one of a plurality of categories of the objects of the visual data, and the at least one synthesis module is used to determine the reconstruction of the visual data.
[0068] Item 6. A method according to any one of Items 1-5, wherein performing the conversion comprises: determining an index of the target module from the bitstream; determining the target module from a plurality of candidate modules based on the index, the plurality of candidate modules being trained based on the plurality of candidate weights; and performing the conversion by determining a reconstruction of the visual data using the target module based on the target weights and a representation of the visual data.
[0069] Item 7. A method according to Item 6, wherein the target weights include a set of weight values, and determining the reconstruction of the visual data includes: determining the reconstruction of the visual data by using the target module based on the set of weight values and multiple sample points representing the visual data.
[0070] Item 8. The method according to Item 6 also includes: determining the index of the target module based on the information associated with the visual data.
[0071] Item 9. A method according to any one of Items 6-8, wherein a first number of the plurality of candidate modules is less than a second number of candidate modules in another codec system that encodes and decodes the visual data without determining the target weight.
[0072] Item 10. A method according to any one of Items 1-9, wherein performing the conversion includes: determining at least one sample point representing the visual data by using a prediction module in the codec system; and determining a reconstruction of the visual data by using the target module based on the target weight and the at least one sample point.
[0073] Item 11. A method according to Item 10, wherein determining the at least one sample point comprises: determining a prediction weight from a plurality of candidate prediction weights based on the information associated with the visual data; and determining the at least one sample point by using the prediction module based on the prediction weight.
[0074] Item 12. The method according to any one of Items 1-11 further includes: updating at least one of the first architecture of the prediction module in the codec system or the second architecture of the entropy codec module in the codec system by correcting at least one of the following: the number of convolutional layers in the at least one architecture, the type of resampling layers in the at least one architecture, or the type of activation layers in the at least one architecture.
[0075] Item 13. A method according to any one of Items 1-12, wherein the codec system includes a decomposition entropy module implemented using the at least one neural network, a variance codec module using super-prior information, a context module and the target module, and wherein performing the conversion includes: determining a first representation of the visual data based on the bit stream by using the decomposition entropy module; determining a first probability parameter of the visual data based on the first representation by using the variance codec module using super-prior information; determining a residual representation of the visual data based on the first probability parameter and the bit stream; determining a second representation of the visual data based on the first representation and the residual representation by using the context module; and determining a reconstruction of the visual data based on the second representation by using the target module based on the target weight.
[0076] Item 14. The method of Item 13, wherein the residual representation is further determined based on a gain module.
[0077] Item 15. A method according to Item 13 or Item 14, wherein the residual representation comprises a quantized residual representation.
[0078] Item 16. A method according to any one of Items 13-16, wherein the context module includes an autoregressive context module and a prediction module, and determining the second representation includes: determining a first intermediate representation based on a first sample point of the second representation by using the autoregressive context module; determining a second probability parameter of the visual data based on at least the first intermediate representation by using the prediction module; and determining a second sample point of the second representation based on the second probability parameter and the residual representation.
[0079] Item 17. A method according to Item 16, wherein the context module also includes a codec module utilizing super-prior information, and determining the second probability parameter includes: determining a second intermediate representation based on the first representation by using the codec module utilizing super-prior information; and determining the second probability parameter based on the first intermediate representation and the second intermediate representation by using the prediction module.
[0080] Item 18. A method according to Item 16 or Item 17, wherein the second probability parameter is also determined by the prediction module based on a prediction weight selected from a plurality of prediction weights.
[0081] Item 19. A method according to any one of Items 1-18, wherein the visual data includes a luminance component and a chrominance component.
[0082] Item 20. A method according to any one of items 1-19, wherein the encoding and decoding system further comprises a scaling module, the scaling module being used to scale an input of the scaling module based on a scaling factor.
[0083] Item 21. The method of Item 20, wherein the scaling factor is included in the bitstream.
[0084] Item 22. A method according to any one of items 1-21, wherein the codec system further comprises an addition module, the addition module being used to add an addition factor to an input of the addition module.
[0085] Item 23. The method of Item 22, wherein the addition factor is included in the bitstream.
[0086] Item 24. A method according to any one of Items 1-23, wherein the coding and decoding system further comprises at least one of the following: an entropy coding and decoding module, a range coding and decoding module or an arithmetic coding and decoding module.
[0087] Item 25. A method according to any one of Items 1-24, wherein the target module comprises a synthesis module for determining a reconstruction of the visual data.
[0088] Item 26. A method according to any of items 1-25, wherein further information about applying the method is included in the bitstream.
[0089] Item 27. A method according to Item 26, wherein the further information indicates at least one of: whether to apply the method, or how to apply the method.
[0090] Item 28. The method according to Item 26 or Item 27 further includes: determining the other information based on the encoding and decoding information of the visual data.
[0091] Item 29. A method according to Item 28, wherein the codec information includes at least one of the following: the size of the visual data, or the color format of the visual data.
[0092] Item 30. A method according to any of Items 1-29, wherein the converting comprises decoding the visual data from the bitstream.
[0093] Item 31. A method according to any of Items 1-29, wherein the converting comprises encoding the visual data into the bitstream.
[0094] Item 32. An apparatus for visual data processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-31.
[0095] Item 33. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of Items 1-31.
[0096] Item 34. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by an apparatus for video processing, wherein the method comprises: determining, based on information associated with the visual data, target weights for use by a target module in a codec system, the codec system being implemented using at least one neural network; and generating the bitstream based on the target weights by using the coding system.
[0097] Item 35. A method for storing a bitstream of visual data, comprising: determining target weights for use by a target module in a codec system based on information associated with the visual data, the codec system being implemented using at least one neural network; generating the bitstream based on the target weights by using the coding system; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device
[0098] Fig.10 A block diagram of a computing device 1000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1000 may be implemented as a source device 100 (or a data encoder 104) or a destination device 120 (or a data decoder 124), or may be included in a source device 100 (or a data encoder 104) or a destination device 120 (or a data decoder 124).
[0099] It should be understood that Fig.10 The computing device 1000 shown in FIG. 1 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0100] like Fig.10 As shown, computing device 1000 includes a general computing device 1000. Computing device 1000 may include at least one or more processors or processing units 1010, memory 1020, storage unit 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060.
[0101] In some embodiments, the computing device 1000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 1000 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0102] The processing unit 1010 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 1020. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the computing device 1000. The processing unit 1010 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0103] The computing device 1000 typically includes various computer storage media. Such media can be any media accessible by the computing device 1000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1020 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 1030 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 1000.
[0104] The computing device 1000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Fig.10 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.
[0105] The communication unit 1040 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 1000 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0106] The input device 1050 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 1060 may be one or more of various output devices, such as a display, a speaker, a printer, etc. With the help of the communication unit 1040, the computing device 1000 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and the computing device 1000 may also communicate with one or more devices that enable a user to interact with the computing device 1000, or if necessary, the computing device 1000 may also communicate with any device (e.g., a network card, a modem, etc.) that enables the computing device 1000 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0107] In some embodiments, some or all components of the computing device 1000 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.
[0108] In an embodiment of the present disclosure, the computing device 1000 may be used to implement visual data encoding / decoding. The memory 1020 may include one or more visual data encoding / decoding modules 1025 having one or more program instructions. These modules are accessible and executable by the processing unit 1010 to perform the functions of the various embodiments described herein.
[0109] In an example embodiment performing visual data encoding, input device 1050 may receive visual data as input to be encoded 1070. The visual data may be processed, for example, by visual data encoding / decoding module 1025 to generate an encoded bitstream. The encoded bitstream may be provided as output 1080 via output device 1060.
[0110] In an example embodiment where visual data decoding is performed, input device 1050 may receive an encoded bitstream as input 1070. The encoded bitstream may be processed, for example, by visual data codec module 1025 to generate decoded visual data. The decoded visual data may be provided as output 1080 via output device 1060.
[0111] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for visual data processing, include: For conversion between visual data and a bitstream of the visual data, determining, based on information associated with the visual data, target weights for use by a target module in a codec system implemented using at least one neural network; as well as The converting is performed based on the target weight by using the codec system.
2. The method of claim 1, wherein the information associated with the visual data comprises at least one of the following: the content of the visual data, or A category of an object of the visual data. The method according to claim 2 , wherein the category of the object comprises at least one of a person, a landscape, or a building. The method of claim 2 , wherein the content of the visual data comprises one of screen content or natural content.
5. A method according to any one of claims 2-4, wherein the target weight is selected from a plurality of candidate weights, and the plurality of candidate weights are determined by training at least one synthesis module based on a training data set associated with a plurality of contents of the visual data or at least one of a plurality of categories of the objects of the visual data, and the at least one synthesis module is used to determine the reconstruction of the visual data.
6. The method according to any one of claims 1 to 5, wherein the conversion is performed include: Determining an index of the target module from the bitstream; determining the target module from a plurality of candidate modules based on the index, the plurality of candidate modules being trained based on the plurality of candidate weights; as well as The transforming is performed by determining a reconstruction of the visual data using the target module based on the target weights and a representation of the visual data.
7. The method of claim 6, wherein the target weights comprise a set of weight values and determining the reconstruction of the visual data include: The reconstruction of the visual data is determined using the target module based on the set of weight values and a plurality of samples of the representation of the visual data.
8. The method according to claim 6, further comprising: include: The index of the target module is determined based on the information associated with the visual data.
9. The method according to any one of claims 6-8, wherein a first number of the plurality of candidate modules is less than a second number of candidate modules in another codec system that encodes and decodes the visual data without determining the target weight.
10. The method according to any one of claims 1 to 9, wherein the conversion is performed include: determining at least one sample of a representation of the visual data by using a prediction module in the codec system; as well as A reconstruction of the visual data is determined using the target module based on the target weight and the at least one sample point.
11. The method according to claim 10, wherein determining the at least one sample point include: determining a prediction weight from a plurality of candidate prediction weights based on the information associated with the visual data; as well as The at least one sample point is determined by using the prediction module based on the prediction weight.
12. The method according to any one of claims 1 to 11, further comprising: include: At least one of the first architecture of the prediction module in the codec system or the second architecture of the entropy codec module in the codec system is updated by modifying at least one of the following: the number of convolutional layers in the at least one architecture, the type of resampling layer in the at least one architecture, or The type of activation layer in the at least one architecture.
13. The method according to any one of claims 1 to 12, wherein the codec system comprises a decomposition entropy module implemented using the at least one neural network, a variance codec module using hyper-prior information, a context module and the target module, and where the conversion is performed include: determining, based on the bitstream, a first representation of the visual data by using the decomposition entropy module; Based on the first representation, determining a first probability parameter of the visual data by using the variance encoding and decoding module using super-prior information; determining a residual representation of the visual data based on the first probability parameter and the bitstream; determining a second representation of the visual data by using the context module based on the first representation and the residual representation; as well as Based on the second representation, a reconstruction of the visual data is determined by using the target module based on the target weight. The method of claim 13 , wherein the residual representation is further determined based on a gain module.
15. A method according to claim 13 or claim 14, wherein the residual representation comprises a quantized residual representation.
16. The method according to any one of claims 13 to 16, wherein the context module comprises an autoregressive context module and a prediction module, and determines the second representation include: determining a first intermediate representation based on a first sample of the second representation using the autoregressive context module; determining, based at least on the first intermediate representation, a second probability parameter for the visual data by using the prediction module; as well as A second sample of the second representation is determined based on the second probability parameter and the residual representation.
17. The method according to claim 16, wherein the context module further comprises a codec module utilizing super prior information and determining the second probability parameter include: Based on the first representation, determine a second intermediate representation by using the encoding and decoding module using the super-prior information; as well as The second probability parameter is determined by using the prediction module based on the first intermediate representation and the second intermediate representation.
18. The method of claim 16 or claim 17, wherein the second probability parameter is further determined by the prediction module based on a prediction weight selected from a plurality of prediction weights.
19. The method of any one of claims 1-18, wherein the visual data comprises a luminance component and a chrominance component.
20. The method according to any one of claims 1-19, wherein the coding system further comprises a scaling module, the scaling module being configured to scale an input of the scaling module based on a scaling factor.
21. The method of claim 20, wherein the scaling factor is included in the bitstream.
22. The method according to any one of claims 1-21, wherein the codec system further comprises an addition module, wherein the addition module is configured to add an addition factor to an input of the addition module.
23. The method of claim 22, wherein the additive factor is included in the bitstream.
24. The method according to any one of claims 1-23, wherein the coding and decoding system further comprises at least one of the following: an entropy coding and decoding module, a range coding and decoding module, or an arithmetic coding and decoding module.
25. The method of any one of claims 1-24, wherein the target module comprises a synthesis module for determining a reconstruction of the visual data.
26. The method according to any of claims 1-25, wherein further information on applying the method is included in the bitstream.
27. The method of claim 26, wherein the further information indicates at least one of: whether to apply the method, or how to apply the method.
28. The method according to claim 26 or claim 27, further comprising: include: The other information is determined based on the codec information of the visual data.
29. The method according to claim 28, wherein the codec information comprises at least one of the following: the size of the visual data, or The color format of the visual data.
30. The method of any one of claims 1-29, wherein the converting comprises decoding the visual data from the bitstream.
31. The method of any one of claims 1-29, wherein the converting comprises encoding the visual data into the bitstream.
32. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-31.
33. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1-31.
34. A non-transitory computer-readable recording medium storing a bit stream of visual data generated by a method performed by an apparatus for video processing, wherein the method include: determining, based on information associated with the visual data, a target weight for use by a target module in a codec system implemented using at least one neural network; as well as The bitstream is generated based on the target weight by using the encoding system.
35. A method for storing a bit stream of visual data, include: determining, based on information associated with the visual data, a target weight for use by a target module in a codec system implemented using at least one neural network; generating the bitstream based on the target weight by using the encoding system; as well as The bit stream is stored in a non-transitory computer-readable recording medium.