Visual data processing method, device and medium
By integrating small wavelet and non-linear transformations through neural networks, the method addresses inefficiencies in existing image and video compression, achieving better encoding efficiency and reduced bit rates.
Patent Information
- Application Number
- CN202380083621.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-10
- Filing Date
- 2023-12-07
- Publication Date
- 2025-07-15
AI Technical Summary
The existing image/video compression technology still has room for improvement in encoding and decoding efficiency, especially the neural network-based methods are not yet mature, and it is difficult to effectively utilize the advantages of wavelet transform and nonlinear transform.
Combining wavelet transform and nonlinear transform, visual data processing is performed through a neural network, including the conversion between the current visual unit of the visual data and the bit stream, the wavelet subband representation is determined, and the image/video encoding and decoding is performed through inverse nonlinear transform and wavelet transform.
Improve the efficiency of image/video encoding and decoding, enhance the performance of learning image compression, and achieve more efficient compression and reconstruction quality.
Smart Images

Figure CN120323028A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to visual data processing technologies, and more particularly, to wavelet transforms and non-linear transforms. Background Art
[0002] Image / video compression is a fundamental technology for reducing the cost of image / video transmission and storage in a lossless or lossy manner. Image / video compression technologies can be divided into two branches, classical video codec methods and neural network-based video compression methods. Classical video codec schemes employ transform-based solutions, where researchers exploit statistical dependencies in latent variables (e.g., wavelet coefficients) through carefully hand-engineered entropy coding / decoding, which models the dependencies in the quantization mechanism. Neural network-based video compression has two forms, neural network-based coding / decoding tools and neural network-based end-to-end video compression. The former is embedded as a coding / decoding tool into existing classical video codecs and only serves as part of the framework, while the latter is a separate framework developed based on neural networks without relying on classical video codecs. It is generally desirable to further improve the coding / decoding efficiency of image / video coding / decoding. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method includes: for the conversion between a current visual unit of visual data and a bitstream of visual data, determining at least one wavelet subband representation of the current visual unit according to a neural network-based non-linear transform or a neural network-based inverse non-linear transform opposite to the non-linear transform, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; and performing the conversion based on the at least one wavelet subband representation. The method according to the first aspect of the present disclosure combines wavelets and non-linear transforms, as well as corresponding processing of visual data. In this way, the performance of learning visual data processing such as neural network-based visual data processing can be improved.
[0005] In a second aspect, a device for visual data processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of visual data, and the bitstream of visual data is generated by a method executed by a device for visual data processing. The method includes: determining at least one wavelet subband representation of a current visual unit of visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, where the wavelet subband representation is associated with a subband of a wavelet of the current visual unit; and generating a bitstream based on the at least one wavelet subband representation.
[0008] In a fifth aspect, a method for storing a bitstream of visual data is proposed. The method includes: determining at least one wavelet subband representation of a current visual unit of visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, where the wavelet subband representation is associated with a subband of a wavelet of the current visual unit; generating a bitstream based on the at least one wavelet subband representation; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] The summary of the invention is provided to introduce a selection of concepts in a simplified form, which will be further described in the detailed description below. The summary of the invention is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram of an exemplary visual data encoding and decoding system according to some embodiments of the present disclosure is shown;
[0012] Figure 2 A typical transform coding scheme is shown;
[0013] Figure 3 An image from the Kodak dataset and different representations of the image are shown;
[0014] Figure 4 The network architecture of an autoencoder implementing a hyperprior model is shown;
[0015] Figure 5 A block diagram of a combined model is shown;
[0016] Figure 6 The encoding process of the combined model is shown;
[0017] Figure 7Shows the decoding process of the combined model;
[0018] Figure 8 Shows an encoder and a decoder with wavelet-based transforms;
[0019] Figure 9 Shows the output of the forward wavelet-based transform;
[0020] Figure 10 Shows the segmentation of the output of the forward wavelet-based transform;
[0021] Figure 11 Shows an example of the decoding process according to an embodiment of the present disclosure;
[0022] Figure 12 Shows an example of the inverse channel network according to an embodiment of the present disclosure;
[0023] Figure 13 Shows an example of the inverse non-linear transform according to an embodiment of the present disclosure;
[0024] Figure 14 Shows another example of the inverse non-linear transform according to an embodiment of the present disclosure;
[0025] Figure 15 Shows a flowchart of a method for visual data processing according to an embodiment of the present disclosure;
[0026] Figure 16 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.
[0027] Throughout the drawings, the same or similar reference numerals generally denote the same or similar elements. Detailed Description
[0028] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustration and to assist those skilled in the art in understanding and implementing the present disclosure, and do not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.
[0029] In the following specification and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0030] As used herein, the terms "one embodiment", "embodiment", "exemplary embodiment", etc. indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes the specific features, structures, or characteristics. In addition, such phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an exemplary embodiment, it is considered within the knowledge of those skilled in the art to affect such features, structures, or characteristics associated with other embodiments, whether or not explicitly described.
[0031] It should be understood that although terms such as "first" and "second" may be used to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0032] As used herein, the term "visual data" may refer to image data or video data. The term "visual data processing" may refer to image processing or video processing. The term "visual data coding / decoding" may refer to image coding / decoding or video coding / decoding. The phrase "coding / decoding visual data" may refer to "encoding visual data (e.g., encoding visual data into a bitstream)" and / or "decoding visual data (e.g., decoding visual data from a bitstream)".
[0033] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that when the terms "comprises", "comprising", "has", "having", "includes", and / or "including" are used herein, they specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Exemplary Environment
[0034] Figure 1FIG. 0 is a block diagram showing an example visual data encoding and decoding system 100 that can utilize the techniques of the present disclosure. As shown, the visual data encoding and decoding system 100 can include a source device 110 and a destination device 120. The source device 110 may also be referred to as a data encoding device or a visual data encoding device, and the destination device 120 may also be referred to as a data decoding device or a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a data source 112, a data encoder 114, and an input / output (I / O) interface 116.
[0035] The data source 112 may include a source such as a data capture device. Examples of data capture devices include, but are not limited to, an interface for receiving data from a data provider, a computer graphics system for generating data, and / or a combination thereof.
[0036] The data may include one or more pictures or one or more images of a video. The data encoder 114 encodes the data from the data source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the data. The bitstream may include encoded pictures and associated data. The encoded pictures are encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded data may be directly transmitted to the destination device 120 via the I / O interface 116 over a network 130A. The encoded data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0037] The destination device 120 may include an I / O interface 126, a data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded data from the source device 110 or the storage medium / server 130B. The data decoder 124 may decode the encoded data. The display device 122 may display the decoded data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, which is configured to interface with an external display device.
[0038] The data encoder 114 and the data decoder 124 may operate according to a data encoding and decoding standard, such as a video encoding and decoding standard or a still picture encoding and decoding standard, and other existing and / or further standards.
[0039] Some example embodiments of the present disclosure will be described in detail below. It should be understood that the chapter headings used in this document are for ease of understanding and do not limit the embodiments disclosed in the chapter to that chapter only. In addition, although some embodiments are described with reference to multi-functional video coding or other specific data codecs, the disclosed techniques are also applicable to other coding and decoding techniques. Further, although some embodiments describe the coding and decoding steps in detail, it can be understood that the corresponding decoding steps will be implemented by the decoder, which undoes the coding and decoding. In addition, the term data processing includes data coding or compression, data decoding or decompression, and data transcoding, where data is represented from one compressed format to another compressed format or at different compression bitrates. 1. Brief Overview A neural network-based image and video compression method is proposed, in which wavelet transform and non-linear transform are combined to improve the coding and decoding efficiency. The present disclosure first addresses the problem of processing sub-bands of wavelet transform and then further improves the performance through non-linear transform. 2. Introduction In the past decade, deep learning has developed rapidly in various fields, especially in the fields of computer vision and image processing. Inspired by the great success of deep learning techniques in the field of computer vision, many researchers have shifted their attention from traditional image / video compression techniques to neural image / video compression techniques. Neural networks were initially proposed in the interdisciplinary research of neuroscience and mathematics. It has shown strong capabilities in the context of non-linear transformation and classification. In the past five years, neural network-based image / video compression techniques have made significant progress. It is reported that the latest neural network-based image compression algorithm achieves R-D performance comparable to that of the multi-functional video coding (VVC) (the latest video coding standard developed by the Joint Video Exploration Team (JVET) composed of experts from MPEG and VCEG). With the continuous improvement of the performance of neural image compression, neural network-based video compression has become an actively developing research field. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its infancy. 2.1. Image / Video Compression Image / video compression generally refers to the computational technique of compressing an image / video into a binary code for convenient storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video, which is called lossless compression and lossy compression. Most work is dedicated to lossy compression because lossless reconstruction is not required in most scenarios. The performance of an image / video compression algorithm is usually evaluated from two aspects, namely, the compression ratio and the reconstruction quality. The compression ratio is directly related to the number of binary codes, the fewer the better; the reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, the higher the better. Image / video compression techniques can be divided into two branches, classical video coding and decoding methods, and neural network-based video compression methods. Classical video coding and decoding schemes adopt transform-based solutions, where researchers utilize the statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) through carefully hand-engineered entropy coding and decoding, which models the dependencies in the quantization mechanism. Neural network-based video compression has two forms, neural network-based coding and decoding tools, and end-to-end neural network-based video compression. The former is embedded as a coding and decoding tool into existing classical video codecs and is only part of the framework, while the latter is a separate framework developed based on neural networks without relying on classical video codecs. Over the past three decades, a series of classical video coding and decoding standards have been developed to accommodate the growing visual content. The International Organization for Standardization ISO / IEC has two expert groups, namely the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), and ITU-T also has its own Video Coding Experts Group (VCEG), which is used for the standardization of image / video coding and decoding technologies. Influential video coding and decoding standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. After H.265 / HEVC, the Joint Video Exploration Team (JVET) composed of MPEG and VCEG has been working on a new video coding and decoding standard, Versatile Video Coding (VVC). The first version of VVC was released in July 2020. Compared with HEVC, VVC reduces the bit rate by an average of 50% at the same visual quality. Neural network-based image / video compression is not new, as many researchers have been working on neural network-based image coding and decoding. However, the network architectures were relatively shallow and the performance was not satisfactory. Thanks to the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, this technology is still far from mature and many challenges need to be addressed. 2.2. Neural Networks A neural network, also known as an artificial neural network (ANN), is a computational model used in machine learning techniques. It typically consists of multiple processing layers, and each layer is composed of multiple simple but non-linear basic computational units. One advantage of such a deep network is considered to be the ability to process data with multiple levels of abstraction and transform the data into different types of representations. Note that these representations are not manually designed; instead, using a general machine learning process, the deep network with processing layers learns from a large amount of data. Deep learning eliminates the need for manually designed representations and is thus considered particularly useful for processing native unstructured data such as acoustic and visual signals, and processing such data has been a long-standing problem in the field of artificial intelligence. 2.3. Neural Networks for Image Compression Existing neural networks for image compression methods can be divided into two categories, namely pixel probability modeling and autoencoders. The former belongs to the prediction coding and decoding strategy, while the latter is a transform-based solution. Sometimes, these two methods are combined together in the literature. 2.3.1. Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless coding and decoding can achieve the minimum coding and decoding rate -log2p(x), where p(x) is the probability of symbol x. Many lossless coding and decoding methods have been proposed in the literature, and arithmetic coding is considered to be one of the optimal methods. Given the probability distribution p(x), arithmetic coding ensures that the coding and decoding rate is as close as possible to its theoretical limit -log2p(x) without considering rounding errors. Therefore, the remaining problem is how to determine the probability. However, due to the curse of dimensionality, this is very challenging for natural images / videos. Following the prediction coding and decoding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is the image. p(x) = p(x1)p(x2|x1)…p(x i |x1,…,x i-1 )…p(x m×n |x1,…,x m×n-1 ) (1) where m and n are the height and width of the image respectively. Previous observations are also called the context of the current pixel. When the image is large, it may be difficult to estimate the conditional probability, so a simplified method is to limit the scope of its context. p(x) = p(x1)p(x2|x1)…p(x i |x i-k ,…,x i-1 )…p(x m×n |x m×n-k ,…,x m×n-1 ) (2) where k is a predefined constant that controls the context scope. It should be noted that the condition can also consider the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on the previously encoded and decoded pixels (including R / G / B samples), the current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample, and for encoding and decoding the current B sample, the previously encoded and decoded pixels and the current R and G samples can also be considered. Neural networks were initially introduced for computer vision tasks and have been proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to estimate the probability p(x i-1 ) given its context x1, x2, …, x i . Pixel probabilities were proposed for binary images, i.e., x i ∈ {-1, +1}. The Neural Autoregressive Distribution Estimator (NADE) was designed for pixel probability modeling and is a feedforward network with a single hidden layer. Similar work has been proposed where the feedforward network also has connections that skip the hidden layer and the parameters are also shared. NADE was extended to the real-valued model RNADE, where the probability p(x i | x1, …, x i-1 ) is derived using Gaussian mixtures. Its feedforward network also has a hidden layer, but the hidden layer is rescaled to avoid saturation and uses the Rectified Linear Unit (ReLU) instead of the Sigmoid. NADE and RNADE have been improved by reordering the pixels and using deeper neural networks. Designing advanced neural networks plays an important role in improving pixel probability modeling. Multidimensional long short-term memory networks (LSTMs) are proposed, which are used for probability modeling together with a mixture of conditional Gaussian scales. LSTM is a special type of recurrent neural network (RNN) that has been shown to be good at modeling sequential data. The spatial domain variant of LSTM is used for images. Several different neural networks are studied, including RNNs and CNNs, namely PixelRNN and PixelCNN. In PixelRNN, two variants of LSTM are proposed, called row LSTM and diagonal BiLSTM, where the latter is specifically designed for images. PixelRNN incorporates residual connections to help train deep neural networks with up to 12 layers. In PixelCNN, masked convolutions are used to fit the shape of the context. Compared with previous work, PixelRNN and PixelCNN are more focused on natural images: they treat pixels as discrete values (e.g., 0, 1, …, 255) and predict a multinomial distribution over the discrete values; they process color images in the RGB color space; they perform well on the large-scale image dataset ImageNet. The gated PixelCNN is proposed to improve PixelCNN and achieves performance comparable to PixelRNN but with much lower complexity. PixelCNN++ is proposed, which makes the following improvements to PixelCNN: using a discretized logistic mixture likelihood instead of a 256-way multinomial distribution; using downsampling to capture structures at multiple resolutions; introducing additional shortcut connections to speed up training; applying dropout for regularization; combining RGB into one pixel. PixelSNAIL is proposed, in which accidental convolutions are combined with self-attention. Most of the above methods directly model the probability distribution in the pixel domain. Some researchers have also tried to model the probability distribution as a conditional distribution based on explicit or latent representations. That is, it can be estimated as where h is the additional condition, and p(x) = p(h)p(x|h), which means that the modeling is divided into unconditional modeling and conditional modeling. The additional condition can be image label information or high-level representations. 2.3.2. Autoencoders Autoencoders are trained for dimensionality reduction and consist of two parts: encoding and decoding. The encoding part converts the high-dimensional input signal into a low-dimensional representation, usually with a reduced spatial size but more channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders can automatically learn representations and eliminate the need for manually designed features, which is also considered one of the most important advantages of neural networks. Figure 2illustrates a typical transform coding scheme 200. The original image x is transformed by the analysis network g a to achieve a latent representation y. The latent representation y is quantized and compressed into bits. The number of bits R is used to measure the coding rate. The quantized latent representation is then inverse-transformed by the synthesis network g s to obtain the reconstructed image The distortion is calculated in the perceptual space by applying the function g p to transform x and for calculation Applying an autoencoder network to lossy image compression is intuitive. It only needs to encode the latent representation learned from a well-trained neural network. However, adapting an autoencoder to image compression is not easy because the original autoencoder is not optimized for compression, resulting in inefficiency by directly using a trained autoencoder. In addition, there are other major challenges: First, the low-dimensional representation should be quantized before being encoded, but quantization is not differentiable, which is required in backpropagation when training a neural network. Second, the objectives in the compression scenario are different because both distortion and rate need to be considered. Estimating the rate is challenging. Third, practical image coding schemes need to support variable rates, scalability, encoding / decoding speed, and interoperability. In response to these challenges, many researchers have been actively contributing to this field. A prototype autoencoder for image compression is shown as Figure 2 follows. The original image x is transformed by the analysis network y = g a (x), where y is the latent representation, which will be quantized and coded. The synthesis network inverse-transforms the quantized latent representation to obtain the reconstructed image The framework is trained using a rate-distortion loss function, i.e., where D is the degree of distortion between x and , R is the rate calculated or estimated from the quantized representation , and λ is the Lagrange multiplier. It should be noted that D can be calculated in the pixel domain or the perceptual domain. All existing research works follow this prototype, and the differences only lie in the network structure or the loss function. In terms of network structure, RNN and CNN are the most widely used architectures. In the RNN-related category, a general framework for variable-ratio image compression using RNN is proposed. It uses binary quantization to generate codes and does not consider the ratio during training. The framework does provide scalable encoding and decoding capabilities, where RNNs with convolutional and deconvolutional layers are reported to perform well. An improved version is proposed by upgrading the encoder with a neural network similar to PixelRNN to compress binary codes. It is reported that the performance on the Kodak image dataset is better than JPEG using the MS-SSIM evaluation metric. The RNN-based solution is further improved by introducing hidden state activation. In addition, an SSIM weighted loss function is designed and a spatial domain adaptive bitrate mechanism is enabled. Using MS-SSIM as the evaluation metric, it achieves better results than BPG on the Kodak image dataset. A general framework for rate-distortion optimized image compression is proposed. The framework uses multi-ary quantization to generate integer codes and considers the ratio during training, i.e., the loss is a joint rate-distortion cost, which can be MSE or other. The framework adds random uniform noise to stimulate quantization during training and uses the differential entropy of the noisy code as a proxy for the ratio. The framework uses generalized divisive normalization (GDN) as the network structure, which consists of a linear mapping followed by a nonlinear parameter normalization. An improved version of GDN is proposed, in which the version uses 3 convolutional layers, each followed by a downsampling layer and a GDN layer as the forward transform. Therefore, the version uses 3 layers of inverse GDN, each followed by an upsampling layer and a convolutional layer to stimulate the inverse transform. In addition, an arithmetic coding method is designed to compress the integer codes. It is reported that the performance on the Kodak dataset is better than JPEG and JPEG 2000 in terms of MSE. In addition, the method is further improved by designing a scale hyper-prior into the autoencoder. The method utilizes the subnetwork h a Transform the latent representation y into z = h a (y), and z will be quantized and transmitted as side information. Therefore, the inverse transformation is achieved using the subnetwork h s The sub-network h s Trying to quantify the side information Decode to quantized The standard deviation of It is further used during the arithmetic coding and decoding of . On the Kodak image set, this method is slightly inferior to BPG in terms of PSNR. The structure in the residual space is further exploited by introducing an autoregressive model to estimate the standard deviation and mean. A Gaussian mixture model is then proposed to further remove the redundancy in the residual. Using PSNR as the evaluation metric, the reported performance is comparable to VVC on the Kodak image set. 2.3.3. Hyper-prior model In the transform coding method for image compression, the encoder sub-network (Section 2.3.2) uses a parametric analysis transform to transform the image vector x into a latent representation y, which is then quantized to form Since are discrete values, entropy coding techniques such as arithmetic coding can be used to losslessly compress them and transmit them as a bit sequence. Figure 3 Shows an example latent representation of an image, including Image 300 from the Kodak dataset, a visualization of the latent representation y of Image 300, the standard deviation σ320 of the latent 310, and the latent y330 after introducing the hyperprior network. The hyperprior network includes an encoder and a decoder that utilize hyperprior information. In the transform coding method for image compression, as Figure 2 shown, the encoder sub-network uses a parametric analysis transform to transform the image vector x into a latent representation y, which is then quantized to form Since are discrete values, they can be losslessly compressed using entropy coding techniques such as arithmetic coding and transmitted as a bit sequence. From Figure 3 the latent 310 and the standard deviation σ320, it can be clearly seen that there is significant spatial dependence among the elements of . An additional set of random variables Figure 4 can be introduced to capture the spatial dependence and further reduce redundancy. In this case, the image compression network is as Figure 4 is a schematic diagram 400 showing an example network architecture of an autoencoder implementing the hyperprior model. The upper side shows the image autoencoder network, and the lower side corresponds to the hyperprior sub-network. The analysis transform and the synthesis transform are denoted as g a and g a . Q represents quantization, and AE and AD represent the arithmetic encoder and the arithmetic decoder respectively. The hyperprior model includes two sub-networks, an encoder (denoted as h a ) that utilizes hyperprior information and a decoder (denoted as h s ) that utilizes hyperprior information. The hyperprior model generates a quantized hyperprior information latent value which includes information related to the probability distribution of the samples of the quantized latent value . is included in the bitstream and transmitted to the receiver (decoder) together with . In the schematic diagram 400, the upper side of the model is the encoder g as discussed above. a and decoder g s The lower side is used to obtain The additional encoder h using hyper-prior information a and a decoder using hyper-prior information h s In this architecture, the encoder subjects the input image x to g a , producing a response y with a spatially varying standard deviation. The response y is fed to h a The distribution of standard deviations in z is summarized in . z is then quantified compressed and transmitted as side information. The encoder then uses the quantized vector To estimate the spatial distribution of the standard deviation σ, and use σ to compress and transmit the quantized image representation The decoder first recovers the compressed signal The decoder then uses h s To obtain σ, this provides the decoder with the correct probability estimate to successfully recover The decoder will then Feed to g s to obtain the reconstructed image. When an encoder utilizing super-prior information and a decoder utilizing super-prior information are added to an image compression network, the quantized latent value The airspace redundancy is reduced. Figure 3 The potential value y 330 in σ corresponds to the quantized potential value when using an encoder / decoder that utilizes super-prior information. Compared with the standard deviation σ 320, the spatial redundancy is significantly reduced due to the lower correlation of the samples of the quantized potential value. 2.3.4. Context Model Although the hyper-prior model improves the quantitative potential value , but further improvement can be obtained by utilizing an autoregressive model that predicts the quantized latent value from its causal context (context model). The term autoregressive means that the output of a process is then used as its input. For example, the context model subnetwork generates a sample of potential values, which is then used as input to get the next sample. Figure 5 5 is a schematic diagram showing an example combined model configured to jointly optimize a context model with a hyper-prior and an autoencoder. The combined model jointly optimizes an autoregressive component that estimates the probability distribution of a latent value from its causal context (context model) with a hyper-prior and an underlying autoencoder. The real-valued latent representation is quantized (Q) to create a quantized latent value and quantify the potential value of hyper-prior information They are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The dashed area corresponds to the components performed by a receiver (e.g., a decoder) to recover the image from the compressed bitstream. A joint architecture is used, where both a hyperprior model subnetwork (an encoder using hyperprior information and a decoder using hyperprior information) and a context model subnetwork are utilized. The hyperprior and context models are combined to learn a probability model on the quantization latent values and then it is used for entropy encoding and decoding. As Figure 5 shown, the outputs of the context subnetwork and the decoder subnetwork using hyperprior information are combined by a subnetwork called the entropy parameter, which generates mean μ and scale (or variance) σ parameters for a Gaussian probability model. Then, the Gaussian probability model is used to encode the samples of the quantization latent values into a bitstream with the help of an arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantization latent values from the bitstream through an arithmetic decoder (AD) module Generally, the latent samples are modeled as a Gaussian distribution or a Gaussian mixture model (not limited to). In subsequent work, the context model and the hyperprior are jointly used to estimate the probability distribution of the latent samples. Since a Gaussian distribution can be defined by the mean and variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ). 2.3.5. Encoding Process Using a Joint Autoregressive Hyperprior Model Figure 5 The design in corresponds to the prior art of the compression method. In this section and the next section, the encoding and decoding processes will be described respectively. Figure 6 An example encoding process 600 is shown. The input image is first processed by an encoder subnetwork. The encoder transforms the input image into a transformed representation called the latent value, denoted as y. y is then input into a quantizer block, denoted as Q, to obtain the quantization latent values and then it is converted into a bitstream (bitstream 1, bits1) using an arithmetic coding module (denoted as AE). The arithmetic coding block sequentially converts each sample of into a bitstream (bits1). The module uses an encoder using hyperprior information, context, a decoder using hyperprior information, and an entropy parameter subnetwork to estimate the probability distribution of the samples of the quantization latent values The latent value y is input into the encoder using hyperprior information, and the encoder using hyperprior information outputs a hyperprior information latent value (denoted as z). The hyperprior information latent value is then quantized And use an arithmetic coding (AE) module to generate a second bitstream (bitstream 1, bits2). The factored entropy module generates a probability distribution that is used to encode the quantized hyperprior information latent values into a bitstream. The quantized hyperprior information latent values include information about the probability distribution of the quantized latent values of the quantized latent values. The entropy parameter subnetwork generates an estimated probability distribution that is used to encode the quantized latent values The information generated by the entropy parameters typically includes the mean μ and scale (or variance) σ parameters, which are together used to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as where the parameter μ is the mean or expectation of the distribution (and also its median and mode), and the parameter σ is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values. The subnetwork uses a decoder of the hyperprior information to generate part of the information used by the entropy parameter subnetwork, and another part of the information is generated by an autoregressive module called the context module. The context module uses the samples that have already been encoded by the arithmetic coding (AE) module to generate information about the probability distribution of the samples of the quantized latent values. The quantized latent values are typically a matrix consisting of many samples. The samples can be indicated using indices, for example or depending specifically on the dimensions of the matrix The samples are encoded one by one by the AE, typically using a raster scan order. In the raster scan order, the rows of the matrix are processed from top to bottom, and the samples in the row are processed from left to right. In such a scenario (where the AE encodes the samples into a bitstream using a raster scan order), the context module uses the samples encoded previously in the raster scan order to generate information related to the sample The information generated by the context module and the decoder using the hyperprior information is combined by the entropy parameter module to generate a probability distribution for encoding the quantized latent values into a bitstream (bits1). Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder. It should be noted that the above modules can also use other names. In the above description, Figure 6 all the elements in are collectively referred to as the encoder. The analysis transformation that converts the input image into a latent representation is also referred to as the encoder (or autoencoder). Figure 7 An example decoding process 700 is shown separately. During the decoding process, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. bits2 is first decoded by an arithmetic decoding (AD) module by utilizing the probability distribution generated by the factored entropy sub-network. The factored entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values in the case of a Gaussian distribution. The output of the arithmetic decoding process of bits2 is which is the quantized hyperprior information latent value. The AD process reverts to the AE process applied in the encoder. The processes of AE and AD are lossless, which means that the quantized hyperprior information latent value generated by the encoder can be reconstructed at the decoder without any change. After obtaining it, it is processed by a decoder that utilizes hyperprior information. The output of the decoder that utilizes hyperprior information is fed into the entropy parameter module. The three sub-networks employed in the decoder, the context, the decoder that utilizes hyperprior information, and the entropy parameter, are the same as those in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder (as in the encoder), which is crucial for losslessly reconstructing the quantized latent value is crucial. As a result, the same version of the quantized latent value obtained in the encoder can be obtained in the decoder. After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter sub-network, the arithmetic decoding module decodes the samples of the quantized latent value one by one from the bitstream bits1. From a practical perspective, the autoregressive model (context model) is inherently serial, so techniques such as parallelization cannot be used to accelerate it. Finally, the fully reconstructed quantized latent value is input into the synthesis transform (represented as the decoder in Figure 7 ) module to obtain the reconstructed image. In the above description, Figure 7 all the elements in are collectively referred to as the decoder. The synthesis transform that converts the quantized latent value into the reconstructed image is also called the decoder (or auto-decoder). Figure 6 The analysis transform (represented as the encoder) in Figure 7 and the synthesis transform (represented as the decoder) in Figure 8 can be replaced by wavelet-based transforms. Figure 8In it, first, the input image is converted from the RGB color format to the YUV color format. This conversion process is optional and may be absent in other implementations. However, if such a conversion is applied to the input image, an inverse conversion (from YUV to RGB) will also be applied before generating the output image. Additionally, in Figure 8 two additional post-processing modules (Post-Processing 1 and 2) are shown. These modules are also optional and thus may be absent in other implementations. The core of an encoder with a wavelet-based transform consists of a wavelet-based forward transform, a quantization module, and an entropy encoding / decoding module. After applying these three modules to the input image, a bitstream is generated. The core of the decoding process consists of entropy decoding, an inverse quantization process, and a wavelet-based inverse transform operation. The decoding process converts the bitstream into an output image. The encoding and decoding processes are as Figure 8 shown. Figure 9 A schematic diagram 900 of the output of the forward wavelet-based transform is shown. After applying the wavelet-based forward transform to the input image, in the output of the wavelet-based forward transform, the image is divided into its frequency components. The output of the two-dimensional forward wavelet transform (depicted as the iWave forward module in Figure 8 ) can take the form depicted in Figure 9 . The input to the transform is an image of a castle. In the example, after the transform, an output with 7 different regions is obtained. The number of different regions depends on the specific implementation of the transform and may be different from 7. The possible number of regions is 4, 7, 10, 13, … In Figure 9 , it can be seen that the input image is transformed into 7 regions, with 3 small images and 4 even smaller images. The transform is based on frequency components. The small image in the lower right quarter includes high-frequency components in both the horizontal and vertical directions. On the other hand, the smallest image in the upper left corner includes the lowest frequency components in both the vertical and horizontal directions. The small image in the upper right quarter includes high-frequency components in the horizontal direction and low-frequency components in the vertical direction. Figure 10 A segmentation 1000 of the output of the forward wavelet-based transform is shown. Figure 10 Depicts the possible partitioning of the latent representation after the 2D forward transform. The latent representation is the samples (latent samples or quantized latent samples) obtained after the 2D forward transform. The latent samples are divided into 7 parts above, denoted as HH1, LH1, HL1, LL2, HL2, LH2, and HH2. HH1 describes that this part includes high-frequency components in the vertical direction, high-frequency components in the horizontal direction, and the partitioning depth is 1. HL2 describes that this part includes low-frequency components in the vertical direction, high-frequency components in the horizontal direction, and the partitioning depth is 2. After obtaining the latent samples by positive wavelet transform at the encoder, the latent samples are transmitted to the decoder using entropy coding. At the decoder, entropy decoding is applied to obtain the latent samples, which are then inverse-transformed (by using the iWave inverse module in Figure 8 ) to obtain the reconstructed image. 2.4 Neural Networks for Video Compression Similar to traditional video coding techniques, neural image compression serves as the basis for intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression techniques lags behind that of neural network-based image compression. However, due to its complexity, more efforts are needed to address the challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared with image compression, video compression requires effective methods to eliminate inter-frame picture redundancy. Inter-frame picture prediction is a key step in these works. Motion estimation and compensation have been widely adopted but have only recently been implemented through trained neural networks. The research on neural network-based video compression can be classified into two categories according to the target scenarios: random access and low latency. In the case of random access, which requires decoding to start from any point in the sequence, the entire sequence is usually divided into multiple individual segments, and each segment can be decoded independently. In the case of low latency, it aims to reduce the decoding time, so usually only the previous frames in the time domain can be used as reference frames to decode subsequent frames. 2.4.1. Low Latency The first method first divides the video sequence frames into blocks, and each block will choose one of two available modes: intra-frame coding or inter-frame coding. If intra-frame coding is selected, there is an associated autoencoder to compress the block. If inter-frame coding is selected, traditional methods are used to perform motion estimation and compensation, and a trained neural network will be used for residual compression. The output of the autoencoder is directly quantized and coded through the Huffman method. Another neural network-based video coding scheme is proposed using PixelMotionCNN. The frames are compressed in temporal order, and each frame is divided into blocks, which are compressed in raster scan order. Each frame will first be inferred using the previous two reconstructed frames. When a block is to be compressed, the inferred frame together with the context of the current block is fed into PixelMotionCNN to derive the latent representation. Then, the residual is compressed through a variable-rate image scheme. The performance of this scheme is comparable to that of H.264. An end-to-end neural-network-based video compression framework with a sense of reality is proposed, where all modules are implemented by neural networks. This solution accepts the current frame and the previously reconstructed frame as inputs, and the optical flow will be derived using a pre-trained neural network as motion information. The motion information will be warped using the reference frame, and then the motion-compensated frame will be generated by the neural network. The residual and the motion information are compressed using two separate neural autoencoders. The entire framework is trained using a single rate-distortion loss function. It achieves better performance than H.264. Then, an advanced neural-network-based video compression solution is proposed. It inherits and extends the traditional video coding and decoding solution using a neural network with the following main features: 1) Only one autoencoder is used to compress the motion information and the residual; 2) Motion compensation using multiple frames and multiple optical flows; 3) Learning the online state and propagating this state over time through subsequent frames. This solution achieves better performance than the HEVC reference software in terms of MS-SSIM. An extended end-to-end neural-network-based video compression framework is proposed. In this solution, multiple frames are used as references. Therefore, by using multiple reference frames and the associated motion information, a more accurate prediction of the current frame can be provided. In addition, motion field prediction is deployed to eliminate the motion redundancy along the temporal channel. A post-processing network is also introduced into this work to eliminate the reconstruction artifacts from the previous process. In terms of PSNR and MS-SSIM, the performance is significantly better than H.265. Then the scaled spatial flow is proposed, which replaces the common optical flow by adding a scaling parameter. It is reported that it achieves better performance than H.264. A multi-resolution representation is proposed for the optical flow. Specifically, the motion estimation network generates multiple optical flows with different resolutions and lets the network learn to select which one under the loss function. The performance is slightly improved and better than H.265. 2.4.2. Random Access The initial method involves a neural-network-based video compression solution with frame interpolation. The key frames are first compressed using a neural image compressor, and the remaining frames are compressed in a hierarchical order. It performs motion compensation in the perceptual domain, that is, the feature maps at multiple spatial scalings of the original frame are derived, and the motion is used to warp the feature maps, which will be used by the image compressor. It is reported that this method is comparable to H.264. Subsequently, interpolation-based video compression is proposed, where the interpolation model combines motion information compression and image synthesis, and the same autoencoder is used for both the image and the residual. Subsequently, a neural network-based video compression method based on a variational autoencoder with a deterministic encoder was proposed. Specifically, the model consists of an autoencoder and an autoregressive prior. Different from previous methods, this method accepts a group of pictures (GOP) as input and incorporates a 3D autoregressive prior by considering temporal correlations when encoding and decoding the latent representation. It provides performance comparable to H.265. 2.5. Preliminary Knowledge Almost all natural images / videos are in digital format. A grayscale digital image can be represented as where is a set of pixel values, m is the image height, and n is the image width. For example, is a common setting, in which case Therefore, pixels can be represented by 8-bit integers. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while the compressed bits are surely less. Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by where the three separate channels store red, green, and blue information. Similar to 8-bit grayscale images, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Most neural network-based video compression schemes are developed in the RGB color space, while traditional codecs usually use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels, namely Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. The benefit comes from the fact that Cb and Cr are usually downsampled for pre-compression because the human visual system is less sensitive to chrominance components. A color video sequence consists of multiple color images (called frames) to record scenes at different timestamps. For example, in the RGB color space, a color video can be represented as X = {x0, x1, …, x t , …, x T-1}, where T is the number of frames in the video sequence, If m = 1080, n = 1920, and the video has 50 frames per second (fps), then the data rate of this uncompressed video is 1920×1080×8×3×50 = 2,488,320,000 bits per second (bps), approximately 2.32 Gbps, which requires a large amount of storage and thus surely needs to be compressed before transmission over the Internet. Typically, lossless methods can achieve a compression ratio of about 1.5 to 3 for natural images, which is significantly lower than the requirements. Therefore, lossy compression has been developed to achieve further compression ratios, but at the cost of introducing distortion. The distortion can be measured by calculating the mean squared difference between the original image and the reconstructed image, i.e., the mean squared error (MSE). For grayscale images, the MSE can be calculated by the following equation. Therefore, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR): where is the maximum value in, e.g., 255 for an 8-bit grayscale image. There are also other quality assessment metrics, such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM). To compare different lossless compression schemes, it is sufficient to compare the compression ratios given the result ratios. Or vice versa, given the compression ratios to compare the ratios. However, to compare different lossy compression methods, both the bit rate and the reconstruction quality must be considered. For example, calculating the relative rates at several different quality levels and then averaging the rates is a commonly used method; the average relative rate is called the Bjontegaard Delta Rate (BD Rate). There are also other important aspects in evaluating image / video codec schemes, including encoding / decoding complexity, scalability, robustness, etc. 3. Problems Most learning-based image compression methods usually utilize non-linear transforms to achieve a compact representation, which has been proven efficient from the perspectives of human vision and objective quality. In addition to non-linear transforms, wavelet transform is also a powerful tool for multi-resolution time-frequency analysis, which has been widely implemented in traditional codecs such as JPEG-2000. Its effectiveness has also been verified in other image processing tasks such as denoising, image enhancement, and fusion. Moreover, in recent work, some wavelet-based learning image frameworks have shown great potential in lossy compression and also support lossless encoding / decoding. To further improve the codec performance and provide more functions for learning image compression, the combination of wavelet transform and non-linear transform has great potential. However, different from directly feeding the image into the transform, the outputs of wavelet transform and non-linear transform may have quite different statistical characteristics, and how to exploit their advantages and combine them remains a key issue. 4. Detailed Solutions The following detailed embodiments should be considered as examples to explain the general concepts. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. 4.1. Objectives of the Embodiment The objective of the embodiment is to discover the advantages of wavelet transform and non - linear transform and combine them to improve the performance of the transform module of the learning image codec, thereby improving the compression efficiency. 4.2 Core of the Embodiment The core of the embodiment lies in controlling the combination of wavelet transform and non - linear transform, as well as the corresponding processing of potential samples. 4.3 Details 4.3.1. Decoding Process According to the embodiment of the present disclosure, decoding the bitstream to obtain the reconstructed picture is performed as follows. An image decoding method includes the following steps: - Using the potential sample reconstruction module, obtaining the quantized potential samples according to the bitstream - Adjusting the information included in the channels of the quantized potential samples through the inverse channel network The channels of which include information. - Based on the output of the inverse channel network, performing an inverse non - linear transform network to obtain the reconstructed sub - bands of the wavelet. - Using the reconstructed sub - bands and the wavelet transform network to obtain the reconstructed image. Figure 11 An example diagram 1100 of the decoding process is shown. During the decoding process, first, based on the bitstream and by using the potential sample reconstruction module, the quantized potential samples are obtained The potential sample reconstruction module may include multiple neural - network - based sub - networks. An example potential sample reconstruction module is as Figure 11 shown, where the potential sample reconstruction module includes a prediction fusion model, a context model, a decoder using hyper - prior information, and a variance decoder using hyper - prior information, all of which are neural - network - based processing units. In Figure 11 the example shown, the potential sample reconstruction module further includes two entropy decoder units (entropy decoder 1 and 2), which are responsible for converting the bitstream (a series of 1s and 0s) into symbols such as integers or floating - point numbers. In Figure 11 the example of, the output of the prediction fusion model is the predicted sample (i.e., the sample expected to be as close as possible to the quantized potential sample), and the output of the entropy decoder 2 is the residual sample (i.e., the sample representing the difference between the predicted sample and the quantized potential sample). Therefore, the unit prediction fusion model, the context model, and the decoder using hyper - prior information are responsible for obtaining the predicted sample, while the variance decoder unit using hyper - prior information is responsible for obtaining the probability parameters used in the decoding of the bitstream 2. It should be noted that Figure 11 the potential sample reconstruction module depicted in is an example, and the present disclosure is applicable to any type of potential sample reconstruction module, unit for obtaining quantized potential samples. The quantized latent representation can be a tensor or matrix including quantized latent samples. After obtaining the quantized latent samples, an inverse channel network is applied to adjust the number of channels of the quantized latent representation. Then, the output of the inverse channel network is fed into an inverse non-linear transformation network to obtain the reconstructed wavelet subbands. Finally, the reconstructed wavelet subbands are used in an inverse wavelet transform network to obtain the reconstructed image. The reconstructed image can be displayed on a display device with or without further processing. The inverse channel network can be used to adjust the number of channels of the input. An example inverse channel network 1200 is as Figure 12 shown. According to an example, the inverse channel network can include operations of BatchNorm, ReLU, and convolutional layers. The input of the inverse channel network is the quantized latent samples which can be a 3D or 4D tensor, where two dimensions can be spatial dimensions such as width or height. The third dimension can be the number of channels or the number of feature maps. The output of the inverse channel network is which includes samples of equally sized reconstructed subbands. - In one example, the inverse channel network can include other activation operations. a) In one example, leaky ReLU can be used inside the network. b) In one example, Sigmoid can be used inside the network. c) In another example, there can be no activation layer. - In one example, the inverse channel network can include more layers. - In one example, BatchNorm can be removed or replaced by other normalization operations. - In one example, the order of operations can be different. - In one example, the stride of the convolution is strictly 1 to ensure that the output spatial size is the same as the input. - In one example, the structure of the network can be a bottleneck, in which case the stride of the convolution can be any number. a) In one example, the input features can be first downsampled and then upsampled to ensure that the output has the same spatial resolution. b) In one example, the input features can be first upsampled and then downsampled to ensure that the output has the same spatial resolution. - The inverse channel network can include fully connected neural network layers. - In one example, the inverse channel network can reduce the number of channels of the input, i.e., the number of channels of can be less than - In one example, the inverse channel network can increase the number of channels of the input, i.e., the number of channels of The output of the inverse channel network is a tensor which includes samples of equal-sized reconstructed subbands. These subbands are then divided into more than one tensor, with each tensor corresponding to a different subband. The division operation is along the channel dimension. For example, the first N channels of can correspond to the first subband, and the subsequent M channels can correspond to the second subband. The function of the inverse channel network is to achieve the distribution of the information included in to each subband. The information included in is usually mixed, and one of the channels (feature maps) of can include information from multiple subbands. Therefore, the function of the inverse channel network is to separate this information into the corresponding subbands. In one implementation of the present disclosure, the inverse channel network does not exist. In such an implementation, the information included in may already have been separated, i.e., the first N channels of Figure 13 can include only information related to the first subband, and the subsequent M channels can include only information related to the second subband, and so on. Figure 13 An example of the inverse nonlinear transformation to obtain the reconstructed wavelet subbands is shown in as shown. According to the example, the input features can be processed using 4 separate branches to obtain 4 sets of information (e.g., subband groups). Each group can include one or more subbands with the same spatial resolution, while the spatial resolutions (dimensions in the spatial domain) between different groups can be different. The input to the inverse nonlinear transformation is which includes subbands of equal size, and its output is which includes one or more groups of subbands. The function of the inverse nonlinear transformation is to adjust the spatial dimension. In one example, the number of subband groups can be greater than one. A part of is adjusted to match the size of the first group of subbands, and a second part of is adjusted to match the size of the second group of subbands. In one example, the inverse nonlinear transformation can include at least two branches, where at least one branch includes a neural network layer that performs size adjustment. The size adjustment layer increases or decreases the spatial size of the input. In one example, the inverse non-linear transformation may include a single branch that includes a neural network layer that performs resizing. The resizing layer may be an upsampling layer or a downsampling layer. The terms resizing, increasing the size, decreasing the size, magnifying, shrinking, upsampling, or downsampling may be used interchangeably. The upsampling layer may be implemented using a transposed convolution layer or a convolution layer. The downsampling layer may be implemented using a transposed convolution layer or a convolution layer. If the inverse non-linear transformation includes more than one branch, the output of each branch may be a set of sub-bands, each set of sub-bands having a different spatial size. If the inverse non-linear transformation includes more than one branch, for each branch, different resizing ratios may be used to obtain different sub-bands. - In one example, the input features may be divided into 4 parts along the channel dimension, and different parts may be used to reconstruct the sub-bands separately. a) In one example, the number of channels in all parts is the same. b) In one example, each part has a different number of channels, depending on the importance of the sub-band. - In one example, all of the above branches may use the same input features as input to obtain different sub-bands. - The resizing operation may be magnifying or upsampling. - The resizing operation may be shrinking or downsampling. - The resizing operation may be implemented using a neural network. Another possible solution for the inverse non-linear transformation is as Figure 14 shown in the schematic diagram 1400. According to the example, each branch may include a convolution layer, a transposed convolution layer, or an activation function. - In one example, the sub-band with the lowest resolution (smallest spatial size) may be obtained without applying resizing. - In one example, the inverse non-linear transformation may include an activation operation. a) In one example, leaky ReLU or ReLU (Rectified Linear Unit) may be used. b) In one example, Sigmoid operation or tanh (hyperbolic tangent function) may be used inside the network. c) Embodiments of the present disclosure are not limited to any particular activation function or the presence of an activation function. - In one example, the inverse non-linear transformation network may include a convolution of a transposed convolution layer. - In one example, the inverse non-linear transformation may include more than 1 layer. - In one example, the inverse non-linear transformation network may include a normalization operation. - In one example, the order of operations may be different. - The resizing operation may be implemented by a transposed convolution layer or a convolution layer with a stride equal to N. a) In one example, N is equal to 2. The wavelet transform network may be implemented as: - In one example, the parameters of the wavelet transform may be fixed parameters, which are the same as those of the traditional wavelet transform. - Alternatively, the parameters of the wavelet transform may be learnable (i.e., based on a neural network), such that the wavelet transform may be optimized jointly with the entire network. The wavelet transform (or inverse wavelet transform) takes the reconstructed subbands as input and applies the transform to convert to the reconstructed image. Another example implementation of the wavelet transform is described in Section 2.3.7. The names of the latent sample reconstruction module, the inverse channel network, and the inverse non-linear transformation network may be different. 4.3.2. Encoding process According to the present disclosure, the encoding process follows a process opposite to that of the decoding process for obtaining the quantized latent samples. The difference is that after obtaining the quantized latent samples, the samples are included in the bitstream using an entropy coding method. According to the present disclosure, encoding the input image to obtain a bitstream is performed as follows. An image encoding method includes the following steps: - Input the image and obtain wavelet subbands by using a wavelet transform. - Process the wavelet subbands by using a non-linear transformation network. - Obtain the latent samples y by using a channel network. - Quantize the latent samples to obtain the quantized latent samples The quantized latent samples may be obtained by a latent sample reconstruction module. Or it may be obtained through a quantization process. - Use the quantized latent values and an entropy coding module to obtain the bitstream. 4.4. Benefits The present disclosure provides a combined method of wavelet transform and non-linear transform to further improve the encoding and decoding efficiency of the transform operations in learning image compression. 5. Embodiments 1. An image or video decoding method, including the following steps: - Use a latent sample reconstruction module to obtain quantized latent samples from the bitstream - Adjusting the quantized latent samples through an inverse channel network for the information included in the channels. - Based on the output of the inverse channel network, performing an inverse non-linear transformation network to obtain the reconstructed wavelet sub-bands. - Using the reconstructed sub-bands and the wavelet transform network to obtain the reconstructed image. 2. An image or video coding method, comprising the following steps: - Using the input image and obtaining wavelet sub-bands by using wavelet transform. - Processing the wavelet sub-bands by using a non-linear transformation network. - Obtaining the latent samples y by using a channel network. - Quantizing the latent samples to obtain the quantized latent samples The quantized latent samples can be obtained by a latent sample reconstruction module. Or it can be obtained through a quantization process. - Using the quantized latent values and an entropy coding module to obtain a bitstream.
[0040] Figure 15 FIG. shows a flowchart of a method 1500 for visual data processing according to an embodiment of the present disclosure. The method 1500 is implemented for the conversion between the current visual unit of the visual data and the bitstream of the visual data.
[0041] At block 1510, at least one wavelet sub-band representation of the current visual unit is determined according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation. The wavelet sub-band representation is associated with the sub-bands of the wavelet of the current visual unit. The sub-bands of the wavelet represent the frequency sub-bands or frequency components of the wavelet.
[0042] At block 1520, the conversion is performed based on at least one wavelet sub-band representation.
[0043] The method 1500 can implement the combination of wavelet transform and non-linear transform, as well as the corresponding visual data processing. For example, the processing of latent samples can be combined. The combination of wavelet transform and non-linear transform can improve the performance of the transform module of the learned image codec, thereby improving the compression efficiency.
[0044] In some embodiments, the conversion includes decoding the current visual unit from the bitstream.
[0045] In some embodiments, determining at least one wavelet subband representation includes: determining, based on a bitstream and a neural network-based latent sample reconstruction module, quantized samples of a current visual unit, the quantized samples being associated with a plurality of channels; determining an intermediate representation of the quantized samples, the intermediate representation being associated with at least one wavelet subband, the wavelet subband being associated with at least one of the plurality of channels; and determining at least one wavelet subband representation based on the intermediate representation and a neural network-based inverse nonlinear transform. For example, the neural network-based inverse nonlinear transform may be performed by an inverse nonlinear transform network, which may be a neural network. As used herein, the term "quantized sample" may be a quantized latent sample, which may be referred to as The term "intermediate representation" may be denoted as and the wavelet subband representation may be denoted as
[0046] In some embodiments, performing the transformation includes: determining a reconstruction of a current visual unit based on at least one wavelet subband representation and an inverse wavelet transform. The inverse wavelet transform may be performed by an inverse wavelet module or an inverse wavelet network.
[0047] For example, as Figure 11 shown, during the decoding process, first, quantized latent samples are obtained based on the bitstream and by using a latent sample reconstruction module The latent sample reconstruction module may include a plurality of neural network-based sub-networks. An example latent sample reconstruction module is as Figure 11 shown, where the latent sample reconstruction module includes a prediction fusion model, a context model, a decoder using hyperprior information, and a variance decoder using hyperprior information, all of which are neural network-based processing units. The latent sample reconstruction module also includes two entropy decoder units (entropy decoder 1 and 2), which are responsible for converting the bitstream (a series of 1s and 0s) into symbols such as integers or floating-point numbers. The output of the prediction fusion model is a predicted sample (i.e., a sample that is expected to be as close as possible to the quantized latent sample), and the output of the entropy decoder 2 is a residual sample (i.e., a sample representing the difference between the predicted sample and the quantized latent sample). Thus, the unit prediction fusion model, context model, and decoder using hyperprior information are responsible for obtaining the predicted sample, while the variance decoder unit using hyperprior information is responsible for obtaining the probability parameters used in the decoding of bitstream 2. It should be noted that Figure 11 the latent sample reconstruction module depicted in
[0048] The quantized latent representation can be a tensor or matrix including quantized latent samples. After obtaining the quantized latent samples, an inverse channel network is applied to adjust the number of channels of the quantized latent representation. Then the output of the inverse channel network is fed into an inverse non-linear transformation network to obtain the reconstructed wavelet subbands. Finally, the reconstructed wavelet subbands are used in an inverse wavelet transformation network to obtain the reconstructed image. The reconstructed image can be displayed on a display device with or without further processing.
[0049] In some embodiments, an intermediate representation of the quantized samples is determined based on an inverse channel network for adjusting the number of channels of the quantized samples. An example of the inverse channel network is as Figure 11 shown.
[0050] In some embodiments, the inverse channel network includes at least one of the following: a batch normalization unit such as BatchNorm, a rectified linear unit (ReLU), or at least one convolutional layer.
[0051] In some embodiments, the quantized samples include at least one of the following: a first dimension of a first spatial dimension of a current visual unit, a second dimension of a second spatial dimension of the current visual unit, a third dimension of the number of multiple channels, or a fourth dimension of the number of feature maps associated with the current visual unit.
[0052] In some embodiments, the activation operation of the inverse channel network includes at least one of the following: a leaky rectified linear unit (ReLU) or a Sigmoid operation.
[0053] In some embodiments, there is no activation layer in the inverse channel network.
[0054] In some embodiments, the stride of the convolution of the inverse channel network is a first predefined value, such as 1.
[0055] In some embodiments, the inverse channel network is a bottleneck, and the stride of the convolution of the inverse channel network is a second value, such as any number.
[0056] In some embodiments, the bottleneck inverse channel network performs a downsampling operation and an upsampling operation.
[0057] In some embodiments, the inverse channel network includes a fully connected neural network layer.
[0058] In some embodiments, the number of channels of the intermediate representation is less than or greater than the number of channels of the quantized samples.
[0059] In some embodiments, the intermediate representation includes samples of the reconstructed wavelet subbands of equal size.
[0060] In some embodiments, a first number of channels of the intermediate representation are associated with a first wavelet subband, and a second number of channels of the intermediate representation are associated with a second wavelet subband.
[0061] In some embodiments, a first network for an inverse non-linear transform includes a plurality of branches associated with multiple sets of wavelet subbands, where a set of wavelet subbands is associated with the same spatial resolution. As used herein, the first network may also be referred to as an inverse non-linear transform network. An example of an inverse non-linear transform network is Figure 13 shown.
[0062] In some embodiments, at least one wavelet subband representation includes multiple sets of subband representations associated with multiple sets of wavelet subbands.
[0063] In some embodiments, the first network adjusts the spatial dimension of the intermediate representation, where a first portion of the intermediate representation is adjusted to match a first size of a first set of wavelet subbands among the multiple sets of wavelet subbands, and a second portion of the intermediate representation is adjusted to match a second size of a second set of wavelet subbands among the multiple sets of wavelet subbands.
[0064] In some embodiments, the plurality of branches includes a first branch that includes a first neural network layer for size adjustment.
[0065] In some embodiments, the first network for an inverse non-linear transform includes a single branch that includes a first neural network layer for size adjustment.
[0066] In some embodiments, the first neural network layer increases or decreases the spatial size of the intermediate representation.
[0067] In some embodiments, the first neural network layer includes an upsampling layer or a downsampling layer.
[0068] In some embodiments, the first neural network layer includes a transposed convolution layer or a convolution layer.
[0069] In some embodiments, a first spatial size of a first set of wavelet subbands among the multiple sets of wavelet subbands is different from a second spatial size of a second set of wavelet subbands among the multiple sets of wavelet subbands.
[0070] In some embodiments, a first size adjustment ratio is used for a first set of wavelet subbands among the multiple sets of wavelet subbands, and a second size adjustment ratio different from the first size adjustment ratio is used for a second set of wavelet subbands among the multiple sets of wavelet subbands.
[0071] In some embodiments, at least one wavelet subband representation includes a plurality of wavelet subband representations, each wavelet subband representation being associated with one of a plurality of sets of wavelet subbands, and wherein determining the plurality of wavelet subband representations includes: determining a plurality of segmented representations of an intermediate representation; and determining the plurality of wavelet subband representations based on the plurality of segmented representations.
[0072] In some embodiments, the plurality of segmented representations are associated with the same number of channels.
[0073] In some embodiments, a first segmented representation among the plurality of segmented representations is associated with a first number of channels, and a second segmented representation among the plurality of segmented representations is associated with a second number of channels different from the first number of channels.
[0074] In some embodiments, the first number of channels and the second number of channels are determined based on the importance of the wavelet subbands associated with the first segmented representation and the second segmented representation.
[0075] In some embodiments, at least one wavelet subband representation includes a plurality of wavelet subband representations, each wavelet subband representation being associated with one of a plurality of sets of wavelet subbands, and the plurality of wavelet subband representations are determined based on an intermediate representation.
[0076] Another example of an inverse non-linear transformation network is as Figure 14 shown. In some embodiments, one of the plurality of branches includes at least one of the following: a convolutional layer, a deconvolutional layer, an activation operation, a rectified linear unit (ReLU), a leaky ReLU, a Sigmoid operation, a hyperbolic tangent operation, or a normalization operation.
[0077] In some embodiments, the wavelet subband representation of the wavelet subband with the lowest resolution among the plurality of sets of wavelet subbands is determined without resizing.
[0078] In some embodiments, the stride of the deconvolutional layer or the stride of the convolutional layer in a branch is a third number, such as 2.
[0079] In some embodiments, the transformation includes encoding the current visual unit into a bitstream.
[0080] In some embodiments, determining at least one wavelet subband representation includes: determining wavelet subband information of the current visual unit based on a wavelet transform; and determining at least one wavelet subband representation based on the wavelet subband information and a non-linear transformation.
[0081] In some embodiments, performing the transformation includes: determining samples of the current visual unit based on at least one wavelet subband representation and a channel network; determining quantized samples of the current visual unit by quantizing the samples of the current visual unit; and determining a bitstream based at least on the quantized samples and an entropy encoding / decoding module.
[0082] In some embodiments, the wavelet transform includes at least one fixed parameter. For example, the wavelet transform may include at least one fixed parameter that is the same as a conventional wavelet transform.
[0083] In some embodiments, the wavelet transform includes at least one learnable parameter, and the at least one learnable parameter is updated together with another neural network for the transformation. For example, the parameters of the wavelet transform may be learnable (e.g., based on a neural network) such that the wavelet transform can be optimized jointly with the entire network.
[0084] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of visual data, and the bitstream of visual data is generated by a method executed by a device for visual data processing. In the method, at least one wavelet subband representation of a current visual unit of the visual data is determined according to a neural network-based non-linear transform or a neural network-based inverse non-linear transform inverse to the non-linear transform. The wavelet subband representation is associated with a subband of a wavelet of the current visual unit. The bitstream is generated based on the at least one wavelet subband representation.
[0085] According to still further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, at least one wavelet subband representation of a current visual unit of the visual data is determined according to a neural network-based non-linear transform or a neural network-based inverse non-linear transform inverse to the non-linear transform. The wavelet subband representation is associated with a subband of a wavelet of the current visual unit. The bitstream is generated based on the at least one wavelet subband representation. The bitstream is stored in a non-transitory computer-readable recording medium.
[0086] Implementations of the present disclosure can be described according to the following items, and the features can be combined in any reasonable manner.
[0087] Item 1. A method for visual data processing, comprising: for the conversion between a current visual unit of visual data and a bitstream of the visual data, determining at least one wavelet subband representation of the current visual unit according to a neural network-based non-linear transform or a neural network-based inverse non-linear transform opposite to the non-linear transform, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; and performing the conversion based on the at least one wavelet subband representation.
[0088] Item 2. The method according to Item 1, wherein the conversion includes decoding the current visual unit from the bitstream.
[0089] Item 3. The method according to Item 2, wherein determining the at least one wavelet subband representation includes: determining, based on the bitstream and a neural network-based latent sample reconstruction module, the quantized samples of the current visual unit, the quantized samples being associated with a plurality of channels; determining an intermediate representation of the quantized samples, the intermediate representation being associated with at least one wavelet subband, the wavelet subband being associated with at least one of the plurality of channels; and determining the at least one wavelet subband representation based on the intermediate representation and the neural network-based inverse nonlinear transform.
[0090] Item 4. The method according to Item 2 or Item 3, wherein performing the transformation includes: determining a reconstruction of the current visual unit based on the at least one wavelet subband representation and an inverse wavelet transform.
[0091] Item 5. The method according to Item 3, wherein the intermediate representation of the quantized samples is determined based on an inverse channel network for adjusting the number of channels of the quantized samples.
[0092] Item 6. The method according to Item 5, wherein the inverse channel network includes at least one of the following: a batch normalization unit, a rectified linear unit (ReLU), or at least one convolutional layer.
[0093] Item 7. The method according to Item 5 or Item 6, wherein the quantized samples include at least one of the following: a first dimension of a first spatial dimension of the current visual unit, a second dimension of a second spatial dimension of the current visual unit, a third dimension of the number of the plurality of channels, or a fourth dimension of the number of feature maps associated with the current visual unit.
[0094] Item 8. The method according to any one of Items 5 to 7, wherein the activation operation of the inverse channel network includes at least one of the following: a leaky rectified linear unit (ReLU) or a Sigmoid operation.
[0095] Item 9. The method according to any one of Items 5 to 8, wherein an activation layer is missing in the inverse channel network.
[0096] Item 10. The method according to any one of Items 5 to 9, wherein the stride of the convolution of the inverse channel network is a first predefined value.
[0097] Item 11. The method according to any one of Items 5 to 9, wherein the inverse channel network is a bottleneck, and the stride of the convolution of the inverse channel network is a second value.
[0098] Item 12. The method according to Item 11, wherein the bottleneck inverse channel network performs a downsampling operation and an upsampling operation.
[0099] Item 13. The method according to any one of Items 5 to 12, wherein the inverse channel network includes a fully connected neural network layer.
[0100] Item 14. The method according to any one of Items 5 to 13, wherein the number of channels of the intermediate representation is less than or greater than the number of channels of the quantization samples.
[0101] Item 15. The method according to any one of Items 3 to 7, wherein the intermediate representation includes samples of reconstructed wavelet subbands of equal size.
[0102] Item 16. The method according to any one of Items 3 to 15, wherein a first number of channels of the intermediate representation is associated with a first wavelet subband, and a second number of channels of the intermediate representation is associated with a second wavelet subband.
[0103] Item 17. The method according to any one of Items 3 to 16, wherein a first network for the inverse non-linear transform includes a plurality of branches associated with multiple sets of wavelet subbands, and one set of wavelet subbands is associated with the same spatial resolution.
[0104] Item 18. The method according to Item 17, wherein the at least one wavelet subband representation includes multiple sets of subband representations associated with the multiple sets of wavelet subbands.
[0105] Item 19. The method according to Item 17 or Item 18, wherein the first network adjusts the spatial dimension of the intermediate representation, a first part of the intermediate representation is adjusted to match a first size of a first set of wavelet subbands in the multiple sets of wavelet subbands, and a second part of the intermediate representation is adjusted to match a second size of a second set of wavelet subbands in the multiple sets of wavelet subbands.
[0106] Item 20. The method according to any one of Items 17 to 19, wherein the plurality of branches includes a first branch, and the first branch includes a first neural network layer for size adjustment.
[0107] Item 21. The method according to any one of Items 3 to 16, wherein a first network for the inverse non-linear transform includes a single branch, and the single branch includes a first neural network layer for size adjustment.
[0108] Item 22. The method according to Item 20 or Item 21, wherein the first neural network layer increases or decreases the spatial size of the intermediate representation.
[0109] Item 23. The method according to any one of Items 20 to 22, wherein the first neural network layer includes an upsampling layer or a downsampling layer.
[0110] Item 24. The method according to any one of Items 20 to 23, wherein the first neural network layer comprises a deconvolution layer or a convolution layer.
[0111] Item 25. The method according to any one of Items 17 to 19, wherein a first spatial domain size of a first group of wavelet subbands among the multiple groups of wavelet subbands is different from a second spatial domain size of a second group of wavelet subbands among the multiple groups of wavelet subbands.
[0112] Item 26. The method according to any one of Items 17 to 19, wherein a first size adjustment ratio is used for a first group of wavelet subbands among the multiple groups of wavelet subbands, and a second size adjustment ratio different from the first size adjustment ratio is used for a second group of wavelet subbands among the multiple groups of wavelet subbands.
[0113] Item 27. The method according to any one of Items 17 to 19, wherein the at least one wavelet subband representation comprises a plurality of wavelet subband representations, each wavelet subband representation being associated with a group of wavelet subbands among the multiple groups of wavelet subbands, and wherein determining the plurality of wavelet subband representations comprises: determining a plurality of segmented representations of the intermediate representation; and determining the plurality of wavelet subband representations based on the plurality of segmented representations.
[0114] Item 28. The method according to Item 27, wherein the plurality of segmented representations are associated with the same number of channels.
[0115] Item 29. The method according to Item 27, wherein a first segmented representation among the plurality of segmented representations is associated with a first number of channels, and a second segmented representation among the plurality of segmented representations is associated with a second number of channels different from the first number of channels.
[0116] Item 30. The method according to Item 29, wherein the first number of channels and the second number of channels are determined based on the importance of the wavelet subbands associated with the first segmented representation and the second segmented representation.
[0117] Item 31. The method according to any one of Items 17 to 19, wherein the at least one wavelet subband representation comprises a plurality of wavelet subband representations, each wavelet subband representation being associated with a group of wavelet subbands among the multiple groups of wavelet subbands, and the plurality of wavelet subband representations are determined based on the intermediate representation.
[0118] Item 32. The method according to Item 17, wherein one of the plurality of branches comprises at least one of the following: a convolution layer, a deconvolution layer, an activation operation, a rectified linear unit (ReLU), a leaky ReLU, a Sigmoid operation, a hyperbolic tangent operation, or a normalization operation.
[0119] Item 33. The method according to Item 32, wherein the wavelet subband representation of the wavelet subband with the lowest resolution among the multiple groups of wavelet subbands is determined without resizing.
[0120] Item 34. The method according to Item 32, wherein the stride of the transposed convolutional layer or the stride of the convolutional layer in the branch is a third number.
[0121] Item 35. The method according to Item 1, wherein the transformation includes encoding the current visual unit into the bitstream.
[0122] Item 36. The method according to Item 35, wherein determining the at least one wavelet subband representation includes: determining wavelet subband information of the current visual unit based on wavelet transform; and determining the at least one wavelet subband representation based on the wavelet subband information and the non - linear transform.
[0123] Item 37. The method according to Item 35 or Item 36, wherein performing the transformation includes: determining samples of the current visual unit based on the at least one wavelet subband representation and a channel network; determining quantized samples of the current visual unit by quantizing the samples of the current visual unit; and determining the bitstream based at least on the quantized samples and an entropy encoding and decoding module.
[0124] Item 38. The method according to Item 36, wherein the wavelet transform includes at least one fixed parameter.
[0125] Item 39. The method according to Item 36, wherein the wavelet transform includes at least one learnable parameter, and the at least one learnable parameter is updated together with another neural network for the transformation.
[0126] Item 40. An apparatus for visual data processing, comprising a processor and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 39.
[0127] Item 41. A non - transitory computer - readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 39.
[0128] Item 42. A non-transitory computer-readable recording medium stores a bitstream of visual data, the bitstream of visual data being generated by a method executed by a device for visual data processing, wherein the method includes: determining at least one wavelet subband representation of a current visual unit of the visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; and generating the bitstream based on the at least one wavelet subband representation.
[0129] Item 43. A method for storing a bitstream of visual data includes: determining at least one wavelet subband representation of a current visual unit of the visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; generating the bitstream based on the at least one wavelet subband representation; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0130] Figure 16 A block diagram of a computing device 1600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1600 may be implemented as the source device 110 (or the data encoder 114) or the destination device 120 (or the data decoder 124), or may be included in the source device 110 (or the data encoder 114) or the destination device 120 (or the data decoder 124).
[0131] It should be understood that Figure 16 the computing device 1600 shown is for illustrative purposes only and does not imply any limitation to the functions and scope of the embodiments of the present disclosure in any way.
[0132] As Figure 16 shown, the computing device 1600 includes a general computing device 1600. The computing device 1600 may include at least one or more processors or processing units 1610, a memory 1620, a storage unit 1630, one or more communication units 1640, one or more input devices 1650, and one or more output devices 1660.
[0133] In some embodiments, computing device 1600 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, and including accessories and peripherals of these devices or any combination thereof. It is contemplated that computing device 1600 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0134] Processing unit 1610 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in memory 1620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 1600. Processing unit 1610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0135] Computing device 1600 generally includes various computer storage media. Such media may be any media accessible by computing device 1600, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 1620 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 1630 may be any removable or non-removable media, and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed in computing device 1600.
[0136] Computing device 1600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 16 A disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0137] The communication unit 1640 communicates with another computing device via a communication medium. Additionally, the functionality of the components in the computing device 1600 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the computing device 1600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0138] The input device 1650 can be one or more of a variety of input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 1660 can be one or more of a variety of output devices, such as a display, speaker, printer, etc. With the aid of the communication unit 1640, the computing device 1600 can also communicate with one or more external devices (not shown), such as storage devices and display devices, and the computing device 1600 can also communicate with one or more devices that enable a user to interact with the computing device 1600, or any device that enables the computing device 1600 to communicate with one or more other computing devices (e.g., network cards, modems, etc.), if needed. Such communication can occur via an input / output (I / O) interface (not shown).
[0139] In some embodiments, some or all of the components of the computing device 1600 can also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which will not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network, such as the Internet. For example, a cloud computing provider provides applications via a wide area network, which can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the location of a remote data center. The cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point for the user. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, they can be provided from a traditional server or installed directly or otherwise on a client device.
[0140] The computing device 1600 can be used to implement visual data encoding / decoding in embodiments of the present disclosure. The memory 1620 can include one or more visual data encoding / decoding modules 1625 having one or more program instructions. These modules can be accessed and executed by the processing unit 1610 to perform the functions of the various embodiments described herein.
[0141] In an example embodiment of performing visual data encoding, the input device 1650 can receive visual data as the input 1670 to be encoded. The visual data can be processed by, for example, the visual data encoding / decoding module 1625 to generate an encoded bitstream. The encoded bitstream can be provided as the output 1680 via the output device 1660.
[0142] In an example embodiment of performing visual data decoding, the input device 1650 can receive the encoded bitstream as the input 1670. The encoded bitstream can be processed by, for example, the visual data encoding / decoding module 1625 to generate decoded visual data. The decoded visual data can be provided as the output 1680 via the output device 1660.
[0143] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes in form and detail can be made without departing from the spirit and scope of the present application as defined by the appended claims. These changes are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for visual data processing, comprising: For the conversion between the current visual unit of visual data and the bitstream of the visual data, determining at least one wavelet subband representation of the current visual unit according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, the wavelet subband representation being associated with the subbands of the wavelet of the current visual unit; and Performing the conversion based on the at least one wavelet subband representation.
2. The method according to claim 1, wherein the conversion includes decoding the current visual unit from the bitstream.
3. The method according to claim 2, wherein determining the at least one wavelet subband representation includes: Determining the quantized samples of the current visual unit based on the bitstream and a neural network-based latent sample reconstruction module, the quantized samples being associated with multiple channels; Determining an intermediate representation of the quantized samples, the intermediate representation being associated with at least one wavelet subband, the wavelet subband being associated with at least one of the multiple channels; and Determining the at least one wavelet subband representation based on the intermediate representation and the neural network-based inverse non-linear transformation.
4. The method according to claim 2 or claim 3, wherein performing the conversion includes: Determining the reconstruction of the current visual unit based on the at least one wavelet subband representation and an inverse wavelet transform.
5. The method according to claim 3, wherein the intermediate representation of the quantized samples is determined based on an inverse channel network for adjusting the number of channels of the quantized samples.
6. The method according to claim 5, wherein the inverse channel network includes at least one of the following: A batch normalization unit, A rectified linear unit (ReLU), or At least one convolutional layer.
7. The method according to claim 5 or claim 6, wherein the quantized samples include at least one of the following: A first dimension of a first spatial dimension of the current visual unit, A second dimension of a second spatial dimension of the current visual unit, A third dimension of the number of the multiple channels, or A fourth dimension of the number of feature maps associated with the current visual unit.
8. The method according to any one of claims 5 to 7, wherein the activation operation of the inverse channel network includes at least one of the following: A leaky rectified linear unit (ReLU), or A Sigmoid operation.
9. The method according to any one of claims 5 to 8, wherein there is no activation layer in the inverse channel network.
10. The method according to any one of claims 5 to 9, wherein the stride of the convolution of the inverse channel network is a first predefined value.
11. The method according to any one of claims 5 to 9, wherein the inverse channel network is a bottleneck, and the stride of the convolution of the inverse channel network is a second value.
12. The method according to claim 11, wherein the bottleneck inverse channel network performs a downsampling operation and an upsampling operation.
13. The method according to any one of claims 5 to 12, wherein the inverse channel network includes a fully connected neural network layer.
14. The method according to any one of claims 5 to 13, wherein the number of channels of the intermediate representation is less than or greater than the number of channels of the quantized samples.
15. The method according to any one of claims 3 to 7, wherein the intermediate representation comprises samples of reconstruction wavelet subbands of equal size.
16. The method according to any one of claims 3 to 15, wherein a first number of channels of the intermediate representation is associated with a first wavelet subband, and a second number of channels of the intermediate representation is associated with a second wavelet subband.
17. The method according to any one of claims 3 to 16, wherein a first network for the inverse non-linear transform comprises a plurality of branches associated with multiple sets of wavelet subbands, and a set of wavelet subbands is associated with the same spatial resolution.
18. The method according to claim 17, wherein the at least one wavelet subband representation comprises multiple sets of subband representations associated with the multiple sets of wavelet subbands.
19. The method according to claim 17 or claim 18, wherein the first network adjusts the spatial dimension of the intermediate representation, a first part of the intermediate representation is adjusted to match a first size of a first set of wavelet subbands among the multiple sets of wavelet subbands, and a second part of the intermediate representation is adjusted to match a second size of a second set of wavelet subbands among the multiple sets of wavelet subbands.
20. The method according to any one of claims 17 to 19, wherein the plurality of branches comprises a first branch, and the first branch comprises a first neural network layer for size adjustment.
21. The method according to any one of claims 3 to 16, wherein a first network for the inverse non-linear transform comprises a single branch, and the single branch comprises a first neural network layer for size adjustment.
22. The method according to claim 20 or claim 21, wherein the first neural network layer increases or decreases the spatial size of the intermediate representation.
23. The method according to any one of claims 20 to 22, wherein the first neural network layer comprises an upsampling layer or a downsampling layer.
24. The method according to any one of claims 20 to 23, wherein the first neural network layer comprises a transposed convolution layer or a convolution layer.
25. The method according to any one of claims 17 to 19, wherein a first spatial size of a first set of wavelet subbands among the multiple sets of wavelet subbands is different from a second spatial size of a second set of wavelet subbands among the multiple sets of wavelet subbands.
26. The method according to any one of claims 17 to 19, wherein a first size adjustment ratio is used for a first set of wavelet subbands among the multiple sets of wavelet subbands, and a second size adjustment ratio different from the first size adjustment ratio is used for a second set of wavelet subbands among the multiple sets of wavelet subbands.
27. The method according to any one of claims 17 to 19, wherein the at least one wavelet subband representation comprises a plurality of wavelet subband representations, each wavelet subband representation is associated with a set of wavelet subbands among the multiple sets of wavelet subbands, and wherein determining the plurality of wavelet subband representations comprises: Determine a plurality of segmented representations of the intermediate representation; and Determine the plurality of wavelet subband representations based on the plurality of segmented representations.
28. The method according to claim 27, wherein the plurality of segmented representations are associated with the same number of channels.
29. The method according to claim 27, wherein a first segmented representation among the plurality of segmented representations is associated with a first number of channels, and a second segmented representation among the plurality of segmented representations is associated with a second number of channels different from the first number of channels.
30. The method according to claim 29, wherein the first number of channels and the second number of channels are determined based on the importance of the wavelet subbands associated with the first segmented representation and the second segmented representation.
31. The method according to any one of claims 17 to 19, wherein the at least one wavelet subband representation includes a plurality of wavelet subband representations, each wavelet subband representation being associated with a group of wavelet subbands among the plurality of groups of wavelet subbands, and the plurality of wavelet subband representations are determined based on the intermediate representation.
32. The method according to claim 17, wherein one of the plurality of branches includes at least one of the following: a convolutional layer, a transposed convolutional layer, an activation operation, a rectified linear unit (ReLU), a leaky ReLU, a sigmoid operation, a hyperbolic tangent operation, or a normalization operation.
33. The method according to claim 32, wherein the wavelet subband representation of the wavelet subband having the lowest resolution among the plurality of groups of wavelet subbands is determined without resizing.
34. The method according to claim 32, wherein the stride of the transposed convolutional layer or the stride of the convolutional layer in the branch is a third number.
35. The method according to claim 1, wherein the transformation includes encoding the current visual unit into the bitstream.
36. The method according to claim 35, wherein determining the at least one wavelet subband representation includes: Determining wavelet subband information of the current visual unit based on wavelet transform; and Determining the at least one wavelet subband representation based on the wavelet subband information and the non-linear transform.
37. The method according to claim 35 or claim 36, wherein performing the transformation includes: Determining samples of the current visual unit based on the at least one wavelet subband representation and a channel network; Determining quantized samples of the current visual unit by quantizing the samples of the current visual unit; and Determining the bitstream based at least on the quantized samples and an entropy encoding and decoding module.
38. The method according to claim 36, wherein the wavelet transform includes at least one fixed parameter.
39. The method according to claim 36, wherein the wavelet transform includes at least one learnable parameter, and the at least one learnable parameter is updated together with another neural network for the transformation.
40. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 39.
41. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 39.
42. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for visual data processing for visual data, wherein the method comprises: determining at least one wavelet subband representation of a current visual unit of the visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; and generating the bitstream based on the at least one wavelet subband representation.
43. A method for storing a bitstream of visual data, comprising: determining at least one wavelet subband representation of a current visual unit of the visual data according to a neural network-based non-linear transformation or a neural network-based inverse non-linear transformation opposite to the non-linear transformation, the wavelet subband representation being associated with a subband of a wavelet of the current visual unit; generating the bitstream based on the at least one wavelet subband representation; and storing the bitstream in a non-transitory computer-readable recording medium.