Method and device for visual data processing and medium
By obtaining the component size relationship and synthesis transformation of visual data in the neural network model, the problem of insufficient encoding and decoding efficiency in the existing technology is solved, and flexible and efficient encoding and decoding adaptability is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing neural network-based image/video encoding and decoding technologies have room for improvement in encoding and decoding efficiency and are difficult to adapt to different application scenarios and needs.
A neural network-based encoding and decoding method is proposed. By obtaining the size relationship between the first and second components of the visual data, the synthesis transformation in the neural network model is used to adapt to different encoding and decoding formats, thereby improving the flexibility and efficiency of encoding and decoding.
It achieves more efficient encoding and decoding, can adapt to different application scenarios and needs, and improves encoding and decoding efficiency.
Smart Images

Figure CN121753335A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to visual data processing techniques, and more particularly, to neural network based visual data coding. BACKGROUND
[0002] In the past decade, deep learning has made rapid progress in various fields, especially in computer vision and image processing. Neural networks were originally invented through interdisciplinary research in neuroscience and mathematics. It shows strong ability in the context of nonlinear transformation and classification. Neural network based image / video compression technology has made significant progress in the past five years. It is reported that the latest neural network based image compression algorithm achieves rate-distortion (R-D) performance comparable to versatile video coding (VVC). With the continuous improvement of the performance of neural image compression, neural network based video compression has become an actively developing research field. However, the coding efficiency of neural network based image / video coding is generally expected to be further improved. SUMMARY
[0003] Embodiments of the present disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is presented. The method comprises: for a conversion between visual data and a bitstream of the visual data with a neural network (NN) based model, obtaining a format used for coding the visual data, the format indicating a relationship between a size of a first component of the coded visual data and a size of a second component of the coded visual data, and a first synthesis transform in the NN based model being used for the first component; based on the format, determining a second synthesis transform in the NN based model used for the second component; and performing the conversion based on the first synthesis transform and the second synthesis transform.
[0005] Based on the method according to the first aspect of the present disclosure, the second synthesis transform used for the second component of the visual data is determined based on the format used for coding the visual data. Compared with the conventional solution with fixed coding format, the proposed solution can advantageously support different coding formats in order to adapt to different applications. In this way, the coding flexibility can be improved, and thus the coding efficiency can be improved.
[0006] In a second aspect, an apparatus for visual data processing is presented. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data processing. The method includes: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; and generating a bitstream based on the first and second synthesis transforms.
[0009] In a fifth aspect, a method for storing a bitstream of visual data is proposed. The method includes: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; generating a bitstream based on the first and second synthesis transforms; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] FIG. 1A A block diagram illustrating an exemplary visual data encoding / decoding system according to some embodiments of the present disclosure is shown; FIG. 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme; FIG. 2 An exemplary potential representation of the image is shown; FIG. 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model; FIG. 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model together with a super-prior and an autoencoder; FIG. 5 An exemplary encoding process is shown; FIG. 6 An exemplary decoding process is shown; FIG. 7 Exemplary decoding processes according to some embodiments of this disclosure are shown; FIG. 8 An example of a learning-based image codec architecture is shown; FIG. 9 An exemplary synthetic transformation for learning-based image encoding and decoding is shown; FIG. 10 An example of a leaky ReLU activation function is shown; FIG. 11 An example ReLU activation function is shown; FIG. 12 The pixel shuffling and unshuffling operations are shown; FIG. 13 This demonstrates a transposed convolution using a 2x2 kernel; FIG. 14 An example of a downsampling encoding / decoding scheme is shown; FIG. 15A to FIG. 15C Examples of the proposed solution are shown, in which different upsampling processes are performed for 4:4:4, 4:2:2 and 4:2:0 encoding and decoding; FIG. 16 An example of the second synthetic transformation is shown; FIG. 17 An exemplary first synthetic transformation is shown; FIG. 18A to FIG. 18C An exemplary implementation of the proposed solution is shown; FIG. 19 An exemplary implementation of a subpixel convolutional unit is shown; FIG. 20A to FIG. 20C An exemplary implementation of the proposed solution is shown; FIG. 21 A flowchart of a method for visual data processing according to embodiments of the present disclosure is shown; and FIG. 22 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0013] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation
[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0015] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0016] The terms "an embodiment," "embodiment," "exemplary embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an exemplary embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0017] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0019] Example Environment FIG. 1AThis is a block diagram illustrating an exemplary visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.
[0020] Visual data source 112 may include sources such as visual data acquisition devices. Examples of visual data acquisition devices include, but are not limited to, interfaces for receiving visual data from visual data content providers, computer graphics systems for generating visual data, and / or combinations thereof.
[0021] Visual data may include one or more images. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the visual data. The bitstream may include encoded images and associated data. The encoded image is an encoded representation of an image. The associated data may include a sequence parameter set, an image parameter set, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data can be directly transmitted to a destination device 120 via network 130A through the I / O interface 116. The encoded visual data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0022] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0023] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or additional standards.
[0024] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Additionally, although some embodiments describe encoding steps in detail, it should be understood that the corresponding decoding steps of the inversion encoding will be implemented by the decoder. Furthermore, the term "visual data processing" includes visual data encoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another or at a different compression bit rate.
[0025] 1. Brief Overview Neural image and video compression methods include encoding and decoding modes. This disclosure relates to an adaptive encoding and decoding method, in which a codec (encoder / decoder) is proposed that can perform 4:4:4 encoding and decoding, 4:2:0 encoding and decoding, and 4:2:2 encoding and decoding.
[0026] 2 Introduction Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression techniques have made significant progress in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its early stages.
[0027] 2.1 Image / Video Compression Image / video compression generally refers to the computational technique of compressing images / videos into binary code for convenient storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video; this is called lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is not always necessary. The performance of image / video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes; the fewer, the better. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; the higher, the better.
[0028] Image / video compression techniques can be divided into two branches: classical video encoding / decoding methods and neural network-based video compression methods. Classical video encoding / decoding schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully hand-designing entropy encoding / decoding to model dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within existing classical video codecs and exists only as part of the framework, while the latter is a separate framework developed based on neural networks, independent of classical video codecs.
[0029] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-growing volume of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups: the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG). The ITU-T also has its own Video Codec Experts Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard: Multi-Functional Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC achieves an average bitrate reduction of 50% while maintaining the same visual quality.
[0030] Neural network-based image / video compression is not a new invention, as many researchers have worked on neural network-based image encoding and decoding. However, the network architectures are relatively shallow, and the performance is not satisfactory. Thanks to the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, the technology is still far from mature and many challenges need to be addressed.
[0031] 2.2 Neural Networks Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. Notably, these representations are not manually designed; instead, deep networks, including processing layers, are learned from large amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations and is therefore considered particularly suitable for processing native unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.
[0032] 2.3 Neural Networks for Image Compression Existing neural networks used for image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature.
[0033] 2.3.1 Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless encoding and decoding can achieve the lowest possible decoding rate. ,in It is a symbol The probability of [the lossless encoding / decoding method]. Many lossless encoding / decoding methods have been proposed in the literature, among which arithmetic encoding / decoding is considered one of the best methods. Given a probability distribution... Arithmetic encoding and decoding ensures that the encoding / decoding rate is as close as possible to its theoretical limit without considering rounding errors. Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality.
[0034] Following the predictive encoding / decoding strategy, for One approach to modeling this is to predict pixel probabilities one by one in raster scan order based on previous observations. It's an image.
[0035] (1) in These are the height and width of the image, respectively. Previous observations are also referred to as the current pixel's... Context When the image is large, estimating the conditional probability can be difficult, so a simplified approach is to limit the scope of its context.
[0036] (2) in It is a predefined constant that controls the scope of the context.
[0037] It should be noted that this condition can also take into account the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on previously encoded and decoded pixels (including R / G / B samples). The current G sample can be encoded and decoded based on previously encoded and decoded pixels and the current R sample. For encoding and decoding the current B sample, previously encoded and decoded pixels as well as the current R sample and the current G sample can also be considered.
[0038] Neural networks were initially introduced for computer vision tasks and have proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to adjust their behavior based on their context. Estimated probability .
[0039] Most methods model the probability distribution directly in the pixel domain. Some researchers have also attempted to model the probability distribution as a conditional distribution based on explicit or latent representations. That is, it is possible to estimate... (3) in It is an additional condition, and This means that modeling is divided into unconditional modeling and conditional modeling. The additional conditions can be image label information or high-level representations.
[0040] 2.3.2 Automatic Encoder Autoencoders originate from the renowned work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part transforms the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders can automatically learn representations and eliminate the need for hand-crafted features, which is considered one of the most significant advantages of neural networks.
[0041] FIG. 1B This is a schematic diagram of a typical transform encoding / decoding scheme. Original image. Analysis network Transformation to achieve latent representation The latent representation y is quantized and compressed into bits. The number of bits... Used to measure codec rate. Quantized latent representation. Then by the synthetic network Inverse transform to obtain the reconstructed image Distortion is achieved through the use of functions. right The transformation is calculated in the perceptual space.
[0042] Applying autoencoder networks to lossy image compression is intuitive. We simply need to encode the learned latent representation from a trained neural network. However, adapting autoencoders to image compression is not straightforward, as the original autoencoders are not optimized for compression, making direct use of the trained autoencoder inefficient. Furthermore, other major challenges exist: First, the low-dimensional representation should be quantized before encoding, but quantization is non-differentiable, which is necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios because both distortion and bit rate need to be considered. Estimating the bit rate is challenging. Third, practical image encoding / decoding schemes need to support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Many researchers have been actively contributing to this field to address these challenges.
[0043] An autoencoder prototype for image compression, such as FIG. 1B As shown, it can be regarded as Transform coding Strategy. Original image use Analysis network Transformed, where It is the potential representation that will be quantized and encoded / decoded. Synthesis The network will quantify the potential representation Perform inverse transform to obtain the reconstructed image Framework utilization distortion loss function (i.e. Training is conducted, among which Distortion between It is a representation based on quantification. The calculated or estimated bit rate, and These are Lagrange multipliers. It should be noted that... It can be computed in the pixel domain or the receptive domain. All existing research follows this prototype, with differences only in network structure or loss function.
[0044] 2.3.3 Super-prior model In transform encoding and decoding methods used for image compression, the encoder subnetwork (Section 2.3.2) uses parametric analysis of the transform. Transform the image vector x into a latent representation Then quantify it to form .because Since they are discrete values, they can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as a bit sequence.
[0045] from FIG. 2 The left and right center images clearly show that Significant spatial dependencies exist among the elements. Notably, their scales (middle right image) appear to be spatially coupled. An additional set of random variables is introduced into the existing design. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as FIG. 3 As shown.
[0046] exist FIG. 3 In the middle, the encoder is on the left side of the model. and decoder (Explained in Section 2.3.2). The right-hand side is used to obtain... Additional encoders utilizing prior information and decoders that utilize prior information Network. In this architecture, the encoder subjectes the input image x to... The response with a standard deviation exhibiting spatial variation is obtained. .response fed to In summary The distribution of standard deviations in the data. Then... Quantified ( The data is compressed and transmitted as side information. The encoder then uses the quantized vector data... To estimate the spatial distribution of standard deviation And use it to compress and transmit quantized image representations. The decoder first recovers the signal from the compressed signal. Then the decoder uses To obtain ,Should To provide the decoder with the correct probability estimate and thus successfully recover the value. Then the decoder will Feed to To obtain a reconstructed image.
[0047] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. FIG. 2 The rightmost image in the middle corresponds to the quantization latent value when using an encoder / decoder that leverages prior information. Compared to the middle right image, spatial redundancy is significantly reduced because the samples of the quantization latent value have lower correlation.
[0048] refer to FIG. 2 Left: Image from the Kodak dataset. Middle left: Visualization of the latent representation y of this image. Middle right: Standard deviation of the latent values. Right side: The latent value y after introducing a super-prior network (an encoder and a decoder that utilize super-prior information).
[0049] FIG. 3 The network architecture of an autoencoder implementing a prior model is shown. The left side shows the image autoencoder network, and the right side corresponds to the prior subnetwork. The analytic transform and the synthetic transform are represented as follows: Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks, utilizing the hyperprior information of the encoder (denoted as...). ) and decoders that utilize prior information (represented as The prior model generates quantified potential values of prior information. The quantified potential value of prior information ( This includes information about quantifying potential values. Information about the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).
[0050] 2.3.4 Context Model Although the prior model improves the quantification of latent values Modeling the probability distribution of quantified potential values is possible, but further improvements can be achieved by utilizing an autoregressive model (context model) that predicts quantified potential values from the causal context of quantified potential values.
[0051] The term autoregressive means that the output of a process is later used as its input. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.
[0052] FIG. 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model with a super-prior and an autoencoder. Table 1 below shows the meaning of the different symbols.
[0053] Table 1 – Symbol Explanation
[0054] The existing design utilizes a joint architecture in which both a super-prior model subnetwork (an encoder and a decoder utilizing super-prior information) and a context model subnetwork are employed. The super-prior and context models are combined to learn about the quantized latent value. The probability model is then used for entropy encoding and decoding. For example... FIG. 4 As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean for the Gaussian probability model. And variance (scale) (or variance) The parameters are then used. The Gaussian probability model is then used by the arithmetic encoder (AE) module to encode the samples of the quantized latent values into the bitstream. In the decoder, the Gaussian probability model is used to obtain the quantized latent values from the bitstream via the arithmetic decoder (AD) module. .
[0055] FIG. 4 A combined model is shown, which jointly optimizes the following: an autoregressive component (context model) that estimates the probability distribution of latent values from the causal context of the latent values, and a super-prior and a low-level autoencoder. Real-valued latent representations are quantized (Q) to create quantized latent values (…). ) and quantified potential value of prior information ( ), this quantified potential value ( ) and quantified potential value of prior information ( The image is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted areas correspond to the components performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream.
[0056] Typically, potential samples are modeled as Gaussian distributions or Gaussian mixture models (not limited to these). In existing designs, according to FIG. 4 The context model and the super-prior are jointly used to estimate the probability distribution of potential samples. Since the Gaussian distribution can be defined by its mean and variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as...). ).
[0057] 2.3.5 Gain Variational Automatic Encoder (G-VAE) Typically, neural network-based image / video compression methods require training multiple models to adapt to different bit rates. Gain variational autoencoders (G-VAEs) are variational autoencoders with a pair of gain units, designed to achieve continuously variable bit rate adaptation using a single model. It consists of a pair of gain units, typically inserted into the encoder's output and the decoder's input. The encoder's output is defined as the latent representation. ,in This represents the number of channels, height, and width of the latent representation. Each channel of the latent representation is represented as... ,in A pair of gain units includes a gain matrix. and the inverse gain matrix, where This is the number of gain vectors. A gain vector can be represented as... , ,in Indicates the index of the gain vector in the gain matrix.
[0058] The motivation behind the gain matrix is similar to the quantization table in JPEG, controlling the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector:
[0059] in It is channel-wise multiplication, that is ,and It is the gain vector The first in There are several gain values. The inverse gain matrix used on the decoder side can be represented as... It includes One inverse gain vector, i.e. , The inverse gain process is represented as:
[0060] in It is the quantized latent representation of the decoded data, and It is the quantized latent representation of the inverse gain, which will be fed into the synthesis network.
[0061] To achieve continuous variable rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors... and The interpolated gain vector can be obtained using the following formula:
[0062] in These are interpolation coefficients, which control the bit rate of the generated gain vector pairs. Because... Since it is a real number, it is possible to achieve any bit rate between two given gain vector pairs.
[0063] 2.3.6 Encoding process using a joint autoregressive hyperprior model FIG. 4 This corresponds to existing compression methods. The encoding and decoding processes will be described in this section and the next section, respectively.
[0064] FIG. 5 The encoding process is described. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transform representation called the latent value, which is... express. It is then fed into the quantizer block, denoted by Q, to obtain the quantization potential value ( ). Then, using an arithmetic coding module (denoted as AE), it is converted into a bitstream (bits1). The arithmetic coding blocks are sequentially... Each sample point is converted into a bit stream (bits1) one by one.
[0065] The module's encoder utilizing prior information, context, decoder utilizing prior information, and entropy parameter subnetwork are used to estimate the quantization potential. The probability distribution of the sample points. Potential The information is input into an encoder that utilizes prior information, and its output is the potential value of the prior information (denoted as...). The latent value of the prior information is then quantified. The second bitstream (bits2) is generated using the arithmetic coding (AE) module. The decompositional entropy module generates the probability distribution used to encode the quantized hyperprior information latent values into a bitstream. The quantized hyperprior information latent values include information about the quantized latent values (…). Information about the probability distribution of ( ).
[0066] Entropy parameter subnetwork generation was used to encode quantized latent values. The probability distribution is estimated. Information generated from the entropy parameter is typically used together to obtain the mean of the Gaussian probability distribution. And variance (scale) (or variance) Parameters. The Gaussian distribution of the random variable x is defined as follows: , where parameters It is the mean or expected value of the distribution (also the median and mode), while the parameter This is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. In the existing design, the entropy parameter module is used to estimate the mean and variance values.
[0067] The decoder of the subnetwork, utilizing prior information, generates part of the information used by the entropy parameter subnetwork, while another part is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the quantized latent values. It is typically a matrix composed of many sample points. Sample points can be indicated using indices, for example... [i,j,k] or [i,j], specifically depends on the matrix. Dimensions. Sample points [i,j] are encoded sequentially by the AE, typically using a raster scan order. In a raster scan order, the matrix rows are processed from top to bottom, with samples in a row processed from left to right. In such a scenario (where the AE encodes samples into a bitstream using a raster scan order), the context module uses samples previously encoded in the raster scan order to generate a sequence with the samples. Information related to [i,j]. The information generated by the context module and the decoder utilizing prior information is combined by the entropy parameter module to generate information used to quantize the latent values. The probability distribution encoded as a bit stream (bits1).
[0068] Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder.
[0069] It is worth noting that other names can be used for the modules mentioned above.
[0070] In the above description, FIG. 5 All elements in the algorithm are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).
[0071] 2.3.7 Decoding process using a joint autoregressive superprior model FIG. 6 The decoding process is shown separately.
[0072] During decoding, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decompositional entropy subnetwork. The decompositional entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This is the quantized potential value of the prior information. The AD process recovers to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.
[0073] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks, context, decoder utilizing prior information, and entropy parameters used in the decoder are the same as the three sub-networks in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder as in the encoder, which is crucial for reconstructing the quantized latent value without any loss. This is essential. As a result, the quantization potential value can be obtained in the decoder as well as in the encoder. Same version.
[0074] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of quantized potential values one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial, and therefore cannot be accelerated using techniques such as parallelization.
[0075] Finally, the fully reconstructed quantized potential value Input into the synthesis transform (in) FIG. 6 The module (represented as decoder) is used to obtain the reconstructed image.
[0076] In the above description, FIG. 6 All elements in the array are collectively referred to as decoders. The synthetic transformation that converts quantized latent values into a reconstructed image is also called a decoder (or autodecoder).
[0077] 2.4 Neural Networks for Video Compression Similar to traditional video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, it requires more effort to overcome its challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Thus, inter-frame prediction is a key step in these works. Motion estimation and compensation have been widely adopted, but only recently have they been implemented using trained neural networks.
[0078] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the case of random access, decoding can begin at any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, the aim is to reduce decoding time, so usually only the earlier frames in the time domain can be used as reference frames to decode subsequent frames.
[0079] 2.5 Prerequisites Almost all natural images / videos are in digital format. Grayscale digital images can be generated by... It means that, among them It is a collection of pixel values. It is the image height. It is the image width. For example, This is a common setting, in this case, Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed images certainly have fewer bits.
[0080] Color images are typically represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by... This means that three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Most neural network-based video compression schemes have been developed in the RGB color space, while traditional codecs typically use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels: Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. The advantage comes from the fact that Cb and Cr are often downsampled for pre-compression, as the human visual system is less sensitive to the chrominance component.
[0081] A color video sequence consists of multiple color images (called frames) to record a scene at different timestamps. For example, in the RGB color space, a color video can be composed of... It means that, among them It is the number of frames in the video sequence. .if If the video has 50 frames per second (fps), then the data rate of the uncompressed video is... Bits per second (bps), approximately 2.32 Gbps, requires a significant amount of storage, so it definitely needs to be compressed before being transmitted over the internet.
[0082] Typically, lossless methods can achieve compression ratios of around 1.5 to 3 for natural images, which is clearly below what is required. Therefore, lossy compression has been developed to achieve further compression ratios, but at the cost of introducing distortion. Distortion can be measured by calculating the mean squared difference (MSE) between the original and reconstructed images. For grayscale images, the MSE can be calculated using the following formula.
[0083]
[0084] Therefore, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR):
[0085] in yes The maximum value in the range is 255, for example, for an 8-bit grayscale image. Other quality assessment metrics include structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).
[0086] To compare different lossless compression schemes, it is sufficient to compare the compression ratio at a given bitrate, or vice versa. However, to compare different lossy compression methods, both bitrate and reconstruction quality must be considered. For example, a common approach is to calculate the relative rates at several different quality levels and then average the bitrates; the average relative bitrate is called the Bjontegaard increment rate (BD rate). Other important aspects for evaluating image / video codec schemes include encoding / decoding complexity, scalability, and robustness.
[0087] 2.6 Separate processing of the luminance and chrominance components of an image FIG. 7 The decoding process according to this disclosure is shown.
[0088] According to one implementation, the luminance and chrominance components of an image can be decoded using separate sub-networks. In the above figures, the luminance component of the image is processed by sub-networks such as "synthesis," "predictive fusion," "masked convolution," "decoder utilizing prior information," and "variance decoder utilizing prior information." The chrominance component is processed by sub-networks such as "synthesized UV," "predictive fusion UV," "masked convolution UV," "decoder UV utilizing prior information," and "variance decoder UV utilizing prior information."
[0089] The advantage of this separate processing is that the computational complexity of image processing is reduced by applying it separately. Typically, in neural network-based image and video decoding, computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is 192, the computational complexity will be proportional to 192 × 192. On the other hand, if the feature maps are divided into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128 × 128 + 64 × 64, which corresponds to a 45% reduction in complexity. Generally, separate processing of the luminance and chrominance components of an image does not lead to an excessive performance degradation because the correlation between the luminance and chrominance components is usually very small.
[0090] FIG. 7 The processing (decoding process) in the code can be explained as follows: 1. First, the decompositional entropy model was used to decode the quantization potential values of luminance and chrominance, i.e. FIG. 7 In and .
[0091] 2. The probability parameters (e.g., variance) generated by the second network are used to generate quantized residual latent values by performing an arithmetic decoding process.
[0092] 3. Quantified residual potential values are utilized as follows: FIG. 7 The orange inverse gain unit (iGain) is inversely gained. The output of the inverse gain unit is expressed for the luminance and chrominance components respectively. and .
[0093] 4. For the luminance component, the following steps are performed repeatedly until the result is obtained. All elements: a. The first subnetwork was used for... The already obtained samples are used to estimate the quantized potential value ( The mean parameter of ).
[0094] b. Quantified residual potential value The mean was used to obtain The next element.
[0095] 5. In obtaining After all the samples are obtained, a synthetic transformation can be applied to obtain the reconstructed image.
[0096] 6. For the chromaticity components, steps 4 and 5 are the same, but with a separate network set.
[0097] 7. The decoded luminance component is used to obtain additional information for the chrominance component. Specifically, an inter-channel related information filter (ICCI) subnetwork is used for chrominance component recovery. The luminance is fed into the ICCI subnetwork as additional information to assist in chrominance component decoding.
[0098] 8. After the luminance and chrominance components are reconstructed, adaptive color transformation (ACT) is performed.
[0099] The module named ICCI is a neural network-based post-processing module. The proposed solution is not limited to the ICCI subnetwork; any other neural network-based post-processing module can also be used.
[0100] An exemplary implementation of the proposed solution is in FIG. 7 The decoding process is illustrated. The framework comprises two branches, one for the luma component and one for the chroma component. Within each branch, the first sub-network includes a context, prediction, and optionally, a decoder module utilizing advanced prior information. The second network includes a variance decoder module utilizing the advanced prior information. The quantized advanced prior information latent value is... and The arithmetic decoding process generates quantized residual latent values, which are further fed into the iGain unit to obtain quantized residual latent values for gain. and .
[0101] After obtaining the residual potential value, a recursive prediction operation is performed to obtain the potential value. and The following steps describe how to obtain potential values. The sample points, and the chromaticity components are processed in the same way but using different networks.
[0102] 1. The autoregressive context module is used when using samples. This is used to generate the first input to the prediction module, where the (m, n) pairs are the indices of the sample points of the obtained potential values.
[0103] 2. Optionally, the second input to the prediction module is obtained by using a decoder that utilizes prior information and quantized potential values of the prior information. Obtained.
[0104] 3. Using the first and second inputs, the prediction module generates the mean. .
[0105] 4. Mean And quantified residual potential value Added together to obtain potential value .
[0106] 5. Steps 1-4 are repeated for the next sample point.
[0107] Whether and / or how at least one of the methods disclosed herein can be applied can, for example, be transmitted from the encoder to the decoder in a bitstream via signal transmission.
[0108] Alternatively, whether and / or how to apply at least one of the methods disclosed herein can be determined by the decoder based on encoding / decoding information such as dimensions, color format, etc.
[0109] Alternative or additional locations, modules named MS1, MS2, or MS3+O (in) FIG. 7 (The input) can be included in the processing stream. The module can perform operations on its input to obtain an output by multiplying the input by a scalar or adding an additional component to the input. The scalar or additional component used by the module can be indicated in the bitstream.
[0110] FIG. 7 The module named RD or AD in the code can be an entropy decoding module. It can be a range decoder or an arithmetic decoder, etc.
[0111] The solutions described in this article are not limited to FIG. 7 The example demonstrates a specific combination of units. Some modules may be missing, and some modules may be replaced in the order they are processed. Additionally, supplementary modules may be included. For example: 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined by another module, which can be based on a neural network.
[0112] 2. One or more modules named MS1, MS2, or MS3+O may be removed. The removal of one or more of the scaling and addition modules does not affect the core of the proposed solution.
[0113] exist FIG. 8The bitstream also uses an asterisk (*) to indicate other operations performed during the processing of the luminance and chrominance components. These operations are denoted as MS1, MS2, MS3+0. These operations can be, but are not limited to, adaptive quantization, latent sample scaling, and latent sample offset operations. For example, in adaptive quantization, this might correspond to scaling the samples with a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. Latent scaling might correspond to scaling the samples with a multiplier after the prediction process, where the multiplier value is predefined or indicated in the bitstream. Offset operations might correspond to adding an appended element to the sample, where the value of the appended element can be indicated, estimated, or predetermined in the bitstream.
[0114] Another operation can be slicing, where the samples are first sliced (grouped) into overlapping or non-overlapping regions, each of which is processed independently. For example, samples corresponding to the luminance component can be divided into slices with a slice height of 20 samples, while the chrominance component can be divided into slices with a slice height of 10 samples for processing.
[0115] Another application is wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the number of samples that can be processed in parallel can be indicated by control parameters. These control parameters can be indicated, estimated, or predetermined in the bitstream. In the case of separate luminance and chrominance processing, the number of samples that can be processed in parallel can be different, so different indicators can be transmitted via signals in the bitstream to control the operation of luminance and chrominance processing separately.
[0116] 2.7 Color Separation and Conditional Encoding / Decoding In one example, such as FIG. 8 As shown, the primary and secondary color components of the image are encoded and decoded separately using networks with similar architectures but different numbers of channels. All boxes with the same name are sub-networks with similar architectures, differing only in input / output tensor sizes and the number of channels. The primary component has the following number of channels: The number of channels for the secondary component is The vertical arrows (pointing downwards) indicate the data flow related to the encoding and decoding of secondary color components. The vertical arrows show the data exchange between the primary and secondary component pipelines.
[0117] The input signal to be encoded is represented as The latent space tensor in the bottleneck of a variational autoencoder is The subscript "Y" indicates the primary component, while the subscript "UV" is used for the secondary components in the stitching, including the chromaticity component.
[0118] FIG. 8A learning-based image codec architecture is shown.
[0119] First, the input image in RGB color format is converted into primary (Y) and secondary (UV) components. Primary component... Independent of secondary components The image is encoded and decoded, and the size of the encoded / decoded image is equal to the input / decoded image size. Secondary components use data from the primary components. As auxiliary information, it is conditionally encoded and decoded for use in encoding. And using auxiliary information from the main component. As a potential tensor, used for decoding Aside from the number of channels, channel size, and several entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost identical. Therefore, the primary and secondary latent tensors will generate two different bitstreams based on two different entropy models. Before encoding, , This module adjusts the sample point position through downsampling (in...) FIG. 8 The above is marked as " This essentially means that the encoded image size of the secondary component differs from that of the encoded image size of the primary component. The scaling factor s is variable, but the default scaling factor is 1. In conditional encoding and decoding, the size of the auxiliary input tensor is adjusted so that the encoder receives primary and secondary component tensors with the same image size. After reconstruction, the secondary component utilizes a neural network-based upsampling filter module (…). FIG. 8 "NN color filter" The image is rescaled to its original size, and the module outputs a factor. The secondary component that is upsampled.
[0120] FIG. 9 The example illustrates an image encoding / decoding system where the input image is first converted into primary (Y) and secondary (UV) components. Output , This corresponds to the reconstructed output of the primary and secondary components. At the end of the processing, , It is converted back to RGB color format. Typically, It is downsampled (resized) before being processed by the encoding and decoding modules (neural network). For example, The size can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Therefore, the processing of minor components involves approximately 50% × 50% = 25% fewer samples, making it computationally less complex.
[0121] 2.8 Cropping Operations in Neural Network-Based Image / Video Encoding and Decoding FIG. 9 An example of synthetic transformation for learning-based image encoding and decoding is shown.
[0122] The above example of a synthetic transform consists of a series of four convolutions and an upsampling with a stride of 2. The synthetic transform subnetwork is in FIG. 9 The tensor dimensions in different parts of the composite transform before the clipping layer are shown in the diagram. FIG. 9 The image above.
[0123] Clipping layers will tensor dimensions Change to ,in ;here This is the depth of the convolution performed in the codec architecture. For the principal components, the synthesized transform receive size is... The input tensor, where The output of the synthesis transform of the principal components is ,in .
[0124] For the secondary components, the synthesized transform receiver size is: The input tensor; For the principal components, the output of the composition transform is ,in For secondary components, the input size is... ,in It is the scaling factor. For example, the scaling factor could be 2, where the minor components are downsampled by a factor of 2.
[0125] Based on the above explanation, the operation of the clipping layer depends on the output dimensions H and W and the depth of the clipping layer. FIG. 10 The leftmost clipping layer has a depth of 0. The output of this clipping layer must be equal to H and W (output dimensions). If the input dimensions of this clipping layer are greater than H or W in the horizontal or vertical dimension, clipping is required in that dimension. The second clipping layer, counting from left to right, has a depth of 1. The output of the second clipping layer must be equal to... This means that if the input to the second clipping layer is greater than h1 or w1 in any dimension, clipping is applied in that dimension. In summary, the operation of the clipping layers is controlled by the output dimensions H and W. In one example, if both H and W are equal to 16, the clipping layer does not perform any clipping. On the other hand, if both H and W are equal to 17, all four clipping layers will perform clipping.
[0126] 2.9 Displacement Bitwise shift operators can use functions is represented, where n is an integer. If n is greater than 0, it corresponds to the right shift operator (>>), which shifts the input bits to the right, and the left shift operator (<<), which shifts the bits to the left. In other words, The operation corresponds to:
[0127] The output of the shift operation is an integer value. In some implementations, the floor() function can be added to the definition. floor(x) is equal to the largest integer less than or equal to x.
[0128] The " / / " operator or integer division operator. It is an operation that includes division and truncating the result towards zero. For example, 7 / / 4 is equivalent to 7 / 4 and is truncated to 1, -7 / / (-4) is equivalent to -7 / (-4) and is truncated to 1. -7 / / 4 is equivalent to -7 / 4 and is truncated to -2; 7 / / (-4) is equivalent to 7 / (-4) and is truncated to -2.
[0129]
[0130] Formula 3: Bitwise shift operators as an alternative implementation of right shift or left shift.
[0131] x >> y The two's complement integer representation of x is arithmetically right-shifted by y binary digits. This function is only defined for non-negative integer values of y. The bit shifted into the most significant bit (MSB) as a result of the right shift has the value equal to the MSB of x before the shift operation.
[0132] x << y The two's complement integer representation of x is arithmetically left-shifted by y binary digits. This function is only defined for non-negative integer values of y. The bit shifted into the least significant bit (LSB) as a result of the left shift has the value equal to 0.
[0133] 2.10 Convolution operation Convolution is essentially a simple operation: starting from the kernel, which is just a small weight matrix. The kernel "slides" over the input data, performs element-wise multiplication with the current part of the input, and then sums the results to a single output pixel. In some cases, the convolution operation can include a "bias", which is added to the output of the element-wise multiplication operation.
[0134] The convolution operation can be described by the following mathematical formula. The output out1 can be obtained as:
[0135] where w1 is the multiplication factor and K1 is called the bias (additive term), Let N be the k-th input, N be the kernel size in one direction, and P be the kernel size in another direction. A convolutional layer can include convolution operations, which can generate more than one output. Other equivalent descriptions of convolution operations can be found below:
[0136] In the above formula, "c" indicates the channel number. It is equivalent to the output number, out[1,x,y] is one output, and out[2,x,y] is the second output. k is the input number. It is an input, and It is the second input.
[0137] w1 or w describes the weights of the convolution operation.
[0138] 2.10.1 Two-dimensional convolution operation Convolution operations can be defined in 1, 2, 3, 4, ... dimensions. For example, a 2D convolution operation can be defined as:
[0139] 2.11 LeakyReLU activation function LeakyReLU activation function FIG. 11 The function is described in the diagram. According to this function, if the input is positive, the output equals the input. If the input (y) is negative, the output equals a. y. a is typically (but not limited to) a value less than 1 and greater than 0. Because the multiplier a is less than 1, it can be implemented as multiplication or division with non-integers. The multiplier a can be referred to as the negative slope of the LeakyReLU function.
[0140] 2.12 ReLU Activation Function The ReLU activation function in FIG. 12 The function is shown in the diagram. According to this function, if the input is positive, the output is equal to the input. If the input (y) is non-positive, the output is 0.
[0141] 2.13 Pixel shuffling and unshuffling functions FIG. 13 The pixel shuffling and unshuffling operations are described.
[0142] PixelShuffle is an operation used in super-resolution models to implement efficient sub-pixel convolutions with a stride of 1 / r. Specifically, it transforms a convolution into a shape of [ The elements in the tensor of ] are rearranged into a shape of [ The tensor of ].
[0143] The pixel deshuffling operation is the inverse of the shuffling operation, wherein [ The input tensor of ] is converted to a shape of [ The tensor of ].
[0144] 2.14 Deconvolution Operation Transposed convolutional (also known as deconvolution) layers are typically used for upsampling, i.e., to generate an output feature map with a spatial dimension larger than the input feature map. The transposed convolution operation... FIG. 13 is exemplified in .
[0145] FIG. 14 The transposed convolution using a 2x2 kernel is shown. The shaded area represents a portion of the intermediate tensor, along with the input and kernel tensor elements used for computation.
[0146] 3 questions 3.1 Core Issues In existing neural network-based image and video codecs, the original input image can first be converted to different color formats (e.g., YCbCr, YUV, etc.), and different color planes can be downsampled to increase the compression ratio. Typically, three input planes are used, such as Y, U, and V corresponding to different color planes. The three different planes can be encoded / decoded with the same spatial size, or some planes can be downsampled before encoding (reducing the spatial size).
[0147] FIG. 14 An example of a downsampling encoding / decoding scheme is shown.
[0148] Such an encoding / decoding scheme is as follows FIG. 15A to FIG. 15C As shown in the figure, the input image has a 4:2:0 format. (4:2:0) is a symbol indicating that the input image has the following dimensions: • The first component (the first plane) has dimensions H and W, where H is the height and W is the width.
[0149] • The second (plane) and third (plane) components have dimensions of H / 2 and W / 2, respectively. The dimensions of the second and third components are reduced because they are chromaticity components and they include less information about the human visual system (e.g., they are less relevant to human perception).
[0150] On the other hand, using 4:2:0 downsampling and compression is not always beneficial. The information included in the second and third components is sometimes important, and it may be desirable to preserve that information. The encoding and decoding process is lossy, especially with 4:2:0 encoding and decoding, resulting in the loss of information that cannot be retrieved later.
[0151] Therefore, the problem lies in the lack of adaptive encoding / decoding methods in neural network-based encoding / decoding systems.
[0152] In the remainder of the document, use the following definitions: 4:4:4 encoding / decoding: The primary and secondary components of an image have the same dimensions as the input to the encoding process and the output to the decoding process.
[0153] 4:2:0 encoding / decoding: The primary and secondary components of an image have different sizes as inputs to the encoding process and outputs to the decoding process. If the size of the first component is... (high (width), then the dimension of the second component is .
[0154] 4:2:2 encoding / decoding: The primary and secondary components of an image have different sizes as inputs to the encoding process and outputs to the decoding process. If the size of the first component is... (high (width), then the dimension of the second component is .
[0155] 4 Specific Solutions This disclosure relates to an adaptive encoding / decoding method, which proposes a codec (encoder / decoder) capable of performing 4:4:4 encoding / decoding, 4:2:0 encoding / decoding, and 4:2:2 encoding / decoding.
[0156] 4.1 Core of the Specific Solution According to some embodiments of this disclosure, a bitstream is generated by an encoder and transmitted to a decoder. The decoder includes two synthesis transforms. A first synthesis transform operates on a first component. And a second synthesis transform operates on a second component. Wherein; • The second synthesis transform consists of three parts: the first part, the upsampler unit, and the third part.
[0157] • The upsampler unit is configured according to the encoding / decoding mode (e.g., 4:4:4, 4:2:0, or 4:2:2).
[0158] • Optionally, the third part of the synthetic transformation can be configured according to the encoding / decoding mode. FIG. 15A Examples of the proposed solution are shown, where different upsampling processes are performed for 4:4:4, 4:2:2, and 4:2:0 encoding and decoding.
[0159] The proposed solution is FIG. 15A , 15B As illustrated in 15C. The upsampler unit is executed according to the encoding / decoding mode.
[0160] • When the encoding / decoding mode is 4:4:4 (corresponding to...) FIG. 15BWhen processing by the upsampler unit, the upsampler unit can be configured to perform upsampling in both the horizontal and vertical directions. The output of the first part has a spatial dimension of h / 2 x w / 2. After upsampling by the upsampler unit, the spatial dimension becomes h x w. Finally, after processing by the third part, the spatial dimension becomes H x W.
[0161] • When the encoding / decoding mode is 4:2:2 (corresponding to...) FIG. 15C When processing by the upsampler unit, the upsampler unit can be configured to upsample only in the horizontal direction or only in the vertical direction. The output of the first part has a spatial dimension of h / 2 × w / 2. After upsampling by the upsampler unit, the spatial dimension becomes h × w / 2. Finally, after processing by the third part, the spatial dimension becomes H × W / 2.
[0162] • When the encoding / decoding mode is 4:2:0 (corresponding to...) FIG. 16 When the first part is processed, the upsampler unit can be configured not to perform upsampling. The output of the first part has spatial dimensions of h / 2 and w / 2. After upsampling by the upsampler unit, the spatial dimensions remain h / 2 and w / 2. Finally, after processing by the third part, the spatial dimensions become h / 2 and w / 2.
[0163] In one example, the second synthetic transform may not include the upsampler unit because it is redundant.
[0164] 4.2 Details of the specific solution • According to some embodiments, the first synthesis transform can be used to process the main components, and its structure can be the same for different encoding / decoding modes (4:4:4, 4:2:0 or 4:2:2).
[0165] • According to some embodiments, the second synthesis transform can be used to process the second component, and its structure can be different for different encoding / decoding modes (4:4:4, 4:2:0 or 4:2:2).
[0166] More specifically, the upsampler unit can be configured differently for different encoding / decoding modes.
[0167] • The vertical and horizontal dimensions of the inputs to the first and second composite transforms can be the same.
[0168] The inputs to the first and second composition transformations can be the same.
[0169] The same input can be input into the first composition transform and the second composition transform.
[0170] • According to some embodiments, the output of the second synthetic transform can be a component of the reconstructed image.
[0171] o This component can be a chromaticity component.
[0172] • Upsampler units can be implemented using deconvolutional layers, transposed convolutional layers, pixel shuffling layers, or nearest-nearest-sample upsampled layers.
[0173] In the most recent sample upsampling layer, the size of the input tensor is increased by copying the most recent sample value or by interpolating the most recent sample to the newly generated sample position (which is generated by increasing the size of the tensor).
[0174] In the case of deconvolution or transposed convolution, the stride can be different for different encoding / decoding modes. The horizontal stride or vertical stride, or both, of deconvolution (or transposed convolution) can be modified based on the encoding / decoding mode.
[0175] • An example of the second composition transformation can be seen as follows: FIG. 17 As shown. In this example, the first part, the third part, and the upsampler unit of the synthesis transform are enclosed in dashed boxes. Based on the example using the encoding / decoding mode, different branches can be used during processing using the synthesis transform.
[0176] For example, if the codec mode is 4:2:0, a bottom branch can be used where upsampling is not performed.
[0177] For example, if the encoding / decoding mode is 4:2:2, then the intermediate branches can be traversed, where only the vertical or horizontal upsampling is performed. This is depicted using deconvolutional layers: In the deconvolutional layer (represented as...) In ), two step size values are defined ( One corresponds to the upsampling ratio in one dimension, and the second corresponds to the upsampling ratio in the second dimension.
[0178] For example, if the encoding / decoding mode is 4:4:4, you can traverse the top branch, where 2x upsampling is performed in both the vertical and horizontal directions.
[0179] • An exemplary first synthetic transformation can be as follows FIG. 17 As shown. In this example, the first synthesis transformation is not changed based on the encoding / decoding mode.
[0180] • The inputs for the first synthesizer and the second synthesizer can be the same. FIG. 16 and FIG. 19 In the middle, one of the inputs (i.e. The two composition transformations are identical. The second composition transformation additionally receives a second input. .
[0181] • The inputs to the first synthesizer and the second synthesizer can be different. The intermediate or final output of the first synthesizer can be used in the second synthesizer.
[0182] • Alternatively or additionally, the second synthetic transform can consist of two parts: a front-end and a final upsampling unit.
[0183] In one example, the upsampling unit can be a subpixel convolutional unit.
[0184] A subpixel convolutional unit can include two layers: a convolutional layer and a pixel shuffling layer. An exemplary implementation of a subpixel convolutional unit is shown below. FIG. 18A As shown.
[0185] The implementation of the pixel shuffling layer can be described as follows: The pixel shuffling layer is represented as ,in It is the upsampling factor. This layer will have a shape of tensor The elements in the text are rearranged into shapes. tensor ,in .
[0186] for , .
[0187] The pixel shuffling layer is represented as ,in It is the upsampling factor. This layer will have a shape of tensor The elements in the text are rearranged into shapes. tensor ,in .
[0188] for , .
[0189] o An exemplary implementation can be FIG. 21 , 18B As shown in 18C.
[0190] o The second synthetic transformation can be configured based on the encoding / decoding mode.
[0191] Subpixel convolutional units can be configured based on encoding / decoding modes.
[0192] Subpixel convolutional units can be controlled by one or two upsampling ratios: • Subpixel conv(s1, s2), where the upsampling ratios are s1 and s2.
[0193] • Subpixel conv(s), where the upsampling ratio is s.
[0194] • The upsampling ratio of subpixel convolution can be controlled according to the encoding / decoding mode.
[0195] • The output channel number of the convolutional layer of a subpixel convolutional unit can be controlled according to the encoding / decoding mode.
[0196] • According to some embodiments, the second synthetic transformation may include four parts: Part One o Upsampling unit (upsampler) o front, o Subpixel convolution part, o An exemplary implementation can be FIG. 21 A, FIG. 21 B and FIG. 21 As shown in C.
[0197] The upsampling unit and subpixel convolution part can be controlled based on the encoding / decoding mode.
[0198] • Signaling The o indicator can be included in the bitstream to indicate the encoding / decoding mode.
[0199] The 'o' indicator can be included in the bitstream to indicate the image format. Control of the second synthesis transformation can be based on such an indicator.
[0200] The o indicator can be included in the bitstream to indicate the image color encoding / decoding format. Control of the second synthesis transform can be based on such an indicator.
[0201] • The upsampling ratio utilized by the upsampling unit can be 2 times.
[0202] 4.3 Explanation and Advantages of the Proposed Solution This disclosure provides a solution for supporting different encoding / decoding modes with very limited changes to the synthesis transform.
[0203] Supporting different codec modes is important because different modes are suitable for different applications. The most typical image / video codec applications require the 4:2:0 codec mode because it is the most efficient in terms of compression. In other words, when the codec mode is 4:2:0, at the same bit rate (using the same number of bits), the image quality is optimal in terms of human visual perception. However, the 4:2:0 codec mode is not always sufficient for every application.
[0204] For example, when the preservation of the original material is important, it is necessary to use a 4:4:4 or 4:2:2 encoding / decoding mode instead of a 4:2:0 encoding / decoding mode. Similarly, if the image to be encoded / decoded is a medical image or has artistic intent, it is necessary to use a 4:4:4 or 4:2:2 encoding / decoding mode.
[0205] Supporting different codec modes typically requires substantial changes to the codecs (decoders and / or encoders). This disclosure provides a solution in which different codec modes can be supported with minimal modifications. More specifically, the proposed solution does not change the input to the second transform. As a result, the entropy decoding pipeline required to generate the input to the second synthetic transform remains unchanged. Furthermore, the first synthetic transform responsible for decoding the first component of the image also remains unchanged.
[0206] Finally, the modifications in the second synthetic transform are very limited. As a result, due to the proposed solution, different encoding and decoding modes can be adapted to different applications and supported with limited modifications.
[0207] Further details of embodiments of this disclosure relating to neural network-based visual data encoding and decoding will now be described. As used herein, the term “visual data” can refer to video, images, pictures in video, or any other visual data suitable for encoding and decoding.
[0208] As discussed above, in existing designs for visual data encoding and decoding based on neural networks (NNs), the output format of the entire synthesis transform is always 4:2:0; that is, it is fixed. When the output format during the entire decoding process is expected to be different from 4:2:0, filters (such as bicubic filters) must be applied to the output of the entire synthesis transform. Therefore, this traditional solution cannot support different output formats for the entire synthesis transform and thus lacks flexibility.
[0209] To address the above-mentioned problems and other issues not mentioned, a visual data processing solution is disclosed as described below. The embodiments of this disclosure should be considered as examples illustrating general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0210] FIG. 22 A flowchart of a method 2100 for visual data processing according to some embodiments of the present disclosure is shown. Method 2100 may be implemented during the conversion between visual data and a bitstream of visual data using a neural network (NN)-based model. As used herein, the NN-based model can be a model based on neural network techniques. For example, the NN-based model may specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module may include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the NN-based model described herein are merely illustrative and should therefore not be construed as limiting the present disclosure in any way.
[0211] like FIG. 22 As shown, method 2100 begins at 2102, where the format for encoding and decoding the visual data is acquired. As used herein, the format for encoding and decoding the visual data can refer to the format of the output of the entire synthetic transformation (which may include one or more synthetic transformations) in a neural network-based model. This format for encoding and decoding the visual data can also be referred to as a “first format,” “encoding / decoding format,” “encoded / decoded format,” “encoded / decoded image format,” “encoding / decoding mode,” and / or similar terms. The input to the analysis transformation on the encoder side can also be in this format. It should be noted that this format is allowed to differ from the format of the output of the transformation (i.e., the output of the entire decoding process). The format of the output of the entire decoding process can also be referred to as a “second format,” “output format,” “output image format,” and / or similar terms. It should also be noted that the term “format” can also be referred to as a “color format” or similar terms. In some embodiments, the first format can be indicated by at least one indicator in the bitstream. In this case, the first format can be acquired by parsing the bitstream. Additionally or alternatively, the second format can be indicated by at least one indicator in the bitstream. Furthermore, the first format may be determined based on the second format or any other suitable codec information for the visual data. In some embodiments, the first format may be the same as the second format. Additionally or alternatively, the first and second formats may be indicated by the same indications in the bitstream(s).
[0212] The format used for encoding and decoding visual data indicates the relationship between the size of a first component and the size of a second component of the encoded and decoded visual data. For example, the format for encoding and decoding visual data may indicate the ratio between the vertical dimension (e.g., height) of the first component and the vertical dimension (e.g., height) of the second component, and / or the ratio between the horizontal dimension (e.g., width) of the first component and the horizontal dimension (e.g., width) of the second component. Similarly, the format of the output of the entire decoding process may indicate the relationship between the size of the first component and the size of the second component of the output of the entire decoding process. For example, the format of the output of the entire decoding process may indicate the ratio between the vertical dimension (e.g., height) of the first component and the vertical dimension (e.g., height) of the second component, and / or the ratio between the horizontal dimension (e.g., width) of the first component and the horizontal dimension (e.g., width) of the second component.
[0213] In some embodiments, the format used for encoding and decoding visual data may be 4:4:4, 4:2:0, 4:2:2, or similar. Additionally, the output format of the entire decoding process may be 4:4:4, 4:2:0, 4:2:2, or similar. For example, in a 4:4:4 format, the vertical dimension of the second component may be the same as the vertical dimension of the first component, and the horizontal dimension of the second component may be the same as the horizontal dimension of the first component. In a 4:2:2 format, the vertical dimension of the second component may be half the vertical dimension of the first component, and the horizontal dimension of the second component may be the same as the horizontal dimension of the first component. Alternatively, the vertical dimension of the second component may be the same as the vertical dimension of the first component, and the horizontal dimension of the second component may be half the horizontal dimension of the first component. In a 4:2:0 format, the vertical dimension of the second component may be half the vertical dimension of the first component, and the horizontal dimension of the second component may be half the horizontal dimension of the first component. In some embodiments, the first component may include one of the following: a principal component, a luminance component, or a Y component, and the second component may include one of the following: a principal component, a chromaticity component, a U component, or a V component. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0214] It can be seen that the ratio between the vertical dimension of the first component and the vertical dimension of the second component can be 1 or 2, and the ratio between the horizontal dimension of the first component and the horizontal dimension of the second component can be 1 or 2. It should be understood that the specific values listed herein are intended to be exemplary and not to limit the scope of this disclosure.
[0215] In a neural network-based model, a first synthesis transform can be used for a first component, and a second synthesis transform can be used for a second component. For example, the first and second synthesis transforms can also be considered as two parts of a single synthesis transform. In this case, the outputs of the first and second synthesis transforms can constitute the output of the entire synthesis transform. In some embodiments, the first synthesis transform may include lightweight residual blocks (LRBs), convolutional layers, pruning layers, residual activation units (ResAUs), shuffling layers, and / or similar elements.
[0216] At 2104, a second synthesis transform used for the second component in the NN-based model is determined based on the format used for encoding and decoding the visual data (hereinafter referred to as the "first format"). In some embodiments, one or more parameters of the second synthesis transform may be configured based on the first format. This will be described in detail below. Additionally or alternatively, the structure of the second synthesis transform may be configured based on the first format. For example, and not limitingly, the structure of the second synthesis transform may be configured differently for different formats.
[0217] In some embodiments, the first composition transformation may be the same for different formats. That is, the first composition transformation may be independent of the first format. Alternatively, the first composition transformation may also be determined based on the first format. The scope of this disclosure is not limited in this respect.
[0218] At 2106, the conversion is performed based on the first and second synthesis transforms. In some embodiments, the conversion may include encoding visual data into a bitstream. Additionally or alternatively, the conversion may include decoding visual data from the bitstream. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0219] In light of the above, the second synthesis transform for the second component of the visual data is determined based on the format used for encoding and decoding the visual data. Compared to conventional solutions with fixed encoding and decoding formats, the proposed solution advantageously supports different encoding and decoding formats to adapt to different applications. In this way, encoding and decoding flexibility can be improved, and therefore encoding and decoding efficiency can be improved.
[0220] In some embodiments, the second synthesis transform may include an upsampling subnetwork (also referred to as an "upsampler unit"). At least one parameter of the upsampling subnetwork may be configured based on a first format. For example, at least one parameter of the upsampling subnetwork may be configured differently for different formats. In one exemplary embodiment, the upsampling subnetwork may include a shuffling layer, and at least one parameter may include at least one scaling factor of the shuffling layer. For example, and not limitingly, the vertical scaling factor of the shuffling layer may be determined based on the ratio between the vertical dimensions of the first component and the vertical dimensions of the second component. Additionally or alternatively, the horizontal scaling factor of the shuffling layer may be determined based on the ratio between the horizontal dimensions of the first component and the horizontal dimensions of the second component. In some additional or alternative exemplary embodiments, the upsampling subnetwork may include a deconvolution layer, a transposed convolution layer, a nearest-nearest-samples upsampling layer, and / or the like.
[0221] For example, rather than being restrictive, if the first format is 4:4:4, the upsampling subnetwork can apply upsampling operations in both the horizontal and vertical directions. If the first format is 4:2:2, the upsampling subnetwork can apply upsampling operations in only one of the horizontal and vertical directions. If the first format is 4:2:0, the upsampling subnetwork does not apply upsampling operations in either the horizontal or vertical directions.
[0222] In some embodiments, the second synthesis transform may further include a first subnetwork (also referred to as the "first part") preceding the upsampling subnetwork. For example, the first subnetwork may include a latent value combination block (LCB). Additionally, the first subnetwork may also include convolutional layers, shuffling layers, pruning layers, residual activation units (ResAU), and / or the like. For example, and not as a limitation, at least one parameter of the first subnetwork may be configured based on a first format.
[0223] Additionally or alternatively, the second synthesis transform may further include a second subnetwork (also referred to as a "third part") following the upsampled subnetwork. For example, the second subnetwork may include a pruning layer. Additionally, the second subnetwork may also include convolutional layers, shuffling layers, pruning layers, residual activation units (ResAU), and / or the like. For example, the second synthesis transform may further include subpixel convolutional units preceding the upsampled subnetwork. For example, subpixel convolutional units may include convolutional layers and shuffling layers. For example, and not limitingly, at least one parameter of the second subnetwork may be configured based on a first format. For example, the parameters of the pruning layer may depend on the first format.
[0224] In some embodiments, the input to the first synthesis transform may be different from the input to the second synthesis transform. For example, and not limited to, the input to the second synthesis transform may include the input to the first synthesis transform and another input. For example, the input to the second synthesis transform can be obtained by concatenating the input to the first synthesis transform and another input. In some embodiments, the vertical dimension of the input to the first synthesis transform may be the same as the vertical dimension of the input to the second synthesis transform, and the horizontal dimension of the input to the first synthesis transform may be the same as the horizontal dimension of the input to the second synthesis transform. Additionally, the number of channels of the input to the first synthesis transform may be different from the number of channels of the input to the second synthesis transform. Alternatively, the vertical dimension of the input to the first synthesis transform may be different from the vertical dimension of the input to the second synthesis transform, and / or the horizontal dimension of the input to the first synthesis transform may be different from the horizontal dimension of the input to the second synthesis transform. In some alternative embodiments, the input to the first synthesis transform may be the same as the input to the second synthesis transform.
[0225] In some embodiments, the output of the second composite transform can be the chromaticity component of the reconstructed visual data. Alternatively, the output of the first composite transform can be the luminance component of the reconstructed visual data. In this case, the output of the entire composite transform, including the first and second composite transforms, can be the reconstructed visual data.
[0226] In some embodiments, the size of the output of the second synthesis transform may be determined based on the first format. For example, the size of the output of the second synthesis transform may include the width and / or height of the output of the second synthesis transform. In one exemplary embodiment, the size of the output of the second synthesis transform may be equal to the size of the output of the first synthesis transform. In another exemplary embodiment, the size of the output of the second synthesis transform may be half the size of the output of the first synthesis transform.
[0227] For example, rather than being restrictive, if the first format is 4:4:4, then the width of the output of the second composition transform can be equal to the width of the output of the first composition transform, and the height of the second composition transform can be equal to the height of the output of the first composition transform. If the first format is 4:2:2, then the width of the output of the second composition transform can be equal to the width of the output of the first composition transform, and the height of the second composition transform can be half the height of the output of the first composition transform. If the first format is 4:2:0, then the width of the output of the second composition transform can be half the width of the output of the first composition transform, and the height of the second composition transform can be half the height of the output of the first composition transform.
[0228] In some embodiments, the size of the output of the second synthetic transform can be determined based on an indication obtained from the bitstream (such as at least one indication of the first format).
[0229] In view of the above, the solutions according to some embodiments of this disclosure can advantageously support different codec formats to adapt to different applications. In this way, codec flexibility can be improved, and therefore codec efficiency can be improved.
[0230] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data processing. The method includes: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; and generating a bitstream based on the first synthesis transform and the second synthesis transform.
[0231] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. The method includes: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; generating a bitstream based on the first and second synthesis transforms; and storing the bitstream in a non-transitory computer-readable recording medium.
[0232] Implementations of this disclosure may be described according to the following entries, and its features may be combined in any reasonable manner.
[0233] Item 1. A method for visual data processing, comprising: for a conversion between visual data and a bitstream of the visual data using a neural network (NN)-based model, obtaining a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in the NN-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; and performing the conversion based on the first synthesis transform and the second synthesis transform.
[0234] Item 2. The method according to Item 1, wherein the first synthetic transformation is the same for different formats.
[0235] Item 3. The method according to any one of Items 1-2, wherein the second synthetic transform includes an upsampled subnetwork, and at least one parameter of the upsampled subnetwork is configured based on the format.
[0236] Item 4. The method according to Item 3, wherein the at least one parameter of the upsampling subnetwork is configured differently for different formats.
[0237] Item 5. The method according to any one of items 3-4, wherein the upsampling subnetwork includes a shuffling layer.
[0238] Item 6. The method according to Item 5, wherein the at least one parameter includes at least one scaling factor of the washing layer.
[0239] Item 7. The method according to any one of items 3-6, wherein the second synthetic transformation further includes a first subnetwork preceding the upsampled subnetwork, and the first subnetwork includes a latent value combination block (LCB).
[0240] Item 8. The method according to any one of items 3-7, wherein the second synthesis transform further comprises a second subnetwork following the upsampled subnetwork, and the second subnetwork comprises a clipping layer.
[0241] Item 9. The method according to Item 8, wherein at least one parameter of the second sub-network is configured based on the format.
[0242] Item 10. The method according to any one of items 1-9, wherein the input of the first synthesis transform is different from the input of the second synthesis transform.
[0243] Item 11. The method according to Item 10, wherein the vertical dimension of the input of the first synthetic transform is the same as the vertical dimension of the input of the second synthetic transform, and the horizontal dimension of the input of the first synthetic transform is the same as the horizontal dimension of the input of the second synthetic transform.
[0244] Item 12. The method according to any one of items 1-11, wherein the output of the second synthetic transformation is the chromaticity component of the reconstructed visual data.
[0245] Item 13. The method according to any one of items 1-12, wherein the size of the output of the second synthetic transformation is determined based on the format.
[0246] Item 14. The method according to Item 13, wherein the size of the output of the second synthetic transformation includes at least one of the width or height of the output of the second synthetic transformation.
[0247] Item 15. The method according to any one of items 13-14, wherein the size of the output of the second synthetic transformation is equal to the size of the output of the first synthetic transformation.
[0248] Item 16. The method according to any one of items 13-14, wherein the size of the output of the second synthetic transform is equal to half the size of the output of the first synthetic transform.
[0249] Item 17. The method according to Item 14, wherein if the format is 4:4:4, the width of the output of the second synthesized transform is equal to the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to the height of the output of the first synthesized transform; or if the format is 4:2:2, the width of the output of the second synthesized transform is equal to the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to half the height of the output of the first synthesized transform; or if the format is 4:2:0, the width of the output of the second synthesized transform is equal to half the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to half the height of the output of the first synthesized transform.
[0250] Item 18. The method according to Items 13-17, wherein the size of the output of the second synthetic transformation is determined based on an indication obtained from the bitstream.
[0251] Item 19. The method according to any one of items 1-18, wherein the format is indicated by at least one indication in the bit stream.
[0252] Item 20. The method according to any one of items 3-19, wherein the ratio between the vertical dimension of the first component and the vertical dimension of the second component is permitted to be 2, and the ratio between the horizontal dimension of the first component and the horizontal dimension of the second component is permitted to be 2.
[0253] Item 21. The method according to any one of items 1-20, wherein the first synthesis transformation comprises at least one of the following: a lightweight residual block (LRB), a convolutional layer, a pruning layer, a residual activation unit (ResAU), and a shuffling layer.
[0254] Item 22. The method according to any one of items 1-21, wherein the first component comprises one of the following: a principal component, a luminance component, or a Y component, and the second component comprises one of the following: a principal component, a chromaticity component, a U component, or a V component.
[0255] Item 23. The method according to any one of items 1-22, wherein the format is permitted to be one of the following: 4:4:4 format, 4:2:0 format, or 4:2:2 format.
[0256] Item 24. The method according to any one of items 3-23, wherein if the format is 4:4:4, the upsampling subnetwork applies upsampling operations in both the horizontal and vertical directions; or if the format is 4:2:2, the upsampling subnetwork applies upsampling operations in only one of the horizontal and vertical directions; or if the format is 4:2:0, the upsampling subnetwork does not apply upsampling operations in either the horizontal or vertical directions.
[0257] Item 25. The method according to any one of items 1-24, wherein the structure of the second synthetic transformation is configured differently for different formats.
[0258] Item 26. The method according to any one of items 1-9, wherein the input of the first synthetic transformation is the same as the input of the second synthetic transformation.
[0259] Item 27. The method according to any one of Items 3-4, wherein the upsampling subnetwork comprises at least one of the following: a deconvolution layer, a transposed convolution layer, or a nearest-nearest-samples upsampling layer.
[0260] Item 28. The method according to any one of items 3-6, wherein the second synthetic transformation further comprises a subpixel convolutional unit preceding the upsampling subnetwork.
[0261] Item 29. The method according to Item 28, wherein the subpixel convolutional unit comprises a convolutional layer and a shuffling layer.
[0262] Item 30. The method according to any one of items 1-29, wherein the visual data includes video, a picture of the video, or an image.
[0263] Item 31. The method according to any one of items 1-30, wherein the conversion includes encoding the visual data into the bitstream.
[0264] Item 32. The method according to any one of items 1-30, wherein the conversion includes decoding the visual data from the bitstream.
[0265] Item 33. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-32.
[0266] Item 34. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-32.
[0267] Item 35. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of visual data processing, wherein the method comprises: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; and generating the bitstream based on the first synthesis transform and the second synthesis transform.
[0268] Item 36. A method for storing a bitstream of visual data, comprising: acquiring a format for encoding and decoding the visual data, the format indicating a relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transform in a neural network (NN)-based model being used for the first component; determining a second synthesis transform in the NN-based model being used for the second component based on the format; generating the bitstream based on the first synthesis transform and the second synthesis transform; and storing the bitstream in a non-transitory computer-readable recording medium.
[0269] Example device FIG. 22 A block diagram of a computing device 2200 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2200 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).
[0270] It should be understood that, FIG. 22 The computing device 2200 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0271] like As shown, computing device 2200 includes general-purpose computing device 2200. Computing device 2200 may include at least one or more processors or processing units 2210, memory 2220, storage unit 2230, one or more communication units 2240, one or more input devices 2250, and one or more output devices 2260.
[0272] In some embodiments, the computing device 2200 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2200 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0273] Processing unit 2210 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 2200. Processing unit 2210 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0274] Computing device 2200 typically includes various computer storage media. Such media can be any media accessible by computing device 2200, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2230 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 2200.
[0275] The computing device 2200 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to the bus (not shown) via one or more visual data media interfaces.
[0276] Communication unit 2240 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 2200 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 2200 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0277] Input device 2250 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2260 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2240, computing device 2200 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2200 can also communicate with one or more devices that enable a user to interact with computing device 2200, or, if needed, with any device that enables computing device 2200 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).
[0278] In some embodiments, some or all components of computing device 2200 may be deployed in a cloud computing architecture rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote visual data center locations. Cloud computing infrastructure may provide services through shared visual data centers, although they appear as a single access point to users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.
[0279] In embodiments of this disclosure, computing device 2200 may be used to implement visual data encoding / decoding. Memory 2220 may include one or more visual data encoding / decoding modules 2225 having one or more program instructions. These modules are accessible and executable by processing unit 2210 to perform the functions of the various embodiments described herein.
[0280] In an exemplary embodiment of performing visual data encoding, input device 2250 may receive visual data as input 2270 to be encoded. The visual data may be processed, for example, by visual data encoding / decoding module 2225 to generate an encoded bitstream. The encoded bitstream may be provided as output 2280 via output device 2260.
[0281] In an exemplary embodiment of performing visual data decoding, input device 2250 may receive an encoded bitstream as input 2270. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 2225 to generate decoded visual data. The decoded visual data may be provided as output 2280 via output device 2260.
[0282] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for visual data processing, comprising: For the conversion between visual data and the bitstream of the visual data using a neural network (NN)-based model, a format for encoding and decoding the visual data is obtained, the format indicating the relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transformation in the NN-based model is used for the first component. Based on the format, the second synthetic transformation used for the second component in the NN-based model is determined; as well as The transformation is performed based on the first and second synthetic transformations.
2. The method of claim 1, wherein the first synthesis transformation is the same for different formats.
3. The method according to any one of claims 1-2, wherein the second synthesis transform comprises an upsampling subnetwork, and at least one parameter of the upsampling subnetwork is configured based on the format.
4. The method of claim 3, wherein the at least one parameter of the upsampling subnetwork is configured differently for different formats.
5. The method according to any one of claims 3-4, wherein the upsampling subnetwork includes a shuffling layer.
6. The method of claim 5, wherein the at least one parameter includes at least one scaling factor of the washing layer.
7. The method according to any one of claims 3-6, wherein the second synthetic transformation further comprises a first subnetwork preceding the upsampled subnetwork, and the first subnetwork comprises a latent value combination block (LCB).
8. The method according to any one of claims 3-7, wherein the second synthesis transform further comprises a second subnetwork following the upsampling subnetwork, and the second subnetwork comprises a clipping layer.
9. The method of claim 8, wherein at least one parameter of the second sub-network is configured based on the format.
10. The method according to any one of claims 1-9, wherein the input of the first synthesis transform is different from the input of the second synthesis transform.
11. The method of claim 10, wherein the vertical dimension of the input of the first synthetic transform is the same as the vertical dimension of the input of the second synthetic transform, and the horizontal dimension of the input of the first synthetic transform is the same as the horizontal dimension of the input of the second synthetic transform.
12. The method according to any one of claims 1-11, wherein the output of the second synthetic transformation is the chromaticity component of the reconstructed visual data.
13. The method according to any one of claims 1-12, wherein the size of the output of the second synthetic transform is determined based on the format.
14. The method of claim 13, wherein the size of the output of the second synthetic transformation includes at least one of the width or height of the output of the second synthetic transformation.
15. The method according to any one of claims 13-14, wherein the size of the output of the second synthetic transform is equal to the size of the output of the first synthetic transform.
16. The method according to any one of claims 13-14, wherein the size of the output of the second synthetic transform is equal to half the size of the output of the first synthetic transform.
17. The method of claim 14, wherein if the format is 4:4:4, the width of the output of the second synthesized transform is equal to the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to the height of the output of the first synthesized transform, or If the format is 4:2:2, then the width of the output of the second synthesized transform is equal to the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to half the height of the output of the first synthesized transform, or If the format is 4:2:0, then the width of the output of the second synthesized transform is equal to half the width of the output of the first synthesized transform, and the height of the second synthesized transform is equal to half the height of the output of the first synthesized transform.
18. The method according to claims 13-17, wherein the size of the output of the second synthetic transform is determined based on an indication obtained from the bitstream.
19. The method according to any one of claims 1-18, wherein the format is indicated by at least one indication in the bit stream.
20. The method according to any one of claims 3-19, wherein the ratio between the vertical dimension of the first component and the vertical dimension of the second component is permitted to be 2, and the ratio between the horizontal dimension of the first component and the horizontal dimension of the second component is permitted to be 2.
21. The method according to any one of claims 1-20, wherein the first synthetic transformation comprises at least one of the following: Lightweight residual block (LRB). Convolutional layer Cut-out layer, Residual Activation Unit (ResAU), or Mixed washing layer.
22. The method according to any one of claims 1-21, wherein the first component comprises one of: a principal component, a luminance component, or a Y component, and The second component includes one of the following: principal component, chromaticity component, U component, or V component.
23. The method according to any one of claims 1-22, wherein the format is permitted to be one of the following: 4:4:4 format, 4:2:0 format, or 4:2:2 format.
24. The method according to any one of claims 3-23, wherein if the format is 4:4:4, the upsampling subnetwork applies an upsampling operation in both the horizontal and vertical directions, or If the format is 4:2:2, then the upsampling subnetwork applies the upsampling operation only in one of the horizontal and vertical directions, or If the format is 4:2:0, then the upsampling subnetwork does not apply upsampling operations in the horizontal and vertical directions.
25. The method according to any one of claims 1-24, wherein the structure of the second synthetic transformation is configured differently for different formats.
26. The method according to any one of claims 1-9, wherein the input of the first synthesis transform is the same as the input of the second synthesis transform.
27. The method according to any one of claims 3-4, wherein the upsampling subnetwork comprises at least one of the following: Deconvolution layer Transposed convolutional layer, or The most recent sample point is upsampled in the layer.
28. The method according to any one of claims 3-6, wherein the second synthesis transform further comprises a subpixel convolutional unit preceding the upsampling subnetwork.
29. The method of claim 28, wherein the subpixel convolutional unit comprises a convolutional layer and a shuffling layer.
30. The method according to any one of claims 1-29, wherein the visual data includes video, a picture of the video, or an image.
31. The method according to any one of claims 1-30, wherein the conversion comprises encoding the visual data into the bitstream.
32. The method according to any one of claims 1-30, wherein the conversion comprises decoding the visual data from the bitstream.
33. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-32.
34. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-32.
35. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of visual data processing, wherein the method includes: A format for encoding and decoding the visual data is obtained, the format indicating the relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transformation in a neural network (NN)-based model is used for the first component; Based on the format, the second synthetic transformation used for the second component in the NN-based model is determined; as well as The bit stream is generated based on the first synthesis transform and the second synthesis transform.
36. A method for storing a bitstream of visual data, comprising: A format for encoding and decoding the visual data is obtained, the format indicating the relationship between the size of a first component of the encoded and decoded visual data and the size of a second component of the encoded and decoded visual data, and a first synthesis transformation in a neural network (NN)-based model is used for the first component; Based on the format, the second synthetic transformation used for the second component in the NN-based model is determined; The bit stream is generated based on the first synthesis transform and the second synthesis transform; as well as The bitstream is stored in a non-transitory computer-readable recording medium.