Neural network-based image and video compression method using integer operation

Through the neural network technology that combines integer value convolution and activation functions with hyper-priority information and context model, the problem of insufficient compression efficiency and reconstruction quality in existing video encoding and decoding technologies is solved, and more efficient video compression and reconstruction are achieved.

CN120513624APending Publication Date: 2025-08-19DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090066.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-28
Filing Date
2023-12-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing neural network-based video encoding and decoding technology has shortcomings in compression efficiency and reconstruction quality, especially when processing video data, it is difficult to effectively utilize computing resources and data characteristics, resulting in poor performance.

Method used

Integer values are used for convolution operations and activation functions, image and video compression is performed through integer weights and biases, data conversion is performed using bit shifting and quantization techniques, avoiding non-integer operations, and optimizing the encoding and decoding process with hyper-priority information and context model.

Benefits of technology

It improves the efficiency and reconstruction quality of video compression, reduces the demand for computing resources, improves the control ability of compressed bit rates, and adapts to different bit rate scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513624A_ABST
    Figure CN120513624A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. It is determined to apply a neural sub-network to perform a convolution operation using only integer values. The neural sub-network also performs an activation function on the output of the convolution operation using only integer values. A conversion between the visual media data and the bitstream is performed based on the convolution operation and the activation function.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefits of International Patent Application No. PCT / CN2022 / 142668, filed on December 28, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the processing of digital images and videos. Background Art

[0004] Digital video consumes the largest amount of bandwidth used for the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is likely to continue to grow. Summary of the Invention

[0005] A first aspect relates to a method for processing video or image data, comprising: applying a neural network to perform a convolution operation on an input and an activation function on an output of the convolution operation, wherein the convolution operation and the activation function are performed using only integer values; and performing conversion between visual media data and a bitstream based on the convolution operation and the activation function.

[0006] Optionally, in any of the above aspects, another embodiment of this aspect provides that the input includes only integer values.

[0007] Optionally, in any of the above aspects, another embodiment of this aspect provides a convolution operation using a convolution kernel with integer weights and an integer bias, wherein the integer weights are element-wise multiplied with the input, and wherein the integer bias includes all integer values.

[0008] Optionally, in any of the above aspects, another embodiment of this aspect provides that the size of the convolution kernel is 3x3.

[0009] Optionally, in any of the above aspects, another embodiment of this aspect provides that the method is applicable to any convolution operation regardless of the dimension, size of the convolution kernel and integer bias.

[0010] Optionally, in any of the above aspects, another embodiment of this aspect provides that the activation function adopts bit shifting.

[0011] Optionally, in any of the above aspects, another embodiment of this aspect provides that non-integer values are not allowed to be used for convolution operations and activation functions.

[0012] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation prohibits non-integer values by employing only multiplications and additions using integer inputs.

[0013] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation is performed using only integer-valued multiplications and integer-valued additions.

[0014] Optionally, in any of the above aspects, another embodiment of this aspect provides that the input only includes integer values to ensure that the output of the neural network only includes integer values.

[0015] Optionally, in any of the above aspects, another embodiment of the aspect provides that the convolution operation is a two-dimensional convolution operation and is performed according to the following formula:

[0016]

[0017] Optionally, in any of the above aspects, another embodiment of this aspect provides that the activation function includes a rectified linear unit (ReLU) or a quantized leaky ReLU.

[0018] Optionally, in any of the above aspects, another embodiment of this aspect provides that the activation function does not use any division operation, any rounding operation or any multiplication operation involving non-integer numbers.

[0019] Optionally, in any of the above aspects, another embodiment of this aspect provides that the activation function is performed according to one or a combination of the following:

[0020] or

[0021] or

[0022] or

[0023] or

[0024]

[0025] The bitshift(x,n) operation corresponds to the bitwise shift operator.

[0026] Optionally, in any of the above aspects, another embodiment of this aspect provides using a bitwise shift operator to implement a division operation followed by a floor() operation.

[0027] Optionally, in any of the above aspects, another embodiment of this aspect provides that the division operation is replaced by a series of operations using at least one lookup table.

[0028] Optionally, in any of the above aspects, another embodiment of the aspect provides that the activation function is performed according to the following formula:

[0029] or

[0030]

[0031] Optionally, in any of the above aspects, another embodiment of this aspect provides that the constants A, B and C are integers that approximate a conventional leaky ReLU function.

[0032] Optionally, in any of the above aspects, another embodiment of this aspect provides that the constants A, B and C are set to approximate the negative slope of the leaky ReLU function.

[0033] Optionally, in any of the above aspects, another embodiment of this aspect provides that the constants A, B and C are set to an arbitrary negative slope that approximates the leaky ReLU function, where the arbitrary negative slope is a floating point number between 0 and 1 or between 0 and -1.

[0034] Optionally, in any one of the above aspects, another embodiment of this aspect provides that when the bit shift is a left shift operation, the value of the constant B is zero, and when the bit shift is a right shift operation, the value of the constant B is one.

[0035] Optionally, in any of the above aspects, another embodiment of the aspect provides that the activation function is performed according to the following formula:

[0036]

[0037] Optionally, in any of the above aspects, another embodiment of the aspect provides that the input is bit-shifted by n bits in both the positive branch and the negative branch.

[0038] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation includes a convolution layer with integer weights and bias values, and the activation function includes a rectified linear unit (ReLU) function.

[0039] Optionally, in any of the above aspects, another embodiment of this aspect provides including an additional layer between the convolution operation and the activation function.

[0040] Optionally, in any of the above aspects, another embodiment of this aspect provides performing a bit shift operation after the convolution operation and before the activation function.

[0041] Optionally, in any of the above aspects, another embodiment of this aspect provides that when the input is an integer value, the negative slope of the activation function is equal to zero, and wherein the output is guaranteed to be an integer value.

[0042] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation includes one or more of integer multiplication, bit shift operation, integer addition and limiting.

[0043] Optionally, in any of the above aspects, another embodiment of this aspect provides that the method is performed without using any division, multiplication or addition with non-integer values, or rounding operations.

[0044] Optionally, in any of the above aspects, another embodiment of this aspect provides that one of the additional layers is a clipping layer.

[0045] Optionally, in any of the above aspects, another embodiment of the aspect provides that the additional layer does not include device-dependent operations using floating point, division or rounding.

[0046] Optionally, in any of the above aspects, another embodiment of this aspect provides that the neural network is a neural sub-network among multiple neural sub-networks, each neural sub-network comprising a convolution operation and an activation function.

[0047] Optionally, in any of the above aspects, another embodiment of this aspect provides performing a quantization operation on the output of the activation function, and wherein the quantization operation uses only integers.

[0048] Optionally, in any of the above aspects, another embodiment of this aspect provides that the integers used for the quantization operation are obtained from a table including quantization levels and quantized output values.

[0049] Optionally, in any of the above aspects, another embodiment of this aspect provides that the quantization operation includes a mapping function between integer inputs and integer outputs.

[0050] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation and activation function are incorporated into a variance decoder module that utilizes super-prior information.

[0051] Optionally, in any of the above aspects, another embodiment of this aspect provides that the output of the variance decoder module using super-prior information includes a probability parameter or a standard deviation.

[0052] Optionally, in any of the above aspects, another embodiment of this aspect provides that one or more of the probability parameters include a standard deviation.

[0053] Optionally, in any of the above aspects, another embodiment of this aspect provides that the output of the activation function is a reconstructed image.

[0054] Optionally, in any of the above aspects, another embodiment of this aspect provides that the convolution operation and activation function are essentially composed of integer-valued multiplication, bit shift operation, integer-valued addition, limiting and combinations thereof.

[0055] Optionally, in any of the above aspects, another embodiment of this aspect provides that the converting includes encoding the visual media data into a bitstream.

[0056] Optionally, in any of the above aspects, another embodiment of this aspect provides that converting includes decoding the visual media data from a bitstream.

[0057] A second aspect relates to an apparatus for processing media data, comprising: one or more processors; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processors, cause the apparatus to perform the method of any disclosed embodiment.

[0058] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any disclosed embodiment.

[0059] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes the method in any disclosed embodiment.

[0060] A fifth aspect relates to a method for storing a bitstream of a video, including the method in any disclosed embodiment.

[0061] A sixth aspect relates to a method, apparatus or system described in the present disclosure.

[0062] For purposes of clarity, any of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments within the scope of the present disclosure.

[0063] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0065] Figure 1 is a schematic diagram illustrating an example transform coding scheme.

[0066] Figure 2 Example latent representations of images are shown.

[0067] Figure 3 is a diagram illustrating an example autoencoder that implements a model that utilizes hyper-prior information.

[0068] Figure 4 is a diagram illustrating an example combined model configured to jointly optimize a context model with a hyper-prior and an autoencoder.

[0069] Figure 5 An example encoding process is shown.

[0070] Figure 6 An example decoding process is shown.

[0071] Figure 7 An example decoding process according to the present disclosure is shown.

[0072] Figure 8 An example learning-based image codec architecture is shown.

[0073] Figure 9 Example synthetic transforms for learning-based image encoding and decoding are shown.

[0074] Figure 10 An example leaky rectified linear unit (ReLU) activation function is shown.

[0075] Figure 11 An example leaky ReLU activation function is shown.

[0076] Figure 12 An example base unit according to the present disclosure is shown.

[0077] Figure 13 Another example base unit according to the present disclosure is shown.

[0078] Figure 14 Another example base unit according to the present disclosure is shown.

[0079] Figure 15 is a block diagram illustrating an example video processing system.

[0080] Figure 16 is a block diagram of an example video processing device.

[0081] Figure 17 is a flow chart of an example method for video processing.

[0082] Figure 18 is a block diagram illustrating an example video encoding and decoding system.

[0083] Figure 19 is a block diagram illustrating an example encoder.

[0084] Figure 20 is a block diagram illustrating an example decoder.

[0085] Figure 21 is a schematic diagram of an example encoder. DETAILED DESCRIPTION

[0086] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and all equivalents thereof.

[0087] 1. Preliminary Discussion

[0088] The present disclosure relates to neural network-based image and video compression methods using separate processing of color components of an image, where control parameters used to process one component are also used for the other component.

[0089] 2. Further Discussion

[0090] Deep learning is gaining momentum in various fields, such as computer vision and image processing. Inspired by the successful application of deep learning techniques in computer vision, neural image / video compression techniques are being investigated for image / video compression. Neural networks are designed based on interdisciplinary research in neuroscience and mathematics. Neural networks have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Exemplary neural network-based image compression algorithms achieve rate-distortion (RD) performance comparable to Versatile Video Codec (VVC), a video codec standard developed by the Joint Video Experts Group (JVET) with experts from the Moving Picture Experts Group (MPEG) and the Video Codec Experts Group (VCEG). Neural network-based video compression is an actively developing research area, resulting in continuous improvements in the performance of neural image compression. However, due to the inherent difficulty of the problems addressed by neural networks, neural network-based video coding and decoding remains a largely unexplored discipline.

[0091] 2.1 Image / Video Compression

[0092] Image / video compression generally refers to computing techniques that compress video images into binary codes for easier storage and transmission. This binary code may or may not support lossless reconstruction of the original image / video. Codecs that achieve no data loss are called lossless compression, while codecs that allow for targeted data loss are called lossy compression. Most codecs employ lossy compression because lossless reconstruction is not necessary in most scenarios. The performance of image / video compression algorithms is typically evaluated based on the resulting compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes generated by the compression; the fewer binary codes, the better the compression. Reconstruction quality is measured by comparing the reconstructed image / video to the original; the greater the similarity, the better the reconstruction quality.

[0093] Image / video compression technology can be divided into video coding methods and neural network-based video compression methods. Video coding schemes adopt transform-based solutions, in which statistical dependencies in latent variables (such as discrete cosine transform (DCT) and wavelet coefficients) are used to carefully hand-design entropy codes to model dependencies in the quantization domain. Neural network-based video compression can be divided into neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded in existing video codecs as a codec tool and is provided only as part of the framework, while the latter is a separate framework developed based on neural networks and does not depend on video codecs.

[0094] A range of video codec standards have been developed to meet the growing demand for visual content delivery. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups: the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG). The International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has the Video Codec Experts Group (VCEG), which is responsible for standardizing image / video codec technologies. Influential video codec standards released by these organizations include the Joint Photographic Experts Group (JPEG), JPEG 2000, H.262, H.264 / Advanced Video Codec (AVC), and H.265 / High Efficiency Video Codec (HEVC). The Joint Video Experts Group (JVET), composed of MPEG and VCEG, developed the Versatile Video Codec (VVC) standard. Compared to HEVC, VVC reduces bitrate by an average of 50% while maintaining the same visual quality.

[0095] Neural network-based image / video compression / encoding is also under development. Exemplary neural network codecs have relatively shallow architectures, and the performance of such networks is unsatisfactory. Neural network-based approaches benefit from abundant data and powerful computing resources, and therefore are better utilized in various applications. Neural network-based image / video compression has shown promising improvements and has been proven feasible. However, this technology is far from mature, and many challenges must be addressed.

[0096] 2.2 Neural Networks

[0097] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. Neural networks typically consist of multiple processing layers, each composed of multiple simple but nonlinear basic computational units. One benefit of such deep networks is their ability to process data at multiple levels of abstraction and transform that data into different kinds of representations. The representations created by neural networks are not manually designed. Instead, the deep networks, including their processing layers, are learned from vast amounts of data using a common machine learning process. Deep learning eliminates the need for handcrafted representations. Consequently, deep learning is considered particularly well-suited for processing raw, unstructured data, such as acoustic and visual signals. Processing such data has been a long-standing challenge in the field of artificial intelligence.

[0098] 2.3 Neural Networks for Image Compression

[0099] Neural networks used for image compression can be divided into two categories: pixel-based probabilistic models and autoencoder models. Pixel-based probabilistic models employ a predictive encoding / decoding strategy, while autoencoder models employ a transform-based solution. Sometimes, these two approaches are combined.

[0100] 2.3.1 Pixel Probabilistic Modeling

[0101] According to Shannon's information theory, the optimal method for lossless coding achieves the minimum coding rate, which is expressed as -log2p(x), where p(x) is the probability of symbol x. Arithmetic coding is a lossless coding method considered to be one of the optimal methods. Given a probability distribution p(x), arithmetic coding achieves a coding rate as close to the theoretical limit -log2p(x) as possible, without considering rounding errors. Therefore, the remaining problem is to determine the probability, which is very challenging for natural images and videos due to the curse of dimensionality. The curse of dimensionality refers to the problem that the increase in dimensionality causes the dataset to become sparse. As the number of dimensions increases, the amount of data required to effectively analyze and organize the data increases rapidly.

[0102] Following the predictive encoding and decoding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is an image, which can be expressed as follows:

[0103] p(x)=p(x1)p(x2|x1)…p(x i |x1,…,x i-1 )…p(x m×n |x1,…,x m×n-1 ) (1)

[0104] Where m and n are the height and width of the image, respectively. The previous observation is also called the context of the current pixel. When the image is large, the estimation of the conditional probability may be difficult. Therefore, a simplified approach is to limit the scope of the context of the current pixel as follows:

[0105] p(x)=p(x1)p(x2|x1)…p(x i |x i-k ,…,x i-1 )…p(x m×n |x m×n-k ,…,x m×n-1 ) (2)

[0106] Where k is a predefined constant that controls the scope of the context.

[0107] It should be noted that this condition can also take into account the sample values of other color components. For example, when encoding and decoding red (R), green (G), and blue (B) (RGB) color components, the R sample depends on the previously encoded and decoded pixels (including R, G, and / or B samples), and the current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample. In addition, when encoding and decoding the current B sample, the previously encoded and decoded pixels as well as the current R and G samples can also be considered.

[0108] Neural networks can be designed for computer vision tasks and can also be effective in regression and classification problems. Therefore, neural networks can be used to estimate the value of a given context x1, x2, ..., x i-1 p(x i ) probability.

[0109] Most methods directly model the probability distribution in the pixel domain. Some designs also model the probability distribution as a conditional probability distribution based on explicit or latent representations. Such a model can be expressed as:

[0110]

[0111] Where h is an additional condition, and p(x)=p(h)p(x|h) indicates that the modeling is divided into an unconditional model and a conditional model. The additional condition can be image label information or a high-level representation.

[0112] 2.3.2 Autoencoder

[0113] Now let's describe autoencoders. Autoencoders are trained for dimensionality reduction and consist of an encoding component and a decoding component. The encoding component converts a high-dimensional input signal into a low-dimensional representation. The low-dimensional representation may have a reduced spatial size but a greater number of channels. The decoding component recovers the high-dimensional input from the low-dimensional representation. The ability of autoencoders to automatically learn representations and eliminate the need for handcrafted features is considered one of the most important advantages of neural networks.

[0114] Figure 1 is a schematic diagram illustrating an example transform coding scheme 100. The original image x is analyzed by the analysis network g a The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the codec rate. Then the synthetic network g s Inverse transform to obtain the reconstructed image The distortion (D) is calculated in the perceptual space by using the function g p Transform x and Get z and Compare them to get D.

[0115] Autoencoder networks can be applied to lossy image compression. The learned latent representation can be encoded from a well-trained neural network. However, adapting autoencoders to image compression is not easy because the original autoencoder is not optimized for compression and is therefore inefficient to use directly as a trained autoencoder. In addition, there are other major challenges. First, the low-dimensional representation should be quantized before being encoded. However, quantization is non-differentiable, which is required in backpropagation while training the neural network. Second, the goals in the compression scenario are different because both distortion and bit rate need to be considered. Estimating the bit rate is challenging. Third, practical image coding and decoding schemes should support variable bitrate, scalability, encoding / decoding speed, and interoperability. Various schemes are under development to address these challenges.

[0116] The example autoencoder for image compression using the example transform coding scheme 100 can be viewed as a transform coding strategy. a (x) transforms the original image x, where y is the latent representation to be quantized and decoded. The synthesis network transforms the quantized latent representation Inverse transform to get the reconstructed image Utilization-distortion loss function Training framework, where D is x and The distortion between them, R is expressed from the quantization The rate is calculated or estimated, and λ is the Lagrange multiplier. D can be calculated in the pixel domain or the receptive domain. Most example systems follow this prototype, and the difference between such systems may only be the network structure or loss function.

[0117] 2.3.3 Models using hyper-prior information

[0118] Figure 2 Example latent representations of images are shown. Figure 2 The image 201 from the Kodak dataset, the visualization of the potential value 202 representing y of the image 201, the standard deviation σ203 of the potential value 202, and the potential value y 204 after the introduction of the super-prior network. The super-prior network includes an encoder that uses super-prior information and a decoder that uses super-prior information. In the transform coding method of image compression, such as Figure 1 As shown, the encoder subnetwork uses parameter analysis transformation The image vector x is transformed into a latent representation y, which is then quantized to form because is a discrete value, so It can be losslessly compressed using entropy coding techniques such as arithmetic coding and transmitted as a bit sequence.

[0119] from Figure 2It can be clearly seen from the potential value 202 and standard deviation σ203 that There is a significant spatial dependence between the elements of . It is noteworthy that their scales (standard deviation σ203) appear to be coupled in the spatial domain. An additional set of random variables can be introduced to capture spatial dependencies and further reduce redundancy. In this case, the image compression network such as Figure 3 shown.

[0120] Figure 3 3 is a diagram illustrating an example network architecture for an autoencoder that implements a model utilizing hyper-prior information. The upper side illustrates the image autoencoder network, and the lower side corresponds to the hyper-prior subnetwork. The analysis and synthesis transforms are denoted as g a and g a Q represents quantization, and AE and AD represent arithmetic encoder and arithmetic decoder respectively. The model using super prior information includes two sub-networks: the encoder using super prior information (using h a denoted) and a decoder using super prior information (denoted by h s The model using super prior information generates a quantitative super prior information potential value. It includes the potential value of quantification Information related to the probability distribution of the sample points. is included in the bitstream and is are transmitted together to the receiver (decoder).

[0121] In the diagram 300, the upper side of the model is the encoder g as discussed above. a and decoder g s The lower side is used to obtain The additional encoder h using super prior information a and decoder h using super prior information s In this architecture, the encoder subjects the input image x to g a , producing a response y with a spatially varying standard deviation. The response y is fed to h a , summarizing the distribution of standard deviations in z. z is then quantified compressed and transmitted as side information. The encoder then uses the quantized vector To estimate the spatial distribution of the standard deviation σ, and the encoder uses σ to compress and transmit the quantized image representation The decoder first recovers the compressed signal The decoder then uses h s to obtain σ, which provides the correct probability estimate for the decoder to successfully recover The decoder will then Feed to gs to obtain the reconstructed image.

[0122] When an encoder utilizing super-prior information and a decoder utilizing super-prior information are added to an image compression network, the quantized latent value The spatial redundancy is reduced. Figure 2 The potential value y 204 in corresponds to the quantized potential value when using an encoder / decoder that utilizes super-prior information. Compared to the standard deviation σ 203 , the spatial redundancy is significantly reduced because the samples of the quantized potential value are less correlated.

[0123] 2.3.4 Context Model

[0124] Although the model using hyper-prior information improves the potential value of quantization , but additional improvement can be obtained by utilizing an autoregressive model that predicts the quantized latent value from the causal context of the quantized latent value, which can be called a context model.

[0125] The term autoregressive indicates that the output of a process is later used as input to that process. For example, the context model subnetwork generates one sample of potential values, which is later used as input to obtain the next sample.

[0126] Figure 4 4 is a diagram illustrating an example combined model configured to jointly optimize a context model and a hyper-prior and an autoencoder. The combined model jointly optimizes an autoregressive component that estimates a probability distribution of latent values from the underlying causal context (context model) and the hyper-prior and underlying autoencoder. The real-valued latent representation is quantized (Q) to create a quantized latent value and quantify the potential value of super prior information They are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD).The dotted area corresponds to the components executed by a receiver (eg, a decoder) to restore the image from the compressed bitstream.

[0127] The example system utilizes a joint architecture where both the model subnetworks utilizing super-prior information (encoder utilizing super-prior information and decoder utilizing super-prior information) and the context model subnetwork are utilized. The super-prior and context models are combined to learn the potential values for the quantization The probability model of the quantized potential value is then used for entropy coding and decoding. As shown in schematic diagram 400, the outputs of the context sub-network and the decoder sub-network that utilizes the super-prior information are combined by a sub-network called the entropy parameter, which generates the mean μ and scale (or variance) σ parameters of the Gaussian probability model. The Gaussian probability model is then used to encode the samples of the quantized potential value into a bitstream with the help of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is utilized to obtain the quantized potential value from the bitstream through the arithmetic decoder (AD) module.

[0128] In the example, the potential sample is modeled as a Gaussian distribution or a Gaussian mixture model (without limitation). In the example according to schematic diagram 400, the context model and the hyper-prior are used jointly to estimate the probability distribution of the potential sample. Since the Gaussian distribution can be defined by a mean and a variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ).

[0129] 2.3.5 Gain Variational Autoencoder (G-VAE)

[0130] In the example, neural network-based image / video compression methodologies require training multiple models to adapt to different bitrates. Gain Variational Autoencoder (G-VAE) is a variational autoencoder with a pair of gain units, which is designed to achieve continuous variable bitrate adaptation using a single model. It includes a pair of gain units that are usually inserted at the output of the encoder and the input of the decoder. The output of the encoder is defined as the latent representation y∈R c*h*w , where c, h, w represent the number of channels, height and width of the potential representation. Each channel of the potential representation is represented as y (i) ∈R h*w , where i = 0, 1, ..., c-1. A pair of gain units includes a gain matrix M∈R c*n and the inverse gain matrix, where n is the number of gain vectors. The gain vector can be expressed as m s ={α s(0) ,α s(1) ,…,α s(c-1)}, α s(i) ∈R, where s represents the index of the gain vector in the gain matrix.

[0131] The motivation of the gain matrix is similar to the quantization table in JPEG, which controls the quantization loss based on the characteristics of different channels.

[0132] To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.

[0133]

[0134] where ⊙ is the channel-wise multiplication, i.e. And α s(i) is the gain vector m s The inverse gain matrix used at the decoder side can be expressed as M′∈R c*n , which consists of n inverse gain vectors, namely M ′ ={δ s(0) ,δ s(1) ,…,δ s(c-1)}, δ s(i) ∈R. The inverse gain process is expressed as

[0135]

[0136] in is the decoded quantized latent representation, and y′ s is the quantized latent representation of the inverse gain, which will be fed into the synthesis network.

[0137] In order to achieve continuous variable bit rate adjustment, interpolation is used between the vectors. Given two pairs of gain vectors {m t ,m′ t} and {m r ,m′ r}, the interpolated gain vector can be obtained by the following equation.

[0138] m v =[(m r ) l ·(m t ) 1-l ]

[0139] m′ v =[(m′ r ) l ·(m′ t ) 1-l ]

[0140] where l∈R is the interpolation coefficient that controls the corresponding bit rate of the generated gain vector pair. Since l is a real number, any bit rate between a given pair of gain vectors can be achieved.

[0141] 2.3.6 Encoding Process of Models Using Joint Autoregression to Exploit Hyper-Prior Information

[0142] Figure 4 The design in corresponds to the example combined compression method. In this section and the next, the encoding and decoding processes are described respectively.

[0143] Figure 5An example encoding process 500 is shown. The input image is first processed by the encoder sub-network. The encoder converts the input image into a transformed representation (called a potential value), denoted by y. y is then input to the quantizer block, denoted by Q, to obtain a quantized potential value Then it is converted into a bit stream (bits1) using the arithmetic coding module (denoted as AE). The arithmetic coding block sequentially converts Each sample point is converted into a bit stream (bits1) one by one.

[0144] The encoder using super prior information, context, decoder using super prior information and entropy parameter sub-network are used to estimate the quantized potential value. The potential value y is input to the encoder using super prior information, which outputs the super prior information potential value (denoted by z). The super prior information potential value is then quantized The arithmetic encoding (AE) module is used to generate a second bit stream (bits2). The decomposition entropy module generates a probability distribution, which is used to encode the quantized super-prior information potential value into a bit stream. The quantized super-prior information potential value includes the potential value of the quantized information about the probability distribution of .

[0145] The entropy parameter subnetwork generates probability distribution estimates, which are used to encode the quantized latent values. The information generated by the entropy parameters usually includes the mean μ and scale (or variance) σ parameters, which together are used to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as The parameter μ is the mean or expectation of the distribution (as well as its median and mode), and the parameter σ is its standard deviation (or variance, or scale). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.

[0146] The decoder subnetwork using the hyper-prior information generates part of the information used by the entropy parameter subnetwork, and the other part of the information is generated by an autoregressive module called the context module. The context module uses the samples that have been encoded by the arithmetic coding (AE) module to generate information about the probability distribution of the samples of the quantized potential value. It is usually a matrix composed of many sample points. Sample points can be indicated by using indices, such as or It depends on the matrix Dimensions of . Sample points Encoded one by one by the AE, usually using raster scan order. In raster scan order, the rows of the matrix are processed from top to bottom, where the samples in the row are processed from left to right. In such a scenario (where raster scan order is used by the AE to encode the samples into the bitstream), the context module uses the samples that were previously encoded in raster scan order to generate the same samples as the ones in the row. The information generated by the context module and the decoder using the hyper-prior information is combined by the entropy parameter module to generate the potential value used to quantize Encoded as a probability distribution of the bitstream (bits1).

[0147] Finally, the first bit stream and the second bit stream are transmitted to a decoder as a result of the encoding process. Note that other names may be used for the above modules.

[0148] In the above description, Figure 5 All elements in are collectively referred to as encoders. The analysis transformation that transforms the input image into the latent representation is also called an encoder (or autoencoder).

[0149] 2.3.7 Decoding Process of Models Using Joint Autoregression and Hyper-Prior Information

[0150] Figure 6 An example decoding process 600 is shown. Figure 6 The decoding process is depicted separately.

[0151] During the decoding process, the decoder first receives the first bit stream (bits1) and the second bit stream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module by using the probability distribution generated by the decomposition entropy sub-network. The decomposition entropy module usually uses a predetermined template to generate the probability distribution, such as a predetermined mean and variance value in the case of a Gaussian distribution. The output of the arithmetic decoding process of bits2 is It is the quantized super prior information potential value. The AD process is restored to the AE process, and the AE process is applied in the encoder. The AE and AD processes are lossless, which means that the quantized super prior information potential value generated by the encoder is Can be reconstructed at the decoder without any changes.

[0152] In obtaining Afterwards, it is processed by the decoder using super prior information, and the output of the decoder using super prior information is fed to the entropy parameter module. The three sub-networks used in the decoder, context, decoder using super prior information, and entropy parameters are the same as those in the encoder. Therefore, exactly the same probability distribution can be obtained in the decoder (as in the encoder), which is necessary for lossless reconstruction of the quantized potential values. As a result, the quantized potential value obtained in the encoder The same version of can be obtained in the decoder.

[0153] After the probability distribution (e.g., mean and variance parameters) is obtained by the entropy parameter sub-network, the arithmetic decoding module decodes the quantized potential value samples one by one from the bitstream bits1. From a practical point of view, the autoregressive model (context model) is serial in nature and therefore cannot be accelerated using techniques such as parallelization. Finally, the fully reconstructed quantized potential value is input to the composite transform (in Figure 6 ) module to obtain the reconstructed image.

[0154] In the above description, Figure 6 All elements in are collectively referred to as decoders. The synthetic transform that converts the quantized latent values into the reconstructed image is also called a decoder (or autodecoder).

[0155] 2.4 Neural Networks for Video Compression

[0156] Similar to video codecs, neural image compression forms the basis for intra-frame compression in neural network-based video compression. Consequently, the development of neural network-based video compression has lagged behind that of neural network-based image compression due to its greater complexity and the associated challenges that require more effort to address. Compared to image compression, video compression requires effective methods to remove inter-frame redundancy. Inter-frame prediction is a primary step in these example systems. Motion estimation and compensation are widely used in video codecs, but are not typically implemented using trained neural networks.

[0157] Neural network-based video compression can be divided into two categories based on the target scenario: random access and low latency. In the case of random access, the system allows decoding to start from any point in the sequence, typically dividing the entire sequence into multiple separate segments and allowing each segment to be decoded independently. In the case of low latency, the system aims to reduce decoding time, whereby previous frames can be used as reference frames to decode subsequent frames.

[0158] 2.5 Preliminary Knowledge

[0159] Almost all natural images and / or videos are in digital format. Grayscale digital images can be obtained by Indicates that is a set of pixel values, m is the image height, and n is the image width. For example, is an example setup, and in this case Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed bits are definitely less.

[0160] Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by Representation, where three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24bpp. Digital images / videos can be represented in different color spaces. Neural network-based video compression schemes are mostly developed in the RGB color space, while video codecs typically use the YUV color space to represent video sequences. In the YUV color space, the image is decomposed into three channels, namely luminance (Y), blue difference chrominance (Cb), and red difference chrominance (Cr). Y is the luminance component, and Cb and Cr are the chrominance components. The compression benefits of YUV arise because Cb and Cr are usually downsampled to achieve pre-compression because the human visual system is not very sensitive to the chrominance components.

[0161] A color video sequence consists of multiple color images (also called frames) to record scenes at different time stamps. For example, in the RGB color space, a color video can be composed of X = {x0, x1, ..., x t ,…,x T-1}, where T is the number of frames in the video sequence, and If m=1080, n=1920, And the video has 50 frames per second (fps), then the data rate of the uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits per second (bps). This results in about 2.32 gigabits per second (Gbps), which uses a lot of storage and should be compressed before transmission over the Internet.

[0162] Typically, lossless methods can achieve a natural image compression ratio of approximately 1.5 to 3, which is clearly lower than streaming requirements. Therefore, lossy compression is employed to achieve a better compression ratio, but at the expense of distortion. Distortion can be measured by calculating the mean squared difference between the original and reconstructed images, for example using the mean squared error (MSE). For grayscale images, the MSE can be calculated using the following equation.

[0163]

[0164] Therefore, the quality of the reconstructed image compared to the original image can be measured by the Peak Signal-to-Noise Ratio (PSNR):

[0165]

[0166] in yes The maximum value in , for example, for an 8-bit grayscale image is 255. There are other quality assessment metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).

[0167] To compare different lossless compression schemes, the compression ratio for a given resulting bit rate can be compared, and vice versa. However, to compare different lossy compression methods, the comparison must take into account both bit rate and reconstruction quality. This can be achieved, for example, by calculating the relative bit rate at several different quality levels and then averaging the bit rates. The average relative rate is called Bjontegaard's delta-rate (BD-rate). There are other aspects to evaluate image and / or video codecs, including encoding / decoding complexity, scalability, robustness, etc.

[0168] 2.6 Separate Processing of Luminance and Chrominance Components of an Image

[0169] Figure 7 An example decoding process 700 according to the present disclosure is shown.

[0170] According to one implementation, the luma and chroma components of an image may be decoded using separate sub-networks. Figure 7 In

[15] , the luminance component of the image is processed by sub-networks such as "synthesis", "prediction fusion", "masked convolution", "decoder using super-prior information", and "variance decoder using super-prior information". The chrominance component is processed by sub-networks such as "synthesis UV", "prediction fusion UV", "masked convolution UV", "decoder using super-prior information", and "variance decoder using super-prior information".

[0171] The benefit of this separate processing is that by applying separate processing, the computational complexity of the processing of the image is reduced. Typically, in neural network-based image and video decoding, the computational complexity is proportional to the square of the number of feature maps. For example, if the number of total feature maps is equal to 192, the computational complexity will be proportional to 192×192. On the other hand, if the feature map is divided into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128×128+64×64, which corresponds to a 45% reduction in complexity. Typically, separate processing of the luminance and chrominance components of an image does not lead to an excessive reduction in performance because the correlation between the luminance and chrominance components is usually very small.

[0172] The processing (decoding process) in the above figure can be explained as follows:

[0173] 1. First, the decomposition entropy model is used to decode the quantized latent values of luminance and chrominance, i.e. Figure 7 in and

[0174] 2. The probability parameters (e.g., variance) generated by the second network are used to generate quantized residual potential values by performing an arithmetic decoding process.

[0175] 3. The quantized residual potential value is inversely gained using the inverse gain unit (iGain), such as Figure 7 The output of the inverse gain unit for the luminance and chrominance components are expressed as and

[0176] 4. For the brightness component, perform the following steps in a loop until you get All elements of:

[0177] a. The first sub-network is used to use The obtained sample points are used to estimate the potential value of quantization The mean parameter of .

[0178] b. Quantified residual potential value and the mean are used to obtain The next element.

[0179] 5. After obtaining After all the samples of , the composite transform can be applied to obtain the reconstructed image.

[0180] 6. For the chroma components, steps 4 and 5 are the same, but with a separate set of networks.

[0181] 7. The decoded luminance component is used as additional information to obtain the chrominance components. Specifically, the inter-channel-related information filter sub-network (ICCI) is used for chrominance component recovery. The luminance is fed into the ICCI sub-network as additional information to assist in chrominance component decoding.

[0182] 8. After the luma and chroma components are reconstructed, adaptive color transform (ACT) is performed.

[0183] The module named ICCI is a post-processing module based on a neural network. The example is not limited to the UCCI sub-network. Any other post-processing module based on a neural network can also be used.

[0184] The exemplary implementation of the present disclosure is as follows Figure 7(Decoding process). The framework includes two branches for luma and chroma components respectively. In each branch, the first sub-network includes context, prediction and optional decoder modules using super-prior information. The second network includes a variance decoder module using super-prior information. The quantized super-prior information potential value is and The arithmetic decoding process generates a quantized residual potential value, which is further fed into the iGain unit to obtain the quantized residual potential value of the gain and

[0185] After obtaining the residual latent value, a recursive prediction operation is performed to obtain the latent value and The following steps describe how to obtain the potential value The chrominance components are processed in the same way but using different networks.

[0186] 1. The autoregressive context module is used to use the sample points The first input to the prediction module is generated by , where the (m, n) pairs are the indices of the samples of the potential values that have been obtained.

[0187] 2. Optionally, the second input of the prediction module is obtained by using a decoder that utilizes the super-prior information and a quantized super-prior information potential value was obtained.

[0188] 3. Using the first and second inputs, the prediction module generates the mean mean[:,i,j].

[0189] 4. Mean mean[:,i,j] and quantized residual potential value are added together to obtain the potential value

[0190] 5. Repeat steps 1-4 for the next sample point.

[0191] Whether and / or how to apply at least one method disclosed in the present disclosure may be signaled from the encoder to the decoder, eg in a bitstream.

[0192] Whether and / or how to apply at least one method disclosed in the present disclosure may be determined by a decoder based on codec information (such as dimension, color format, etc.).

[0193] In addition, modules named MS1, MS2, or MS3+O (in Figure 7 The modules may be included in the processing flow. The modules may perform operations on their inputs by multiplying the inputs by a scalar or adding an additional component to the inputs to obtain an output. The scalars or additional components used by the modules may be indicated in the bitstream.

[0194] Figure 7 The module named RD or the module named AD in the example may be an entropy decoding module, such as a range decoder or an arithmetic decoder.

[0195] The examples described herein are not limited to Figure 7 Some modules may be missing, and some modules may be repositioned in the processing order. In addition, additional modules may be included. For example:

[0196] 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined through another module, which can be based on a neural network.

[0197] 2. One or more modules named MS1, MS2 or MS3+0 may be removed. The core of the present disclosure is not affected by removing one or more of the scaling and adding modules.

[0198] exist Figure 7 In the example, other operations performed during the processing of the luma and chroma components are also indicated using asterisks. These processes are denoted as MS1, MS2, MS3+0. These processes may be, but are not limited to, adaptive quantization, potential sample scaling, and potential sample offset operations. For example, the adaptive quantization process may correspond to scaling the samples by a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. The potential scaling process may correspond to scaling the samples by a multiplier after the prediction process, where the value of the multiplier is predefined or indicated in the bitstream. The offset operation may correspond to adding an additional element to the sample, where again the value of the additional element may be indicated in the bitstream or inferred or predetermined.

[0199] Another operation may be a slice operation, in which samples are first sliced (grouped) into overlapping or non-overlapping regions, where each region is processed independently. For example, samples corresponding to the luma component may be divided into slices of 20 samples in height, while chroma components may be divided into slices of 10 samples in height for processing.

[0200] Another operation may be the application of wavefront parallel processing. In wavefront parallel processing, multiple samples may be processed in parallel, and the number of samples that may be processed in parallel may be indicated by a control parameter. The control parameter may be indicated in the bitstream, inferred, or predetermined. In the case of separate luma and chroma processing, the number of samples that may be processed in parallel may be different, so different indicators may be signaled in the bitstream to separately control the operation of the luma and chroma processing.

[0201] 2.7 Color Separation and Conditional Encoding and Decoding

[0202] Figure 8 An example learning-based image encoding and decoding architecture 800 is shown.

[0203] In one example, if Figure 8 As shown in Figure 1, the primary and secondary color components of an image are encoded and decoded separately using networks with similar architecture but different number of channels. All boxes with the same name are subnetworks with similar architecture, differing only in input and output tensor sizes and number of channels. The number of channels for the primary component is C p =128, the number of channels for the secondary component is C s = 64. Vertical arrows (arrows pointing downward) indicate the data flow related to the encoding and decoding of the secondary color component. Vertical arrows show the data exchange between the primary component pipeline and the secondary component pipeline.

[0204] The input signal to be encoded is denoted as x, and the latent space tensor in the bottleneck of the variational autoencoder is y. The subscript “Y” indicates the primary component, and the subscript “UV” is used for the concatenated secondary components, including the chrominance components.

[0205] First, the input image with RGB color format is converted into the primary component (Y) and secondary component (UV). Y Independent of the minor component x UV is coded and the coded picture size is equal to the input / decoded picture size. The secondary component is conditionally coded and decoded using x Y As auxiliary information from the main component used to encode x UV , and use as a latent tensor with auxiliary information from the main components for decoding Reconstruction. Except for the number of channels, the channel dimensions, and the multiple entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost the same, so the primary latent tensor and the secondary latent tensor will generate two different bitstreams based on two different entropy models. Y 、x UV Before, x Y 、x UV By module, the module adjusts the sample point position by downsampling (in Figure 8The scaling factor s is variable, but the default scaling factor is s=2. The size of the auxiliary input tensor in the conditional encoding and decoding is adjusted so that the encoder receives the primary component tensor and the secondary component tensor with the same picture size. After reconstruction, the secondary component is upsampled using a neural network based upsampling filter module ( Figure 8 The “NN color filter s↑” on is rescaled to the original image size, and the module outputs the secondary component upsampled by a factor s.

[0206] Figure 8 The example in illustrates an image encoding and decoding system where the input image is first converted into a primary component (Y) and a secondary component (UV). Output is the reconstructed output corresponding to the primary and secondary components. At the end of the processing, is converted back to RGB color format. Typically, x UV It is downsampled (resized) before being processed by the encoding and decoding modules (neural networks). For example, x UV The size of can be reduced by 50% in each of the vertical and horizontal dimensions.Thus, the processing of the secondary component involves approximately 50% x 50% = 25% fewer samples and is therefore less computationally complex.

[0207] 2.8 Pruning Operations in Neural Network-Based Encoding and Decoding

[0208] Figure 9 An example synthetic transform for learning-based image encoding and decoding 900 is shown.

[0209] The example synthetic transformation above consists of a sequence of 4 convolutions followed by upsampling with a stride of 2. The synthetic transformation subnetwork is as follows Figure 9 The dimensions of the tensors in different parts of the composite transformation before the clipping layer are as follows Figure 9 As shown in the figure.

[0210] The cropping layer reduces the tensor size h d ×w d Change to h d-1 ×w d-1 , where h d =2·ceil(H / 2 d );w d =2·ceil(w / 2 d); where d is the depth of the convolutions being performed in the codec architecture. For the principal component, the composition transform receives an input tensor of size h × w; h = ceil(H / 16); w = ceil(W / 16). The output of the composition transform for the principal component is 1 × h0 × w0, where h0 = H; h0 = W.

[0211] For the secondary component, the composite transform receives a size of h UV ×w UV The input tensor of h UV =ceil(ceil(H / s) / 16); w UV =ceil(ceil(W / s) / 16). The output of the composite transform of the main components is 2×h UV0 ×w UV0 , where h UV0 =ceil(H / s);h UV0 =ceil(W / s). For the secondary component, the input dimensions are h0=ceil(H / s); w0=ceil(W / s), where s is the scaling factor. The scaling factor can be, for example, 2, where the secondary component is downsampled by a factor of 2.

[0212] Based on the above explanation, the operation of the cropping layer depends on the output size H, W and the depth of the cropping layer. Figure 9 The leftmost cropping layer has a depth of 0. The output of this cropping layer must be equal to H, W (output size). If the size of the input to this cropping layer is larger than H or W in the horizontal or vertical dimension, then cropping needs to be performed in that dimension. The second cropping layer from the left to the right has a depth of 1. The output of the second layer must be equal to h1 = 2·ceil(H / 2 1 ); w1=2·ceil(W / 2 1 ), which means that if the input to this second cropping layer is larger than h1, w1 in any dimension, cropping is applied in that dimension. In summary, the operation of the cropping layer is controlled by the output dimensions H and W. In one example, if H and W are both equal to 16, the cropping layer does not perform any cropping. On the other hand, if H and W are both equal to 17, all four cropping layers will perform cropping.

[0213] 2.9 Bitwise Shift

[0214] The bitwise shift operator can be represented using the function bitshift(x,n), where n is an integer. If n is greater than 0, it corresponds to the right shift operator (>>), which shifts the bits of the input to the right, and the left shift operator (<<), which shifts the bits to the left. In other words, the bitshift(x,n) operation corresponds to:

[0215] bitshift(x,n) = x * 2 n

[0216] or

[0217] bitshift(x,n) = floor(x * 2 n )

[0218] or

[0219] bitshift(x,n) = x / / 2 n

[0220] The output of the bit shift operation is an integer value. In some embodiments, the floor() function may be added to the definition.

[0221] Floor(x) is equal to the largest integer less than or equal to x.

[0222] The " / / " operator or integer division operator: It is an operation that includes division and truncating the result to zero. For example, 7 / 4 and -7 / -4 are truncated to 1, and -7 / 4 and 7 / -4 are truncated to -1.

[0223] rightshift(x,n) = x >> n or

[0224] leftshift(x,n) = x << n

[0225] Equation 3: The bitwise shift operator as an alternative implementation of right shift (rightshift) or left shift (leftshift).

[0226] x >> y: The two's complement integer representation of x is arithmetically right-shifted by y bits. This function is only defined for non-negative integer values of y. The bits shifted into the most significant bit (MSB) due to the right shift have the value equal to the MSB of x before the shift operation.

[0227] x << y: The two's complement integer representation of x is arithmetically left-shifted by y bits. This function is only defined for non-negative integer values of y. The bits shifted into the least significant bit (LSB) due to the left shift have the value equal to 0.

[0228] 2.10 Convolution operation

[0229] The convolution operation starts with a convolution kernel, which is a small matrix of weights. This convolution kernel "slides" over the input data, performs element-wise multiplication with the part of the input it is currently on, and then sums the results into a single output pixel. In some cases, the convolution operation may include a "bias", which is added to the output of the element-wise multiplication operation.

[0230] 2.11Leaky_Relu activation function

[0231] Figure 10 An example leaky ReLU activation function 1000 is shown. The leaky_ReLU activation function is as follows Figure 10 As shown. According to this function, if the input is positive, the output is equal to the input. If the input (y) is negative, the output is equal to a*y. a is usually (but not limited to) a value less than 1 and greater than 0. Since the multiplier a is less than 1, it can be implemented as a multiplication with a non-integer, or as a division operation. The multiplier a can be called the negative slope of the leaky ReLU function.

[0232] 2.12ReLU activation function

[0233] Figure 11 An example ReLU activation function 1100 is shown. A ReLU activation function is as follows Figure 11 According to this function, if the input is positive, the output is equal to the input. If the input (y) is negative, the output is equal to 0.

[0234] 3. Technical problems solved by the disclosed technical solutions

[0235] In image or video compression systems, arithmetic coding or other forms of entropy coding are often used to convert symbols into strings of bits. The process of converting symbols into bits is sequential, which means that small errors (for example, a single bit interpreted as a "0" instead of a "1") can cause the entire bit stream to be corrupted, making decoding of the image impossible.

[0236] The layers used in the example neural network implementation include the following operations:

[0237] Multiplication of floating-point numbers,

[0238] Division operation,

[0239] Floating point addition,

[0240] Rounding operation.

[0241] ·..

[0242] Such operations are not well defined, so the output values can vary from one device to another. For most applications, such small differences are negligible, however, in the case of image and video compression, such small differences can render the bitstream undecodable. The reason is that a single error in the decoding of a single bit will affect the interpretation of all subsequent bits, thus corrupting the entire bitstream.

[0243] Therefore, neural network-based image or video encoding and decoding systems are prone to decoding errors: a bitstream encoded in one device may not be decodable in another.

[0244] 4. List of solutions and implementation examples

[0245] According to the present disclosure, a basic unit including a convolutional layer and an activation layer is provided. The basic unit is designed in such a way that all operations performed by the basic unit are performed using integers. In addition, any functions whose results may depend on the device (e.g., rounding or division operations) are not used. Instead, rounding or division operations are replaced by bitwise shifts.

[0246] 4.1 Core Examples

[0247] Figure 12 An example base unit 1200 according to the present disclosure is shown. Figure 13 Another example base unit 1300 according to the present disclosure is shown.

[0248] Convolutional layers are frequently used in image and video encoding and decoding using neural networks. Examples are disclosed in Figure 12 and 13 Depicted in . According to the example, the basic convolution and activation layers are presented. If the input of the unit is an integer, then the basic unit guarantees that all operations performed to obtain the output can be completed using integers.

[0249] The basic unit (basic convolution and activation unit) is Figure 12 and 13 Depicted in.

[0250] According to the example, first, all parameters of the convolution layer (the first function that operates on the input) are integers. The convolution operation parameters can include the convolution kernel (the number multiplied by the input) and the bias (the number added to the output of the multiplication). For example, if the size of the convolution kernel is 3x3, the two-dimensional (2D) convolution operation can be expressed using the following equation:

[0251]

[0252] The above mathematical formula describes a 2D convolution operation with a convolution kernel size of 3x3. The present disclosure is applicable to any convolution operation, regardless of the dimension, the size of the convolution kernel, or the bias (which can be a vector, scalar, or matrix).

[0253] According to the example, first, the weights of the convolution kernel (used for element-by-element multiplication with the input) and the bias values are all integers. No non-integer values are allowed. Therefore, since the convolution operation consists of multiplication and addition, if the input of the convolution operation is an integer, the output of the convolution operation is guaranteed to be an integer. Secondly, a novel activation layer ("quantized leaky ReLU") is adopted, which does not require any division operations, any rounding operations, and any multiplication with non-integer numbers. The "quantized leaky ReLU" layer corresponds to the activation function of the basic unit, which can be described by an equation with only integer operations such as the following equation:

[0254]

[0255] or

[0256]

[0257] or

[0258]

[0259] or

[0260]

[0261] or

[0262]

[0263] The bitshift(x,n) operation corresponds to the bitwise shift operator.

[0264] Note that the division operation followed by the floor() operation is just a different representation of the bit shift operation. In other words, it can be implemented using the bit shift() operator.

[0265] In another embodiment, the division operation may be replaced by a series of operations that depend on at least one lookup table.

[0266] 4.2 Example Details

[0267] A "quantized leaky ReLU" layer is an activation layer where all processing is performed using integers. The general form of the quantized leaky ReLU function can be found in the example equation in Section 4.1.

[0268] In one example, n is equal to zero. In this case, the “quantized leaky ReLU” function can be described as:

[0269]

[0270] or

[0271]

[0272] …

[0273] In the above formula, constants A, B, and C are integers and can be determined in such a way as to approximate the regular leaky_ReLU function. For example, if the negative slope of the leaky ReLU is 0.01, then A, B, and C can be set to round(2 m *0.01), 0 and m. For example, m can be 13. In this case, A is equal to 82, and C is equal to 13. In this case, A / 2 C is equal to 82 / 8192, which is a number close to 0.01. In other words, by setting A, B, and C, the negative slope of the leaky ReLU function can be approximated using only integer and bit shift operations.

[0274] In another example, A, B, and C can be set to round (2 m *0.01),2 m-1 and m. Compared to the previous example (where B is equal to 0), B is set equal to 2 m-1 As a result, the output of the quantized leaky ReLU function becomes more accurate because the floor operation becomes more accurate. If B is set equal to 2 m-1 and is added to the numerator before the floor operation, the output of the floor() operation becomes the nearest integer value. Note that in this example and the previous examples, the round() operation is only used once during the determination of the number A. In other words, once the value of A is calculated, the computing device does not need to perform the rounding operation each time; the computing device can use the pre-calculated value.

[0275] In other words, A, B, and C can be chosen in such a way as to approximate any negative slope of the leaky ReLU function, which is typically a floating point number between 0 and 1 or between 0 and -1. A, B, and C can be chosen in such a way that the negative slope can be approximated arbitrarily closely.

[0276] Alternative embodiments of the present disclosure may be described as:

[0277]

[0278] In the above example, the floor() operation and the division operation (division by 2 C ) is replaced by a bit shift operation. This is just an alternative representation of the same mathematical formula. In computer science, a bit shift operation can be represented by a division operation (or multiplication operation) followed by the floor function.

[0279] In another example, A, B, and C can be set to another value that approximates the slope of the leaky ReLU function in such a way that, since A, B, and C are all integer-valued numbers and since bit shift operations are used, the approximation is very friendly to implementation in a computing device.

[0280] In an example, if the bit shift is a left shift operation (e.g., the bit shift operation is equivalent to a multiplication by a number greater than 1), B is generally equal to 0. If the bit shift is a right shift operation, B can be greater than 0 to increase the precision of the right shift operation.

[0281] Multiplication and addition operations with integers, as well as bit shift operations, are well-defined in computer hardware. On the other hand, operations with floating-point numbers and division are not as well defined, so the output values can vary from one device to another. For most applications, these small differences are negligible. However, in the case of image and video compression, they can render the bitstream undecodable. This is because a single error in the decoding of a single bit will affect the interpretation of all subsequent bits, corrupting the entire bitstream.

[0282] The implementation of the negative slope of the leaky ReLU function using integers A, B, and C and a bit shift operation ensures that the result of the quantized leaky ReLU function is the same in any computing device. Therefore, a bitstream encoded using one device (e.g., a computer) can be decoded by another device (e.g., a mobile computing device).

[0283] Another embodiment of the present disclosure is:

[0284]

[0285] In this embodiment, the input is bit-shifted by n bits in both the positive and negative branches. In this case, D may be equal to 0. Alternatively, it may be an integer value greater than 0, for example, it may be equal to 2 n-1 .

[0286] In another example, a basic convolution and activation layer including a convolution layer with integer weights and bias values and a ReLU function is presented. Two alternative implementations of the basic convolution and activation layer can be depicted in the following figure:

[0287] Figure 14 Another example base unit 1400 according to the present disclosure is shown.

[0288] The above example is a special version of the present disclosure in which the quantized leaky ReLU function is replaced by a ReLU function. This is because, since the negative slope of the ReLU function is equal to 0, if its input is an integer value, its output is also an integer value. Furthermore, the ReLU function does not include any multiplication or division with floating-point numbers, making it a function suitable for implementation in computing hardware.

[0289] According to the present disclosure, if the input of a basic unit (basic convolution and activation layer) is an integer value, the output is guaranteed to be an integer value. In addition, the operations that the unit can perform include:

[0290] Integer multiplication,

[0291] Bit shift operations,

[0292] Integer addition,

[0293] Limiting,

[0294] Rounding.

[0295] Therefore, there is no need for division, multiplication / addition, or rounding operations with non-integer values, which are device-dependent operations (their output may change when executed on different devices).

[0296] According to the present disclosure, additional layers may be included between the convolutional layer and the activation layer. Possible layers that may be included between the convolutional layer and the activation layer may be clipping layers as described in Section 2.8. In general, any layer that does not include any device-dependent operations (such as operations using floating point or division or rounding, etc.) may be included between the convolutional and activation layers.

[0297] According to the present disclosure, a sub-network can be designed using the invented basic units (basic convolution and activation units). One or more basic units can be used to construct a neural sub-network.

[0298] Alternatively or additionally, a quantization operation may be performed on the output of the elementary unit. For example, a table comprising only integer values may be used for the quantization of the output of the elementary unit.

[0299] An example table could be as follows:

[0300]

[0301]

[0302] According to the example table above, if the input value is between 0 and 110, it is quantized to assume an output value of 0. If the quantized input value is between 228 and 448, the output value is equal to 2.

[0303] Since the output of the basic unit is an integer value, a quantization table such as the one described above (which only includes integer quantization values) can be used to quantize the output of the basic unit. The quantization operation is essentially a mapping function between integer input and integer output. As a result, the implementation of this quantizer is also device-independent.

[0304] According to an example, a variance decoder module based on a neural network and utilizing super-prior information is constructed using basic units. The output of the variance decoder utilizing super-prior information is a probability parameter, such as a Gaussian variance parameter.

[0305] 4.3 Benefits of Examples

[0306] According to the present disclosure, if the input of a basic unit (basic convolution and activation layer) is an integer value, then the output is guaranteed to be an integer value. In addition, the operation performed by the unit is:

[0307] Integer multiplication,

[0308] Bit shift operations,

[0309] Integer addition.

[0310] Therefore, there is no need for division, multiplication / addition, or rounding operations with non-integer values, which are device-dependent operations (their output may change when executed on different devices).

[0311] The basic units (basic convolution and activation layers) presented in this disclosure ensure that the neural sub-networks implemented using such basic units are device-independent. In other words, the output of the sub-network implemented using such basic units is guaranteed to provide the same output on any computing device. Therefore, a bitstream encoded using one device (e.g., a computer) can be decoded by another device (e.g., a mobile computing device), which is a basic requirement for image encoding / decoding.

[0312] 5. Examples

[0313] 1. Decoder item:

[0314] A method for decoding an image or video including a neural subnetwork, comprising the following steps:

[0315] - perform convolution operations,

[0316] - Perform an activation function on the output of the convolution operation,

[0317] A reconstructed image is obtained based on the output of the neural sub-network, wherein the convolution and activation functions are performed using only integer values.

[0318] 6. References

[0319] [1] Z.Cheng, H.Sun, M.Takeuchi and J.Katto, "Learned image compression with discretized gaussian mixture likelihoods and attention modules", in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 7939-7948, 2020.

[0320] [2] B. Bross, J. Chen, S. Liu, and Y.-K. Wang, “Versatile Video Draft (Draft 10)”, JVET-S2001, July 2020.

[0321] [3] R. D. Dony and S. Haykin, “Neural network approaches to image compression”, Proceedings of the IEEE, vol. 83, no. 2, pp. 288–303, 1995.

[0322] [4] Z. Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600–612, 2004.

[0323] [5] G. Bjontegaard, “Calculation of average PSNR differences between RD-curves”, VCEG, Technical Report VCEG-M33, 2001.

[0324] [6] C.E. Shannon, “A mathematical theory of communication”, Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948.

[0325] [7] G.E. Hinton and R.R. Salakhutdinov, “Reducing the dimensionality of data with neural networks”, Science, vol. 313, no. 5786, pp. 504–507, 2006.

[0326] [8] J. Ballé, D. Minnen, S. Singh, S. Hwang and N. Johnston, "Variational image compression with a scale hyperprior", International Conference on Learning Representations, 2018.

[0327] [9] D.Minnen, J.Ballé, G.Toderici, "Joint Autoregressive and HierarchicalPriors for Learned Image Compression", arXiv.1809.02736.1, 2, 3, 4, 7.

[0328] Figure 15 4 is a block diagram illustrating an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as wireless fidelity (Wi-Fi) or a cellular interface.

[0329] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this disclosure. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to produce a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or transmitted via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video received at input 4002 or the communication transmission can be used by component 4008 to generate pixel values or transmit to a displayable video of display interface 4010. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it is understood that the codec tools or operations are used by the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0330] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The technology described in this disclosure may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0331] Figure 16 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this disclosure. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, for example, a graphics coprocessor.

[0332] Figure 174 is a flow chart of an example method 4200 for video processing. At step 402, a determination is made to apply a neural subnetwork to perform a convolution operation using only integer values and to perform an activation function on the output of the convolution operation using only integer values. At step 4204, conversion between visual media data and a bitstream is performed based on the convolution operation and the activation function. The conversion at step 4204 may include encoding at an encoder or decoding at a decoder, depending on the example.

[0333] It should be noted that method 4200 can be implemented in an apparatus for processing video data, such as video encoder 4400, video decoder 4500, and / or encoder 4600, that includes a processor and non-transitory memory having instructions stored thereon. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use with a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that, when executed by the processor, the video codec device performs method 4200. Furthermore, a non-transitory computer-readable recording medium may store a bitstream of a video generated by performing method 4200 by the video processing device. Furthermore, method 4200 can be performed by an apparatus for processing video data, including a processor and non-transitory memory having instructions stored thereon. The instructions, when executed by the processor, cause the processor to perform method 4200.

[0334] Figure 18 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where destination device 4320 can be referred to as a video decoding device.

[0335] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to target device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0336] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or may be external to target device 4320, which may be configured to interface with an external display device.

[0337] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or further standards.

[0338] Figure 19 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 18 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0339] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406.

[0340] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0341] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but for purposes of explanation, these components are shown separately in the example of the video encoder 4400 .

[0342] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.

[0343] The mode selection unit 4403 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0344] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 4413 (other than the picture associated with the current video block).

[0345] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0346] In some examples, motion estimation unit 4404 may perform unidirectional prediction on the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0347] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block. The motion estimation unit 4404 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 4404 may then generate a reference index and a motion vector, where the reference index indicates the reference pictures in list 0 and list 1 that contain the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0348] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal motion information for the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0349] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.

[0350] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0351] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0352] The intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, the intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0353] The residual generation unit 4407 can generate residual data for the current video block by subtracting the (one or more) prediction video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0354] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform a subtraction operation.

[0355] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0356] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0357] Inverse quantization unit 4410 and inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in buffer 4413.

[0358] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0359] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0360] Figure 20 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 18 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0361] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the inverse of the encoding process described for video encoder 4400.

[0362] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.

[0363] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0364] The motion compensation unit 4502 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information, and the motion compensation unit 4502 may use the interpolation filters to generate a prediction block.

[0365] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0366] The intra-frame prediction unit 4503 can form a prediction block from spatially adjacent blocks using, for example, an intra-frame prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0367] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If desired, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.

[0368] Figure 21 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and by applying a finite impulse response (FIR) filter, respectively, where the offset and filter coefficients are transmitted by the encoded side information. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts created by previous stages.

[0369] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed to the transform (T) component 4614 and the quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed to the entropy codec component 4618. The entropy codec component 4618 performs entropy coding and decoding on the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed to the inverse quantization (IQ) component 4620, the inverse transform component 4622, and the reconstruction (REC) component 4624. The REC component 4624 is capable of outputting images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these images are stored in the reference picture cache 4612 .

[0370] A list of some example preferred solutions is provided below.

[0371] The following solutions illustrate examples of the techniques discussed herein.

[0372] 1. A method for processing video data, comprising: determining to apply a neural subnetwork to perform a convolution operation using only integer values and to perform an activation function on an output of the convolution operation using only integer values; and performing conversion between visual media data and a bitstream based on the convolution operation and the activation function.

[0373] 2. The method according to solution 1, wherein the convolution operation uses a convolution kernel with integer weights and an integer bias.

[0374] 3. A method according to any one of solutions 1-2, wherein the activation function uses bit shifting.

[0375] 4. The method according to any one of solutions 1-3, wherein the convolution operation is performed according to the following formula:

[0376]

[0377] 5. A method according to any of solutions 1-4, wherein the convolution operation prohibits non-integer values by employing multiplication and addition only on integer inputs.

[0378] 6. A method according to any one of solutions 1-5, wherein the activation function includes a quantized leaky ReLU function or a ReLU function that does not include division and does not include non-integer multiplication.

[0379] 7. The method according to any one of solutions 1-6, wherein the activation function is performed according to the following formula:

[0380]

[0381]

[0382]

[0383] or

[0384]

[0385] 8. The method according to any one of solutions 1-7, wherein the activation function is performed according to the following formula:

[0386] or

[0387]

[0388] 9. The method according to any one of solutions 1-8, wherein the activation function is performed according to the following formula:

[0389]

[0390] 10. A method according to any one of solutions 1-9, wherein the convolution operation and activation function essentially consist of integer-valued multiplication, bit shift operations, integer-valued addition, clipping, rounding and their combinations.

[0391] 11. A method according to any one of solutions 1-10, wherein the neural sub-network further includes a pruning layer.

[0392] 12. A method according to any one of solutions 1-11, wherein the neural sub-network also includes a quantization operation.

[0393] 13. A method according to any of solutions 1-12, wherein the neural sub-network is included in a variance decoder that utilizes super-prior information of the output probability parameters.

[0394] 14. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-13.

[0395] 15. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of solutions 1-13.

[0396] 16. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to apply a neural subnetwork to perform a convolution operation using only integer values and to perform an activation function on an output of the convolution operation using only integer values; and generating a bitstream based on the determination.

[0397] 17. A method for storing a bitstream of a video, comprising: determining to apply a neural subnetwork to perform a convolution operation using only integer values and to perform an activation function on an output of the convolution operation using only integer values; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0398] 18. A method, apparatus or system as described in the present disclosure.

[0399] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can parse syntax elements in the codec representation according to the format rules using known information about the presence and absence of syntax elements to produce decoded video.

[0400] In the present disclosure, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may correspond, for example, to bits that are co-located or scattered at different locations as defined by the syntax within the bitstream. For example, a macroblock may be encoded based on a transformed and coded error residual value and also using bits in the header and other fields in the bitstream. In addition, during conversion, the decoder may parse the bitstream based on determining known information that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or not include particular syntax fields and generate the codec representation accordingly by including or excluding the syntax fields from the codec representation.

[0401] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this disclosure may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0402] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file preserving other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers, which are located at a site or distributed across multiple sites and interconnected by a communication network.

[0403] The processes and logic flows described in this disclosure may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0404] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor that executes instructions and one or more memory devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices to receive data from them or transfer data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0405] Although this disclosure contains many details, these details should not be interpreted as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features that are unique to particular embodiments of particular technologies. In this disclosure, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable subcombination. In addition, although features may function in certain combinations as described above, and may even be initially claimed in this manner, in some cases, one or more features in a claimed combination may be omitted from that combination, and a claimed combination may be directed to a subcombination or a variation of a subcombination.

[0406] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed sequentially in the particular order or sequence shown, or that all illustrated operations be performed to achieve desired results. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be understood as requiring such partitioning in all embodiments.

[0407] Only a few implementations and examples are described, and other implementations, improvements, and variations can be made based on what is described and illustrated in this disclosure.

[0408] A first component is directly coupled to a second component when there are no intervening components other than a line, trace, or other medium between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than a line, trace, or other medium between the first and second components. The term "coupled" and its variations encompass both direct and indirect couplings. Unless otherwise specified, the use of the term "about" is intended to encompass a range of ±10% of the subsequent figure.

[0409] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples should be considered illustrative rather than restrictive, and the present invention is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0410] In addition, the techniques, systems, subsystems, and methods described and illustrated in the various embodiments, as discrete or separate, may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as coupled may be directly connected, or may be indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and alterations can be determined by those skilled in the art and may be made without departing from the spirit and scope of the present disclosure.

Claims

1. A method for processing video or image data, comprising: applying a neural network to perform a convolution operation on an input and an activation function on an output of the convolution operation, wherein the convolution operation and the activation function are performed using only integer values; and Conversion between visual media data and a bitstream is performed based on the convolution operation and the activation function.

2. The method according to claim 1, wherein The input consists only of integer values.

3. The method according to any one of claims 1 to 2, wherein The convolution operation employs a convolution kernel of integer weights and an integer bias, wherein the integer weights are element-wise multiplied with the input, and wherein the integer bias comprises all integer values.

4. The method according to claim 3, wherein: The size of the convolution kernel is 3x3.

5. The method according to any one of claims 1 to 4, wherein The method is applicable to any convolution operation regardless of the dimensionality, the size of the convolution kernel and the integer bias.

6. The method according to any one of claims 1 to 5, wherein The activation function uses bit shifting.

7. The method according to any one of claims 1 to 6, wherein Non-integer values are not allowed for the convolution operation and the activation function.

8. The method according to any one of claims 1 to 6, wherein The convolution operation prohibits non-integer values by employing only multiplications and additions using integer inputs.

9. The method according to any one of claims 1 to 6, wherein The convolution operation is performed using only integer-valued multiplications and integer-valued additions.

10. The method according to any one of claims 1 to 6, wherein The input comprises only integer values to ensure that the output of the neural network comprises only integer values.

11. The method according to any one of claims 1 to 10, wherein The convolution operation is a two-dimensional convolution operation and is performed according to the following formula:

12. The method according to any one of claims 1 to 11, wherein The activation function includes a rectified linear unit (ReLU) or a quantized leaky ReLU.

13. The method according to claim 12, wherein: The activation function does not use any division operations, any rounding operations, or any multiplication operations involving non-integer numbers.

14. The method according to any one of claims 1 to 13, wherein: The activation function is performed according to one or a combination of the following: or or or or The bitshift(x,n) operation corresponds to a bitwise shift operator.

15. The method according to claim 14, wherein The bitwise shift operator is used to implement a division operation followed by the floor() operation.

16. The method according to claim 14, wherein The division operation is replaced by a series of operations using at least one lookup table.

17. The method according to any one of claims 1 to 13, wherein: The activation function is performed according to the following formula: or 18. The method according to claim 17, wherein: The constants A, B, and C are integers that approximate the regular leaky ReLU function.

19. The method according to claim 18, wherein The constants A, B, and C are set to approximate the negative slope of the leaky ReLU function.

20. The method according to claim 18, wherein The constants A, B, and C are set to approximate any negative slope of the leaky ReLU function, where the arbitrary negative slope is a floating point number between 0 and 1 or between 0 and -1.

21. The method according to any one of claims 1 to 20, wherein When the bit shift is a left shift operation, the value of the constant B is zero, and wherein when the bit shift is a right shift operation, the value of the constant B is one.

22. The method according to any one of claims 1 to 13, wherein: The activation function is performed according to the following formula:

23. The method according to claim 22, wherein The input is bit-shifted by n bits in both the positive branch and the negative branch.

24. The method according to claim 1, wherein The convolution operation includes a convolution layer having integer weights and bias values, and the activation function includes a rectified linear unit (ReLU) function.

25. The method according to any one of claims 1 to 24, wherein Additional layers are included between the convolution operation and the activation function.

26. The method according to any one of claims 1 to 25, wherein The bit shift operation is performed after the convolution operation and before the activation function.

27. The method according to claim 26, wherein When the input is an integer value, the negative slope of the activation function is equal to zero, and wherein the output is guaranteed to be an integer value.

28. The method according to claim 24, wherein The convolution operation includes one or more of integer multiplication, bit shift operation, integer addition and clipping.

29. The method according to any one of claims 1 to 28, wherein The method is performed without using any division, multiplication or addition operations with non-integer values, or rounding operations.

30. The method according to any one of claims 25 to 29, wherein One of the additional layers is a clipping layer.

31. The method according to claim 30, wherein The additional layer does not include device-dependent operations using floating point, division, or rounding.

32. The method according to any one of claims 1 to 31, wherein The neural network is a neural sub-network among multiple neural sub-networks, and each neural sub-network includes a convolution operation and an activation function.

33. The method according to any one of claims 1 to 32, wherein A quantization operation is performed on an output of the activation function, and wherein the quantization operation uses only integers.

34. The method according to claim 33, wherein The integers used for the quantization operation are obtained from a table comprising quantization levels and quantized output values.

35. The method of claim 33, wherein: The quantization operation includes a mapping function between integer inputs and integer outputs.

36. The method according to any one of claims 1 to 35, wherein The convolution operation and the activation function are incorporated into a variance decoder module that utilizes super-prior information.

37. The method according to any one of claims 1 to 36, wherein The output of the variance decoder module utilizing super-prior information includes the probability parameter or standard deviation.

38. The method according to any one of claims 1 to 37, wherein One or more of the probability parameters include a standard deviation.

39. The method according to any one of claims 1 to 38, wherein The output of the activation function is the reconstructed image.

40. The method of claim 1, wherein The convolution operation and the activation function essentially consist of integer-valued multiplication, bit shift operation, integer-valued addition, clipping, and combinations thereof.

41. The method according to any one of claims 1 to 37, wherein The converting includes encoding the visual media data into the bitstream.

42. The method according to any one of claims 1 to 37, wherein The converting includes decoding the visual media data from the bitstream.

43. An apparatus for processing media data, comprising: one or more processors; as well as A non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the apparatus to perform the method of any one of claims 1-42.

44. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method according to any one of claims 1-42.

45. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein: The method comprises the method according to any one of claims 1-42.

46. A method for storing a bitstream of a video, comprising the method according to any one of claims 1-42.

47. A method, apparatus or system as described in the present disclosure.