Method and device for visual data processing and medium
By dividing the tensors in the adaptive filter of the neural network model into slices that are multiples of a predetermined value, the problem of insufficient encoding and decoding efficiency in the existing technology is solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202480068348.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-23
- Filing Date
- 2024-10-22
- Publication Date
- 2026-05-26
AI Technical Summary
Existing neural network-based image/video encoding and decoding technologies still have room for improvement in encoding and decoding efficiency, especially when dealing with large-scale data, where traditional methods struggle to effectively utilize the potential of neural networks.
Encoding and decoding efficiency is improved by dividing the tensors used in the adaptive filter of the neural network-based model into multiples of predetermined values and performing transformations based on these pieces.
This ensures that the size of most chips is a multiple of the predetermined value, improving encoding and decoding efficiency and enabling more efficient data processing.
Smart Images

Figure CN122095374A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to visual data processing techniques, and more specifically, to visual data encoding and decoding based on neural networks. Background Technology
[0002] Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. In the past five years, neural network-based image / video compression technologies have made significant strides. It has been reported that the latest neural network-based image compression algorithms have achieved rate-distortion (RD) performance comparable to that of Multifunctional Video Coding (VVC). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, the encoding and decoding efficiency of neural network-based image / video codecs is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method includes: for a conversion between visual data utilizing a neural network (NN)-based model and a bitstream of visual data, dividing a first tensor used in an adaptive filter of the NN-based model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; and performing a conversion based on the first set of slices.
[0005] Based on the method according to the first aspect of this disclosure, the first tensor used in the adaptive filter of the NN-based model is divided into a first group of pieces based on a first piece size, and the first piece size is a multiple of a first predetermined value. Compared with conventional solutions that do not restrict the piece size, the proposed method can advantageously ensure that the size of most pieces is a multiple of the predetermined value. In this way, encoding and decoding efficiency can be improved.
[0006] In a second aspect, an apparatus for visual data processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that enable a processor to execute the method according to the first aspect of this disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by means of a device for visual data processing. The method includes: dividing a first tensor used in an adaptive filter of a neural network (NN)-based model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; and generating a bitstream based on the first set of slices using the NN-based model.
[0009] In a fifth aspect, a method for storing bitstreams of visual data is proposed. The method includes: dividing a first tensor used in an adaptive filter of a neural network (NN)-based model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; generating a bitstream based on the first set of slices using the NN-based model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0011] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.
[0012] Figure 1A A block diagram of an example visual data encoding / decoding system according to some embodiments of the present disclosure is shown; Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme; Figure 2 An example potential representation of the image is shown; Figure 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model; Figure 4 This is a schematic diagram illustrating an example combined model configured to jointly optimize the context model, as well as the super-prior and autoencoder; Figure 5 An example encoding process is shown; Figure 6 An example decoding process is shown; Figure 7 Example decoding processes according to some embodiments of this disclosure are shown; Figure 8 An example of a learning-based image codec architecture is shown; Figure 9 This demonstrates enhancement filter techniques; Figure 10 An example implementation of a primary component-guided adaptive upsampling filter is shown; Figure 11 An example implementation of an EFE nonlinear filter is shown; Figure 12 A simplified example implementation of an EFE nonlinear filter is shown; Figure 13 An example implementation of an adaptive upsampling filter guided by the main quantization component is shown; Figure 14 An example implementation of a quantized EFE nonlinear filter is shown; Figure 15 A simplified example implementation of a principal component-guided adaptive upsampling filter is shown; Figure 16 A simplified example implementation of an EFE nonlinear filter is shown; Figure 17 A flowchart of a method for visual data processing according to embodiments of the present disclosure is shown; and Figure 18 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0013] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0015] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0016] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0017] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0019] Example Environment Figure 1A This is a block diagram illustrating an example visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.
[0020] Visual data source 112 may include sources such as visual data capture devices. Examples of visual data capture devices include, but are not limited to, interfaces for receiving visual data from visual data providers, computer graphics systems for generating visual data, and / or combinations thereof.
[0021] Visual data may include one or more pictures or images from a video. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded / decoded representation of the visual data. The bitstream may include encoded / decoded pictures and associated visual data. An encoded / decoded picture is an encoded / decoded representation of a picture. Associated visual data may include sequence parameter sets, picture parameter sets, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data may be transmitted directly to a destination device 120 via network 130A through I / O interface 116. Encoded visual data may also be stored on storage medium / server 130B for access by the destination device 120.
[0022] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0023] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or further standards.
[0024] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Additionally, although some embodiments describe encoding steps in detail, it should be understood that corresponding decoding steps will be implemented by a decoder that undoes the encoding. Furthermore, the term visual data processing includes visual data encoding / decoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another or at different compression bitrates.
[0025] 1. Brief Overview This disclosure relates to image and video encoding and decoding based on neural networks (NNs). Specifically, it relates to improvements in enhancement filters.
[0026] 2 Introduction Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Significant progress has been made in neural network-based image / video compression techniques over the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its early stages.
[0027] 2.1 Image / Video Compression Image / video compression (also known as image / video encoding / decoding) generally refers to the computational techniques used to compress images / videos into binary code to facilitate storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video, a process known as lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is not always necessary. The performance of image / video compression algorithms is typically evaluated in two ways: compression ratio and reconstruction quality. Compression ratio is directly related to the number of binary codes, and lower is better; reconstruction quality is measured by comparing the reconstructed image / video to the original image / video, and higher is better.
[0028] Image / video compression techniques can be divided into two branches: classical video codec methods and neural network-based video compression methods. Classical video codec schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully hand-designing entropy codes that model dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded as a codec tool within existing classical video codecs, serving only as part of the framework; the latter is a separate framework developed based on neural networks without relying on classical video codecs.
[0029] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-increasing amount of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups, the Joint Picture Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), and the ITU-T also has its own Video Codec Experts Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard, Multifunction Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC reduces the bit rate by an average of 50% while maintaining the same visual quality.
[0030] Neural network-based image / video compression is not a new invention, as many researchers have worked on neural network-based image encoding and decoding. However, the network architecture is relatively shallow, and the performance is not entirely satisfactory. Thanks to the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. Nevertheless, this technology is still far from mature and many challenges need to be addressed.
[0031] 2.2 Neural Networks Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different types of representations. Note that these representations are not manually designed; instead, they are learned from massive amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations and is therefore considered particularly suitable for processing natively unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.
[0032] 2.3 Neural Networks for Image Compression Existing neural network methods for image compression can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature.
[0033] 2.3.1 Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless encoding and decoding can achieve the lowest possible decoding rate. ,in It is a symbol The probability of [the lossy encoding / decoding method]. Many lossless encoding / decoding methods have been developed in the literature, among which arithmetic encoding / decoding is considered one of the best methods. Given a probability distribution... Arithmetic encoding and decoding ensures that the encoding / decoding rate is as close as possible to its theoretical limit without considering rounding errors. Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality.
[0034] Following the predictive encoding / decoding strategy, for One approach to modeling this is to predict pixel probabilities one by one in raster scan order based on previous observations. It's an image.
[0035] (1) in These are the height and width of the image, respectively. Previous observations are also referred to as the current pixel's... Context When the image is large, estimating the conditional probability can be difficult, so a simplified approach is to limit the scope of its context.
[0036] (2) in It is a predefined constant that controls the scope of the context.
[0037] It should be noted that this condition can also take into account the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on previously encoded and decoded pixels (including R / G / B samples). The current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample. For encoding and decoding the current B sample, the previously encoded and decoded pixels as well as the current R and G samples can also be considered.
[0038] Neural networks were initially introduced for computer vision tasks and have proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to estimate values given their context. The probability of.
[0039] Most methods directly model the probability distribution in the pixel domain. Some researchers have also attempted to model the probability distribution as a conditional distribution based on explicit or latent representations. That is, we can estimate... (3) in It is an additional condition, and This means that modeling is divided into unconditional modeling and conditional modeling. The additional conditions can be image label information or high-level representations.
[0040] 2.3.2 Automatic Encoder Autoencoders originate from the renowned work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part transforms the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders enable the automatic learning of representations and eliminate the need for hand-crafted features, which is considered one of the most significant advantages of neural networks.
[0041] Figure 1B A typical transform encoding / decoding scheme is shown. Original image. Analysis network Transformation to achieve latent representation The latent representation y is quantized and compressed into bits. The number of bits... Used to measure codec rate. Latent representation of quantization. Then by the synthetic network Inverse transform to obtain the reconstructed image By using functions Transformation Calculate distortion in the perceptual space.
[0042] Applying autoencoder networks to lossy image compression is intuitive. We simply need to encode the latent representations learned from a well-trained neural network. However, adapting autoencoders to image compression is not straightforward, as the original autoencoders are not optimized for compression, making direct use of trained autoencoders inefficient. Other major challenges include: First, low-dimensional representations should be quantized before encoding, but quantization is non-differentiable, necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios, as both distortion and bit rate need to be considered. Estimating the bit rate is challenging. Third, practical image encoding / decoding schemes need to support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Many researchers have been actively contributing to this field to address these challenges.
[0043] Prototype autoencoders for image compression, such as Figure 1B As shown, it can be regarded as Transform encoding and decoding Strategy. Original image It is to utilize analyze network The transformation is used, among which It is the potential representation that will be quantized and encoded / decoded. synthesis The network quantizes the latent representation of the inverse transform. To obtain reconstructed images The framework is trained using a rate-distortion loss function, i.e. ,in Distortion between From quantification The calculated or estimated bit rate, and These are Lagrange multipliers. It should be noted that... It can be computed in the pixel domain or the receptive domain. All existing research follows this prototype, with differences only in network structure or loss function.
[0044] 2.3.3 Super-prior model In transform encoding and decoding methods used for image compression, the encoder subnetwork (Section 2.3.2) uses parametric analysis of the transform. Transform the image vector x into a latent representation Then quantify it to form .because Since they are discrete values, they can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as bit sequences.
[0045] from Figure 2 The left and right center images clearly show this. Significant spatial dependencies exist among the elements. Notably, their scales (right center image) appear to be spatially coupled. An additional set of random variables is introduced. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as Figure 3 As shown.
[0046] exist Figure 3 In the middle, the encoder is on the left side of the model. and decoder (Explained in Section 2.3.2). The right-hand side is used to obtain... Additional encoders utilizing prior information and decoders that utilize prior information Network. In this architecture, the encoder subjectes the input image x to... This produces a response with a standard deviation that varies in the spatial domain. .response fed to In summary The distribution of standard deviations in the data. Then... Quantified ( The data is compressed and transmitted as side information. The encoder then uses the quantization vector... To estimate the spatial distribution of standard deviation And use it to compress and transmit quantized image representations. The decoder first recovers from the compressed signal. Then it uses To obtain This provided it with the correct probability estimate, which also enabled successful recovery. Then it will Feed to To obtain a reconstructed image.
[0047] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. Figure 2 The rightmost image in the diagram corresponds to the quantization latent value when using an encoder / decoder that leverages prior information. Compared to the right-middle image, spatial redundancy is significantly reduced because the samples of the quantization latent value have lower correlation.
[0048] Figure 2 The image from the Kodak dataset is shown (left), a visualization of the latent representation y of the image (middle left), the latent standard deviation σ (middle right), and the latent y (right) after the introduction of a super-prior network (an encoder and decoder network that utilizes super-prior information).
[0049] Figure 3 The network architecture of an autoencoder implementing a priori model is shown. The left side shows the image autoencoder network, and the right side corresponds to the priori subnetwork. The analysis and synthesis transforms are represented as follows: Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks, utilizing the hyperprior information of the encoder (denoted as...). ) and decoders that utilize prior information (represented as The advanced prior model generates quantified advanced prior information latent values (). ), which includes information on quantifying potential values Information about the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).
[0050] 2.3.4 Context Model Although the prior model improves the quantification of latent values Modeling the probability distribution of quantified potential values is possible, but further improvements can be achieved by utilizing an autoregressive model (context model) that predicts quantified potential values from the causal context of quantified potential values.
[0051] The term autoregressive means that the output of a process is later used as its input. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.
[0052] Figure 4 This is a schematic diagram illustrating an example combined model configured to jointly optimize the context model, as well as the super-prior and autoencoder. The meanings of the different symbols are shown below.
[0053] Table - Symbol Explanation
[0054] A joint architecture is employed, utilizing both a hyperprior model subnetwork (an encoder and a decoder utilizing hyperprior information) and a context model subnetwork. The hyperprior and context models are combined to learn about the quantized latent value. The probability model is then used for entropy encoding and decoding. For example... Figure 4 As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean for the Gaussian probability model. And scaling (or variance) The parameters are then used. A Gaussian probability model is then used to encode the samples of the quantized latent values into a bitstream with the aid of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized latent values from the bitstream via the arithmetic decoder (AD) module. .
[0055] Figure 4 This illustrates the combined model's joint optimization of the autoregressive component (context model) to estimate the probability distribution of latent values from their causal context, along with the super-prior and low-level autoencoder. Real-valued latent representations are quantized (Q) to create quantized latent values (…). ) and quantified potential value of prior information ( The image is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted area corresponds to the component performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream.
[0056] Typically, latent samples are modeled as Gaussian distributions or Gaussian mixture models (not limited to). According to Figure 4 The context model and the super-prior are jointly used to estimate the probability distribution of potential samples. Since the Gaussian distribution can be defined by the mean and variance (also known as sigma or scaling), the joint model is used to estimate the mean and variance (denoted as...). ).
[0057] 2.3.5 Gain Variational Automatic Encoder (G-VAE) Typically, neural network-based image / video compression methods require training multiple models to adapt to different bit rates. A gain variational autoencoder (G-VAE) is a variational autoencoder with a pair of gain units, designed to achieve continuously variable bit rate adaptation using a single model. It consists of a pair of gain units, typically inserted into the encoder's output and the decoder's input. The encoder's output is defined as the latent representation. ,in This represents the number of channels, height, and width of the latent representation. Each channel of the latent representation is represented as... ,in A pair of gain units includes a gain matrix. and the inversion gain matrix, where This is the number of gain vectors. A gain vector can be represented as... ,in Indicates the index of the gain vector in the gain matrix.
[0058] The motivation behind the gain matrix is similar to the quantization table in JPEG, controlling the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.
[0059]
[0060] in It is channel multiplication, that is ,and It is the gain vector The first in There are several gain values. The inversion gain matrix used on the decoder side can be expressed as: It is made of Composed of inversion gain vectors, i.e. The inversion gain process is expressed as follows:
[0061] in It is the quantized latent representation of the decoded data, and It is a quantized potential representation of the inversion gain, which will be fed into the synthesis network.
[0062] To achieve continuous variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors... The interpolation gain vector can be obtained using the following equation.
[0063]
[0064]
[0065] in These are interpolation coefficients, which control the bit rate of the generated gain vector pairs. Because... Since it is a real number, it is possible to achieve any bit rate between two given gain vector pairs.
[0066] 2.3.6 Encoding process using a joint autoregressive hyperprior model Figure 4 This corresponds to the existing compression method proposed. The encoding and decoding processes will be described in this section and the next section, respectively.
[0067] Figure 5 The encoding process is described. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transform representation called the latent value, which is... Indicate. Then It is input into the quantizer block represented by Q to obtain the quantized potential value ( Then, the arithmetic coding module (represented as AE) is used to... Converted to a bitstream (bits1). Arithmetic-coded blocks are sequentially... Each sample point is converted into a bit stream (bits1) one by one.
[0068] The module utilizes an encoder with prior information, context, a decoder with prior information, and an entropy parameter subnetwork to estimate the quantized latent value. The probability distribution of the sample points. Latent value The information is input into an encoder that utilizes prior information, and the encoder outputs the prior information latent value (represented as...). Then the latent value of the prior information is quantized (). The second bitstream (bits2) is generated using the Arithmetic Encoding (AE) module. The Decompositional Entropy module generates a probability distribution used to encode the quantized prior information latent values into a bitstream. The quantized prior information latent values include information about the quantized latent values (…). Information about the probability distribution of ( ).
[0069] Entropy parameter subnetwork generation is used for quantizing potential values The probability distribution is estimated for encoding. Information generated from the entropy parameter typically includes the mean of the Gaussian probability distribution. And scaling (or variance) Parameters. The Gaussian distribution of a random variable x is defined as follows: , where parameters It is the mean or expected value of the distribution (also the median and mode), while the parameter This is its standard deviation (or variance or scaling). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.
[0070] The subnetwork uses a decoder with prior information to generate part of the information used by the entropy parameter subnetwork, while the other part is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the quantized latent values. It is usually a matrix composed of many sample points. According to the matrix The dimension, the sample points can use such as [i,j,k] or The index of [i,j] is used to indicate the sample points. [i,j] are encoded sequentially by the AE, typically using a raster scan order. In a raster scan order, the matrix rows are processed from top to bottom, with samples within each row processed from left to right. In such a scenario (where the AE encodes the samples into a bitstream using a raster scan order), the context module uses samples previously encoded in a raster scan order to generate a bitstream with the samples. Information related to [i,j]. The information generated by the context module and the decoder utilizing prior information is combined by the entropy parameter module to generate information for quantizing the latent values. The probability distribution encoded as a bit stream (bits1).
[0071] Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder.
[0072] It should be noted that the above modules can also use other names.
[0073] In the above description, Figure 5 All elements in the encoder are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).
[0074] 2.3.7 Decoding process using a joint autoregressive superprior model Figure 6 Depicting and Figure 5 The encoding process shown corresponds to a separate decoding process.
[0075] During decoding, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decompositional entropy subnetwork. The decompositional entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This refers to the quantized potential value of the prior information. The AD process reverts to the AE process applied in the encoder. Both the AE and AD processes are lossless, meaning that the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.
[0076] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks used in the decoder (context, the decoder utilizing prior information, and the entropy parameter) are the same as those in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder as in the encoder, which is crucial for lossless reconstruction of the quantized latent value. This is crucial. As a result, the quantized latent values can be obtained in the decoder as well as those obtained in the encoder. Same version.
[0077] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of quantized latent values one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial and therefore cannot be accelerated using techniques such as parallelization.
[0078] Finally, the fully reconstructed quantized potential value Input to the synthesis transformation (in) Figure 6 The module (represented as decoder in the image) is used to obtain the reconstructed image.
[0079] In the above description, Figure 6 All elements in the array are collectively referred to as the decoder. The synthetic transformation that converts quantized latent values into a reconstructed image is also called a decoder (or automatic decoder).
[0080] 2.4 Neural Networks for Video Compression Similar to traditional video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, more effort is needed to address the challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to remove inter-frame redundancy. Inter-frame prediction is a key step in these works. Motion estimation and compensation have been widely adopted, but only recently have they been implemented using trained neural networks.
[0081] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the case of random access, decoding can begin from any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, the aim is to reduce decoding time; therefore, usually only the earlier frames in the temporal domain can be used as reference frames to decode subsequent frames.
[0082] 2.5 Prerequisites Almost all natural images / videos are in digital format. Grayscale digital images can be generated by... It means that, among them It is a set of pixel values. It is the image height. It is the image width. For example, This is a common setting, in this case Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed images certainly have fewer bits.
[0083] Color images are typically represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by... This means that three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Most neural network-based video compression schemes are developed in the RGB color space, while traditional codecs typically use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels: Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. The advantage is that Cb and Cr are often downsampled for pre-compression because the human visual system is less sensitive to the chrominance component.
[0084] A color video sequence consists of multiple color images (called frames) to record a scene at different timestamps. For example, in the RGB color space, a color video can be represented as... ,in It is the frame number in the video sequence. .if If the video has 50 frames per second (fps), then the data rate of the uncompressed video is... Bits per second (bps), approximately 2.32 Gbps, requires a significant amount of storage, so compression is definitely necessary before transmission over the internet.
[0085] Typically, lossless methods can achieve compression ratios of around 1.5 to 3 for natural images, which is clearly below what is required. Therefore, lossy compression has been developed to achieve higher compression ratios, but at the cost of introducing distortion. Distortion can be measured by calculating the mean squared difference between the original and reconstructed images, known as the mean squared error (MSE). For grayscale images, the MSE can be calculated using the following equation.
[0086] (4) Accordingly, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR): (5) in The maximum value in the range is, for example, 255 for an 8-bit grayscale image. There are other quality assessment metrics, such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).
[0087] To compare different lossless compression schemes, it is sufficient to compare the compression ratio of a given rate, or vice versa. However, to compare different lossy compression methods, both rate and reconstruction quality must be considered. For example, a common approach is to calculate the relative rates at several different quality levels and then average the rates; the average relative rate is called Bjontegaard's incremental rate (BD rate). Other important aspects for evaluating image / video encoding / decoding schemes include encoding / decoding complexity, scalability, robustness, and more.
[0088] 2.6 Separate processing of the luminance and chrominance components of an image Figure 7 The decoding process according to this disclosure is shown.
[0089] Depending on one implementation, the luminance and chrominance components of an image can be decoded using separate sub-networks. Figure 7In this process, the luminance component of the image is processed by sub-networks such as "synthesis", "predictive fusion", "mask convolution", "decoder using prior information", and "variance decoder using prior information". The chrominance component is processed by sub-networks such as "synthesis UV", "predictive fusion UV", "mask convolution UV", "decoder UV using prior information", and "variance decoder UV using prior information".
[0090] The advantage of this separate processing is that it reduces the computational complexity of image processing. Typically, in neural network-based image and video decoding, computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is 192, the computational complexity will be proportional to 192 × 192. On the other hand, if the feature maps are divided into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128 × 128 + 64 × 64, which corresponds to a 45% reduction in complexity. Generally, separate processing of the luminance and chrominance components of an image does not lead to a significant performance degradation because the correlation between the luminance and chrominance components is usually very small.
[0091] Figure 7 The processing (decoding process) in the code can be explained as follows: 1. First, the decompositional entropy model is used to decode the quantization potential values for luminance and chrominance, i.e. Figure 7 In .
[0092] 2. The probability parameters (e.g., variance) generated by the second network are used to generate quantized residual latent values by performing an arithmetic decoding process.
[0093] 3. For example Figure 7 As shown in orange, the quantized residual latent value is inverted and gained using an inversion gain unit (iGain). The output of the inversion gain unit is expressed for the luminance and chrominance components as follows: .
[0094] 4. For the luminance component, the following steps are performed repeatedly until the result is obtained. All elements: a. The first subnetwork is used for... The already obtained samples are used to estimate the quantized potential value ( The mean value parameter of ).
[0095] b. Quantified residual potential value The mean value is used to obtain The next element. 5. In obtaining After all the samples are obtained, a synthetic transformation can be applied to obtain the reconstructed image.
[0096] 6. For the chromaticity components, steps 4 and 5 are the same, but with a separate network set.
[0097] 7. The decoded luminance component is used to obtain additional information for the chrominance component. Specifically, the Inter-Frame Channel Related Information Filter (ICCI) subnetwork is used for chrominance component recovery. Luminance is fed as additional information into the ICCI subnetwork to assist in chrominance component decoding.
[0098] 8. After the luminance and chrominance components are reconstructed, perform adaptive color transformation (ACT).
[0099] The module named ICCI is a neural network-based post-processing module. This solution is not limited to the UCCI subnetwork; any other neural network-based post-processing module can also be used.
[0100] An exemplary implementation of this solution is in Figure 7 The decoding process is described in the diagram. The framework comprises two branches, one for the luma and one for the chroma components. Within each branch, the first sub-network includes context, prediction, and optionally a decoder module utilizing advanced prior information. The second network includes a variance decoder module utilizing the advanced prior information. The quantized advanced prior information latent value is... The arithmetic decoding process generates quantized residual latent values, which are further fed into the iGain unit to obtain quantized residual latent values for gain. .
[0101] After obtaining the residual potential value, a recursive prediction operation is performed to obtain the potential value. The following steps describe how to obtain potential values. The sample points, and the chromaticity components are processed in the same way but using different networks.
[0102] 1. The autoregressive context module is used when using samples. This is used to generate the first input to the prediction module, where the (m, n) pairs are the indices of the sample points of the obtained potential values.
[0103] 2. Optionally, the second input to the prediction module is obtained by using a decoder that utilizes prior information and quantized potential values of the prior information. It was obtained.
[0104] 3. Using the first and second inputs, the prediction module generates the mean. .
[0105] 4. Mean And quantified residual potential value Added together to obtain potential value .
[0106] 5. Steps 1-4 are repeated for the next sample point.
[0107] Whether and / or how at least one method disclosed in the document can be applied, for example, in a bitstream transmitted from the encoder to the decoder via signal transmission.
[0108] Alternatively, whether and / or how to apply at least one method disclosed in the document can be determined by the decoder based on encoding / decoding information such as dimensions, color format, etc.
[0109] Alternative or additional locations, named MS1, MS2, or MS3+O (in Figure 7 Modules (in the process stream) can be included in the processing stream. A module can perform operations on its input by multiplying it by a scalar or adding an additional component to the input to obtain an output. The scalar or additional component used by the module can be indicated in the bitstream.
[0110] Figure 7 The module named RD or AD in the code can be an entropy decoding module. It can be a range decoder or an arithmetic decoder, etc.
[0111] The solutions described in this article are not limited to Figure 7 The diagram shows a specific combination of units. Some modules may be missing, and some modules may be replaced in the order they are processed. Additional modules may also be included. For example: 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined by another module, which can be based on a neural network.
[0112] 2. One or more modules named MS1, MS2, or MS3+O can be removed. The removal of one or more of these scaling and adding modules does not affect the core of the solution.
[0113] exist Figure 7 In the bitstream, star symbols are also used to indicate other operations performed during the processing of the luminance and chrominance components. These operations are denoted as MS1, MS2, MS3+0. These operations can be, but are not limited to, adaptive quantization, latent sample scaling, and latent sample offset operations. For example, in adaptive quantization, this can correspond to scaling the samples using a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. Latent scaling can correspond to scaling the samples using a multiplier after the prediction process, where the multiplier value is predefined or indicated in the bitstream. Offset operations can correspond to adding an appended element to the sample, again where the value of the appended element can be indicated, estimated, or predetermined in the bitstream.
[0114] Another operation can be a slice operation, where samples are first sliced (grouped) into overlapping or non-overlapping regions, where each region is processed independently. For example, samples corresponding to the luminance component can be divided into slices with a slice height of 20 samples, while chrominance components can be divided into slices with a slice height of 10 samples for processing.
[0115] Another application is wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the number of samples that can be processed in parallel can be indicated by control parameters. These control parameters can be indicated in the bitstream, estimated, or predetermined. In the case of separate luma and chroma processing, the number of samples that can be processed in parallel can be different, so different indicators can be transmitted via signals in the bitstream to individually control the operation of luma and chroma processing.
[0116] 2.7 Color Separation and Conditional Encoding / Decoding In one example, such as Figure 8 As shown, using networks with similar architectures but different numbers of channels, the first and second color components of an image are encoded and decoded separately. All boxes with the same name are sub-networks with similar architectures, differing only in input / output tensor sizes and the number of channels. The number of channels for the principal components is... The number of channels for the secondary components is The vertical arrows (pointing downwards) indicate the data flow related to the encoding and decoding of the second color component. The vertical arrows also show the data exchange between the primary and secondary component pipelines.
[0117] The input signal to be encoded is represented as The latent space tensor in the bottleneck of a variational autoencoder is The subscript "Y" indicates the primary component, and the subscript "UV" is used for the secondary components in series, including the chromaticity component.
[0118] Figure 8 A learning-based image codec architecture is shown.
[0119] First, the input image in RGB color format is converted into a first (Y) component and a second (UV) component. (Main component) Independent of secondary components The image is encoded and decoded, and the size of the encoded / decoded image is equal to the input / decoded image size. Secondary components are conditionally encoded and decoded, using data from the primary components. As used for encoding The auxiliary information, and the auxiliary information from the main component. Used as a potential tensor for decoding Reconstruction. Aside from the number of channels, channel sizes, and several entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost identical. Therefore, the first and second latent tensors will generate two different bitstreams based on two different entropy models. In encoding... Before, This module adjusts the sample point position through downsampling (in...) Figure 8 The superscript "s" indicates that the encoded image size for the minor component differs from the encoded image size for the major component. The scaling factor s is variable, but the default scaling factor is 0. In conditional encoding and decoding, the size of the auxiliary input tensor is adjusted so that the encoder receives primary and secondary component tensors with the same image size. After reconstruction, the secondary component utilizes a neural network-based upsampling filter module (…). Figure 8 The "NN color filters" on the image are rescaled to the original image size, and the module output utilizes a factor. Secondary components of upsampling.
[0120] Figure 8 The example illustrates an image encoding / decoding system where the input image is first converted into a first (Y) component and a second (UV) component. Output This corresponds to the reconstructed output of the primary and secondary components. At the end of the processing, It is converted back to RGB color format. Typically, this is done before processing using encoding and decoding modules (neural networks). Downsampled (resized). For example, The size can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Therefore, the processing of minor components involves approximately 50% × 50% = 25% fewer samples, making it computationally less complex.
[0121] 2.8 JPEG AI Image Codec Standard At the time of writing, the JPEG AI image codec standard is an image codec standard standardized by the JPEG Working Group (WG), JPEG WG being WG 1 of ISO / IEC JTC 1 SC 29. The ISO / IEC number for the JPEG AI standard is ISO / IEC 6048. The latest JPEG AI draft specification is included in the JPEG output document WG1N100602.
[0122] The latest JPEG AI draft specification utilizes some of the neural network-based image encoding and decoding methods mentioned above. The following describes or outlines some features of the latest JPEG AI specification.
[0123] 2.8.1 Quantized Convolution Two-dimensional quantized convolution is represented as The convolution process receives data of size 1. of Integer Tensor, and output size is of Integer Tensor, where ,factor Known as stride In the absence of stride In the case of this parameter, its default value is 1, which means that no spatial resolution change is performed. `is` is a non-negative integer, defining the maximum amplitude of the input tensor elements after amplitude limiting. Tensor Includes descaling shift for each channel of the output tensor.
[0124] The operation in quantized convolution can be described as a three-step process:
[0125]
[0126] in" "is the core size is" 2D cross-correlation operators.
[0127]
[0128] Shape tensor Includes learnable Integer Weights, shape is tensor Includes learnable Integer Bias. All parameters These are all part of a learnable quantization model.
[0129] Quantization model parameter limit d Descaling and shifting The combination of amplitude allows the control register Bit depth.
[0130] 2.8.2 Inverse Integer Functions It is represented as deinterger(wP, A). The inputs to the function are the bit depth wP and the integer input value A.
[0131] The function outputs a floating-point value, out.
[0132] The variables add and max are assigned as: add = 2(wP-5) , max = 2 wP – 1, The output 'out' is set to equal to (A – max / 2) ÷ add.
[0133] 2.8.3 Bitwise Operators The following bitwise operators are defined as follows:
[0134] 2.8.4 Decoder-side upsampling for minor components Based on the output image and the ratio between the size of the primary component and the size of the secondary component in the encoded and decoded image format, the scaling factor is applied to the vertical direction of the secondary component. ) or horizontal direction (with scaling factor) Perform upsampling. Output image format and corresponding principal components ( ) and secondary components ( The supported combinations of sizes and component scaling factors are listed in Table 1.
[0135] Table 1. Supported combinations of output image formats and scaling factors
[0136] If scaling factor Then upsampling is performed on the minor components. By default, upsampling is bicubic: For c=0..1, i=0, ... -1, j=0, ... -1:
[0137] If the adaptive upsampler is enabled (EFE_enabled_flag is true), the primary component-guided adaptive upsampler is executed.
[0138] If the adaptive upsampler is enabled (EFE_enabled_flag is false), then after bicubic upsampling... Switch to the inter-frame channel related information filter.
[0139] 2.8.5 bitstream structure The structure of a JPEG AI bitstream (also known as a codec stream or codec stream) consists of six parts with byte boundaries: 1. SOC - Start of Stream Marker; 2. PIH (Picture Header Marker), followed by the picture header; 3. TOH (Tool Head Marker), followed by tool information; 4. SOZ (Start of Z-stream marker), followed by the encoded / decoded stream of the super-prior information tensor z, including... ; 5. SORp (Start of Major Component Residual Stream), followed by the encoded / decoded stream of the major component residual, which includes... ; 6. SORs (Small Component Orders) followed by the codec stream of minor component residuals, which includes... ; 7. EOC - End of Encoder / Decoder Stream Marker.
[0140] The overall grammatical structure of an image is as follows:
[0141] Each codec stream begins with a tag. All tags used in this specification are as follows:
[0142] 2.8.6 Image Header This substream contains information about the image height. ,width Potential space piece location and size, control flags for each tool, scaling factors for primary and secondary components, - Learnable model index and displacement for rate control parameters (for principal components) and for secondary components (information).
[0143] The syntax and semantics are as follows: syntax table
[0144] Resolution changer syntax
[0145] Model head
[0146]
[0147]
[0148]
[0149]
[0150]
[0151] The following service information is transmitted via signaling: PIH is a 32-bit tag that includes the type (first two bytes) and the image header size (last two bytes). Adding 64 to img_width specifies the width of the input image (from 64 to 65600). img_height plus 64 specifies the height of the input image (from 64 to 65600); picture_format is the data format of the output image (YUV420 = 0, YUV444 = 1, sRGB = 2, YUV444 = 3). bit_depth is the bit depth of the output image ("0" corresponds to 8, "1" corresponds to 10); res_changer_enable is an enable flag for the resolution changer tool.
[0152] scale_comp[2] is the size of the primary and secondary components of the encoded and decoded image. The vertical and horizontal ratios between them; if they do not exist (res_changer_enable = false), then .
[0153] independent_beta_uv is a rate control parameter indicating the primary and secondary components. Whether the signs are the same (false / true).
[0154] beta_displacement_log_y – A parameter indicating the ratio between the rate control parameter beta selected by the encoder for the main components and the rate control parameter beta used in model training.
[0155] betaDisplacementLogY = beta_displacement_log_y – 2 11 beta_displacement_log_uv – A parameter indicating the ratio between the rate control parameter beta selected by the encoder for minor components and the rate control parameter beta used in model training.
[0156] betaDisplacementLogUV = beta_displacement_log_uv – 2 11 model_id is an identifier for a pre-stored checkpoint with the model's weights; model_id = 0, 1, 2, 3, or 4.
[0157] opIdx is an identifier for the operation point; 0 means "basic" and 1 means "high" operation point.
[0158] tile_enable_Luma and tile_enable_Chroma are enable flags for tiles of the primary and secondary components.
[0159] tile_size_Luma and tile_size_Chroma are the tile sizes for the primary and secondary components.
[0160] tile_overlap_Luma and tile_overlap_Luma are the dimensions of the tile overlap area for the primary and secondary components.
[0161] `cube_group_flag` is a 1-bit unsigned integer. `cube_group_flag=0` indicates that no `cube_flag` is signaled in a group, and the cube flag of a group is set to 1. `cube_group_flag=1` indicates that the cube flag of a group is signaled.
[0162] cube_luma_flag is a size of A 1D array containing cube flags for the principal components. A 1 indicates that a skip mode is applied to a cube of the residual tensor of the principal components. 0 indicates that a cube of the residual tensor of the principal components disables skip mode. .
[0163] cube_chroma_flag is a value of (( A 1D array containing cube flags for minor components. A 1 indicates that a skip mode is applied to a cube of the residual tensor of the minor component. 0 indicates that a cube of the residual tensor for minor components disables skip mode. .
[0164] color_transform_enable is an enable flag for the color transformation module.
[0165] `color_transform_matrix[i][j]` is the color transformation matrix. If it does not exist (`color_transform_enable` is false), the default ITU-R BT.709 color transformation is used.
[0166] `color_transform_offset[i]` is the offset for the color transformation. If it does not exist (`color_transform_enable` is false), the default ITU-R BT.709 color transformation is used.
[0167] 2.8.7 Tool Head This optional substream contains information about the tool.
[0168] Tool Header Syntax Table
[0169]
[0170]
[0171]
[0172]
[0173] On the first call to decodeFilters(), the values of fl[0], fl[1], best_cand_idx[0], and best_cand_idx[1] are restricted to 1.
[0174] TOH is a 32-bit tag that includes the type (first two bytes) and the tool header size (last two bytes). EFE_upsampler_enabled_flag – A flag indicating whether the EFE luminance-assisted upsampling process is enabled.
[0175] EFE_nonlinear_filter_enabled_flag – A flag indicating whether the EFE nonlinear filtering process is enabled.
[0176] best_cand_idx[0] – Specifies a 4-bit non-negative integer value corresponding to the candidate index of the u-component (the first of the minor components), indicating the number of slices and slice coordinates. It is used as input to the cand[X][Y] table.
[0177] best_cand_idx[1] – Specifies a 4-bit non-negative integer value corresponding to the candidate index of the v component (the second of the minor components), indicating the number of slices and slice coordinates. It is used as input to cand[X][Y]..
[0178] fl[0] – A non-negative integer value of 9 that specifies the core size. The value of fl[0] is restricted to be less than 4 and greater than 0.
[0179] fl[1] – A non-negative integer value of 9 that specifies the core size. The value of fl[1] is restricted to be less than 4 and greater than 0.
[0180] W 1A - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights).
[0181] W 1B - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights).
[0182] W 4A - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights).
[0183] W 4B - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights).
[0184] bS – A 10-bit non-negative integer value that specifies the block size of the EFE output adjustment subprocess.
[0185] minSymbol – Specifies the minimum 16-bit non-negative integer value to be used in the deinteger() function.
[0186] maxSymbol – A 16-bit non-negative integer value that specifies the maximum coefficient value.
[0187] mask1_enabled_flag – Specifies whether the values of len_mask_1_x and len_mask_1_y are zero or a 1-bit non-negative integer greater than zero.
[0188] mask2_enabled_flag – Specifies whether the values of len_mask_2_x and len_mask_2_y are zero or a 1-bit non-negative integer greater than zero.
[0189] B1 – Specifies the 16-bit value of the bias (additional component).
[0190] nonLinear_enabled_U_flag – An on / off switch for the nonlinear filtering process of the U component.
[0191] nonLinear_enabled_V_flag – An on / off switch for the nonlinear filtering process of the V component.
[0192] nonlinear_width – The width of the weight tensor in the nonlinear filtering process.
[0193] nonlinear_height – The height of the weight tensor in the nonlinear filtering process.
[0194] 3. Enhancement Filter Techniques in Existing Designs In total, it includes four different enhancement filter techniques with different functions: - Adaptive upsampling, – Inter-frame channel related information filter – Nonlinear chroma enhancement filter, – Luminance edge filter.
[0195] Data streams such as Figure 9 As shown. The process starts from targeting the main components. and secondary components The output of the synthesis transform begins. The output of the filter processing block is the reconstructed color components. It performs inverse color transformation.
[0196] Set to equal to Set to equal to ,and They were set to equal to They are used in the remaining sub-parts.
[0197] Note – When the encoded and decoded image format is 4:2:0 and the output image format is 4:4:4, The values are 1, 1, 2, 2, 1 and 1 respectively.
[0198] Figure 9 The enhancement filter technique is illustrated.
[0199] The size of the tensor, which depends on the output and the encoded / decoded image format, is defined in Table 1.
[0200] 3.1 Adaptive Upsampler This section details the primary component-guided adaptive upsampling process. This process utilizes information from the primary component to enhance the secondary components (color information plane) of the image.
[0201] If EFE_upsampler_enabled_flag is true, the process is enabled.
[0202] The input to this process is (Output of the synthesis transformation for the main components). (Output of the synthesis transformation for minor components). The output of this process is Enhanced secondary components ,in Go to the ICCI filter block, and Go to the nonlinear filter block.
[0203] Figure 10 An example implementation of a primary component-guided adaptive upsampling filter is shown.
[0204] If EFE_upsampler_enabled_flag equals 0, then Upsampling is performed via bicubic interpolation. Otherwise, the following ordered steps are executed: Invoke the parsing process based on a parsing table such as a predefined parsing table in the standard to obtain... .
[0205] Invoke the slice processing as described, where the parsed syntax elements are taken as input and the Tile1 tensor is taken as output.
[0206] Invoke the update procedure as specified in the parameters, where As input, and modified As output.
[0207] [2] Vector from [2, The middle channel has been removed.
[0208] They were set to equal to .
[0209] For x = 0.. (W+1) / 2-1, y = 0…(H+1) / 2-1, k = 0..1, i=0… -1, j=0… -1, perform the following operations:
[0210]
[0211]
[0212]
[0213] [k] + [k] They were set to equal to .
[0214] 3.1.1 Adaptive Upsampler Parameter Update Process The input to this process is a weight tensor. .
[0215] The output of this process is the modified weight tensor. .
[0216] Weight Tensor The following changes have been made: For i = 0..3, j = 0..3, ch = 0..3, comp = 0…1, cand = 0..5;
[0217]
[0218] Among them 3D The tensor is obtained as follows: Tensors are initialized by setting all their elements to zero.
[0219] The tensor is set to .
[0220] For i = 0…3 and j = 0…3, o , o , o .
[0221] 3.1.2 Adaptive Upsampler Chip Processing The input to this process is the width and height of the brightness reconstruction tensor, and the output is the image header parsing process.
[0222] The output of this process is the Tile1 tensor.
[0223] For comp = 0…1, y = 0… (H+1) / 2 , x = 0…( W+1) / 2 X is set to be equal to best_cand_idx2[comp].
[0224] For tileIdx = 0…cand[X][1]; Perform the following assignment.
[0225]
[0226]
[0227]
[0228]
[0229] if Then Tile1[comp,y,x] is set to equal tileIdx.
[0230] The cand[X][Y][4] table includes the number of slices and the coordinates of the slices.
[0231]
[0232] 3.2 Nonlinear Chromaticity Enhancement Filter This section details nonlinear filters used for minor component enhancement.
[0233] If EFE_nonlinear_filter_enabled_flag equals 1, then this procedure is invoked.
[0234] – The input to this process is the output of the ICCI process. and as the second output of the adaptive upsampling process as well as (Output synthesis transformation).
[0235] – The output of this process is .
[0236] Figure 11 An example implementation of an EFE nonlinear filter is shown.
[0237] Perform the following sequential steps: As described in Section 3.2.1, the nonlinear filter parameter slice is invoked, with the parsed syntax elements as input and the Tile2 tensor as output.
[0238] Addition bias parameters [8] was obtained as follows:
[0239] NLEnable[0] is set to equal nonLinear_enabled_U_flag, and NLEnable[1] is set to equal nonLinear_enabled_V_flag.
[0240] It was obtained using the nearest neighbor downsampling process, where As input, and As the downsampling ratio.
[0241] For x = 0... -1, y is at 0... In -1, and k is in 0..1, perform the following operations:
[0242] For x in the range of 0... In -1, y is at 0... In -1, perform the following operations:
[0243]
[0244] 3.2.1 Nonlinear Filter Parameter Processing The input to this process is nonlinear H, nonlinear W, H UV and W UV .
[0245] The output of this process is a Tile2 tensor.
[0246] The Tile2 tensor is obtained as follows: NumVerSplit = ( H UV + nonlinearH -1) / nonlinearH, .
[0247] 4 questions Existing designs for JPEG AI include enhancement filter techniques as described in Section 3, which improve the quality of reconstructed images. However, for practical applications, there is still room for further performance improvements, which are listed below: 1) The entire enhancement filter is not integerized, which may cause device interoperability issues.
[0248] 2) In the adaptive upsampler, the number of Cands can be further increased to improve performance.
[0249] 3) In addition, different filter sizes can be selected in the adaptive upsampler to improve encoding and decoding performance.
[0250] 4) Furthermore, in nonlinear filters, there is still room for reducing the complexity of the structure.
[0251] 5) In addition, the chip size used in the adaptive upsampler and nonlinear filter is not a multiple of 64.
[0252] 6) To improve the performance of the data format upsampling process (e.g., YUV420 to YUV444, YUV420 to YUV422, YUV422 to YUV444, etc.), an adaptive upsampler with fixed upsampling weights will be used.
[0253] 5 Detailed Solutions To address the above-mentioned problems and some other issues not mentioned, a method outlined below is disclosed. The detailed embodiments below should be considered as examples illustrating general concepts. These embodiments should not be interpreted narrowly. Furthermore, these embodiments can be combined in any manner.
[0254] 1) To address the first problem, a quantization strategy is applied to the adaptive upsampler and the nonlinear filter. All operations in the adaptive upsampler are integers, all accumulators in the computation are in the range of 32-bit integers, and the input is... Bits, the weights of the convolution operation are quantized to A bit integer, and the output is Bits. This ensures the precise bit-based behavior of the neural network. Quantized convolutions are applied in the adaptive upsampling unit.
[0255] a. In one example, the weights of the convolution transmitted in the bitstream are integers with bias. In order to recover the weights after decoding, a bias shift function is applied in these cases.
[0256] i. In one example, the bias shift function is expressed as The function's input is bits. deep Spend and Integer input value The function outputs an integer value, out.
[0257] - The variable max is assigned as: max = 2 wP – 1, The output 'out' is set to equal to .
[0258] 2. In one example, It can be any integer between 1 and 32.
[0259] b. In one example, It is quantized to support the entire integerization process of the adaptive upsampler.
[0260] i. In one example, let B_bit be used for storage. The bits of the elements. In Figure 10 Before using subtraction, first, scale using bitwise operations. It is expressed as: B1[2]<<( ) 1. In one example, It can be 16, and B_bit can be 15.
[0261] c. In one example, instead of using Figure 10 In the convolutions, quantized convolutions (as mentioned in Section 2.8.1) are used, making... This represents the input bits.
[0262] i. In one example, the threshold value d can be 2 << ( -1).
[0263] ii. In one example, such that The bits representing the quantized model parameters, and The bits representing the unquantized model parameters, downscaling shift p can be obtained by... Sure.
[0264] 1. In one example, It can be 2, and It can be 11, and p is 9.
[0265] 2. In one example, It can be an integer value between 1 and 32.
[0266] 3. In one example, It can be an integer value between 1 and 32.
[0267] iii. In one example, such that Indicates that it was used for storage The bits of the elements, the bias of the convolution can be determined by... <<( Scaling.
[0268] 1. In one example, It's 16. It's 15. It is 11, and It is 2, therefore Will be scaled to <<10.
[0269] d. In one example, the adaptive upsampler parameter update process needs to be integerized, and the modification process is provided as follows: Weight Tensor The following changes have been made: For i = 0..3, j = 0..3, ch = 0..3, comp = 0…1, cand = 0..5;
[0270]
[0271] Among them 3D The tensor is obtained as follows: Tensors are initialized by setting all their elements to zero.
[0272] The tensor is set to .
[0273] For i = 0…3; o 512.
[0274] if Not equal to and equal For i = 0…3; o , o .
[0275] if Not equal to and equal For i = 0…3; o , o .
[0276] if Not equal to and Not equal to For i = 0…3 and j = 0…3, o , o , o .
[0277] 2) To address the second issue, the Cand table (in Section 3.1.2) can be supplemented with more potential slice candidates to further improve performance.
[0278] a. In one example, cand[X][1] can be 8.
[0279] i. In one example, cand[X][2..9][4] could be [[0, 0.25, 0, 0.5], [0.25, 0.5, 0, 0.5], [0.5, 0.75, 0, 0.5], [0.75, 1, 0, 0.5], [0, 0.25, 0.5, 1], [0.25, 0.5, 0.5, 1], [0.5, 0.75, 0.5, 1], [0.75, 1, 0.5, 0.25]].
[0280] ii. In one example, cand[X][2..9][4] could be [[0, 0.5, 0, 0.25], [0, 0.5, 0.25, 0.5], [0, 0.5, 0.5, 0.75], [0, 0.5, 0.75, 1], [0.5, 1, 0, 0.25], [0.5, 1, 0.25, 0.5], [0.5, 1, 0.5, 0.75], [0.5, 1, 0.75, 1]].
[0281] b. In one example, cand[X][1] can be 9.
[0282] i. In one example, cand[X][2..10][4] could be [0.00, 0.33, 0.00, 0.33], [0.00, 0.33, 0.33, 0.67], [0.00, 0.33, 0.67, 1.00], [0.33, 0.67, 0.00, 0.33], [0.33, 0.67, 0.33, 0.67], [0.33, 0.67, 0.67, 1.00], [0.67, 1.00, 0.00, 0.33], [0.67, 1.00, 0.33, 0.67], [0.67, 1.00, 0.67, 1.00].
[0283] c. In one example, cand[X][1] can be 10.
[0284] i. In one example, `cand[X][2..11][4]` could be `[0.00, 0.50, 0.00, 0.20]`, `[0.00, 0.50, 0.20, 0.40]`, `[0.00, 0.50, 0.40, 0.60]`, `[0.00, 0.50, 0.60, 0.80]`, `[0.00, 0.50, 0.80, 1.00]`, `[0.50, 1.00, 0.00, 0.20]`, `[0.50, 1.00, 0.20, 0.40]`, `[0.50, 1.00, 0.40, 0.60]`, `[0.50, 1.00, 0.60, 0.80]`. [0.50, 1.00, 0.80, 1.00].
[0285] ii. In one example, `cand[X][2..11][4]` could be `[0.00, 0.20, 0.00, 0.50]`, `[0.00, 0.20, 0.50, 1.00]`, `[0.20, 0.40, 0.00, 0.50]`, `[0.20, 0.40, 0.50, 1.00]`, `[0.40, 0.60, 0.00, 0.50]`, `[0.40, 0.60, 0.50, 1.00]`, `[0.60, 0.80, 0.00, 0.50]`, `[0.60, 0.80, 0.50, 1.00]`, `[0.80, 1.00, 0.00, 0.50]`. [0.80, 1.00, 0.50, 1.00].
[0286] d. In one example, cand[X][1] can be 16.
[0287] i. In one example, `cand[X][2..17][4]` could be `[0.00, 0.50, 0.00, 0.12]`, `[0.00, 0.50, 0.12, 0.25]`, `[0.00, 0.50, 0.25, 0.38]`, `[0.00, 0.50, 0.38, 0.50]`, `[0.00, 0.50, 0.50, 0.62]`, `[0.00, 0.50, 0.62, 0.75]`, `[0.00, 0.50, 0.75, 0.88]`, `[0.00, 0.50, 0.88, 1.00]`, `[0.50, 1.00, 0.00, 0.12]`. [0.50, 1.00, 0.12, 0.25],[0.50, 1.00, 0.25, 0.38], [0.50, 1.00, 0.38, 0.50], [0.50, 1.00, 0.50, 0.62],[0.50, 1.00, 0.62, 0.75], [0.50, 1.00, 0.75, 0.88], [0.50, 1.00, 0.88, 1.00].
[0288] ii. In one example, cand[X][2..17][4] could be [0.00, 0.12, 0.00, 0.50], [0.00, 0.12, 0.50, 1.00], [0.12, 0.25, 0.00, 0.50], [0.12, 0.25, 0.50, 1.00], [0.25, 0.38, 0.00, 0.50], [0.25, 0.38, 0.50, 1.00], [0.38, 0.50, 0.00, 0.50], [0.38, 0.50, 0.50, 1.00], [0.50, 0.62, 0.00, 0.50], [0.50, 0.62, 0.50, 1.00],[0.62, 0.75, 0.00, 0.50], [0.62, 0.75, 0.50, 1.00], [0.75, 0.88, 0.00, 0.50],[0.75, 0.88, 0.50, 1.00], [0.88, 1.00, 0.00, 0.50], [0.88, 1.00, 0.50, 1.00].
[0289] iii. In one example, `cand[X][2..17][4]` could be `[0.00, 0.25, 0.00, 0.25]`, `[0.00, 0.25, 0.25, 0.50]`, `[0.00, 0.25, 0.50, 0.75]`, `[0.00, 0.25, 0.75, 1.00]`, `[0.25, 0.50, 0.00, 0.25]`, `[0.25, 0.50, 0.25, 0.50]`, `[0.25, 0.50, 0.50, 0.75]`, `[0.25, 0.50, 0.75, 1.00]`, `[0.50, 0.75, 0.00, 0.25]`. [0.50, 0.75, 0.25, 0.50],[0.50, 0.75, 0.50, 0.75], [0.50, 0.75, 0.75, 1.00], [0.75, 1.00, 0.00, 0.25],[0.75, 1.00, 0.25, 0.50], [0.75, 1.00, 0.50, 0.75], [0.75, 1.00, 0.75, 1.00].
[0290] 3) To address the third issue, the syntax `fl` will be replaced with `fId` to support decoding different shapes. The modified syntax parsing table is shown below.
[0291]
[0292]
[0293] a. In one example, fId_fl_table is a predefined table.
[0294] i. In one example, fId_fl_table is a subset of [(1,1), (2,2), (3,3), (4,4), (2,1), (1,2), (3,1), (1,3), (2,3), (3,2), (1,4), (4,1), (2,4), (4,2), (3,4), (4,3)].
[0295] 1. In one example, fId_fl_table is [(1,1), (2,2), (3,3), (4,4)].
[0296] 2. In one example, fId_fl_table is [(1,1), (2,2), (2,1), (1,2)].
[0297] 3. In one example, fId_fl_table is [(1,1), (2,2), (4,1), (1,4)].
[0298] 4. In one example, fId_fl_table is [(1,1), (2,2), (2,1), (1,2), (4,1), (1,4)].
[0299] 5. In one example, fId_fl_table is [(1,1), (2,2), (3,3), (4,4), (2,1), (1,2), (3,1), (1,3), (2,3), (3,2), (1,4), (4,1), (2,4), (4,2), (3,4), (4,3)].
[0300] b. In one example, the elements of fId are 9-valued non-negative integers specifying the kernel size. This value is restricted to be less than 4 and greater than 0.
[0301] i. In one example, to support more possible filter sizes, this value can be limited to less than 16 and greater than 0.
[0302] ii. In one example, to support more possible filter sizes, this value can be limited to less than 6 and greater than 0.
[0303] iii. In one example, to support more possible filter sizes, this value can be limited to less than 8 and greater than 0.
[0304] iv. In one example, the maximum value of fId can be a value between 1 and 16.
[0305] c. In one example, to save on the bit cost of fId, the elements of fId are non-negative integer values of N, which are restricted to being less than N and greater than 0.
[0306] 4) To address the fourth problem, different modules will be removed to reduce the complexity of the adaptive upsampler and the nonlinear chroma enhancement filter.
[0307] a. In one example, the ReLU() operation included in nonlinear chroma enhancement will be removed to make it more quantization-friendly for the network.
[0308] b. In one example, It will be used directly as the output of nonlinear chroma enhancement.
[0309] i. In one example, It will not be calculated in the upsampler, and related operations will be disabled.
[0310] ii. In one example, the first branch of the non-linear chroma enhancement will be removed, and the framework of the non-linear chroma enhancement can be... Figure 12 The structure shown illustrates an exemplary implementation of a simplified EFE nonlinear filter.
[0311] 5) To address the fifth issue, the chip processing used in the adaptive upsampler and nonlinear chroma enhancement filter needs to be modified to ensure that the chip size is a multiple of 64.
[0312] a. In one example, in adaptive upsampler slice processing (Section 3.1.2), for tileIdx = 0…cand[X][1], the following assignment is made: , , , .
[0313] b. In one example, in a nonlinear chroma enhancement filter (Section 3.2.1), It must be a multiple of 64.
[0314] 6) To address the sixth problem, an adaptive upsampler with fixed weights was applied as a replacement for the upsampling algorithm.
[0315] a. In one example, referring to Table 1, if the scaling factor ,but Instead of using upsampling to scale the shape of the input, an adaptive upsampler with fixed weights is applied if the result is false.
[0316] i. In one example, an adaptive upsampler with fixed weights can implement the upsampling process in an adaptive upsampler using predefined weights, and does not need to transmit the adaptive weights over a bitstream.
[0317] 1. In one example, [1,2,4,4]、 The samples in [1,2,4,4] were set to equal to .
[0318] 2. In one example, 3D The tensor is obtained as follows: • Tensors are initialized by setting all their elements to zero.
[0319] • The tensor is set to .
[0320] • For i = 0…3; o 512.
[0321] • if Not equal to and equal For i = 0…3; o , o .
[0322] • if Not equal to and equal For i = 0…3; o , o .
[0323] • if Not equal to and Not equal to For i = 0…3 and j = 0…3, o , o , o .
[0324] b. In one example, this process also requires an adaptive convolution process with quantization. Detailed quantization solutions can be found in Project 1).
[0325] 7) All of the above examples can be combined in any way to achieve a better trade-off between performance and complexity.
[0326] 6 Examples The following are some example implementations of the detailed solutions outlined in Section 5 of the previous article.
[0327] Figure 13 An example implementation of an adaptive upsampling filter guided by the quantized primary component is shown. Figure 13 An adaptive upsampling filter guided by the main quantized component is shown.
[0328] Figure 14 An example implementation of a quantized EFE nonlinear filter is shown. Figure 14 A quantized EFE nonlinear filter is shown.
[0329] Figure 15 A simplified example implementation of a principal component-guided adaptive upsampling filter is shown. Figure 15 A simplified main component-guided adaptive upsampling filter is shown.
[0330] Figure 16 A simplified example implementation of an EFE nonlinear filter is shown. Figure 16 A simplified EFE nonlinear filter is shown.
[0331] Further details of embodiments of this disclosure, relating to neural network-based visual data encoding and decoding, will be described below. As used herein, the term "visual data" may refer to images, pictures in videos, or any other visual data suitable for encoding and decoding.
[0332] As discussed above, in existing designs for visual data encoding and decoding based on neural networks (NNs), the chip size used in adaptive upsamplers and nonlinear filters is not required to be a multiple of a predetermined value. This makes chip processing of adaptive upsamplers and nonlinear filters poorly controlled and increases computational complexity. Consequently, encoding and decoding efficiency is reduced.
[0333] To address the above-mentioned problems and other issues not mentioned, a visual data processing solution described below is disclosed. The embodiments of this disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0334] Figure 17 A flowchart of a method 1700 for visual data processing according to some embodiments of the present disclosure is shown. Method 1700 may be implemented during the conversion between visual data and a bitstream of visual data utilizing a neural network (NN)-based model. As used herein, the NN-based model can be a model based on neural network techniques. For example, the NN-based model may specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module may include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the NN-based model described herein are merely illustrative and should not be construed as limiting the present disclosure in any way.
[0335] like Figure 17 As shown, method 1700 begins at 1702, where a first tensor used in the adaptive filter of the NN-based model is divided into a first set of pieces based on a first piece size. The first piece size is a multiple of a first predetermined value. For example, the first tensor may be a brightness reconstruction tensor, an input tensor for the adaptive filter, etc. The first piece size may include the width and / or height of the piece. By way of example and not limitation, the first predetermined value is 32, 64, etc. For example, both the width and height of the first piece size must be multiples of 64. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect. Furthermore, the adaptive filter may be an adaptive upsampler, an adaptive linear filter, etc.
[0336] In some embodiments, the boundary of each piece in the first set of pieces can be determined based on a first predetermined value. For example, with the first predetermined value being 64, the position of the boundary of each piece with index tileIdx = 0…cand[X][1] is calculated as follows: , , , , Where lowH represents the lower boundary in the vertical direction, upperH represents the upper boundary in the vertical direction, lowW represents the lower boundary in the horizontal direction, upperW represents the upper boundary in the horizontal direction, the value of cand[X][1] is defined in the table in Section 3.1.2, H represents the height of the first tensor, and W represents the width of the first tensor.
[0337] Alternatively, the floor() function can be used to determine the location of the boundary, as shown below: 32, , , .
[0338] It should be noted that the boundary of each piece in the first set of pieces can be determined based on the first predetermined value in any other suitable manner. The scope of this disclosure is not limited in this respect.
[0339] It should be understood that only a majority of the slices in the first set can be guaranteed to have the first slice size. For the slice(s) located at the boundaries of the first tensor, their size may differ from the first slice size, depending on the original size of the first tensor. For example, if the size of the input tensor is a multiple of a first predetermined value, then each slice in the first set can have the first slice size.
[0340] At 1704, the conversion is performed based on the first set of slices. In some embodiments, the conversion may include encoding visual data into a bitstream. Additionally or alternatively, the conversion may include decoding visual data from the bitstream. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0341] In light of the above, the first tensor used in the adaptive filter of the NN-based model is divided into a first group of pieces based on a first piece size, and the first piece size is a multiple of a first predetermined value. Compared to traditional solutions that do not restrict the piece size, the proposed method can advantageously ensure that the size of most pieces is a multiple of the predetermined value. In this way, encoding and decoding efficiency can be improved.
[0342] In some embodiments, the second tensor used in the nonlinear filter of the NN-based model may be divided into a second set of pieces based on a second piece size, and the second piece size may be a multiple of a second predetermined value. For example, the second tensor may include the input tensor for the nonlinear filter, etc. The second piece size may include the width and / or height of the piece. By way of example and not limitation, the second predetermined value is 32, 64, etc. Additionally, the nonlinear filter may be a nonlinear chroma enhancement filter, etc.
[0343] In some embodiments, the size of the weight tensor for at least one of the second set of slices can be indicated in the bitstream, and the size of the weight tensor can be a multiple of a second predetermined value. For example, the size of the weight tensor can include the width and / or the height of the weight tensor. By way of example and not limitation, a first syntax element (e.g., denoted as nonlinearH or nonlinear_height) can indicate the height of the weight tensor of the nonlinear filtering process, and the value of the first syntax element must be a multiple of a second predetermined value (such as 64). Additionally or alternatively, a second syntax element (e.g., denoted as nonlinearW or nonlinear_width) can indicate the width of the weight tensor of the nonlinear filtering process, and the value of the second syntax element must be a multiple of a second predetermined value (such as 64).
[0344] In some embodiments, adaptive filters can be integerized. For example, all operations of an adaptive filter can be integer operations, and all values involved in the adaptive filter can be integer values. Additionally or alternatively, nonlinear filters can be integerized. For example, all operations of a nonlinear filter based on a neural network model can be integer operations, and all values involved in the nonlinear filter can be integer values.
[0345] In some embodiments, the weights of convolution operations in adaptive filters and / or nonlinear filters can be indicated in the bitstream by integer values with bias. For example, the weights can be determined by applying a bias shift function to the integer values and the bias. As an example and not a limitation, the output of the bias shift function can be equal to (A - (2 wP - 1) / 2), where A represents an integer value, and wP represents the bias (also known as the bit depth). The bias wP can be an integer between 1 and 32. It should be noted that the bias shift function can also be called the inverse integer function.
[0346] In some embodiments, the offset term to be subtracted from the input of the adaptive filter can be quantized. As an example, and not a limitation, the offset term can be scaled as follows: B1[2]<<( ), in Indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the element in the offset entry. This is for illustrative purposes only and not as a limitation. It can be 16, and It can be 15. It should be understood that the specific values listed herein are intended to be exemplary and not to limit the scope of this disclosure.
[0347] In some embodiments, the adaptive filter and / or nonlinear filter may include a quantized convolution operation. For example, a quantized convolution operation may be used to replace the original convolution operation included in the adaptive filter and / or nonlinear filter. The quantized convolution operation has been described in detail in Section 2.8.1 above. By way of example, the value of the limiting parameter d used for the quantized convolution operation can be equal to 2 << ( -1), where This represents the number of bits in the input of the quantized convolution operation. Additionally or alternatively, the value of the downscaling shift parameter p used for the quantized convolution operation can be equal to... ,in This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters. For example, It can be an integer between 1 and 32, and / or It can be an integer between 1 and 32. In one example embodiment, It can be 2, It can be 11, therefore p can be 9.
[0348] In some embodiments, the offset term to be subtracted from the input of the adaptive filter can be scaled as follows: <<( ), in Indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the element of the offset. This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters. This is for illustrative purposes only and not as a limitation. It can be 16. It can be 15. It can be 11, and It can be 2. In this case, Will be scaled to <<10.
[0349] In some embodiments, the parameter update process for an adaptive filter can be integerized. In one example embodiment, each element of the Discrete Cosine Transform (DCT) tensor can be an integer. For example, as described above, A tensor can be set as .
[0350] In some embodiments, the number of slices included in at least one candidate segmentation pattern may be greater than 6. By way of example and not limitation, the number of slices included in at least one candidate segmentation pattern may be 8, 9, 10, 16, etc. In one example, the number of slices included in at least one candidate segmentation pattern may be 8. In this case, a possible set of coordinates for a slice candidate may be [[0, 0.25, 0, 0.5], [0.25, 0.5, 0, 0.5], [0.5, 0.75, 0, 0.5], [0.75, 1, 0, 0.5], [0, 0.25, 0.5, 1], [0.25, 0.5, 0.5, 1], [0.5, 0.75, 0.5, 1], [0.75, 1, 0.5, 0.25]]. Furthermore, another possible set of coordinates for a slice candidate could be [[0, 0.5, 0, 0.25], [0, 0.5, 0.25, 0.5], [0, 0.5, 0.5, 0.75], [0, 0.5, 0.75, 1], [0.5, 1, 0, 0.25], [0.5, 1, 0.25, 0.5], [0.5, 1, 0.5, 0.75], [0.5, 1, 0.75, 1]]. For illustrative purposes, possible coordinates for other slice candidates are listed in Section 5 above. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0351] In some embodiments, the bitstream may include a first indicator that indicates the index of the kernel size of the adaptive or nonlinear filter of the NN-based model within a set of kernel sizes. For example, the first indicator may be represented as fId. In an example embodiment, the set of kernel sizes may be predetermined and stored in a table. By way of example, the set of kernel sizes may be a subset of [(1,1), (2,2), (3,3), (4,4), (2,1), (1,2), (3,1), (1,3), (2,3), (3,2), (1,4), (4,1), (2,4), (4,2), (3,4), (4,3)].
[0352] In some embodiments, the elements of the core size can be 9-valued non-negative integers, and the element values must be less than 4 and greater than 0. Alternatively, the element values of the core size must be less than 16 and greater than 0. In another embodiment, the element values must be less than 6 and greater than 0. Alternatively, the element values must be less than 8 and greater than 0. In some other embodiments, the value of the first indication is required to be less than a predetermined value in order to save the bit cost of the first indication.
[0353] In some embodiments, the corrected linear unit (ReLU) may not be present in the nonlinear filter. Additionally or alternatively, the initial enhancement result (e.g., in...) Figure 11 The Chinese character is represented as This can be directly used as the output of a nonlinear filter. Additionally or alternatively, the enhanced secondary components of the inter-frame channel correlation information (ICCI) filter output from the NN-based model (e.g., in...) Figure 11 The Chinese character is represented as It can be excluded from input to the nonlinear filter, and in this case, the operations related to the enhanced minor component can be disabled.
[0354] In some other embodiments, the enhanced secondary components from the adaptive filter output (e.g., in...) Figure 11 The Chinese character is represented as This component may not be input to the nonlinear filter. In this case, the first branch used to process the enhanced minor component can be removed, and the resulting structure of the nonlinear filter is as follows: Figure 12 As shown.
[0355] In some embodiments, if image format upsampling is required and enhancement filter extension (EFE) luminance-assisted upsampling is disabled, an adaptive upsampler with fixed weights can be applied. For example, the fixed weights may be predetermined and not in the bitstream, i.e., not transmitted through the signal in the bitstream. Additionally, the adaptive upsampler may include a quantized adaptive convolution process as described in detail above.
[0356] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding quality.
[0357] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by means of an apparatus for visual data processing. The method includes: partitioning a first tensor used in an adaptive filter of a neural network (NN)-based model into a first set of pieces based on a first piece size, the first piece size being a multiple of a first predetermined value; and generating a bitstream based on the first set of pieces using the NN-based model.
[0358] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. The method includes: dividing a first tensor used in an adaptive filter of a neural network (NN)-based model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; generating a bitstream based on the first set of slices using the NN-based model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0359] Implementations of this disclosure may be described in accordance with the following entries, and its features may be combined in any reasonable manner.
[0360] Item 1. A method for visual data processing, comprising: for a conversion between visual data utilizing a neural network (NN)-based model and a bitstream of the visual data, dividing a first tensor used in an adaptive filter of the NN-based model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; and performing the conversion based on the first set of slices.
[0361] Item 2. The method according to Item 1, wherein the first tensor includes a brightness reconstruction tensor or an input tensor for the adaptive filter.
[0362] Item 3. The method according to any one of items 1 to 2, wherein the boundary of each piece in the first set of pieces is determined based on the first predetermined value.
[0363] Item 4. The method according to any one of items 1 to 3, wherein the first piece size includes at least one of the following: the width of the piece, or the height of the piece.
[0364] Item 5. The method according to any one of items 1 to 4, wherein each piece in the first set of pieces has the first piece size if the size of the input tensor is a multiple of the first predetermined value.
[0365] Item 6. The method according to any one of items 1 to 5, wherein the first predetermined value is 64.
[0366] Item 7. The method according to any one of items 1 to 6, wherein the adaptive filter is an adaptive upsampler.
[0367] Item 8. The method according to any one of Items 1 to 7, wherein the second tensor used in the nonlinear filter of the NN-based model is divided into a second set of pieces based on a second piece size, and the second piece size is a multiple of a second predetermined value.
[0368] Item 9. The method according to Item 8, wherein the second tensor includes the input tensor for the nonlinear filter.
[0369] Item 10. The method according to any one of Items 8 to 9, wherein the second piece size includes at least one of the following: the width of the piece, or the height of the piece.
[0370] Item 11. The method according to any one of Items 8 to 10, wherein the size of the weight tensor for at least one of the second set of slices is indicated in the bitstream, and the size of the weight tensor is a multiple of the second predetermined value.
[0371] Item 12. The method according to Item 11, wherein the size of the weight tensor includes at least one of the following: the width of the weight tensor, or the height of the weight tensor.
[0372] Item 13. The method according to any one of items 8 to 12, wherein the second predetermined value is 64.
[0373] Item 14. The method according to any one of items 8 to 13, wherein the nonlinear filter is a nonlinear chroma enhancement filter.
[0374] Item 15. The method according to any one of items 1 to 14, wherein all operations of the adaptive filter are integer operations and all values involved in the adaptive filter are integer values, and / or all operations of the nonlinear filter of the NN-based model are integer operations and all values involved in the nonlinear filter are integer values.
[0375] Item 16. The method according to Item 15, wherein the weights of the convolution operations in the adaptive filter and / or the nonlinear filter are indicated in the bitstream by integer values having bias.
[0376] Item 17. The method according to Item 16, wherein the weight is determined by applying a bias shift function to the integer value and the bias.
[0377] Item 18. The method according to Item 17, wherein the output of the bias shift function is equal to (A - (2 wP -1) / 2), where A represents the integer value and wP represents the bias.
[0378] Item 19. The method according to Item 18, wherein the bias wP is an integer between 1 and 32.
[0379] Item 20. The method according to any one of items 15 to 19, wherein the offset term to be subtracted from the input of the adaptive filter is quantized.
[0380] Item 21. According to the method described in Item 20, wherein the offset term is scaled as follows: B1[2]<<( ), in This indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the elements of the offset item.
[0381] Item 22. The method according to Item 21, wherein It is 16, and It is 15.
[0382] Item 23. The method according to any one of items 15 to 22, wherein the adaptive filter and / or the nonlinear filter comprises a quantized convolution operation.
[0383] Item 24. The method according to Item 23, wherein the value of the limiting parameter d for the convolution operation used in the quantization is equal to 2<<( -1), where This represents the number of bits in the input of the quantized convolution operation.
[0384] Item 25. The method according to any one of items 23 to 24, wherein the value of the downscaling shift parameter p for the convolution operation used in the quantization is equal to ,in This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters.
[0385] Item 26. The method according to Item 25, wherein It is an integer between 1 and 32, or It is an integer between 1 and 32.
[0386] Item 27. The method according to any one of items 25 to 26, wherein It is 2. It is 11, and p is 9.
[0387] Item 28. The method according to any one of items 23 to 27, wherein the offset term to be subtracted from the input of the adaptive filter is scaled as follows: <<( ), in This indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the elements of the offset item. This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters.
[0388] Item 29. The method according to Item 28, wherein It's 16. It's 15. It is 11, and It is 2.
[0389] Item 30. The method according to any one of items 15 to 29, wherein the parameter update process for the adaptive filter is integerized.
[0390] Item 31. The method according to Item 30, wherein each element of the Discrete Cosine Transform (DCT) tensor is an integer.
[0391] Item 32. The method according to any one of items 1 to 31, wherein at least one candidate segmentation pattern includes a number of slices greater than 6.
[0392] Item 33. The method according to Item 32, wherein the number of said slices included in at least one candidate segmentation pattern is 8, 9, 10 or 16.
[0393] Item 34. The method according to any one of items 1 to 33, wherein the bitstream includes a first indication for indicating an index of the kernel size of the adaptive filter or the nonlinear filter of the NN-based model within a set of kernel sizes.
[0394] Item 35. The method according to Item 34, wherein the set of core sizes is predetermined and stored in a table.
[0395] Item 36. The method according to Item 35, wherein the set of kernel sizes is a subset of [(1,1), (2,2), (3,3), (4,4), (2,1), (1,2), (3,1), (1,3), (2,3), (3,2), (1,4), (4,1), (2,4), (4,2), (3,4), (4,3)].
[0396] Item 37. The method according to any one of items 34 to 36, wherein the elements of the kernel size are non-negative integer values with a value of 9, and the value of the elements is less than 4 and greater than 0.
[0397] Item 38. The method according to any one of items 34 to 36, wherein the value of an element of the kernel size is less than 16 and greater than 0, or the value of the element is less than 6 and greater than 0, or the value of the element is less than 8 and greater than 0.
[0398] Item 39. The method according to any one of items 34 to 38, wherein the value of the first indication is less than a predetermined value.
[0399] Item 40. The method according to any one of items 8 to 39, wherein the modified linear unit (ReLU) is not in the nonlinear filter.
[0400] Item 41. The method according to any one of items 8 to 40, wherein the result of the initial enhancement is directly used as the output of the nonlinear filter.
[0401] Item 42. The method according to any one of items 8 to 41, wherein the enhanced secondary component output from the inter-frame channel correlation information (ICCI) filter in the NN-based model is not input to the nonlinear filter.
[0402] Item 43. The method according to any one of items 8 to 42, wherein the enhanced minor component output from the adaptive filter is not input to the nonlinear filter.
[0403] Item 44. The method according to any one of items 1 to 43, wherein if image format upsampling is required and the enhancement filter extension (EFE) luminance-assisted upsampling process is disabled, an adaptive upsampler with fixed weights is applied.
[0404] Item 45. The method according to Item 44, wherein the fixed weight is predetermined and is not in the bitstream.
[0405] Item 46. The method according to any one of items 44 to 45, wherein the adaptive upsampling includes a quantized adaptive convolution process.
[0406] Item 47. The method according to any one of items 1 to 46, wherein the visual data includes video, a picture of the video, or an image.
[0407] Item 48. The method according to any one of items 1 to 47, wherein the conversion includes encoding the visual data into the bitstream.
[0408] Item 49. The method according to any one of items 1 to 47, wherein the conversion includes decoding the visual data from the bitstream.
[0409] Item 50. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 49.
[0410] Item 51. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 49.
[0411] Item 52. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of an apparatus for visual data processing, wherein the method comprises: dividing a first tensor used in an adaptive filter of a neural network (NN) model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; and generating the bitstream based on the first set of slices using the NN-based model.
[0412] Item 53. A method for storing a bitstream of visual data, comprising: dividing a first tensor used in an adaptive filter of a neural network (NN) model into a first set of slices based on a first slice size, the first slice size being a multiple of a first predetermined value; generating the bitstream based on the first set of slices using the NN-based model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0413] Example device Figure 18 A block diagram of a computing device 1800 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1800 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).
[0414] It should be understood that, Figure 18 The computing device 1800 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0415] like Figure 18As shown, computing device 1800 includes general-purpose computing device 1800. Computing device 1800 may include at least one or more processors or processing units 1810, memory 1820, storage unit 1830, one or more communication units 1840, one or more input devices 1850, and one or more output devices 1860.
[0416] In some embodiments, the computing device 1800 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 1800 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0417] Processing unit 1810 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 1820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 1800. Processing unit 1810 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0418] Computing device 1800 typically includes various computer storage media. Such media can be any media accessible by computing device 1800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 1820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 1830 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or visual data and can be accessed within computing device 1800.
[0419] The computing device 1800 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 18 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to the bus (not shown) via one or more visual data media interfaces.
[0420] The communication unit 1840 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 1800 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, the computing device 1800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0421] Input device 1850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1860 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 1840, computing device 1800 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 1800 can also communicate with one or more devices that enable a user to interact with computing device 1800, or any device that enables computing device 1800 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0422] In some embodiments, some or all components of computing device 1800 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (WAN) such as the Internet using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations within remote visual data centers. Cloud computing infrastructure can provide services through shared visual data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0423] In embodiments of this disclosure, computing device 1800 may be used to implement visual data encoding / decoding. Memory 1820 may include one or more visual data encoding / decoding modules 1825 having one or more program instructions. These modules can be accessed and executed by processing unit 1810 to perform the functions of the various embodiments described herein.
[0424] In an example embodiment of performing visual data encoding, input device 1850 may receive visual data as input 1870 to be encoded. The visual data may be processed, for example, by a visual data encoding / decoding module 1825 to generate an encoded bitstream. The encoded bitstream may be provided as output 1880 via output device 1860.
[0425] In an example embodiment of performing visual data decoding, input device 1850 may receive an encoded bitstream as input 1870. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 1825 to generate decoded visual data. The decoded visual data may be provided as output 1880 via output device 1860.
[0426] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for visual data processing, comprising: For the conversion between visual data using a neural network (NN)-based model and the bitstream of said visual data, a first tensor used in the adaptive filter of the NN-based model is divided into a first set of pieces based on a first piece size, the first piece size being a multiple of a first predetermined value; and The conversion is performed based on the first set of slices.
2. The method of claim 1, wherein the first tensor comprises a brightness reconstruction tensor or an input tensor for the adaptive filter.
3. The method according to any one of claims 1 to 2, wherein the boundary of each piece in the first group of pieces is determined based on the first predetermined value.
4. The method according to any one of claims 1 to 3, wherein the size of the first piece comprises at least one of the following: The width of the piece, or The height of the piece.
5. The method according to any one of claims 1 to 4, wherein if the size of the input tensor is a multiple of the first predetermined value, then each piece in the first set of pieces has the first piece size.
6. The method according to any one of claims 1 to 5, wherein the first predetermined value is 64.
7. The method according to any one of claims 1 to 6, wherein the adaptive filter is an adaptive upsampler.
8. The method according to any one of claims 1 to 7, wherein the second tensor used in the nonlinear filter of the NN-based model is divided into a second set of pieces based on a second piece size, and the second piece size is a multiple of a second predetermined value.
9. The method of claim 8, wherein the second tensor comprises an input tensor for the nonlinear filter.
10. The method according to any one of claims 8 to 9, wherein the second piece size comprises at least one of the following: The width of the piece, or The height of the piece.
11. The method according to any one of claims 8 to 10, wherein the size of the weight tensor for at least one of the second set of slices is indicated in the bitstream, and the size of the weight tensor is a multiple of the second predetermined value.
12. The method of claim 11, wherein the size of the weight tensor comprises at least one of the following: The width of the weight tensor, or The height of the weight tensor.
13. The method according to any one of claims 8 to 12, wherein the second predetermined value is 64.
14. The method according to any one of claims 8 to 13, wherein the nonlinear filter is a nonlinear chroma enhancement filter.
15. The method according to any one of claims 1 to 14, wherein all operations of the adaptive filter are integer operations, and all values involved in the adaptive filter are integer values, and / or All operations of the nonlinear filter based on the NN model are integer operations, and all values involved in the nonlinear filter are integer values.
16. The method of claim 15, wherein the weights of the convolution operations in the adaptive filter and / or the nonlinear filter are indicated in the bitstream by integer values having bias.
17. The method of claim 16, wherein the weight is determined by applying a bias shift function to the integer value and the bias.
18. The method of claim 17, wherein the output of the bias shift function is equal to (A - (2 wP - 1) / 2), where A represents the integer value and wP represents the bias.
19. The method of claim 18, wherein the bias wP is an integer between 1 and 32.
20. The method according to any one of claims 15 to 19, wherein the offset term to be subtracted from the input of the adaptive filter is quantized.
21. The method of claim 20, wherein the offset term is scaled as follows: B1[2]<<( ), in This indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the elements of the offset item.
22. The method of claim 21, wherein It is 16, and It is 15.
23. The method according to any one of claims 15 to 22, wherein the adaptive filter and / or the nonlinear filter comprises a quantized convolution operation.
24. The method of claim 23, wherein the value of the limiting parameter d for the convolution operation used in the quantization is equal to 2 << ( -1), where This represents the number of bits in the input of the quantized convolution operation.
25. The method according to any one of claims 23 to 24, wherein the value of the downscaling shift parameter p for the convolution operation used in the quantization is equal to ,in This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters.
26. The method of claim 25, wherein It is an integer between 1 and 32, or It is an integer between 1 and 32.
27. The method according to any one of claims 25 to 26, wherein It is 2. It is 11, and p is 9.
28. The method according to any one of claims 23 to 27, wherein the offset term to be subtracted from the input of the adaptive filter is scaled as follows: <<( ), in This indicates the offset item. This indicates the number of bits in the input of the adaptive filter, and This indicates the number of bits used to store the elements of the offset item. This represents the number of bits in the quantized model parameters, and This represents the number of bits in the unquantized model parameters.
29. The method of claim 28, wherein It's 16. It's 15. It is 11, and It is 2.
30. The method according to any one of claims 15 to 29, wherein the parameter update process for the adaptive filter is integerized.
31. The method of claim 30, wherein each element of the discrete cosine transform (DCT) tensor is an integer.
32. The method according to any one of claims 1 to 31, wherein the number of slices included in at least one candidate segmentation pattern is greater than 6.
33. The method of claim 32, wherein the number of slices included in at least one candidate segmentation pattern is 8, 9, 10 or 16.
34. The method according to any one of claims 1 to 33, wherein the bitstream includes a first indication for indicating an index of the kernel size of the adaptive filter or the nonlinear filter of the NN-based model within a set of kernel sizes.
35. The method of claim 34, wherein the set of core sizes is predetermined and stored in a table.
36. The method of claim 35, wherein the set of kernel sizes is a subset of [(1,1), (2,2), (3,3), (4,4), (2,1), (1,2), (3,1), (1,3), (2,3), (3,2), (1,4), (4,1), (2,4), (4,2), (3,4), (4,3)].
37. The method according to any one of claims 34 to 36, wherein the elements of the kernel size are non-negative integer values with a value of 9, and the value of the elements is less than 4 and greater than 0.
38. The method according to any one of claims 34 to 36, wherein the value of the element of the kernel size is less than 16 and greater than 0, or The value of the element is less than 6 and greater than 0, or The value of the element is less than 8 and greater than 0.
39. The method according to any one of claims 34 to 38, wherein the value of the first indication is less than a predetermined value.
40. The method according to any one of claims 8 to 39, wherein the modified linear unit (ReLU) is not in the nonlinear filter.
41. The method according to any one of claims 8 to 40, wherein the result of the initial enhancement is directly used as the output of the nonlinear filter.
42. The method according to any one of claims 8 to 41, wherein the enhanced secondary component output from the inter-frame channel correlation information (ICCI) filter in the NN-based model is not input to the nonlinear filter.
43. The method according to any one of claims 8 to 42, wherein the enhanced minor component output from the adaptive filter is not input to the nonlinear filter.
44. The method according to any one of claims 1 to 43, wherein if image format upsampling is required and enhancement filter extension (EFE) luminance-assisted upsampling is disabled, an adaptive upsampler with fixed weights is applied.
45. The method of claim 44, wherein the fixed weight is predetermined and is not in the bitstream.
46. The method according to any one of claims 44 to 45, wherein the adaptive upsampling includes a quantized adaptive convolution process.
47. The method according to any one of claims 1 to 46, wherein the visual data includes video, a picture of the video, or an image.
48. The method of any one of claims 1 to 47, wherein the conversion comprises encoding the visual data into the bitstream.
49. The method according to any one of claims 1 to 47, wherein the conversion comprises decoding the visual data from the bitstream.
50. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 49.
51. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 49.
52. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of visual data processing, wherein the method includes: The first tensor used in the adaptive filter of the neural network (NN) based on the first piece size is divided into a first set of pieces, the first piece size being a multiple of a first predetermined value; as well as The bitstream is generated based on the first set of slices using the NN-based model.
53. A method for storing a bitstream of visual data, comprising: The first tensor used in the adaptive filter of the neural network (NN) based on the first piece size is divided into a first set of pieces, the first piece size being a multiple of a first predetermined value; The bitstream is generated based on the first set of slices using the NN-based model; as well as The bitstream is stored in a non-transitory computer-readable recording medium.