Method and device for visual data processing and medium
By using two different filtering processes to generate candidate reconstructions in image/video encoding and decoding, and combining them with a neural network model to generate a bitstream, the problem of insufficient encoding and decoding quality in existing technologies is solved, achieving higher encoding and decoding adaptability and quality.
Patent Information
- Application Number
- CN202480019805.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-29
- Filing Date
- 2024-03-22
- Publication Date
- 2025-11-04
AI Technical Summary
Existing neural network-based image/video encoding and decoding technologies still have room for improvement in encoding and decoding quality, especially in their ability to adapt to different content.
Two different filtering processes are used to generate candidate reconstructions of visual data, and a bitstream is generated by combining target reconstruction with a neural network model to improve encoding and decoding quality.
By utilizing candidate reconstructions generated through two different filtering processes, the content of the visual data can be better adapted, thereby improving the encoding and decoding quality.
Smart Images

Figure CN120898429A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to visual data processing techniques, and more specifically, to visual data encoding and decoding based on neural networks. Background Technology
[0002] Over the past decade, deep learning has seen rapid development across various fields, particularly in computer vision and image processing. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. In the past five years, neural network-based image / video compression techniques have made significant progress. It has been reported that the latest neural network-based image compression algorithms have achieved rate-distortion (RD) performance comparable to Multifunctional Video Coding (VVC). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, the encoding and decoding quality of neural network-based image / video codecs is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for visual data processing.
[0004] In a first aspect, a method for visual data processing is proposed. The method includes: for a conversion between visual data utilizing a neural network (NN)-based model and one or more bitstreams of visual data; determining a target reconstruction of a first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; and performing a conversion based on the target reconstruction.
[0005] According to the method of the first aspect of this disclosure, two different filtering processes are used to generate two candidate reconstructions of a component, and the two candidate reconstructions are further used to generate the target reconstruction of that component. Compared with conventional solutions that use only a single filtering process to generate the component reconstruction, the proposed method can advantageously utilize two different filtering processes to generate the component reconstruction. In this way, the encoding and decoding process can be adapted to the content of the visual data, thereby improving the encoding and decoding quality.
[0006] In a second aspect, an apparatus for visual data processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data processing. The method includes: determining a target reconstruction of the first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; and generating a bitstream based on the target reconstruction using a neural network (NN)-based model.
[0009] In a fifth aspect, a method for storing bitstreams of visual data is proposed. The method includes: determining a target reconstruction of a first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; generating a bitstream based on the target reconstruction using a neural network (NN)-based model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1A A block diagram of an example visual data encoding / decoding system according to some embodiments of the present disclosure is shown.
[0013] Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme.
[0014] Figure 2 An example potential representation of the image is shown.
[0015] Figure 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model.
[0016] Figure 4This is a schematic diagram illustrating an example combined model configured to jointly optimize the context model along with a super-prior and an autoencoder. Table 1 below shows the meaning of the different symbols.
[0017] Table 1 - Symbol Explanation
[0018]
[0019]
[0020] Figure 5 An example encoding process is shown.
[0021] Figure 6 An example decoding process is shown.
[0022] Figure 7 An example decoding process according to this disclosure is shown.
[0023] Figure 8 An example of a learning-based image codec architecture is shown.
[0024] Figure 9 An example synthetic transform for learning-based image encoding and decoding is shown. Figure 10 An example LeakyReLU activation function is shown.
[0025] Figure 11 An example ReLU activation function is shown.
[0026] Figure 12 Examples of pixel shuffling and unshuffling operations are shown.
[0027] Figure 13 An example of a transposed convolution with a 2x2 kernel is shown.
[0028] Figure 14 An example subnetwork of a neural network is shown.
[0029] Figure 15 An example subnetwork of a neural network is shown.
[0030] Figure 16 An example subnetwork of a neural network is shown.
[0031] Figure 17 An example subnetwork of a neural network is shown.
[0032] Figure 18 An example of the base weight (W_base[i,j]) value is shown.
[0033] Figure 19 An example of the value of W_base[i,j] is shown.
[0034] Figure 20 An example of the value of W_base[i,j] is shown.
[0035] Figure 21 An example of the value of W_base[i,j] is shown.
[0036] Figure 22 An example of the value of W_base[i,j] is shown.
[0037] Figure 23 An example of the value of W_base[i,j] is shown.
[0038] Figure 24 An example flowchart for executing the disclosed example is shown.
[0039] Figure 25 An example neural network configured to perform the disclosed example is shown.
[0040] Figure 26 An example neural network configured to perform the disclosed example is shown.
[0041] Figure 27 An example neural network configured to perform the disclosed example is shown.
[0042] Figure 28 An example implementation of an embodiment according to this disclosure is shown.
[0043] Figure 29 An example convolution process for obtaining component 1 is shown.
[0044] Figure 30 An example convolution process for obtaining component 1 is shown.
[0045] Figure 31 An example convolution process for obtaining component 1 is shown.
[0046] Figure 32 An example convolution process for obtaining component 1 is shown.
[0047] Figure 33 An example convolution process is shown to obtain component 1 and component 2.
[0048] Figure 34 An example convolution process is shown to obtain component 1 and component 2.
[0049] Figure 35 An example convolution process for obtaining component 1 is shown.
[0050] Figure 36 An example convolution process for obtaining component 1 is shown.
[0051] Figure 37 An example convolution process for obtaining component 1 is shown.
[0052] Figure 38 An example layer structure for EFE is shown.
[0053] Figure 39 An example layer structure for EFE is shown.
[0054] Figure 40 A flowchart of a method for visual data processing according to an embodiment of the present disclosure is shown.
[0055] Figure 41 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0056] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0057] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0058] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0059] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0060] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0061] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0062] Example Environment
[0063] Figure 1A This is a block diagram illustrating an example visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.
[0064] Visual data source 112 may include sources such as visual data capture devices. Examples of visual data capture devices include, but are not limited to, interfaces for receiving visual data from visual data providers, computer graphics systems for generating visual data, and / or combinations thereof.
[0065] Visual data may include one or more pictures or images from a video. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded / decoded representation of the visual data. The bitstream may include encoded / decoded pictures and associated visual data. An encoded / decoded picture is an encoded / decoded representation of a picture. Associated visual data may include sequence parameter sets, picture parameter sets, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data may be transmitted directly to a destination device 120 via network 130A through I / O interface 116. Encoded visual data may also be stored on storage medium / server 130B for access by the destination device 120.
[0066] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0067] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or further standards.
[0068] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-function video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Furthermore, while some embodiments describe encoding steps in detail, it should be understood that the corresponding inverse encoding decoding steps will be implemented by the decoder. Additionally, the term visual data processing encompasses visual data encoding / decoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another compressed format or at a different compression bitrate.
[0069] 1. Preliminary Discussion
[0070] This document relates to a neural network-based image and video compression method employing an output adjustment unit. The output of the neural network-based encoder / decoder is processed by two different upsampling units that generate two intermediate reconstructions. The final reconstruction is obtained by combining the two intermediate reconstructions. Additionally, this document relates to a neural network-based image and video compression method employing explicit control of the output of the processing layers. Indicators are included in the bitstream to indicate how the output of the processing layers is modified. As a result, the encoder and decoder can better adapt to unprecedented content (e.g., images not present in the training dataset), thereby improving compression performance. Furthermore, this patent document relates to a neural network-based image and video compression method employing convolutional layers to modify the components of an image. The weights of (multiple) convolutional layers are included in the bitstream.
[0071] 2. Further discussion
[0072] Deep learning is advancing across various fields, such as computer vision and image processing. Inspired by the successful applications of deep learning in computer vision, neural image / video compression techniques are being researched for application in image / video compression. Neural networks are designed based on interdisciplinary research in neuroscience and mathematics. Neural networks demonstrate powerful capabilities in the context of nonlinear transformations and classification. Example neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), a video coding standard developed by the Joint Video Experts Group (JVET), comprised of experts from the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). Neural network-based video compression is a rapidly developing research area, leading to continuous improvements in the performance of neural image compression. However, due to the inherent difficulty of the problems addressed by neural networks, neural network-based video coding remains a largely unexplored discipline.
[0073] 2.1 Image / Video Compression
[0074] Image / video compression generally refers to the computational technique of compressing video images into binary code for easier storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video. Codecs that do not lose data are called lossless compression, while those that allow targeted data loss are called lossy compression. Most codec systems use lossy compression because lossless reconstruction is not always necessary. Typically, the performance of image / video compression algorithms is evaluated based on the resulting compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes produced by compression; fewer binary codes result in better compression. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; higher similarity indicates better reconstruction quality.
[0075] Image / video compression techniques can be categorized into video encoding / decoding methods and neural network-based video compression methods. Video encoding / decoding schemes employ transform-based solutions, where statistical dependencies in latent variables (such as Discrete Cosine Transform (DCT) and wavelet coefficients) are utilized to carefully hand-design entropy encoding / decoding to model dependencies in the quantization domain. Neural network-based video compression can be further divided into neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within existing video codecs and serves only as part of the framework, while the latter is a separate framework developed based on neural networks, independent of the video codec.
[0076] A range of video codec standards have been developed to meet the growing demand for visual content transmission. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups: the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG). The International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has a Video Codecs Expert Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards published by these organizations include JPEG, JPEG 2000, H.262, H.264 / Advanced Video Codec (AVC), and H.265 / High-Efficiency Video Codec (HEVC). The Joint Video Experts Group (JVET), comprised of MPEG and VCEG, developed the Multi-Functional Video Codec (VVC) standard. Compared to HEVC, VVC reduces the bit rate by an average of 50% while maintaining the same visual quality.
[0077] Neural network-based image / video compression / encoding / decoding is also under development. However, example neural network encoding / decoding architectures are relatively shallow, and the performance of such networks is unsatisfactory. Neural network-based methods benefit from the support of abundant data and powerful computing resources, and are therefore better utilized in a variety of applications. Neural network-based image / video compression has shown promising improvements and has been proven feasible. However, the technology is far from mature, and many challenges need to be addressed.
[0078] 2.2 Neural Networks
[0079] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. Neural networks typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. The representations created by neural networks are not manually designed. Instead, deep networks, including processing layers, learn from massive amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations. Therefore, deep learning is considered particularly suitable for processing natural, unstructured data, such as acoustic and visual signals. The processing of such data has been a long-standing challenge in the field of artificial intelligence.
[0080] 2.3 Neural Networks for Image Compression
[0081] Neural networks used for image compression can be divided into two categories: pixel probabilistic models and autoencoder models. Pixel probabilistic models employ predictive encoding / decoding strategies. Autoencoder models use transform-based solutions. Sometimes, these two approaches are combined.
[0082] 2.3.1 Pixel Probability Modeling
[0083] According to Shannon's information theory, the best methods for lossless encoding and decoding can achieve the minimum decoding rate, denoted as -log2 p(x), where p(x) is the probability of symbol x. Arithmetic encoding and decoding is a lossless encoding and decoding method and is considered one of the best methods. Given a probability distribution p(x), arithmetic encoding and decoding makes the encoding and decoding rate as close as possible to the theoretical limit -log2 p(x) without considering rounding errors. Therefore, the remaining problem is to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality. The curse of dimensionality refers to the problem that the increase in dimensionality causes the dataset to become sparse, thus requiring a rapidly increasing amount of data to effectively analyze and organize the data as the number of dimensions increases.
[0084] Following a predictive encoding / decoding strategy, one way to model p(x) is based on predicting pixel probabilities one by one in raster scan order according to previous observations, where x is the image, which can be represented as follows:
[0085] p(x)=p(x1)p(x2|x1)...p(x i |x1, ...,x i-1 ...p(x) m×n |x1, ...,x m×n-1 (1)
[0086] Where m and n are the height and width of the image, respectively. Previous observations are also called the context of the current pixel. When the image is large, estimating the conditional probability can be difficult. Therefore, a simplified approach is to restrict the context of the current pixel as follows:
[0087] p(x)=p(x1)p(x2|x1)...p(x i |x i-k , ..., x i-1 ...p(x) m×n |x m×n-k , ..., x m×n-1 (2)
[0088] Where k is a predefined constant that controls the scope of the context.
[0089] It should be noted that this condition can also consider the sample values of other color components. For example, when encoding and decoding the red (R), green (G), and blue (B) (RGB) color components, the R sample depends on previously encoded pixels (including R, G, and / or B samples), and the current G sample can be encoded and decoded based on previously encoded pixels and the current R sample. Furthermore, when encoding and decoding the current B sample, previously encoded pixels, as well as the current R and G samples, can also be considered.
[0090] Neural networks can be designed for computer vision tasks and are also effective in regression and classification problems. Therefore, neural networks can be used to solve problems given a context x1, x2, ..., x... i-1 Estimate p(x) under the following circumstances i The probability of ).
[0091] Most methods directly model the probability distribution in the pixel domain. Some designs also model the probability distribution as a conditional probability distribution based on explicit or latent representations. Such models can be represented as:
[0092]
[0093] Where h is an additional condition, and p(x) = p(h)p(x|h) indicates that the modeling is divided into unconditional and conditional models. The additional condition can be image label information or high-level representation.
[0094] 2.3.2 Automatic Encoder
[0095] Now, let's describe autoencoders. Autoencoders are trained for dimensionality reduction and consist of an encoding component and a decoding component. The encoding component transforms a high-dimensional input signal into a low-dimensional representation. The low-dimensional representation can have a reduced spatial size but a greater number of channels. The decoding component recovers the high-dimensional input from the low-dimensional representation. Autoencoders enable the automatic learning of representations and eliminate the need for hand-crafted features, which is considered one of the most important advantages of neural networks.
[0096] Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme. The original image x is processed by the analysis network g. a A transformation is performed to achieve the latent representation y. The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the encoding / decoding rate. Quantized latent representation Then the synthesis network g s Inverse transform to obtain the reconstructed image Distortion (D) is achieved by using the function g p Transform x and Calculated in the perceptual space, resulting in z and It is compared to obtain D.
[0097] Autoencoder networks can be applied to lossy image compression. The learned latent representations can be encoded from a well-trained neural network. However, applying autoencoders to image compression is not straightforward because the original autoencoder is not optimized for compression and therefore cannot be effectively used directly as a trained autoencoder. Furthermore, other major challenges exist. First, the low-dimensional representation should be quantized before encoding. However, quantization is non-differentiable, which is necessary during backpropagation when training the neural network. Second, the objectives differ in compression scenarios because both distortion and rate need to be considered. Estimating the rate is challenging. Third, practical image encoding / decoding schemes should support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Various solutions are under development to address these challenges.
[0098] An example autoencoder for image compression using the example transform encoding / decoding scheme can be considered as a transform encoding / decoding strategy. The original image x is processed using the analysis network y = ... ga (x) is transformed, where y is the latent representation to be quantized and encoded / decoded. The synthetic network inversely transforms the quantized latent representation. To obtain reconstructed images Framework utilization distortion loss function Trained, where D is the relationship between x and The distortion between them, R is from the quantization representation The rate is calculated or estimated, and λ is a Lagrange multiplier. D can be computed in the pixel domain or the receptive domain. Most example systems follow this prototype, and the differences between such systems are likely only in the network structure or the loss function.
[0099] 2.3.3 Advanced Prior Model
[0100] Figure 2 An example potential representation of the image is shown. Figure 2 This includes an image 201 from the Kodak dataset, a visualization of the latent value 202 representing y from image 201, the standard deviation σ of the latent value 202 203, and the latent value y204 after introducing a super-prior network. The super-prior network consists of an encoder and a decoder that utilize super-prior information. In... Figure 1B In the transform encoding / decoding method for image compression shown, the encoder subnetwork uses parameter analysis transform. The image vector x is transformed into a latent representation y, which is then quantized to form because It is a discrete value, so It can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as a bit sequence.
[0101] from Figure 2The potential value 202 and the standard deviation σ203 can be clearly seen. Significant spatial dependencies exist among the elements. Notably, their scales (standard deviation σ²⁰³) appear to be spatially coupled. An additional set of random variables can be introduced. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as Figure 3 As shown.
[0102] Figure 3 This is a schematic diagram illustrating an example network architecture for an autoencoder that implements a super-prior model. The top side shows the image autoencoder network, and the bottom side corresponds to the super-prior subnetwork. The analysis and synthesis transformations are represented as g. a and g a Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks: an encoder utilizing hyperprior information (using h...). a (representation) and decoders utilizing prior information (using h) s (Representation). The prior model generates quantified potential values of prior information. It includes quantifying potential values. Information related to the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).
[0103] exist Figure 3 In the middle, the upper part of the model is the encoder g, as discussed above. a and decoder g s The lower side is used to obtain... The additional encoder h that utilizes prior information a and decoder h that utilizes prior information s Network. In this architecture, the encoder subjects the input image x to g. a This produces a response y with a standard deviation that varies in the spatial domain. The response y is fed into h. a In the middle, the distribution of the standard deviation in z is summarized. z is then quantized. The data is compressed and transmitted as side information. The encoder then uses the quantization vector... To estimate the spatial distribution σ of the standard deviation, and to use σ to compress and transmit the quantized image representation. The decoder first recovers from the compressed signal Then the decoder uses h s To obtain σ, which provides the decoder with the correct probability estimate so as to successfully recover the original value. Then the decoder will Feed to g sTo obtain a reconstructed image.
[0104] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. Figure 2 The latent value y204 in the model corresponds to the quantized latent value when using an encoder / decoder that leverages prior information. Compared to the standard deviation σ203, spatial redundancy is significantly reduced because the samples of the quantized latent value have lower correlation.
[0105] 2.3.4 Context Model
[0106] Although the prior model improves the quantification of latent values Modeling the probability distribution of the quantified potential value is possible, but additional improvements can be obtained by using an autoregressive model that predicts the quantified potential value from the causal context of the quantified potential value, which can be called a context model.
[0107] The term autoregressive indicates that the output of a process is later used as the input to that process. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.
[0108] Figure 4 This is a schematic diagram illustrating an example combined model configured to jointly optimize a context model with a super-prior and an autoencoder. The combined model jointly optimizes an autoregressive component (context model) that estimates the probability distribution of latent values from the causal context of the latent values, along with the super-prior and the underlying autoencoder. The real-valued latent representation is quantized (Q) to create quantized latent values. Quantization of the latent value of prior information It is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The dashed areas correspond to components performed by the receiver (e.g., the decoder) to recover the image from the compressed bitstream.
[0109] The example system utilizes a joint architecture, where both a super-prior model subnetwork (an encoder and a decoder utilizing super-prior information) and a context model subnetwork are leveraged. The super-prior and context models are combined to learn about the quantized latent value. The probabilistic model was then used for entropy encoding and decoding. For example... Figure 4As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean μ and scale (or variance) σ parameters for the Gaussian probability model. The Gaussian probability model is then used to encode samples of the quantized latent values into a bitstream with the aid of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized latent values from the bitstream via the arithmetic decoder (AD) module.
[0110] In the example, the potential samples are modeled as a Gaussian distribution or a Gaussian mixture model (not limited to this). Based on... Figure 4 In the example, the context model and the super-prior are used together to estimate the probability distribution of the potential samples. Since the Gaussian distribution can be defined by the mean and variance (also known as sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ).
[0111] 2.3.5 Gain Variational Automatic Encoder (G-VAE)
[0112] In the example, neural network-based image / video compression methods require training multiple models to adapt to different rates. A gain variational autoencoder (G-VAE) is a variational autoencoder with a pair of gain units, designed to achieve continuous variable bit rate adaptation using a single model. It consists of a pair of gain units, typically inserted into the encoder's output and the decoder's input. The encoder's output is defined as the latent representation y∈R. c*h*w , where c, h, and w represent the number, height, and width of the latent representation. Each channel of the latent representation is represented as y. (i) ∈R h*w Where i = 0, 1, ..., c-1. A pair of gain units includes a gain matrix M ∈ R. c*n And the inversion gain matrix, where n is the number of gain vectors. The gain vector can be represented as m s ={α s(0) α s(1) , ..., α s(c-1)}, α s(i) ∈R, where s represents the index of the gain vector in the gain matrix.
[0113] The motivation for the gain matrix is similar to the quantization table in JPEG, controlling the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.
[0114]
[0115] Where ⊙ represents channel multiplication, i.e. And α s(i) This is the i-th gain value in the gain vector ms. The inversion gain matrix used on the decoder side can be represented as M′∈R c*n It includes n inversion gain vectors, i.e., M′={δ s(0) δ s(1) ,...,δ s(c-1)}, δ s(i) ∈R. The inversion gain process is represented as:
[0116]
[0117] in It is the quantized latent representation of the decoded data, and y′ s It is a quantized potential representation of the inversion gain, which will be fed into the synthesis network.
[0118] To achieve continuous variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors {m t ,m′ t} and {m r ,m′ r The interpolation gain vector can be obtained through the following equation.
[0119] m v =[(m r ) l ·(m t ) 1-l ]
[0120] m′ v =[(m′ r ) l ·(m′ t ) 1-l ]
[0121] Where l∈R are interpolation coefficients, which control the corresponding bit rate of the generated gain vector pairs. Since l is a real number, any bit rate between two given gain vector pairs can be achieved.
[0122] 2.3.6 Encoding process using a joint autoregressive hyperprior model
[0123] Figure 4 The design described here corresponds to the example combined compression method. The encoding and decoding processes are described separately in this and the next section.
[0124] Figure 5 An example encoding process is illustrated. The input image is first processed using an encoder subnetwork. The encoder transforms the input image into a transform representation called the latent value, denoted by y. y is then fed into a quantizer block, denoted by Q, to obtain the quantized latent value. Then, it is converted into a bitstream (bits1) using an arithmetic coding module (denoted as AE). The arithmetic coding block will... Each sample point is sequentially converted into a bit stream (bits1).
[0125] The module utilizes an encoder with prior information, context, a decoder with prior information, and an entropy parameter subnetwork to estimate the quantization latent value. The probability distribution of the sample points. The latent value y is input to an encoder that utilizes prior information, and its output is the prior information latent value (denoted by z). The prior information latent value is then quantized. Furthermore, the second bitstream (bits2) is generated using the arithmetic coding (AE) module. The decompositional entropy module generates the probability distribution used to encode the quantized prior information latent values into a bitstream. The quantized prior information latent values include information about the quantized latent values. Information about the probability distribution.
[0126] Entropy parameter subnetwork generation was used to encode quantized latent values. The probability distribution estimation of a random variable x. Information generated from the entropy parameter typically includes the mean μ and the scale (or variance) σ parameter, which together are used to obtain the Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as... Here, parameter μ is the mean or expected value (also the median and mode) of the distribution, while parameter σ is its standard deviation (or scale or variance). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.
[0127] The subnetwork utilizes a decoder with prior information to generate part of the information used by the entropy parameter subnetwork; the other part of the information is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the samples in the quantized latent value. (Quantized latent value) It is typically a matrix composed of many sample points. Sample points can be indicated using indices, for example... or Specifically depends on the matrix Dimensions. Sample points Encoding is performed sequentially by the AE (Automatic Image Processor), typically using a raster scan order. In this raster scan order, the matrix rows are processed from top to bottom, with samples within each row processed from left to right. In scenarios where the AE encodes samples into a bitstream using a raster scan order, the context module uses previously encoded samples in raster scan order to generate a bitstream with the samples. The relevant information, generated by the context module and the decoder utilizing prior information, is combined by the entropy parameter module to generate information used to quantize the latent value. The probability distribution of encoding as a bit stream (bits1).
[0128] Finally, the first and second bitstreams, as the result of the encoding process, are transmitted to the decoder. It should be noted that other names can be used for the modules described above.
[0129] In the above description, Figure 5 All elements in the encoding are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).
[0130] 2.3.7 Decoding process using a joint autoregressive superprior model
[0131] Figure 6 An example decoding process is shown. Figure 6 The decoding process is described separately.
[0132] During decoding, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decompositional entropy subnetwork. The decompositional entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This refers to the quantized potential value of the prior information. The AD process reverts to the AE process applied in the encoder. Both the AE and AD processes are lossless, meaning the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.
[0133] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks employed in the decoder—the context, the decoder utilizing prior information, and the entropy parameter—are the same as the three sub-networks in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder (as in the encoder), which is crucial for lossless reconstruction of the quantized latent value. This is essential. As a result, the quantization potential value can be obtained in the decoder as well as in the encoder. Same version.
[0134] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter sub-network, the arithmetic decoding module decodes the samples of the quantized latent values one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial, and therefore cannot be accelerated using techniques such as parallelization. Finally, the fully reconstructed quantized latent values... Input into the synthesis transform (in) Figure 6 The module (represented as decoder) is used to obtain the reconstructed image.
[0135] In the above description, Figure 6 All elements in the array are collectively referred to as the decoder. The synthetic transformation that converts the quantized latent values into the reconstructed image is also called the decoder (or autodecoder).
[0136] 2.4 Neural Networks for Video Compression
[0137] Similar to video encoding and decoding techniques, neural image compression is the foundation of intra-frame compression in video compression based on neural networks. Therefore, the development of neural network-based video compression technology has lagged behind that of neural network-based image compression, as neural network-based video compression is more complex and requires greater effort to address its challenges. Compared to image compression, video compression requires effective methods to remove inter-frame redundancy. Inter-frame prediction is a key step in these example systems. Motion estimation and compensation are widely used in video codecs, but are typically not implemented by trained neural networks.
[0138] Neural network-based video compression can be categorized into two types based on the target scenario: random access and low latency. In the case of random access, the system allows decoding to begin at any point in the sequence, typically dividing the entire sequence into multiple individual segments, and allowing each segment to be decoded independently. In the case of low latency, the system aims to reduce decoding time, allowing previous frames in the temporal domain to be used as reference frames for decoding subsequent frames.
[0139] 2.5 Prerequisites
[0140] Almost all natural images and / or videos are in digital format. Grayscale digital images can be represented as... in It is a set of pixel values, where m is the image height and n is the image width. For example, This is the example setting, in which case... Therefore, a pixel can be represented as an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed images certainly have fewer bits.
[0141] Color images are typically represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by... This means that three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Most neural network-based video compression schemes have been developed in the RGB color space, while video codecs typically use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels: luminance (Y), blue chrominance (Cb), and red chrominance (Cr). Y is the luminance component, and Cb and Cr are the chrominance components. The compression benefits of YUV arise because Cb and Cr are often downsampled for pre-compression, as the human visual system is less sensitive to the chrominance component.
[0142] A color video sequence consists of multiple color images (also called frames) to record a scene at different timestamps. For example, in the RGB color space, a color video can be composed of x = {x0, x1, ..., x...} t , ..., x T-1} represents, where T is the number of frames in the video sequence, and If m = 1080 and n = 1920, And if the video has 50 frames per second (fps), then the data rate of the uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits per second (bps). This results in approximately 2.32 gigabits per second (Gbps), which uses a lot of storage and should be compressed before being transmitted over the Internet.
[0143] Typically, lossless methods can achieve natural image compression ratios of around 1.5 to 3, which is clearly below streaming requirements. Therefore, lossy compression is employed to achieve better compression ratios, but at the cost of distortion. Distortion can be measured by calculating the mean squared difference between the original and reconstructed images, for example, based on MSE. For grayscale images, MSE can be calculated using the following equation.
[0144]
[0145] Therefore, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR):
[0146]
[0147] in yes The maximum value in the range is 255, for example, for an 8-bit grayscale image. Other quality assessment metrics include structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).
[0148] To compare different lossless compression schemes, one can compare compression ratios that yield a given rate, and vice versa. However, to compare different lossy compression methods, the comparison must consider both rate and reconstruction quality. This can be achieved, for example, by calculating the relative rates at several different quality levels and then averaging the rates. The average relative rate is called Bjontegaard's incremental rate (BD rate). Other aspects to consider when evaluating image and / or video codec schemes include encoding / decoding complexity, scalability, robustness, and more.
[0149] 2.6 Separate processing of the luminance and chrominance components of an image
[0150] Figure 7 An example decoding process according to this disclosure is shown.
[0151] Depending on one implementation, the luminance and chrominance components of an image can be decoded using separate sub-networks. Figure 7 In this process, the luminance component of the image is processed by sub-networks such as "synthesis", "predictive fusion", "mask convolution", "decoder using prior information", and "variance decoder using prior information". The chrominance component is processed by sub-networks such as "synthesis UV", "predictive fusion UV", "mask convolution UV", "decoder UV using prior information", and "variance decoder UV using prior information".
[0152] The advantage of this separate processing is that the computational complexity of image processing is reduced by applying it separately. Typically, in neural network-based image and video decoding, computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is 192, the computational complexity will be proportional to 192 × 192. On the other hand, if the feature maps are divided into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128 × 128 + 64 × 64, which corresponds to a 45% reduction in complexity. Generally, separate processing of the luminance and chrominance components of an image does not lead to an excessive performance degradation because the correlation between the luminance and chrominance components is usually very small.
[0153] Figure 7 The processing (decoding process) in the code can be explained as follows:
[0154] 1. First, a decompositional entropy model was used to decode the quantization potential values for luminance and chrominance, i.e. Figure 7 In and
[0155] 2. The probability parameters (e.g., variance) generated by the second network are used to generate quantized residual latent values by performing an arithmetic decoding process.
[0156] 3. The quantized residual potential value is inverted using the inversion gain unit (iGain), such as... Figure 7 The orange color is shown in the image. The output of the inversion gain unit is expressed for the luminance and chrominance components as follows: and
[0157] 4. For the luminance component, the following steps are performed repeatedly until the result is obtained. All elements:
[0158] a. The first subnetwork was used for... The obtained samples are used to estimate the quantized potential value. The mean parameter.
[0159] b. Quantified residual potential value The mean was used to obtain The next element.
[0160] 5. In obtaining After all the samples are processed, a synthetic transformation can be applied to obtain the reconstructed image.
[0161] 6. For the chromaticity components, steps 4 and 5 are the same, but with a separate set of networks.
[0162] 7. The decoded luminance component is used to obtain additional information for the chrominance component. Specifically, the Inter-Frame Channel Related Information Filter (ICCI) subnetwork is used for chrominance component recovery. Luminance is fed as additional information into the ICCI subnetwork to assist in chrominance component decoding.
[0163] 8. After the luminance and chrominance components are reconstructed, adaptive color transformation (ACT) is performed.
[0164] The module named ICCI is a neural network-based post-processing module. Examples are not limited to the UCCI subnetwork. Any other neural network-based post-processing module can also be used.
[0165] Exemplary implementations of publicly available content are in Figure 7 The decoding process is described in the diagram. The framework comprises two branches, one for the luma and one for the chroma components. Within each branch, the first sub-network includes context, prediction, and optionally a decoder module utilizing advanced prior information. The second network includes a variance decoder module utilizing the advanced prior information. The quantized advanced prior information latent value is... and The arithmetic decoding process generates quantized residual latent values, which are then fed into the iGain unit to obtain quantized residual latent values for gain. and
[0166] After obtaining the residual potential value, a recursive prediction operation is performed to obtain the potential value. and The following steps describe how to obtain potential values. The sample points, and the chromaticity components are processed in the same way but using different networks.
[0167] 1. The autoregressive context module is used when using samples. This is used to generate the first input to the prediction module, where the (m, n) pairs are the indices of the sample points of the obtained potential values.
[0168] 2. Optionally, the second input to the prediction module is obtained by using a decoder that utilizes prior information and quantized potential values of the prior information. Obtained.
[0169] 3. Using the first and second inputs, the prediction module generates the mean [:, i, j].
[0170] 4. Mean [i, j] and quantified residual potential values Added together to obtain potential value
[0171] 5. Steps 1-4 are repeated for the next sample point.
[0172] Whether and / or how at least one method disclosed in the document can be applied, for example, in a bitstream transmitted from the encoder to the decoder via signal transmission.
[0173] Whether and / or how to apply at least one of the methods disclosed in the document can be determined by the decoder based on encoding and decoding information (such as dimensions, color format, etc.).
[0174] In addition, modules named MS1, MS2, or MS3+O (in) Figure 7 (The input) can be included in the processing stream. This module can perform operations on its input to obtain the output by multiplying the input by a scalar or adding an accumulated component to the input. The scalar or accumulated component used by this module can be indicated in the bitstream.
[0175] Figure 7 The module named RD or AD in the code can be an entropy decoding module. It can be a range decoder or an arithmetic decoder, etc.
[0176] The examples described in this article are not limited to Figure 7The example demonstrates a specific combination of units. Some modules may be missing, and some modules may be shifted according to the processing order. Additionally, supplementary modules may be included. For example:
[0177] 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined in one module, which can be based on a neural network.
[0178] 2. One or more modules named MS1, MS2, or MS3+o can be removed. The core of the public content is unaffected by the removal of one or more modules during scaling and adding.
[0179] exist Figure 7 The code also uses an asterisk to indicate other operations performed during the processing of the luminance and chrominance components. These operations are denoted as MS1, MS2, MS3+0. These operations can be, but are not limited to, adaptive quantization, latent sample scaling, and latent sample offset operations. For example, in adaptive quantization, this might correspond to scaling the samples using a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. Latent scaling might correspond to scaling the samples using a multiplier after the prediction process, where the multiplier's value is predefined or indicated in the bitstream. Offset operations might correspond to adding an accumulated element to the sample, again where the value of the accumulated element can be indicated, estimated, or predetermined in the bitstream.
[0180] Another operation can be slicing, where the samples are first sliced (grouped) into overlapping or non-overlapping regions, each of which is processed independently. For example, samples corresponding to the luminance component can be divided into slices with a slice height of 20 samples, while the chrominance component can be divided into slices with a slice height of 10 samples for processing.
[0181] Another application is wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the number of samples that can be processed in parallel can be indicated by control parameters. These control parameters can be indicated, estimated, or predetermined in the bitstream. In the case of separate luma and chroma processing, the number of samples that can be processed in parallel can be different, so different indicators can be transmitted via signals in the bitstream to control the operation of luma and chroma processing separately.
[0182] 2.7 Color Separation and Conditional Encoding / Decoding
[0183] Figure 8 An example of a learning-based image codec architecture is shown.
[0184] In one example, such as Figure 8As shown, the primary and secondary color components of the image are encoded and decoded separately using networks with similar architectures but different numbers of channels. All boxes with the same name are subnetworks with similar architectures, differing only in input / output tensor sizes and the number of channels. The number of channels for the primary component is C. p =128, the number of channels for the secondary components is C s =64. The vertical arrows (pointing downwards) indicate the data flow related to the encoding and decoding of secondary color components. The vertical arrows show the data exchange between the primary and secondary component pipelines.
[0185] The input signal to be encoded is represented as x, and the latent space tensor in the bottleneck of the variational autoencoder is y. The subscript "Y" indicates the primary component, and the subscript "UV" is used for the cascaded secondary components, including the chromaticity component.
[0186] First, the input image in RGB color format is converted into primary (Y) and secondary (UV) components. Primary component x Y Independent of minor component x UV The image is encoded and decoded, and the size of the encoded / decoded image is equal to the input / decoded image size. The secondary component uses x from the primary component. Y As auxiliary information, it is conditionally encoded and decoded to encode x. UV And using auxiliary information from the main component. As a potential tensor, used for decoding Reconstruction. Aside from the number of channels, channel sizes, and several entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost identical. Therefore, the primary and secondary latent tensors will generate two different bitstreams based on two different entropy models. In encoding x... Y Previously, x UV This module adjusts the sample point position through downsampling (in...) Figure 8 The superscript "s↓" indicates that the encoded image size for the minor component differs from the encoded image size for the major component. The scaling factor s is variable, but the default scaling factor is s=2. The size of the auxiliary input tensor in conditional encoding is adjusted so that the encoder receives major and minor component tensors with the same image size. After reconstruction, the minor component utilizes a neural network-based upsampling filter module ( Figure 8 The “NN color filter s↑” is rescaled to the original image size, and the module outputs the secondary components that are upsampled by the factor s.
[0187] Figure 8The example illustrates an image encoding / decoding system where the input image is first transformed into a primary (Y) component and a secondary (UV) component. The output... This corresponds to the reconstructed output of the primary and secondary components. At the end of the processing, It is converted back to RGB color format. Typically, x UV It is downsampled (resized) before being processed by the encoding and decoding modules (neural network). For example, x UV The size can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Therefore, the processing of minor components involves approximately 50% × 50% = 25% fewer samples, making it computationally less complex.
[0188] 2.8 Pruning Operations in Neural Network-Based Encoding and Decoding
[0189] Figure 9 An example synthetic transform for learning-based image encoding and decoding is shown.
[0190] The example synthetic transform above consists of a sequence of four convolutions, each with an upsampling stride of 2. The synthetic transform subnetwork in... Figure 9 The tensor dimensions in different parts of the composite transform before the clipping layer are depicted as follows: Figure 9 As shown in the diagram.
[0191] The clipping layer will have tensor size h d ×w d Change to h d-1 ×w d-1 , where h d =2·ceil(H / 2) d );w d =2·ceil(W / 2) d Here, d is the depth of the previous convolution in the codec architecture. For the principal components, the synthesis transform receives an input tensor of size h×w, where h = ceil(H / 16); w = ceil(W / 16). The output of the synthesis transform for the principal components is 1×h0×w0, where h0 = H; h0 = W.
[0192] For the secondary components, the synthesized transform receiver size is h. UV ×w UV The input tensor; h UV =ceil(ceil(H / s) / 16); w UV =ceil(ceil(W / s) / 16). The output of the synthesis transform of the principal components is 2×h. UV ×W UV0 , where h UV0 =ceil(H / s); hUV0 =ceil(W / s). For the minor component, the input size is h0 = ceil(H / s); w0 = ceil(W / s), where s is the scaling factor. For example, the scaling factor can be 2, where the minor component is downsampled by a factor of 2.
[0193] Based on the above explanation, the operation of the clipping layer depends on the output dimensions H and W and the depth of the clipping layer. Figure 9 The leftmost clipping layer has a depth of 0. The output of this clipping layer must be equal to H and W (output dimensions). If the input dimensions of this clipping layer are greater than H or W in the horizontal or vertical dimension, respectively, then clipping needs to be performed in that dimension. The second clipping layer, counting from left to right, has a depth of 1. The output of the second clipping layer must be equal to h1 = 2·ceil(H / 2). 1 w1 = 2·ceil(W / 2) 1 This means that if the input of the second clipping layer is greater than h1 or w1 in any dimension, clipping is applied to that dimension. In summary, the operation of the clipping layers is controlled by the output dimensions H and W. In one example, if both H and W are equal to 16, the clipping layer does not perform any clipping. On the other hand, if both H and W are equal to 17, all four clipping layers will perform clipping.
[0194] 2.9 Displacement
[0195] Bitwise shift operators can be represented using the function `bitshift(x, n)`, where `n` is an integer. If `n` is greater than 0, it corresponds to the right shift operator (`>>`), which shifts the input bits to the right, and the left shift operator (`<<`), which shifts the bits to the left. In other words, the `bitshift(x, n)` operation corresponds to:
[0196] bitshift(x, n) = x * 2 n ,
[0197] or
[0198] bitshift(x, n) = floor(x * 2) n ),
[0199] or
[0200] bitshift(x, n) = x / / 2 n .
[0201] The output of a shift operation is an integer value. In some implementations, the floor() function can be added to the definition.
[0202] floor(x) is equal to the largest integer less than or equal to x.
[0203] The " / / " operator is the integer division operator. It includes division and truncation of the result towards zero. For example, 7 / 4 and -7 / -4 are truncated to 1, and -7 / 4 and 7 / -4 are truncated to -1.
[0204] rightshift(x, n) = x >> n or
[0205] leftshift(x, n) = x << n
[0206] Equation 3: Alternative implementations of shift operators for right or left shifting
[0207] The arithmetic right shift of the two's complement integer representation of x >> y, shifting y bits. This function is defined only for non-negative integer values of y. The bits shifted into the most significant bit (MSB) as a result of the right shift have the same MSB value as x before the shift operation.
[0208] This function performs an arithmetic left shift of the two's complement integer representation of x << y, shifting y bits. It is defined only for non-negative integer values of y. The bits shifted into the least significant bit (LSB) as a result of the left shift have a value equal to 0.
[0209] 2.10 Convolution Operations
[0210] A convolution operation begins with a kernel, which is a small matrix of weights. This kernel "slides" across the input data, performing element-wise multiplication with its current portion of the input, and then sums the results into a single output pixel. In some cases, a convolution operation may include a "bias," which is added to the output of the element-wise multiplication operation.
[0211] The convolution operation can be described by the following mathematical formula. The output out1 can be obtained as:
[0212]
[0213] Where w1 is the multiplication factor, K1 is called the bias (accumulation term), and I k Let N be the k-th input, N be the kernel size in one direction, and P be the kernel size in another direction. A convolutional layer can include convolution operations, where more than one output can be generated. Other equivalent descriptions of convolution operations can be found below:
[0214]
[0215] In the above equation, "c" indicates the channel number. It is equivalent to the output number, out[1, x, y] is one output, and out[2, x, y] is the second output. k is the input number, I[1, x, y] is one input, and I[2, x, y] is the second input. w1 or w describes the weights of the convolution operation.
[0216] 2.10.1 Two-dimensional convolution operation
[0217] Convolution operations can be defined in 1, 2, 3, 4, ... dimensions. For example, a 2D convolution operation can be defined as:
[0218]
[0219] 2.11 LeakyReLU activation function
[0220] Figure 10 An example LeakyReLU activation function is shown. The LeakyReLU activation function in... Figure 10 The function is described in the diagram. According to this function, if the input is positive, the output is equal to the input. If the input (y) is negative, the output is equal to a*y. a is typically (but not limited to) a value less than 1 and greater than 0. Since the multiplier a is less than 1, it can be implemented as multiplication or division with non-integers. The multiplier a can be referred to as the negative slope of the LeakyReLU function.
[0221] 2.12 ReLU Activation Function
[0222] Figure 11 An example ReLU activation function is shown. The ReLU activation function is in Figure 11 The function is described in the diagram. According to this function, if the input is positive, the output equals the input. If the input (y) is non-positive, the output equals 0.
[0223] 2.13 Pixel shuffling and unshuffling functions
[0224] Figure 12 Examples of pixel shuffling and deshuffling operations are shown. PixelShuffle is an operation used in super-resolution models to implement efficient subpixel convolutions with a stride of 1 / r. Specifically, it shuffles the elements in a tensor of shape [Cxr2, W, H] into a tensor of shape [C, Wxr, Hxr]. Pixel deshuffling is the inverse operation of shuffling, where the input tensor of shape [C, Wxr, Hxr] is transformed into a tensor of shape [Cxr2, W, H].
[0225] 2.14 Deconvolution Operation
[0226] Transposed convolutional (also known as deconvolution) layers are typically used for upsampling, i.e., to generate an output feature map with a spatial dimension larger than the input feature map. The transposed convolution operation is performed in... Figure 13 The example is shown in the text. Figure 13 An example of a transposed convolution with a 2x2 kernel is shown. The shaded area represents a portion of the intermediate tensor and the input and kernel tensor elements used for computation.
[0227] 3. The technical problem solved by the disclosed technical solution
[0228] In some codecs based on example neural networks, the input image is first converted to YUV420 format. This instructs the image to be decomposed into three components (e.g., "Y", "U", and "V"), and then the chroma components are downsampled by a factor of 2. If the width and height of the luminance component are W and H, respectively, then the width and height of the chroma components "U" and "V" are W / 2 and H / 2, respectively. And at the end of the decoding process, the chroma components are upsampled back to their original size using an upsampling filter.
[0229] There are different types of upsampling methods, each with its own advantages and disadvantages. In particular, different upsampling methods may perform better than others for different content types (different images). Furthermore, different parts of the content (e.g., the image) may favor different upsampling methods.
[0230] 4. List of solutions and implementation examples
[0231] 4.1 Core Example
[0232] The objective of this disclosure is to improve the quality of upsampled reconstructions by combining two different upsampled reconstructions. Side information can be included in the bitstream to control the combination method.
[0233] Decoder operation:
[0234] Based on some examples, using a neural network to convert a bitstream into a reconstructed image includes the following operations:
[0235] - Use the first upsampler to obtain the first upsampled reconstruction (e.g., the components of the image).
[0236] - Use a second upsampler to obtain a second upsampled reconstruction (e.g., the components of the image).
[0237] - [Optional] Obtain the side information of the control combination unit from the bit stream.
[0238] - The first upsampled reconstruction and the second upsampled reconstruction are combined by combining the combination unit to obtain the reconstructed components.
[0239] - Obtain the reconstructed image based on the reconstructed components.
[0240] Encoder operation:
[0241] Based on some examples, using a neural network to convert an image into a bitstream includes the following operations:
[0242] - Use the first upsampler to obtain the first upsampled reconstruction (e.g., the components of the image).
[0243] - Use a second upsampler to obtain a second upsampled reconstruction (e.g., the components of the image).
[0244] - [Optional] Obtain the edge information of the control combination unit.
[0245] - The first upsampled reconstruction and the second upsampled reconstruction are combined by combining the combination unit to obtain the reconstructed components.
[0246] - [Optional] Include side information in the bitstream.
[0247] 4.2 Details of the Example
[0248] Figure 14 An example subnetwork of a neural network is shown. Figure 15 An example subnetwork of a neural network is shown. Figure 16 An example subnetwork of a neural network is shown. Figure 17 An example subnetwork of a neural network is shown. Figure 18 An example of the value of W_base[i,j] is shown. Figure 19 An example of the value of W_base[i,j] is shown. Figure 20 An example of the value of W_base[i,j] is shown. Figure 21 An example of the value of _base[i, j] is shown.
[0249] Figure 22 An example of the value of W_base[i,j] is shown. Figure 23 An example of the value of W_base[i,j] is shown.
[0250] ● Figure 24 An example flowchart is shown for performing the disclosed example. Image components. The output is processed by upsampling unit 1 and upsampling unit 2. The output of each unit is combined by a combination unit to obtain the upsampled output.
[0251] ● Figure 25 An example neural network configured to perform the disclosed example is shown. Figure 26An example neural network 26 configured to perform the disclosed examples is shown. In the figure, the first and second upsampler units are enclosed within dashed boxes. A combination unit is also shown. In this figure, the upsampler units and combination units are implemented using neural network layers, such as convolutional, cascaded, and pixel shuffling units, which are the basic elements of the neural network.
[0252] ● Figure 27 An example neural network configured to perform the disclosed example is shown. Another processing unit or augmentation unit may be applied to one or both of the first upsampled reconstruction or the second upsampled reconstruction. This is in Figure 27 The example is shown below. The processing unit is applied after the first upsampler and before the combining unit.
[0253] ●Details of the first or second upsampling unit:
[0254] The first or second upsampling unit can be an adaptive filter. In other words, the multiplication or accumulation parameters controlling the upsampling process can be obtained from the bitstream.
[0255] ■ The parameters of upsampling units 1 and 2 can be different. In other words, two different sets of parameters can be obtained from the bitstream to control the first upsampling unit and the second upsampling unit.
[0256] ○ The upsampling unit can be a fixed filter.
[0257] ■ In one example, the upsampler could be an interpolation filter based on Discrete Cosine Transform (DCT-IF).
[0258] ■ The coefficient of the upsampler can be compared with... Figures 18 to 23 Any one of them is the same.
[0259] ■ The coefficient of the upsampler can be close to Figures 18 to 23 Any one of them. The coefficients can have higher or lower numerical precision.
[0260] ■ The upsampling unit can be a bicubic filter, a Lanzos filter, or a bilinear filter.
[0261] The first upsampling unit can be an adaptive filter. The second upsampling unit can be a fixed (predefined) filter.
[0262] ○ The upsampling unit can be implemented as a convolution operation or a deconvolution operation.
[0263] ○ The upsampling unit can be implemented using pixel shuffling or pixel deshuffling operations.
[0264] ○ The upsampling unit can have an M×M kernel size (M samples in width and M samples in height).
[0265] ■M can be equal to 5, 4, 3 or 2.
[0266] ■M can be adjusted. The value of M can be obtained from the bitstream.
[0267] ■ M can be different for the first upsampling unit and the second upsampling unit.
[0268] ●Details of the combined unit:
[0269] ○ The combination unit can be selected between the first reconstruction and the second reconstruction.
[0270] The combined unit can include at least one sample from the first reconstruction and at least one sample from the second reconstruction in the final reconstruction.
[0271] The combined unit can obtain the final reconstructed sample points by taking the average of the sample points in the first reconstruction and the sample points in the second reconstruction.
[0272] ○ The combination unit can be controlled by side information obtained from the bit stream.
[0273] ■ Edge information can have 3 states (3 possible indications):
[0274] ● The indicator can indicate whether the samples in the final reconstruction are obtained by setting them to be equal to the samples in the first reconstruction or the samples in the second reconstruction, or the average of both.
[0275] ■ Edge information can have 2 states (2 possible indications):
[0276] ● The indicator can indicate whether the samples in the final reconstruction are obtained by setting them to be equal to the samples in the first reconstruction or the samples in the second reconstruction.
[0277] The final reconstruction can be divided into rectangular pieces of size M×M. For each piece, different methods can be obtained from combining units.
[0278] ■ The sample points of a patch can be set to be equal to the first reconstruction, the sample points of a second patch can be set to be equal to the second reconstruction, and the sample points of a third patch can be set to be equal to the average of the first and second reconstructions.
[0279] ■M can be included in the bitstream (or obtained from the bitstream).
[0280] The example implementation of the disclosed example is as follows Figure 27 As shown. The output of the first upsampling unit is The output of the second upsampling unit is Output The processing unit further processes the data, and its output is... Finally, the combination unit was used for combination. and This is to obtain the final reconstructed output. An example of the process details is provided below.
[0281] EFE luminance-assisted adaptive upsampling process
[0282] The input to this process is and The output of this process is Multiplication weight parameter W 1A [8, 4, 4] and W 1B [8, 4, 4] is used. The accumulated bias parameter B1[2] is used. For x in 0..W, y in 0..H, and k in 0..1, the following is performed:
[0283] -
[0284]
[0285] Where fi = 4*k + 2*(y%2) + x(%2).
[0286] EFE output adjustment
[0287] The input to this process is and The output of this process is The multiplication weight parameter is W 4A [8, 4, 4], W 4B [8, 4, 4] and W5[2]. The accumulated bias parameter B1[2] is used. First, for x in 0..W, y in 0..H, and k in 0..1, the following is performed;
[0288]
[0289] Where fi = 4*k + 2*(y%2) + x(%2).
[0290] Then, for x in 0..W, y in 0..H, and k in 1..2, the following is executed:
[0291]
[0292] or
[0293] -
[0294]
[0295] or
[0296] --
[0297]
[0298] Finally, for x in 0..W and y in 0..H, the following assignments are made:
[0299] -
[0300] The first component, the second component, or any of the components mentioned above can be components of the image.
[0301] - It can be a chromaticity component or a luminance component.
[0302] - Before applying the proposed solution, the mean can be subtracted from any component.
[0303] - After applying the proposed solution, the mean can be added to the upsampled components.
[0304] 4.3. Explanation and advantages of the example
[0305] The example uses parameters obtained from the bitstream to improve the quality of the reconstructed image. The example is designed to achieve the following advantages:
[0306] 1. Some parameters used in the equation are obtained from the bitstream. This provides the possibility of content adaptation. In neural network-based image compression networks, the network can be pre-trained using a very large dataset. After training, the network parameters (e.g., weights and / or biases) cannot be adjusted. However, when the network is used, it is applied to entirely new images that are not part of the training dataset. Therefore, there is a difference between the training dataset and real-world images. To address this issue, a small set of parameters optimized for the new images is passed to the decoder to improve adaptation to new content.
[0307] The second advantage of including parameters in the bitstream is that a much shorter network can be used to serve the same purpose when the parameters are transmitted. In other words, a much longer neural network (containing many more convolutional and activation layers) might be necessary to achieve the same goal if the parameters were not transmitted as side information.
[0308] 2. The examples can be implemented using the most basic neural network layers. The equations used to explain the examples are designed so that they can be implemented using the most basic processing layers (i.e., convolution and ReLU operations) found in neural network literature. This intentional choice is made because image encoders / decoders are expected to be implemented in a wide variety of devices, including mobile phones. Importantly, an image encoded in one device can be decoded in almost any device. Although the neural processing chipsets or GPUs in such devices are becoming increasingly complex, implementing arbitrary functions on such processing units remains impossible. As a simple example, the function f(x) = x 2 Although it appears very simple, it cannot be implemented efficiently in a neural processing unit (NNU), but only in a general-purpose processing unit such as a CPU. If the function cannot be implemented in a NNU, processing speed and battery consumption will increase significantly.
[0309] The example eliminates the aforementioned problems by using the most basic processing layers from neural network literature. Convolutions and ReLU (as well as some other activation functions such as leaky ReLU, sigmoid, etc.) are almost guaranteed to be implemented in a neural processing unit or GPU. Therefore, mobile phones with neural processing units or GPUs can be expected to perform the defined operations efficiently.
[0310] 3. The example utilizes at least two different upsampling methods. Different upsampling methods can exhibit different characteristics in different images (content). Furthermore, different upsampling methods can exhibit different characteristics in different parts of the image. This disclosure utilizes two upsampling methods and adaptively combines them to achieve superior reconstruction quality.
[0311] Further solutions
[0312] 5.1. Technical problem solved by the disclosed technical solution
[0313] In image compression, the images to be compressed can have very different statistical properties. For example, natural images depicting natural scenes can differ significantly in statistical properties from screen content (e.g., computer-generated images). Therefore, some layers of a neural network, including dozens or hundreds of processing layers, are not always able to improve compression performance.
[0314] 5.2. List of Solutions and Implementation Examples
[0315] In one example, a neural network-based image and video compression method is provided that modifies the output of the processing layer. Indicators are included in the bitstream to explicitly control the output of the processing layer.
[0316] 5.2.1 Core Example
[0317] Example decoder operation:
[0318] Example 1: According to this disclosure, a bitstream is converted into a reconstructed image using a neural network, including the following operations:
[0319] ● The first intermediate output is processed using a processing layer to obtain the second intermediate output.
[0320] ● Obtain the indicator from the bitstream.
[0321] ● The output samples are obtained based on the indicator and at least two of the following;
[0322] ○ The first intermediate output sample, or,
[0323] ○ The second intermediate output sample, or
[0324] ○Samples of the first intermediate output and sample of the second intermediate output.
[0325] ● Obtain the reconstructed image based on the output.
[0326] Figure 28 An example implementation of this disclosure is illustrated. In the example, a first intermediate output is obtained using a neural subnetwork. The first intermediate output is fed to a processing layer to obtain a second intermediate output. Furthermore, an indicator is obtained from a bitstream. The indicator, the first intermediate output, and the second intermediate output are fed to a decision unit. The decision unit obtains at least two candidates from the first and second intermediate outputs. The output of the decision unit is determined based on the value of the indicator and the at least two candidates. The value of the indicator determines which candidate is selected. For example, the output of the decision may be obtained based on at least two of the following candidates;
[0327] ●First intermediate output,
[0328] ●Second intermediate output
[0329] ●The combination of the first intermediate output and the second intermediate output.
[0330] For example, the combination can be obtained as (FIO1+FIO2) / 2, where FIO1 is the first intermediate output and FIO2 is the second intermediate output.
[0331] For example, the combination can be obtained as (FIO1*K+FIO2*M) / (M+K), where FIO1 is the first intermediate output, FIO2 is the second intermediate output, and K and M are scalars.
[0332] ■K and M can be pre-ordered.
[0333] ■ At least one of K or M can be transmitted as a signal in a bit stream.
[0334] ■ A relationship between K and M, such as K = 1 - M, can exist.
[0335] For example, the combination can be obtained as f(FIO1, FIO2), where f() is a function.
[0336] ○ Combinations can be obtained based on limiting or clamping operations.
[0337] Combinations can be obtained based on the following:
[0338] ■Clip(FIO2-FIO1, maximum value)+FIO1.
[0339] ■Clamp(FIO2-FIO1, maximum value, minimum value)+FIO1.
[0340] The limiting operation selects the minimum value of the input parameter, and the clamping operation can be described as min(max(FIO2-FIO1, minimum value), maximum value). The min() and max() operations output the minimum and maximum values of the input parameters, respectively.
[0341] Encoder operation:
[0342] Example 1: In this example, the reconstructed image is converted into a bitstream using a neural network, including the following operations:
[0343] ● The first intermediate output is processed using a processing layer to obtain the second intermediate output.
[0344] ● Obtain at least two candidates based on the following:
[0345] ○ The first intermediate output sample, or,
[0346] ○ The second intermediate output sample, or
[0347] ○Samples of the first intermediate output and sample of the second intermediate output.
[0348] ● Select the best candidate from at least two candidates.
[0349] ● Include the indicator corresponding to the best candidate in the bitstream.
[0350] 5.2.2 Details of the Example
[0351] In the example, the first intermediate output and the second intermediate output can be divided into blocks of size N×N.
[0352] For each block, the indicator can be transmitted via signaling in the bitstream.
[0353] The N×N blocks of output can be obtained based on the corresponding N×N blocks of the first intermediate output and / or the corresponding N×N blocks of the second intermediate output.
[0354] ○ Block size can be included in the bitstream.
[0355] Typical values for ○N can be 32, 48, 64, 80, 96, 112, 128, etc.
[0356] ● The indicator value can be 0 or 1. 0 indicates that the corresponding output is obtained from the first intermediate output, and 1 indicates that the corresponding output is obtained from the second intermediate output (and vice versa).
[0357] ● The indicator value can be 0, 1, or 2. 0 indicates that the corresponding output is obtained based on the first intermediate output, 2 indicates that the corresponding output is obtained based on the second intermediate output (and vice versa). A value of 1 indicates that the corresponding output is obtained as a combination of the first and second intermediate outputs.
[0358] The combination can be a linear combination of the first intermediate output and the second intermediate output.
[0359] ● In the example, at least two indicators are included in the bitstream, where the first indicator controls the sample group of the output, and the second indicator controls the different sample groups of the output.
[0360] ●Details of the processing layer:
[0361] ○ The processing layer can be a neural network layer.
[0362] It can include convolutional layers.
[0363] It can include one or more convolutional layers.
[0364] It can include an activation layer.
[0365] The processing layer may include filters.
[0366] The processing layer may include adding offset values to the input.
[0367] 5.2.3. Advantages of the Example
[0368] According to the example, the difference between training time and application time is reduced. In a neural network (NN)-based image encoding and decoding system, the encoder and decoder consist of neural network layers. These neural network layers are trained using a training dataset. After training is complete, the encoder and decoder are subjected to images not present in the training dataset. Therefore, the results obtained by the encoder and decoder may not be optimal for new images.
[0369] This disclosure improves the ability to increase the adaptability of the encoder and decoder. Indicators can be included in the bitstream to modify the output of some processing layers of the encoder and decoder. The encoder can select the value of the indicator in a way that increases compression performance. The decoder obtains the indicator from the bitstream and applies it in the same way as the encoder. As a result, the encoder and decoder can better adapt to unprecedented content (e.g., images not present in the training dataset), thereby improving compression performance.
[0370] 6. Further Solutions
[0371] 6.1. Technical problem solved by the disclosed technical solution
[0372] When image components (e.g., luminance and chrominance components) are processed using different synthetic subnetworks, the correlation between the different components is not fully utilized. In other words, information that might be important for the reconstruction of one component may also be relevant to the reconstruction of the second component. This joint information is not fully utilized when two different synthetic transforms are used to reconstruct two different components.
[0373] 6.2. List of Solutions and Implementation Examples
[0374] According to this disclosure, a subnetwork including convolutional layers is included at the end of two synthetic transforms. The first synthetic transform processes a first component of the image, and the second synthetic transform processes a second component. The subnetwork takes the outputs of the two subnetworks as input and refines at least one component. It should be noted that the second component mentioned in Section 6 may include a minor component, and the first component mentioned in Section 6 may include a major component. Alternatively, the second component mentioned in Section 6 may include a chroma component, and the first component mentioned in Section 6 may include a luminance component. In another example, the second component mentioned in Section 6 may include a U component and / or a V component, and the first component mentioned in Section 6 may include a Y component. This correspondence in Section 6 can be reversed compared to the remainder of this disclosure.
[0375] 6.2.1 Core Example
[0376] Decoder operations can be performed as follows.
[0377] The bitstream is converted into a reconstructed image, including the following operations:
[0378] ● The composite transformation is used to obtain the first and second components of the image.
[0379] ● The first and second components are input into the convolutional layer.
[0380] ● Modify at least one component of the convolutional layer.
[0381] ● The reconstructed image (decoded image) is obtained from two components.
[0382] In one example, the composition transformation consists of two composition transformations, where the first component is obtained using the first composition transformation and the second component is obtained using the second composition transformation.
[0383] 6.2.2 Details of the Example
[0384] In some examples, convolutional layers may have the following details:
[0385] ● A convolutional layer can have at least two inputs.
[0386] ○ An input can be a luminance component.
[0387] ○The second input can be a chromaticity component.
[0388] ● A convolutional layer can have 3 inputs: 1 luminance component and 2 chrominance components.
[0389] ● A convolutional layer can have one output, a chroma component.
[0390] ● A convolutional layer can have two outputs and two chroma components.
[0391] ● A convolutional layer can have two outputs: a luminance component and a chrominance component.
[0392] ● A convolutional layer can have three outputs: a luminance component and two chrominance components.
[0393] In some examples, the operations performed by the convolutional layers may have the following details:
[0394] ● In one example, the mean of the first component can be calculated by subtracting it from the first component before it is input into the convolutional layer.
[0395] ● The mean of the second component can be calculated by subtracting it from the second component before it is input into the convolutional layer.
[0396] ○ In one example, the mean can be obtained from the bitstream.
[0397] In another example, the mean can be predefined.
[0398] In another example, the mean can be calculated by adding the samples of the first or second component and dividing the result by the number of samples.
[0399] ● The calculated mean can be added to the output of the convolutional layer.
[0400] ● The output of a convolutional layer can be one of the components.
[0401] ● At least one component of the image is modified by a convolutional layer.
[0402] ● The output of a convolutional layer can be added to the output of a synthesis transform to obtain a processed component.
[0403] ● Component 1 (i.e., the output of the convolutional layer can be obtained according to any of the following formulas:)
[0404] ○Component1=conv(in2-E(in2), in1-E(in1))+in1+K
[0405] ○Component1=conv(in2,in1-E(in1))+in1+K
[0406] ○Component1=conv(in2-E(in1),in1)+K
[0407] ○Component1=conv(in2,in1)+K
[0408] ○Component1=conv(in2,in1-E(in1))+E(in1)+K
[0409] ○Component1=conv(in2-E(in2),in1-E(in1))+E(in1)+K
[0410] ○Component1=conv(in2-E(in2))+in1+K
[0411] ○Component1=conv(in2-E(in2))+in1
[0412] Where in1 and in2 are the two components of the image obtained as the output of the synthesis transform, E(in1) is the mean of in1, and K is the addition parameter. In one example, K equals 0. In another example, K is a scalar whose value is transmitted as a signal in the bitstream.
[0413] ○ In a specific example, the chromaticity U component can be obtained from the chromaticity U input (component) and the luminance input (component).
[0414] In another specific example, the chromaticity V component can be obtained from the chromaticity V input (component) and the luminance input (component).
[0415] In another specific example, the luminance component can be obtained solely from the luminance input (component).
[0416] ●According to this disclosure, different modified components of an image can be obtained using convolutional layers with different numbers of inputs:
[0417] ○ In a specific example, the chromaticity U component can be obtained from the chromaticity U input and the luminance input.
[0418] In another specific example, the chromaticity V component can be obtained from the chromaticity V input and the luminance input.
[0419] In another specific example, the luminance component can be obtained solely from the luminance input.
[0420] The number of inputs used can be indicated in the bitstream. For example, to obtain the chroma U component, one input (e.g., luma component only) or two inputs (e.g., luma component and chroma U component) can be used. The choice can be indicated in the bitstream.
[0421] An indicator can be included in the bitstream to indicate which input is used to obtain the output. For example, depending on the value of the indicator, the luminance component or the chrominance U component can be used as input to obtain the chrominance U output.
[0422] ● The formula used to obtain the components can be indicated in the bitstream. For example, depending on the indicator, one or both outputs of the two synthesis transforms can be used. More specifically, if the output of synthesis transform 1 is out1 and the output of synthesis transform 2 is out2, then depending on the value of the indicator, only out1 or both out1 and out2 can be used as input to the convolutional layer.
[0423] ○ In one example, component 1 can be obtained based on the value of the indicator obtained from the bitstream according to Component1 = Conv(in2 - E(in2)) + in1 + K or according to conv(in2 - E(in2), in1 - E(in1)) + in1 + K.
[0424] In one example, an indicator is included in the bitstream to indicate how many inputs are used to obtain a component. For example, the chroma U component can be obtained with 1 input, and the chroma V component can be obtained with 2 inputs. The indicator indicates how many inputs are used in obtaining the output component.
[0425] ● The kernel size for a convolution operation can be indicated in the bitstream.
[0426] ● The weights (multiplier parameters) of the convolution operation can be included in the bitstream (and obtained from the bitstream).
[0427] In one example, the weights of the convolution can be included in the bitstream using N bits.
[0428] ■ N can be adjustable, and the indication controlling N can be included in the bitstream. For example, based on the indication in the bitstream, the value of N can be presumed to be equal to 16. Or the value of N can be presumed to be equal to 12.
[0429] The output of the synthesized transform can be sliced into multiple slices. Different convolution weights can be applied to different slices. In other words, different convolution weights corresponding to different slices can be obtained from the bitstream.
[0430] In one example, the number of slices can be transmitted via signals in the bitstream.
[0431] The number of ○ pieces can vary for each component.
[0432] Examples of operations performed by convolutional layers are depicted in the following examples.
[0433] Figure 29 An example convolution process for obtaining component 1 is shown.
[0434] Figure 30 An example convolution process for obtaining component 1 is shown.
[0435] Figure 31 An example convolution process for obtaining component 1 is shown.
[0436] Figure 32 An example convolution process for obtaining component 1 is shown.
[0437] Figure 33 An example convolution process for obtaining component 1 and component 2 is shown.
[0438] Figure 34 An example convolution process for obtaining component 1 and component 2 is shown.
[0439] Figure 35 An example convolution process for obtaining component 1 is shown.
[0440] Figure 36 An example convolution process for obtaining component 1 is shown.
[0441] exist Figure 29 In the example depicted, the mean is first calculated based on the output (out1) of Synthetic Transformation 1. The mean (mean1) is subtracted from the output of Synthetic Transformation 1. The output (out2) of Synthetic Transformation 2 and (out1-mean1) are fed into a convolutional layer. The output of the convolutional layer is added to out1 to obtain component 1. The reconstructed image (decoded image) is obtained based on component 1. In this example, out1 and out2 represent the outputs of Synthetic Transformations 1 and 2.
[0442] exist Figure 31 In addition to Figure 29 In addition to the example in the example, the second mean (mean2) is calculated based on the output (out2) of the synthetic transform 2. (Out1-mean1) and (out2-mean2) are fed into the convolution. Out1 is added to the output of the convolutional layer to obtain component 1.
[0443] Figure 30 Examples and Figure 31 The examples in [the example] are similar. The difference between the two examples is that, in [the example]... Figure 30 In this approach, the mean is either obtained from the bitstream or predefined. Using a predefined mean or one obtained from the bitstream has the advantage of reducing computational complexity because the mean calculation does not need to be performed. When the mean is obtained from the bitstream, it means that the mean is calculated at the encoder and included in the bitstream. Therefore, the decoder can obtain the mean from the bitstream and perform the convolution operation.
[0444] Figure 33 and Figure 34 An example is depicted where the output of the convolutional layer is component 1 and component 2.
[0445] Figure 32 The examples depicted in Figure 31 Similarly. In Figure 32 In the example, component 1 is obtained by adding the calculated mean (instead of the output of synthetic transformation 1) to the output of the convolutional layer.
[0446] In some examples, the details of the components may be as follows.
[0447] ● One component can be a chromaticity component, and another component can be a luminance component.
[0448] ● The output of the first composite transform can be the luminance component. The output of the second composite transform can be the chrominance U and chrominance V components. In another example, the output of the second composite transform can be the chrominance Cb and chrominance Cr components.
[0449] ● In another example, the components can be R, G, and B components (e.g., red, green, and blue).
[0450] Figure 37 An example convolution process for obtaining component 1 is shown.
[0451] Figure 37 An aspect of this disclosure is illustrated, in which an intermediate module is placed between the convolution operation and the composition transformation. In any of the examples above, conv(A, B) is equivalent to conv(A) + conv(B). According to one example, the components are modified according to one of the following formulas:
[0452] ○Component1=(in1-mean)*r+mean,
[0453] ○Component1=(in1-mean) / r+mean,
[0454] ○Component1=(in1-mean)*r+mean+K,
[0455] ○Component1=(in1-mean) / r+mean+K,
[0456] ○Component1=conv(in1-mean)+mean+K.
[0457] Here, mean and r are the mean and scaling factor, which can be obtained from the bitstream. In the decoder, the values of mean and r can be obtained from the bitstream. The weights (coefficients) of the convolution can be obtained from the bitstream.
[0458] At the encoder, the mean can be calculated as the mean of one of the components of the input image. And at the encoder, r can be chosen as a scaling factor. The scaling factor helps stretch the histogram of the input components, allowing more detail to be preserved after the quantization process during encoding. Depending on the value of r, more information can be preserved after quantization at the encoder, at the cost of an increased bit rate. The encoder can choose r in a way that achieves a desired balance between the bit rate and the amount of information retained after quantization.
[0459] At the decoder, the histogram stretching performed by the encoder is inverted based on the values of the mean and r. The mean and r values are determined by the encoder and included in the bitstream. These values are obtained by the decoder from the bitstream to perform the inversion operation.
[0460] Section A below provides an example implementation of the proposed solution. In the example, Figure 37 An example network structure is depicted, and Section A.2 provides details about each processing layer. Section A.3 describes an example method for transmitting parameters via signaling in a bitstream. Section A.4 describes the semantics of the parameters corresponding to those in Section A.3. Finally, Section A.5 describes an example method for slicing an input image into multiple rectangular regions (patches) for processing. When slicing is applied, different weights and bias parameters can be used in different parts of the input.
[0461] Section A Enhanced Filter Extension (EFE) Layer
[0462] A.1 Overview
[0463] This appendix describes in detail the Enhancement Filter Extension (EFE) process. This process provides enhancement to the color information plane (second component) of the image, utilizing information from the luminance (first component).
[0464] A. 2-layer structure
[0465] EFE subnetwork module receives and As input, and output full-size enhanced ( Figure 38 First component) The components are processed sequentially through the bicubic 2x↑, CONV1(1×1,1,1), CONV3(M×M,2,1), Mask&Offset1, and OutputAdjust1 layers. The second component... The processing layers proceed sequentially through bicubic 2x↑, CONV2(1×1, 1, 1), CONV4(MxM, 2, 1), Mask&Offset2, and OutputAdjust2. Figure 38 The details of the layered structure are depicted.
[0466] Figure 38 An example layer structure for an EFE is shown. Details of each layer are as follows:
[0467] -CONV1(1×1,1,1): The weight tensor is set to W1 and the bias tensor is set to B[1].
[0468] -CONV2(1×1,1,1): The weight tensor is set to W2 and the bias tensor is set to B[2].
[0469] -CONV3(M×M,2,1), the weight tensor is set to W3, and the bias tensor is set to all zeros.
[0470] -CONV4(N×N, 2, 1), the weight tensor is set to W4, and the bias tensor is set to all zeros.
[0471] -Mask&Offset Z where z has possible values {1, 2}:
[0472]
[0473] -output adjust1:
[0474]
[0475] -output adjust2:
[0476]
[0477] in;
[0478] -mask[n, x, y]:
[0479]
[0480] and
[0481] -mean(.): Outputs the average sample value of the input tensor.
[0482]
[0483] -subtract: The input from the side branch is subtracted from the input from the main branch, as shown in the example below.
[0484] out[1,x,y]=in[1,x,y]-B[1]
[0485] -concatenation: The two inputs are concatenated along the channel dimension.
[0486] A.3 Parameters are transmitted via signal.
[0487] To perform the processing steps described in Sections I.1 and I.2, adjustable weights, biases, and offset parameters are transmitted via signals in the image header. The parameters transmitted via signals in the image header are:
[0488] Weights and biases of the CONV1(1×1,1,1) and CONV1(1×1,1,1) operations: W1[1], W2[1], B[2].
[0489] - Kernel size and weights for CONV3(M×M, 2, 1) and CONV4(N×N, 2, 1) operations: N, M, W3[2, M, M], W4[2, N, N].
[0490] - The number and offset values for the Mask&Offset1 and Mask&Offset2 operations: Q, C1[Q], C2[Q].
[0491] -Block size and adjustment weights for output adjust1 and output adjust2 operations: Bs,
[0492] wP is set to equal to 17.
[0493]
[0494]
[0495] A.4 Parameter Semantics
[0496] `best_cand_u_idx` - Specifies a 4-bit non-negative integer value corresponding to the candidate index of the `u` component (the first component in the minor components), indicating the number of slices and slice coordinates. It is used as input to the `cand[X][Y]` table in Section I.5. `best_cand_v_idx` - Specifies a 4-bit non-negative integer value corresponding to the candidate index of the `v` component (the second component in the minor components), indicating the number of slices and slice coordinates. It is used as input to the `cand[X][Y]` table in Section I.5. `fl_U` - Specifies a 6-value non-negative integer value for the kernel size of the CONV3 (M×M, 2, 1) processing layer, i.e., M = `fl_U`.
[0497] fl_V - Specifies a non-negative integer value of 6 for the kernel size of the CONV4(N×N, 2, 1) processing layer, i.e., N = fl_V.
[0498] WU - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights) of the CONV3(M×M, 2, 1) processing layer.
[0499] WV - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights) of the CONV4(N×N, 2, 1) processing layer.
[0500] bS - A 10-bit non-negative integer value specifying the block size of the output adjust1 and output adjust2 processing layers.
[0501] len_mask_1_x - A 10-bit non-negative integer value that specifies the number of elements in the vertical direction of the S1 tensor.
[0502] len_mask_1_y - A 10-bit non-negative integer value specifying the number of elements in the horizontal direction of the S1 tensor.
[0503] len_mask_2_x - A 10-bit non-negative integer value that specifies the number of elements in the vertical direction of the S2 tensor.
[0504] len_mask_2_y - A 10-bit non-negative integer value specifying the number of elements in the horizontal direction of the S2 tensor.
[0505] B[1] - Specifies the 16-bit value of the bias (accumulated component) of the CONV1(1×1,1,1) processing layer.
[0506] B[2] - Specifies the 16-bit value of the bias (accumulated component) of the CONV2(1×1,1,1) processing layer.
[0507] W1 - Specifies the 16-bit value of the weights (multiplication components) of the CONV1 (1×1, 1, 1) processing layer.
[0508] W2 - Specifies the 16-bit value of the weights (multiplication components) of the CONV2 (1×1, 1, 1) processing layer.
[0509] S1 - Specifies the 3 non-negative integer values for the multiplication coefficients of the output adjust1 processing layer.
[0510] S2 - Specifies that the 3 values of the multiplication coefficients of the output adjust2 processing layer are non-negative.
[0511] C1 - Specifies the wP bit value of the accumulated offset parameter used in the Mask & Offset1 processing layer.
[0512] C2 - Specifies the wP bit value of the accumulated offset parameter used in the Mask&Offset2 processing layer.
[0513] A.5 zoning
[0514] The weights (i.e., W3[2, M, M] and W4[2, N, N]) of the CONV3(M×M, 2, 1) and CONV4(N×N, 2, 1) operations are set based on the spatial coordinates of the processed sample. In other words, rectangular swatches can be used in the processing of input samples. If the spatial coordinates of the processed sample are in (x, y), the weight parameters are set as follows:
[0515]
[0516] The cand[X][Y][4] table referenced in Sections I.3 and I.4 includes the number of slices and the coordinates of the slices.
[0517]
[0518] Figure 39 An example layer structure for EFE is shown.
[0519] Another example implementation of the proposed solution is in Figure 39 It is depicted in the middle. With Figure 38 Compared to the example in the previous example, the subtraction operation has been removed in this example.
[0520] Section B provides an example of signal transmission as an alternative.
[0521] B.1 Parameters are transmitted via signal.
[0522] To perform the processing steps described in Sections I.1 and I.2, adjustable weights, biases, and offset parameters are transmitted via signals in the image header. The parameters transmitted via signals in the image header are:
[0523] Weights and biases of the CONV1(1×1,1,1) and CONV1(1×1,1,1) operations: W1[1], W2[1], B[2].
[0524] - Kernel size and weights for CONV3(M×M, 2, 1) and CONV4(N×N, 2, 1) operations: N, M, W3[2, M, M], W4[2, N, N].
[0525] - The number and offset values for the Mask&Offset1 and Mask&Offset2 operations: Q, C1[Q], C2[Q].
[0526] -Block size and adjustment weights for output adjust1 and output adjust2 operations: Bs,
[0527]
[0528]
[0529]
[0530] B.2 Parameter Semantics
[0531] `best_cand_u_idx` - Specifies a 4-bit non-negative integer value corresponding to the candidate index of the `u` component (the first component among the minor components), indicating the number of slices and slice coordinates. It is used as input to the `cand[X][Y]` table in Section I.5.
[0532] `best_cand_v_idx` - Specifies a 4-bit non-negative integer value corresponding to the candidate index of the `v` component (the second component in the minor components), indicating the number of slices and slice coordinates. It is used as input to the `cand[X][Y]` table in Section I.5.
[0533] fl_U - Specifies a 6-valued non-negative integer for the kernel size of the CONV3 (M×M, 2, 1) processing layer, i.e., M = fl_U.
[0534] fl_V - Specifies a non-negative integer value of 6 for the kernel size of the CONV4(N×N, 2, 1) processing layer, i.e., N = fl_V.
[0535] WU - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights) of the CONV3(M×M, 2, 1) processing layer.
[0536] WV - A 4-dimensional tensor that specifies the multiplier coefficients (e.g., weights) of the CONV4(N×N, 2, 1) processing layer.
[0537] bS - A 10-bit non-negative integer value specifying the block size of the output adjust1 and output adjust2 processing layers.
[0538] minSymbol - Specifies a 17-bit non-negative integer value to be added to the multiplier coefficients WU, WV and the offset parameters C1 and C2.
[0539] maxSymbol - Specifies a 17-bit non-negative integer value that is used during the uf() decoding of multiplier coefficients WU, WV, and offset parameters C1 and C2.
[0540] mask1_enabled_flag - Specifies whether the values of len_mask_1_x and len_mask_1_y are zero or a 1-bit non-negative integer value greater than zero.
[0541] mask2_enabled_flag - Specifies whether the values of len_mask_2x and len_mask_2y are zero or a 1-bit non-negative integer greater than zero.
[0542] len_mask_1_x - A 10-bit non-negative integer value that specifies the number of elements in the vertical direction of the S1 tensor.
[0543] len_mask_1_y - A 10-bit non-negative integer value specifying the number of elements in the horizontal direction of the S1 tensor.
[0544] len_mask_2_x - A 10-bit non-negative integer value that specifies the number of elements in the vertical direction of the S2 tensor.
[0545] len_mask_2_y - A 10-bit non-negative integer value specifying the number of elements in the horizontal direction of the S2 tensor.
[0546] B[1] - Specifies the 16-bit value of the bias (accumulated component) of the CONV1(1×1,1,1) processing layer.
[0547] B[2] - Specifies the 16-bit value of the bias (accumulated component) of the CONV2(1×1,1,1) processing layer.
[0548] W1 - Specifies the 16-bit value of the weights (multiplication components) of the CONV1 (1×1, 1, 1) processing layer.
[0549] W2 - Specifies the 16-bit value of the weights (multiplication components) of the CONV2 (1×1, 1, 1) processing layer.
[0550] S1 - Specifies the 3 non-negative integer values for the multiplication coefficients of the output adjust1 processing layer.
[0551] S2 - Specifies the 3 non-negative integer values for the multiplication coefficients of the output adjust2 processing layer.
[0552] C1 - Specifies the wP bit value of the accumulated offset parameter used in the Mask & Offset1 processing layer.
[0553] C2 - Specifies the wP bit value of the accumulated offset parameter used in the Mask&Offset2 processing layer.
[0554] In Section B, the parameters for convolution and filtering operations are presented using alternative methods of signal transmission. The uf() operator is described in the syntax table. The definition of the uf() operation is as follows:
[0555] uf(x): Syntax elements are encoded and decoded using a uniform probability distribution. The minimum value of the distribution is 0, and its maximum value is x.
[0556] Based on the proposed solution,
[0557] ● First, the maximum and / or minimum values are included in (or decoded from) the bitstream. These are described as minSymbol and maxSymbol in Section B1 above (lines 9 and 10). These values are first encoded into (or decoded from) the bitstream to indicate the range of values that some subsequent syntax elements can assume.
[0558] ●After minSymbol and maxSymbol in the encoding / decoding order, syntax elements can be encoded / decoded based on the value of maxSymbol.
[0559] ●The value of minSymbol can be added to the encoded / decoded value (decoded value).
[0560] For example, in Section B1, on line 16, the weights WU[0, idx, i, j] of the convolution operation are obtained as follows:
[0561] ● First, the syntax element A1 is obtained based on the value of maxSymbol. This is described in uf(maxSymbol). The syntax element A1 is obtained based on the maximum value of maxSymbol. The maximum value of A1 can be assumed to be maxSymbol.
[0562] ●Additionally or alternatively, minSymbol is added to A1, and the weights of the convolution parameters are obtained based on A1 + minSymbol. This is depicted in line 16.
[0563] In another example in Section B1, on line 58, the offset parameter C1[i] is obtained as follows:
[0564] ● First, the syntax element A1 is obtained based on the value of maxSymbol. This is depicted in uf(maxSymbol) on line 57. The syntax element A1 is obtained based on the maximum value of maxSymbol. The maximum value of A1 can be assumed to be maxSymbol.
[0565] ●Additionally or alternatively, minSymbol is added to A1, and the offset parameter C1[i] is obtained according to A1+minSymbol. This is depicted in line 18.
[0566] The advantage of using minimum (minSymbol) and / or maximum (maxSymbol) values lies in the encoding and decoding of convolution weights or offset parameters, allowing for content adaptation. In some compressed images, the values of convolution and offset parameters can fall within a small range. In these cases, the value of maxSymbol can be very small. In encoding and decoding syntax elements that can assume a small range of values, very little side information needs to be transmitted. On the other hand, there may exist images where the values of syntax elements do not fall within a small range, so maxSymbol can be increased. In these cases, more side information needs to be transmitted. The proposed solution allows for bit rate savings when the image to be encoded or decoded results in syntax elements whose values fall within a small range of values.
[0567] On the encoder side, the encoder can estimate the values of minSymbol and / or maxSymbol by calculating the minimum and maximum values of all syntax elements encoded and decoded according to minSymbol and / or maxSymbol. For example, in Section B1, minSymbol can be obtained as the minimum of all values of WU[0, idx, i, j] or WV[0, idx, i, j]. Or it can be obtained as the minimum of all values of C1[i]. Similarly, maxSymbol can be obtained as the maximum of all values of WU[0, idx, i, j] or WV[0, idx, i, j]. Or it can be obtained as the maximum of all values of C1[i].
[0568] According to the proposed solution, flags are included in the bitstream to indicate whether a masking operation is performed on the components. For example, `mask1_enabled_flag` (line 25 in Section B1) is included in the bitstream to indicate whether the masking process is enabled. If `mask1_enabled_flag` is true, the number of mask samples in the horizontal and vertical directions (lines 30 and 31) can be included in the bitstream. Alternatively or additionally, if `mask1_enabled_flag` is true, the block size (bS, e.g., line 29) can be included in the bitstream.
[0569] In another example, two flags, mask1_enabled_flag and mask2_enabled_flag, can be included in the bitstream to indicate whether masking operations are enabled for the first and second components, respectively. If at least one of the flags is true (e.g., the check in line 28), one of the following is included in the bitstream:
[0570] ● Block size,
[0571] ● The number of mask elements in the vertical direction (e.g., len mask y).
[0572] ● The number of mask elements in the horizontal direction (e.g., len mask x).
[0573] 6.2.3. Example
[0574] 1. Decoder / encoder example:
[0575] An image or video decoding or encoding method includes a neural subnetwork comprising the following:
[0576] ● Obtain the first and second components of the image based on the bitstream.
[0577] ● The two components are processed using a convolutional layer that modifies at least one of the components, wherein the weights of the convolutional layer are obtained from the bitstream.
[0578] ● The reconstructed image is obtained based on at least one of the two modified components.
[0579] 2. According to the above embodiments;
[0580] ●The first and second components are obtained by composition transformation.
[0581] 3. According to the second embodiment;
[0582] ●The first component is obtained using the first composition transformation, and the second component is obtained using the second composition transformation.
[0583] 4. According to any of the embodiments described above, processing two components using a convolutional layer includes:
[0584] ● First, obtain the mean of at least one of the first or second components.
[0585] ●The mean is subtracted from the at least one component before processing with the convolutional layer.
[0586] 5. According to any of the embodiments described above, processing two components using a convolutional layer to obtain modified component 1 includes any of the following:
[0587] ●Component1=conv(in2-E(in2), in1-E(in1))+in1+K
[0588] ●Component1=conv(in2,in1-E(in1))+in1+K
[0589] ●Component1=conv(in2-E(in1),in1)+K
[0590] ●Component1=conv(in2,in1)+K
[0591] ●Component1=conv(in2,in1-E(in1))+E(in1)+K
[0592] ●Component1=conv(in2-E(in2), in1-E(in1))+E(in1)+K
[0593] ●Component1=conv(in2-E(in2))+in1+K
[0594] ●Component1=conv(in2-E(in2))+in1
[0595] Where in1 is the unmodified component 1 before the convolutional layer, in2 is the unmodified component 2, conv() describes the convolution operation, K is a scalar, and E() is the mean operation.
[0596] 6. According to Example 5, K equals 0.
[0597] 7. According to any of the embodiments described above, the convolutional layer may have two or more outputs (e.g., modified component 1, modified component 2, modified component 3, etc.).
[0598] Further details of embodiments of this disclosure will be described below, relating to neural network-based visual data encoding and decoding. As used herein, the term "visual data" can refer to video, images, pictures in video, or any other visual data suitable for encoding and decoding.
[0599] As discussed above, in existing designs for visual data encoding and decoding based on neural networks (NNs), only a single filtering process is used to generate the reconstructed visual data. This can degrade the quality of the reconstructed visual data if the content of the visual data is diverse.
[0600] To address the above-mentioned problems and other issues not mentioned, a visual data processing solution is disclosed as described below. The embodiments of this disclosure should be considered as examples illustrating general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.
[0601] Figure 40 A flowchart of a method 4000 for visual data processing according to some embodiments of the present disclosure is shown. Method 4000 may be implemented during a conversion between visual data and a bitstream of visual data, the conversion being performed using a neural network (NN)-based model. As used herein, the NN-based model can be a model based on neural network techniques. For example, the NN-based model may specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module may include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the NN-based model described herein are merely illustrative and should therefore not be construed as limiting the present disclosure in any way.
[0602] like Figure 40 As shown, method 4000 begins at 4002, where a target reconstruction of a first component of the visual data is determined based on a first candidate reconstruction and a second candidate reconstruction of the first component. The first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. By way of example and not limitation, the first filtering process may include a first upsampling process. Additionally or alternatively, the second filtering process may include a second upsampling process.
[0603] In some embodiments, both the first and second filtering processes are adaptive filtering processes. For example, one or more filtering parameters (such as multiplication parameters, addition parameters, etc.) for the first and second filtering processes can be obtained based on information indicated in the bitstream. Furthermore, one or more filtering parameters are different for the first and second filtering processes. Alternatively, the filtering parameters(s) for the first and second filtering processes can be fixed and different.
[0604] In some embodiments, the first filtering process may be a Discrete Cosine Transform-based interpolation filter (DCT-IF), while the second filtering process may be a bicubic filter. Alternatively, the first filtering process may be a Lanzos filter, while the second filtering process may be a bilinear filter. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0605] In some other embodiments, the first filtering process and / or the second filtering process may be applied to the first component based on a second component different from the first component. By way of example, and not limitation, the second component may be scaled down, for example, using a downsampling operation or a de-shuffling operation, to obtain a downsized second component. Furthermore, the first component and the downsized second component (which may be) may be processed based on the following:
[0606] Conv1(recU)+Conv2(recY d ),
[0607] Where recU represents the first component, recY d The second component is represented by the scaled-down convolution function, Convl() represents the first convolution function, and Conv2() represents the second convolution function. Since the outputs of the two convolution functions are summed, this process can also be implemented using a single convolution function. In some embodiments, the result of this process can be scaled up (e.g., using a shuffling operation) to obtain the filtered first component. For example, the parameter values of the first and / or second convolution functions can be different for the first and second filtering processes, thus allowing different candidate reconstructions of the first component to be obtained.
[0608] In one example, the first component may include a secondary component, and the second component may include a primary component. Alternatively, the first component may include a chromaticity component, and the second component may include a luminance component. In another example, the first component may include a U component and / or a V component, and the second component may include a Y component. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0609] In some embodiments, the first component can be reconstructed using a synthetic transformation in a neural network-based model. By way of example, and not limitation, the synthetic transformation can be a neural network used to transform a latent representation of visual data from the transform domain to the pixel domain. In one example, the first component can be directly output by the synthetic transformation. Alternatively, the first component can be obtained by further processing the output of the synthetic transformation.
[0610] At 4004, the conversion is performed based on the target reconstruction. By way of example, and not limitation, the visual data can be reconstructed based on the target reconstruction. In some embodiments, the conversion may include encoding the visual data into a bitstream. Additionally or alternatively, the conversion may include decoding the visual data from the bitstream. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0611] In light of the above, two distinct filtering processes are used to generate two candidate reconstructions of the component, and these two candidate reconstructions are further used to generate the target reconstruction of that component. Compared to conventional solutions that use only a single filtering process to generate the component reconstruction, the proposed method advantageously utilizes two different filtering processes to generate the component reconstruction. In this way, the encoding and decoding process can be adapted to the content of the visual data, thereby improving the encoding and decoding quality.
[0612] In some embodiments, at 4002, the first candidate reconstruction and the second candidate reconstruction can be combined based on edge information to obtain the target reconstruction. The combination result can be one of the following: the first candidate reconstruction itself, the second candidate reconstruction itself, or a mixture of the first candidate reconstruction and the second candidate reconstruction.
[0613] For ease of discussion, the first sample point in the first candidate reconstruction, the second sample point in the second candidate reconstruction, and the third sample point in the target reconstruction will be illustrated as examples. The first and second sample points correspond to the third sample point. In one example, the first, second, and third sample points may share the same coordinates. In another example, both the first and second sample points may share the same coordinates, and the coordinates of the third sample point may be determined based on the coordinates of the first and second sample points, for example, depending on the color format, etc.
[0614] In some embodiments, the third sample point can be obtained by determining a weighted sum of the first and second sample points based on edge information. For example, the first weight used to weight the first sample point and / or the second weight used to weight the second sample point can be determined based on edge information. Furthermore, the sum of the first and second weights can be equal to a predetermined value, such as 1.
[0615] By way of examples rather than restrictions, the weighted sum can be determined based on the following:
[0616] A×W+B×(1-W),
[0617] Where A represents the first sample point, B represents the second sample point, and W represents the first weight. The first weight can be determined based on edge information.
[0618] In some embodiments, the third sample point can be equal to one of the following: the first sample point, the second sample point, or the average of the first and second sample points. For example, side information can indicate one of three possible values. If the side information indicates a first value, the third sample point can be equal to the first sample point. If the side information indicates a second value, the third sample point can be equal to the second sample point. If the side information indicates a third value, the third sample point can be equal to the average of the first and second sample points. In some embodiments, the side information can be indicated in one or more bitstreams.
[0619] As an example result of the combination process, each sample point of the target reconstruction is set equal to the sample point from the first candidate reconstruction, which can be determined as the target reconstruction. As another example result of the combination process, each sample point of the target reconstruction is set equal to the sample point from the second candidate reconstruction, which can be determined as the target reconstruction. As yet another example result of the combination process, the target reconstruction may include at least one sample point from the first candidate reconstruction and at least one sample point from the second candidate reconstruction. As yet another example result of the combination process, at least one sample point of the target reconstruction can be determined by averaging the sample points from the first and second candidate reconstructions. In this way, the reconstruction of visual data can be adapted to the content of the visual data, thereby improving the encoding and decoding quality.
[0620] In some embodiments, the target reconstruction can be divided into multiple patches. For example, a patch can be a rectangular sub-block corresponding to a component. It should be understood that patches can also have any other suitable shape. At least two of the multiple patches can be determined based on different combinations of a first candidate reconstruction and a second first candidate reconstruction. For example, different weight values can be used for samples from different patches.
[0621] By way of example, one of at least two slices can be determined based on a first set of weights used to weight the samples of the first and second candidate reconstructions. The other slice can be determined based on a second set of weights used to weight the samples of the first and second candidate reconstructions, and the second set of weights may differ from the first set of weights. In this way, the encoding and decoding process can be adapted to the content of visual data with smaller granularity, thereby further improving the encoding and decoding quality.
[0622] In some embodiments, all samples within one of a plurality of slices can be determined based on the same combination of a first candidate reconstruction and a second candidate reconstruction. For example, the same weight pair (i.e., the first weight and the second weight) can be used to determine all samples within the slice.
[0623] As an example result of the combination process, the sample points of the first piece of at least two pieces can be determined from the sample points of the first candidate reconstruction. Additionally, the sample points of the second piece of at least two pieces can be determined from the sample points of the second candidate reconstruction. Alternatively, the sample points of the third piece of at least two pieces can be determined by averaging the sample points of the first and second candidate reconstructions.
[0624] In some embodiments, the size of one or more of the multiple slices can be M×M, and M can be an integer. For example, the slice(s) at the boundary can have a size different from M×M. In some other embodiments, the size of each of the multiple slices can be M×M. In one example, M can be indicated in the bitstream. In another example, M can be predetermined.
[0625] In some embodiments, the size of one of a plurality of slices can be indicated as a block size or a slice size. Additionally or alternatively, the number of slices can be indicated in one or more bitstreams. For example, the number of slices in the horizontal direction and / or the number of slices in the vertical direction can be indicated in one or more bitstreams. By way of example, and not limitation, the number of slices in the horizontal direction can be indicated by the syntax element len_mask_y. Additionally or alternatively, the number of slices in the vertical direction can be indicated by the syntax element len_mask_x.
[0626] In some other embodiments, at least one of the following may be included in one or more bitstreams based on the flags: the number of slices, the number of slices in the horizontal direction, the number of slices in the vertical direction, the block size, or the slice size. Flags may be included in one or more bitstreams. By way of example, and not limitation, the flag may be an enable flag, such as mask1_enabled_flag or mask2_enabled_flag.
[0627] Additionally or alternatively, the number of samples in a slice can be indicated in one or more bitstreams. For example, the number of samples in the horizontal direction and / or the number of samples in the vertical direction can be indicated in one or more bitstreams. By way of example, and not limitation, the number of samples in the horizontal direction can be indicated by the syntax element len_mask_y. Additionally or alternatively, the number of samples in the vertical direction can be indicated by the syntax element len_mask_x.
[0628] In some other embodiments, at least one of the following may be included in one or more bitstreams based on the flags: the number of samples in a slice, the number of samples in the horizontal direction, the number of samples in the vertical direction, the block size, or the sample size. Flags may be included in one or more bitstreams. By way of example, and not limitation, the flag may be an enabling flag, such as mask1_enabled_nag or mask2_enabled_flag.
[0629] In some embodiments, one or more parameters (such as at least one weight, bias, etc.) used for the first filtering process and / or the second filtering process may be obtained based on information indicated in one or more bitstreams. For example, one or more parameters used for the first filtering process may be different from one or more parameters used for the second filtering process.
[0630] In some embodiments, the first filtering process and / or the second filtering process may include at least one of a convolution operation or a shuffling operation. For example, the kernel size of the convolution operation may be N×N, and N may be an integer. In one example, the kernel size of the convolution operation may be indicated in the bitstream. Alternatively, the kernel size may be predetermined.
[0631] In some embodiments, the kernel size of the convolution operation in the first filtering process is different from the kernel size of the convolution operation in the second filtering process. In some other embodiments, the first filtering process and / or the second filtering process may include at least one of a deconvolution operation or a deshuffling operation.
[0632] In view of the above, the solutions according to some embodiments of the present disclosure can advantageously enable the encoding and decoding process to adapt to the content of the visual data, thereby improving the encoding and decoding quality.
[0633] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by a visual data processing apparatus. In this method, a target reconstruction of a first component of the visual data is determined based on a first candidate reconstruction and a second candidate reconstruction of the first component. The first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. Furthermore, the bitstream is generated based on the target reconstruction using a neural network (NN)-based model.
[0634] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. In this method, a target reconstruction of a first component of the visual data is determined based on a first candidate reconstruction and a second candidate reconstruction of the first component. The first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. Furthermore, the bitstream is generated based on the target reconstruction using a neural network (NN)-based model and is stored in a non-transitory computer-readable recording medium.
[0635] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0636] Item 1. A method for visual data processing, comprising: for a conversion between visual data utilizing a neural network (NN)-based model and one or more bitstreams of the visual data, determining a target reconstruction of a first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; and performing the conversion based on the target reconstruction.
[0637] Item 2. The method according to Item 1, wherein the first filtering process includes a first upsampling process.
[0638] Item 3. The method according to any one of items 1-2, wherein the second filtering process includes a second upsampling process.
[0639] Item 4. The method according to any one of items 1-3, wherein determining the target reconstruction comprises: combining the first candidate reconstruction and the second candidate reconstruction based on edge information to obtain the target reconstruction.
[0640] Item 5. The method according to Item 4, wherein the first sample point in the first candidate reconstruction and the second sample point in the second candidate reconstruction correspond to the third sample point in the target reconstruction, and combining the first candidate reconstruction and the second candidate reconstruction includes: obtaining the third sample point by determining a weighted sum of the first sample point and the second sample point based on the edge information.
[0641] Item 6. The method according to Item 5, wherein at least one of the following is determined based on the edge information: a first weight for weighting the first sample point, or a second weight for weighting the second sample point.
[0642] Item 7. The method according to Item 6, wherein the sum of the first weight and the second weight is equal to a predetermined value.
[0643] Item 8. The method according to Item 7, wherein the predetermined value is 1.
[0644] Item 9. The method according to any one of Items 5-8, wherein the third sample point is equal to one of the following: the first sample point, the second sample point, or the average of the first sample point and the second sample point.
[0645] Item 10. The method according to Item 9, wherein if the edge information indicates a first value, then the third sample point is equal to the first sample point; if the edge information indicates a second value, then the third sample point is equal to the second sample point; and if the edge information indicates a third value, then the third sample point is equal to the average of the first sample point and the second sample point.
[0646] Item 11. The method according to any one of items 5-10, wherein the coordinates of the first sample point, the second sample point, and the third sample point are the same.
[0647] Item 12. The method according to any one of Items 4-11, wherein the side information is indicated in the one or more bit streams.
[0648] Item 13. The method according to any one of items 1-12, wherein the first candidate reconstruction is determined as the target reconstruction, or wherein the second candidate reconstruction is determined as the target reconstruction, or wherein the target reconstruction includes at least one sample from the first candidate reconstruction and at least one sample from the second candidate reconstruction, or wherein at least one sample of the target reconstruction is determined by averaging the sample from the first candidate reconstruction and the sample from the second candidate reconstruction.
[0649] Item 14. The method according to any one of items 1-13, wherein the target reconstruction is divided into a plurality of slices, and at least two of the plurality of slices are determined based on different combinations of the first candidate reconstruction and the second first candidate reconstruction.
[0650] Item 15. The method according to Item 14, wherein one of the at least two slices is determined based on a first set of weights for weighting the samples of the first candidate reconstruction and the second candidate reconstruction, and the other slice of the at least two slices is determined based on a second set of weights for weighting the samples of the first candidate reconstruction and the second first candidate reconstruction, and the second set of weights is different from the first set of weights.
[0651] Item 16. The method according to any one of Items 14-15, wherein all samples of one of the plurality of slices are determined based on the same combination scheme of the first candidate reconstruction and the second candidate reconstruction.
[0652] Item 17. The method according to any one of items 14-16, wherein the sample points of the first piece of the at least two pieces are determined from the sample points of the first candidate reconstruction, or the sample points of the second piece of the at least two pieces are determined from the sample points of the second candidate reconstruction, or the sample points of the third piece of the at least two pieces are determined by averaging the sample points of the first candidate reconstruction and the second candidate reconstruction.
[0653] Item 18. The method according to any one of items 14-17, wherein the size of one of the plurality of pieces is M×M, and M is an integer.
[0654] Item 19. The method according to any one of items 14-19, wherein the size of each of the plurality of pieces is M×M, and M is an integer.
[0655] Item 20. The method according to any one of items 18-19, wherein M is indicated in the bit stream.
[0656] Item 21. The method according to any one of Items 14-20, wherein the size of one of the plurality of sheets is indicated as a block size or a sheet size.
[0657] Item 22. The method according to any one of items 14-21, wherein the number of slices is indicated in the one or more bit streams.
[0658] Item 23. The method according to any one of items 14-22, wherein the number of slices in the horizontal direction or the number of slices in the vertical direction is indicated in the one or more bit streams.
[0659] Item 24. The method according to any one of items 14-23, wherein at least one of the following is included in the one or more bit streams based on a flag included in the one or more bit streams: the number of slices, the number of slices in the horizontal direction, the number of slices in the vertical direction, the block size, or the slice size.
[0660] Item 25. The method according to Item 14, wherein the flag is an enable flag.
[0661] Item 26. The method according to any one of items 1-25, wherein at least one of the first filtering process or the second filtering process is adaptive.
[0662] Item 27. The method according to any one of items 1-26, wherein one or more parameters used in at least one of the first filtering process or the second filtering process are obtained based on information indicated in the one or more bit streams.
[0663] Item 28. The method according to Item 27, wherein the one or more parameters include at least one of the following: at least one weight or bias.
[0664] Item 29. The method according to any one of Items 27-28, wherein one or more parameters used for the first filtering process are different from one or more parameters used for the second filtering process.
[0665] Item 30. The method according to any one of items 1-29, wherein the first filtering process or at least one of the second filtering processes includes at least one of the following: a convolution operation or a shuffling operation.
[0666] Item 31. The method according to Item 30, wherein the kernel size of the convolution operation is N×N, and N is an integer.
[0667] Item 32. The method according to any one of items 30-31, wherein the kernel size of the convolution operation is indicated in the bitstream.
[0668] Item 33. The method according to any one of items 30-32, wherein the kernel size of the convolution operation in the first filtering process is different from the kernel size of the convolution operation in the second filtering process.
[0669] Item 34. The method according to any one of items 1-33, wherein at least one of the first filtering process or the second filtering process includes at least one of the following: a deconvolution operation or a deshuffling operation.
[0670] Item 35. The method according to any one of items 1-34, wherein the first component is reconstructed using a synthetic transformation in the NN-based model.
[0671] Item 36. The method according to any one of items 1-35, wherein the first component includes a secondary component, or wherein the first component includes a chromaticity component, or wherein the first component includes at least one of a U component or a V component.
[0672] Item 37. The method according to any one of items 1-36, wherein performing the transformation comprises: reconstructing the visual data based on the target reconstruction.
[0673] Item 38. The method according to any one of items 1-37, wherein the visual data includes video, pictures or images of the video.
[0674] Item 39. The method according to any one of items 1-38, wherein the conversion includes encoding the visual data into the one or more bit streams.
[0675] Item 40. The method according to any one of items 1-38, wherein the conversion includes decoding the visual data from the one or more bitstreams.
[0676] Item 41. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-40.
[0677] Item 42. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-40.
[0678] Item 43. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of means for visual data processing, wherein the method includes: determining a target reconstruction of the first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; and generating the bitstream based on the target reconstruction using a neural network (NN)-based model.
[0679] Item 44. A method for storing a bitstream of visual data, comprising: determining a target reconstruction of a first component based on a first candidate reconstruction and a second candidate reconstruction of a first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process; generating the bitstream based on the target reconstruction using a neural network (NN)-based model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0680] Example device
[0681] Figure 41 A block diagram of a computing device 4100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4100 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).
[0682] It should be understood that, Figure 41 The computing device 4100 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0683] like Figure 41 As shown, computing device 4100 includes general-purpose computing device 4100. Computing device 4100 may include at least one or more processors or processing units 4110, memory 4120, storage unit 4130, one or more communication units 4140, one or more input devices 4150, and one or more output devices 4160.
[0684] In some embodiments, the computing device 4100 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4100 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0685] Processing unit 4110 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 4120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 4100. Processing unit 4110 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0686] Computing device 4100 typically includes various computer storage media. Such media can be any media accessible by computing device 4100, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 4120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 4130 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or visual data and can be accessed within computing device 4100.
[0687] The computing device 4100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 41 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more visual data media interfaces.
[0688] Communication unit 4140 communicates with another computing device via a communication medium. Furthermore, the functionality of components in computing device 4100 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 4100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0689] Input device 4150 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4160 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 4140, computing device 4100 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 4100 can also communicate with one or more devices that enable a user to interact with computing device 4100, or any device that enables computing device 4100 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0690] In some embodiments, instead of being integrated into a single device, some or all components of computing device 4100 may be arranged in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (e.g., the Internet) using suitable protocols. For example, a cloud computing provider provides applications via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations within remote visual data centers. Cloud computing infrastructure may be provided through shared visual data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0691] In embodiments of this disclosure, computing device 4100 can be used to implement visual data encoding / decoding. Memory 4120 may include one or more visual data encoding / decoding modules 4125 having one or more program instructions. These modules can be accessed and executed by processing unit 4110 to perform the functions of the various embodiments described herein.
[0692] In an example embodiment of performing visual data encoding, input device 4150 may receive visual data as input 4170 to be encoded. The visual data may be processed, for example, by visual data encoding / decoding module 4125 to generate an encoded bitstream. The encoded bitstream may be provided as output 4180 via output device 4160.
[0693] In an example embodiment of performing visual data decoding, input device 4150 may receive an encoded bitstream as input 4170. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 4125 to generate decoded visual data. The decoded visual data may be provided as output 4180 via output device 4160.
[0694] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for visual data processing, comprising: For the conversion between visual data using a neural network (NN)-based model and one or more bitstreams of the visual data, a target reconstruction of the first component is determined based on a first candidate reconstruction and a second candidate reconstruction of the first component of the visual data, wherein the first candidate reconstruction is generated based on a first filtering process and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. as well as The transformation is performed based on the target reconstruction.
2. The method according to claim 1, wherein the first filtering process includes a first upsampling process.
3. The method according to any one of claims 1-2, wherein the second filtering process includes a second upsampling process.
4. The method according to any one of claims 1-3, wherein determining the target reconstruction comprises: The first candidate reconstruction and the second candidate reconstruction are combined based on edge information to obtain the target reconstruction.
5. The method of claim 4, wherein the first sample point in the first candidate reconstruction and the second sample point in the second candidate reconstruction correspond to the third sample point in the target reconstruction, and combining the first candidate reconstruction and the second candidate reconstruction comprises: The third sample point is obtained by determining the weighted sum of the first and second sample points based on the edge information.
6. The method of claim 5, wherein at least one of the following is determined based on the edge information: The first weight used to weight the first sample points, or The second weight is used to weight the second sample point.
7. The method of claim 6, wherein the sum of the first weight and the second weight is equal to a predetermined value.
8. The method of claim 7, wherein the predetermined value is 1.
9. The method according to any one of claims 5-8, wherein the third sample point is equal to one of the following: The first sample point, The second sample point, or The average value of the first sample point and the second sample point.
10. The method of claim 9, wherein if the edge information indicates a first value, then the third sample point is equal to the first sample point, or If the edge information indicates a second value, then the third sample point is equal to the second sample point, or If the edge information indicates a third value, then the third sample point is equal to the average value of the first sample point and the second sample point.
11. The method according to any one of claims 5-10, wherein the coordinates of the first sample point, the second sample point, and the third sample point are the same.
12. The method according to any one of claims 4-11, wherein the side information is indicated in the one or more bit streams.
13. The method according to any one of claims 1-12, wherein the first candidate reconstruction is determined as the target reconstruction, or Wherein the second candidate reconstruction is determined as the target reconstruction, or The target reconstruction includes at least one sample from the first candidate reconstruction and at least one sample from the second candidate reconstruction, or At least one sample of the target reconstruction is determined by averaging the sample of the first candidate reconstruction and the sample of the second candidate reconstruction.
14. The method according to any one of claims 1-13, wherein the target reconstruction is divided into a plurality of slices, and at least two of the plurality of slices are determined based on different combinations of the first candidate reconstruction and the second first candidate reconstruction.
15. The method of claim 14, wherein one of the at least two slices is determined based on a first set of weights for weighting the samples of the first candidate reconstruction and the second candidate reconstruction, and the other of the at least two slices is determined based on a second set of weights for weighting the samples of the first candidate reconstruction and the second first candidate reconstruction, and the second set of weights is different from the first set of weights.
16. The method according to any one of claims 14-15, wherein all samples of one of the plurality of slices are determined based on the same combination scheme of the first candidate reconstruction and the second first candidate reconstruction.
17. The method according to any one of claims 14-16, wherein the sample points of the first piece of the at least two pieces are determined from the sample points of the first candidate reconstructed piece, or The sample points of the second of the at least two slices are determined from the sample points of the second candidate reconstruction, or The samples of the third slice of the at least two slices are determined by averaging the samples of the first candidate reconstruction and the second candidate reconstruction.
18. The method according to any one of claims 14-17, wherein the size of one of the plurality of slices is M×M, and M is an integer.
19. The method according to any one of claims 14-19, wherein the size of each of the plurality of slices is M×M, and M is an integer.
20. The method according to any one of claims 18-19, wherein M is indicated in the bit stream.
21. The method according to any one of claims 14-20, wherein the size of one of the plurality of sheets is indicated as a block size or a sheet size.
22. The method according to any one of claims 14-21, wherein the number of slices is indicated in the one or more bit streams.
23. The method according to any one of claims 14-22, wherein the number of slices in the horizontal direction or the number of slices in the vertical direction is indicated in the one or more bit streams.
24. The method according to any one of claims 14-23, wherein at least one of the following is included in the one or more bitstreams based on a flag included in the one or more bitstreams: The number of pieces The number of pieces in the horizontal direction, The number of slices in the vertical direction, Block size, or Piece size.
25. The method of claim 14, wherein the flag is an enable flag.
26. The method according to any one of claims 1-25, wherein at least one of the first filtering process or the second filtering process is adaptive.
27. The method according to any one of claims 1-26, wherein one or more parameters used in at least one of the first filtering process or the second filtering process are obtained based on information indicated in the one or more bit streams.
28. The method of claim 27, wherein the one or more parameters include at least one of the following: At least one weight, or Bias.
29. The method according to any one of claims 27-28, wherein one or more parameters used for the first filtering process are different from one or more parameters used for the second filtering process.
30. The method according to any one of claims 1-29, wherein at least one of the first filtering process or the second filtering process comprises at least one of the following: Convolution operation, or Mixed washing operation.
31. The method of claim 30, wherein the kernel size of the convolution operation is N×N, and N is an integer.
32. The method according to any one of claims 30-31, wherein the kernel size of the convolution operation is indicated in the bitstream.
33. The method according to any one of claims 30-32, wherein the kernel size of the convolution operation in the first filtering process is different from the kernel size of the convolution operation in the second filtering process.
34. The method according to any one of claims 1-33, wherein at least one of the first filtering process or the second filtering process comprises at least one of the following: Deconvolution operation, or Reverse washing operation.
35. The method according to any one of claims 1-34, wherein the first component is reconstructed using a synthetic transformation in the NN-based model.
36. The method according to any one of claims 1-35, wherein the first component includes a secondary component, or The first component includes a chromaticity component, or The first component includes at least one of the U component or the V component.
37. The method according to any one of claims 1-36, wherein performing the conversion comprises: The visual data is reconstructed based on the target reconstruction.
38. The method according to any one of claims 1-37, wherein the visual data includes video, a picture of the video, or an image.
39. The method according to any one of claims 1-38, wherein the conversion comprises encoding the visual data into the one or more bitstreams.
40. The method of any one of claims 1-38, wherein the conversion comprises decoding the visual data from the one or more bitstreams.
41. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-40.
42. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-40.
43. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of visual data processing, wherein the method includes: Based on the first candidate reconstruction and the second candidate reconstruction of the first component of the visual data, the target reconstruction of the first component is determined, wherein the first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. as well as The target reconstruction utilizes a neural network (NN)-based model to generate the bitstream.
44. A method for storing a bitstream of visual data, comprising: Based on the first candidate reconstruction and the second candidate reconstruction of the first component of the visual data, the target reconstruction of the first component is determined, wherein the first candidate reconstruction is generated based on a first filtering process, and the second candidate reconstruction is generated based on a second filtering process different from the first filtering process. The bitstream is generated based on the target reconstruction using a neural network (NN)-based model; as well as The bitstream is stored in a non-transitory computer-readable recording medium.