Method and device for visual data processing and medium

By employing a multi-synthesis transformation method in visual data processing, and selecting appropriate synthesis transformations to process the reconstructed latent representation of visual data, the problem of insufficient encoding and decoding efficiency in existing technologies is solved, achieving greater flexibility and compatibility.

CN121866767APending Publication Date: 2026-04-14DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing neural network-based image/video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in terms of flexibility and compatibility when processing the same bitstream.

Method used

By employing multiple synthetic transformations, and selecting appropriate synthetic transformations to process the reconstructed latent representation of visual data, the flexibility and efficiency of encoding and decoding are improved.

Benefits of technology

By selecting multiple synthesis transformations, the flexibility and efficiency of encoding and decoding are enhanced, and the compatibility and adaptability for processing the same bit stream are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121866767A_ABST
    Figure CN121866767A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for visual data processing. A method for visual data processing is presented. The method comprises: for a conversion between visual data and a bitstream of visual data using a neural network (NN)-based model, selecting a synthetic transformation from a plurality of synthetic transforms in the NN-based model, the plurality of synthetic transforms being configured for processing a reconstructed potential representation of visual data derived from the same bitstream; and performing a conversion based on the selected composite conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to visual data processing techniques, and more specifically, to visual data encoding and decoding based on neural networks. Background Technology

[0002] Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression technology has made significant strides in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved rate-distortion (RD) performance comparable to that of Multifunctional Video Coding (VVC). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, the encoding and decoding efficiency of neural network-based image / video codecs is generally expected to be further improved. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for visual data processing.

[0004] In a first aspect, a method for visual data processing is proposed. The method includes: for a conversion between visual data and a bitstream of visual data using a neural network (NN)-based model, selecting a synthetic transformation from a plurality of synthetic transformations in the NN-based model, the plurality of synthetic transformations being configured to process a reconstructed latent representation of visual data derived from the same bitstream; and performing the conversion based on the selected synthetic transformation.

[0005] According to the method of the first aspect of this disclosure, multiple synthesis transforms in an NN-based model can process the same bitstream, and one of the multiple synthesis transforms is selected to perform the conversion. Compared with the conventional solution of designing two synthesis transforms to process different bitstreams, the proposed method can advantageously improve the compatibility of multiple synthesis transforms, thereby providing more options for processing the same bitstream to cater to different applications. In this way, encoding and decoding flexibility can be improved, and therefore encoding and decoding efficiency can be enhanced.

[0006] In a second aspect, an apparatus for visual data processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method described according to the first aspect of this disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data processing. The method includes: selecting a synthetic transform from a plurality of synthetic transforms in a neural network (NN)-based model, the plurality of synthetic transforms being configured to process a reconstructed latent representation of visual data derived from the same bitstream; and generating a bitstream using the NN-based model based on the selected synthetic transform.

[0009] In a fifth aspect, a method for storing bitstreams of visual data is proposed. The method includes: selecting a synthetic transform from a plurality of synthetic transforms in a neural network (NN)-based model, the plurality of synthetic transforms being configured to process a reconstructed latent representation of visual data derived from the same bitstream; generating a bitstream using the NN-based model based on the selected synthetic transform; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1A A block diagram illustrating an exemplary visual data encoding / decoding system according to some embodiments of the present disclosure is shown; Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme; Figure 2 An exemplary potential representation of the image is shown; Figure 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model; Figure 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model together with a super-prior and an autoencoder; Figure 5 An exemplary encoding process is shown; Figure 6 An exemplary decoding process is shown; Figure 7 The composite transformation of the high operating point (OP) and the basic OP operating point is shown; Figure 8 An exemplary convolution-based attention block (CAB) is shown. Figure 9 An exemplary transformer-based attention module (TAM) is shown. Figure 10 An exemplary transformer-based attention block (TAB) is shown. Figure 11 An exemplary multi-branch decoder utilizing high-OP synthesis transform is shown according to an embodiment of the present disclosure; Figure 12 An exemplary multi-branch decoder implemented using multiple synthesis transformations according to embodiments of the present disclosure is shown; Figure 13 A flowchart of a method for visual data processing according to embodiments of the present disclosure is shown; and Figure 14 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0013] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation

[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0015] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0016] The terms "an embodiment," "embodiment," "exemplary embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an exemplary embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0017] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0019] Exemplary Environment Figure 1A This is a block diagram illustrating an exemplary visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.

[0020] Visual data source 112 may include sources such as visual data acquisition devices. Examples of visual data acquisition devices include, but are not limited to, interfaces for receiving visual data from visual data content providers, computer graphics systems for generating visual data, and / or combinations thereof.

[0021] Visual data may include one or more images. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the visual data. The bitstream may include encoded images and associated data. The encoded image is an encoded representation of an image. The associated data may include a sequence parameter set, an image parameter set, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data can be directly transmitted to a destination device 120 via network 130A through the I / O interface 116. The encoded visual data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0022] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0023] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or additional standards.

[0024] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Additionally, although some embodiments describe encoding steps in detail, it should be understood that the corresponding decoding steps of the inversion encoding will be implemented by the decoder. Furthermore, the term "visual data processing" includes visual data encoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another or at a different compression bit rate.

[0025] 1. Brief Overview A neural network-based image and video compression method with multiple decoder branches is disclosed. This disclosure relates to a multi-branch decoder, i.e., multiple synthesis transforms. The same bitstream can be decoded using one of the multiple synthesis transforms to extract the latent representation. Obtain the reconstructed image. Multi-branch decoders can be implemented in different ways, such as a single synthesis transform with one or more modules / layers that can be individually turned on or off; multiple synthesis transforms, etc.

[0026] 2. Introduction Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression techniques have made significant progress in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its early stages.

[0027] 2.1 Image / Video Compression Image / video compression generally refers to the computational technique of compressing images / videos into binary code for convenient storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video; this is called lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is not always necessary. The performance of image / video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes; the fewer, the better. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; the higher, the better.

[0028] Image / video compression techniques can be divided into two branches: classical video encoding / decoding methods and neural network-based video compression methods. Classical video encoding / decoding schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully hand-designing entropy encoding / decoding to model dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within existing classical video codecs and exists only as part of the framework, while the latter is a separate framework developed based on neural networks, independent of classical video codecs.

[0029] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-growing volume of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups: the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG). The ITU-T also has its own Video Codec Experts Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard: Multi-Functional Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC achieves an average bitrate reduction of 50% while maintaining the same visual quality.

[0030] Neural network-based image / video compression is not a new invention, as many researchers have worked on neural network-based image encoding and decoding. However, the network architectures are relatively shallow, and the performance is not satisfactory. Thanks to the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, the technology is still far from mature and many challenges need to be addressed.

[0031] 2.2 Neural Networks Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. Notably, these representations are not manually designed; instead, deep networks, including processing layers, are learned from large amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations and is therefore considered particularly suitable for processing native unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.

[0032] 2.3 Neural Networks for Image Compression Existing neural networks used for image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature.

[0033] 2.3.1 Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless encoding and decoding can achieve the lowest possible decoding rate. ,in It is a symbol The probability of [the lossless encoding / decoding method]. Many lossless encoding / decoding methods have been proposed in the literature, among which arithmetic encoding / decoding is considered one of the best methods. Given a probability distribution... Arithmetic encoding and decoding ensures that the encoding / decoding rate is as close as possible to its theoretical limit without considering rounding errors. Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality.

[0034] Following the predictive encoding / decoding strategy, for One approach to modeling this is to predict pixel probabilities one by one in raster scan order based on previous observations. It's an image.

[0035] (1) in These are the height and width of the image, respectively. Previous observations are also referred to as the current pixel's... Context When the image is large, estimating the conditional probability can be difficult, so a simplified approach is to limit the scope of its context.

[0036] (2) in It is a predefined constant that controls the scope of the context.

[0037] It should be noted that this condition can also take into account the sample values ​​of other color components. For example, when encoding and decoding RGB color components, the R sample depends on previously encoded and decoded pixels (including R / G / B samples). The current G sample can be encoded and decoded based on previously encoded and decoded pixels and the current R sample. For encoding and decoding the current B sample, previously encoded and decoded pixels as well as the current R sample and the current G sample can also be considered.

[0038] Neural networks were initially introduced for computer vision tasks and have proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to adjust their behavior based on their context. Estimated probability For binary images, pixel probabilities are proposed, namely... The Neural Autoregressive Distribution Estimator (NADE) is designed for pixel probability modeling, with its feedforward network having a single hidden layer. Similar work has been proposed where the feedforward network also has connections that skip hidden layers, and the parameters are shared. Experiments have been conducted on the binarized MNIST dataset. NADE is extended to a real-valued model, RNADE, where the probability... It is derived using Gaussian mixture derivation. Their feedforward network also has a single hidden layer, but the hidden layer is rescaled to avoid saturation and uses a modified linear unit (ReLU) instead of a sigmoid. NADE and RNADE are improved by reorganizing the pixel order and using a deeper neural network.

[0039] Designing advanced neural networks plays a crucial role in improving pixel probabilistic modeling. Multidimensional Long Short-Term Memory (LSTM) was proposed, which, along with a mixture of conditional Gaussian scaling mixtures, is used for probabilistic modeling. LSTM is a special type of recurrent neural network (RNN) that has proven adept at modeling sequential data. Spatial variants of LSTM have been later applied to images. Several different neural networks, including RNNs and CNNs, namely PixelRNN and PixelCNN, were investigated. In PixelRNN, two variants of LSTM were proposed, called row LSTM and diagonal bidirectional LSTM (BiLSTM), the latter specifically designed for images. PixelRNN incorporates residual connections to support training deep neural networks with up to 12 layers. In PixelCNN, masked convolutions are used to adapt to the shape of the context. Compared to previous work, PixelRNN and PixelCNN are more focused on natural images: they treat pixels as discrete values ​​(e.g., 0, 1, ..., 255) and predict a multinomial distribution over these discrete values; they handle color images in the RGB color space; and they work well on the large-scale image dataset ImageNet. Gated PixelCNN was proposed to improve upon PixelCNN, achieving performance comparable to PixelRNN but with significantly lower complexity. Building upon PixelCNN, PixelCNN++ was proposed with the following improvements: using discretized logistic mixture likelihood instead of a 256-way multinomial distribution; using downsampling to capture structure at multiple resolutions; introducing additional shortcut connections to accelerate training; employing dropout for regularization; and combining RGB values ​​into a single pixel. PixelSNAIL was also proposed, combining causal convolution with self-attention.

[0040] Most of the methods described above model the probability distribution directly in the pixel domain. Some researchers have also attempted to model the probability distribution as a conditional distribution based on explicit or latent representations. That is, it is possible to estimate... (3) in It is an additional condition, and This means that modeling is divided into unconditional modeling and conditional modeling. The additional conditions can be image label information or high-level representations.

[0041] 2.3.2. Automatic Encoder Autoencoders originate from the renowned work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part transforms the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders can automatically learn representations and eliminate the need for hand-crafted features, which is considered one of the most significant advantages of neural networks.

[0042] Figure 1B This diagram illustrates a typical transform encoding / decoding scheme. Original image. Analysis network Transformation to achieve latent representation The latent representation y is quantized and compressed into bits. The number of bits... Used to measure codec rate. Quantized latent representation. Then by the synthetic network Inverse transform to obtain the reconstructed image Distortion is achieved through the use of functions. right The transformation is calculated in the perceptual space.

[0043] Applying autoencoder networks to lossy image compression is intuitive. It simply requires encoding the learned latent representation from a trained neural network. However, adapting autoencoders to image compression is not straightforward, as the original autoencoders are not optimized for compression, making direct use of the trained autoencoder inefficient. Furthermore, other major challenges exist: First, the low-dimensional representation should be quantized before encoding, but quantization is non-differentiable, which is necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios because both distortion and bit rate need to be considered. Estimating the bit rate is challenging. Third, practical image encoding / decoding schemes need to support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Many researchers have been actively contributing to this field to address these challenges.

[0044] An autoencoder prototype for image compression, such as Figure 1B As shown, it can be viewed as a transformation encoding / decoding strategy. Original image Using analysis networks Transformed, where This is the latent representation that will be quantized and encoded / decoded. The synthetic network will process the quantized latent representation... Perform inverse transform to obtain the reconstructed image Framework utilization distortion loss function (i.e. Training is conducted, among which Distortion between It is a representation based on quantification. The calculated or estimated bit rate, and These are Lagrange multipliers. It should be noted that... It can be computed in the pixel domain or the receptive domain. All existing research follows this prototype, with differences only in network structure or loss function.

[0045] In terms of network architecture, RNNs and CNNs are the most widely used architectures. Within the RNN-related category, a general framework for variable-rate image compression using RNNs has been proposed. They use binary quantization to generate the encoding and decoding, and do not consider the rate during training. This framework does indeed provide scalable encoding and decoding capabilities, with RNNs having both convolutional and deconvolutional layers reportedly performing well. An improved version is then proposed by upgrading the encoder to compress the binary encoding and decoding using a neural network similar to PixelRNN. Using the MS-SSIM evaluation metric, its performance on the Kodak image dataset is reported to outperform JPEG. The RNN-based solution is further improved by introducing hidden-state start-up. Furthermore, an SSIM-weighted loss function is designed, and a spatial adaptive bitrate mechanism is enabled. Using MS-SSIM as the evaluation metric, they achieve better results than BPG on the Kodak image dataset.

[0046] A general framework for rate-distortion optimized image compression is designed. They use multivariate quantization to generate integer codes and consider rate during training, i.e., the loss is the joint rate-distortion cost, which can be MSE or others. They add random uniform noise to simulate quantization during training and use the differential entropy of the noisy code as a surrogate for rate. They use Generalized Divisibility Normalization (GDN) as the network structure, which consists of a linear mapping and nonlinear parameter normalization. The effectiveness of GDN on image codes and decoders is validated. An improved version is proposed, where they use three convolutional layers, each followed by a downsampling layer and a GDN layer as the forward transform. Therefore, they use three layers of inverse GDN, each followed by an upsampling layer and a convolutional layer to simulate the inverse transform. Furthermore, an arithmetic codec method is designed to compress integer codes and decoders. In terms of MSE, performance on the Kodak dataset is reported to outperform JPEG and JPEG 2000. Furthermore, the method is further improved by incorporating a variance super-prior into the autoencoder. They use latent representations... With subnet Transform into ,and This will be quantized and transmitted as side information. Therefore, the inverse transform utilizes the subnet. This was achieved; the subnet attempted to extract quantized edge information. Decoding into quantized The standard deviation, which will be... The arithmetic encoding and decoding process was further utilized. On the Kodak image set, their method slightly outperformed BPG in terms of PSNR. These structures were further explored in the residual space by introducing an autoregressive model to estimate the standard deviation and mean. In the latest work, a Gaussian mixture model was used to further remove redundancy in the residuals. Using PSNR as the evaluation metric, the reported performance on the Kodak image set was comparable to VVC.

[0047] 2.3.3. Pre-Prior Model In transform encoding and decoding methods used for image compression, the encoder subnetwork (Section 2.3.2) uses parametric analysis of the transform. Transform the image vector x into a latent representation Then quantify it to form .because Since they are discrete values, they can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as a bit sequence.

[0048] from Figure 2 The left and right center images clearly show this. Significant spatial dependencies exist among the elements. Notably, their scales (middle right image) appear to be spatially coupled. An additional set of random variables was introduced. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as Figure 3 As shown.

[0049] exist Figure 3 In the middle, the encoder is on the left side of the model. and decoder (Explained in Section 2.3.2). The right-hand side is used to obtain... Additional encoders utilizing prior information and decoders that utilize prior information Network. In this architecture, the encoder subjectes the input image x to... The response with a standard deviation exhibiting spatial variation is obtained. .response fed to In summary The distribution of standard deviations in the data. Then... Quantified ( The data is compressed and transmitted as side information. The encoder then uses the quantized vector data... To estimate the spatial distribution of standard deviation And use it to compress and transmit quantized image representations. The decoder first recovers the signal from the compressed signal. Then the decoder uses To obtain ,Should To provide the decoder with the correct probability estimate and thus successfully recover the value. Then the decoder will Feed to To obtain a reconstructed image.

[0050] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. Figure 2 The rightmost image in the middle corresponds to the quantization latent value when using an encoder / decoder that leverages prior information. Compared to the middle right image, spatial redundancy is significantly reduced because the samples of the quantization latent value have lower correlation.

[0051] Figure 2 Left side: Image from the Kodak dataset. Middle left: Visualization of the latent representation y of this image. Middle right: Standard deviation of the latent values. Right side: The latent value y after introducing a super-prior network (an encoder and a decoder that utilize super-prior information).

[0052] Figure 3 The network architecture of an autoencoder implementing a prior model is shown. The left side shows the image autoencoder network, and the right side corresponds to the prior subnetwork. The analytic transform and the synthetic transform are represented as follows: Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks, utilizing the hyperprior information of the encoder (denoted as...). ) and decoders that utilize prior information (represented as The prior model generates quantified potential values ​​of prior information. The quantified potential value of prior information ( This includes information about quantifying potential values. Information about the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).

[0053] 2.3.4. Context Model Although the prior model improves the quantification of latent values Modeling the probability distribution of quantified potential values ​​is possible, but further improvements can be achieved by utilizing an autoregressive model (context model) that predicts quantified potential values ​​from the causal context of quantified potential values.

[0054] The term autoregressive means that the output of a process is later used as its input. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.

[0055] Figure 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model with a super-prior and an autoencoder. Table 1 below shows the meaning of the different symbols.

[0056]

[0057] The joint architecture is used in scenarios where both the advanced prior model subnetwork (an encoder utilizing advanced prior information and a decoder utilizing advanced prior information) and the context model subnetwork are utilized. The advanced prior and context models are combined to learn about the quantized latent value. The probability model is then used for entropy encoding and decoding. For example... Figure 4 As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean for the Gaussian probability model. And variance (scale) (or variance) The parameters are then used. The Gaussian probability model is then used by the arithmetic encoder (AE) module to encode the samples of the quantized latent values ​​into the bitstream. In the decoder, the Gaussian probability model is used to obtain the quantized latent values ​​from the bitstream via the arithmetic decoder (AD) module. .

[0058] Figure 4 A combined model is shown, which jointly optimizes the following: an autoregressive component (context model) that estimates the probability distribution of latent values ​​from the causal context of the latent values, and a super-prior and a low-level autoencoder. Real-valued latent representations are quantized (Q) to create quantized latent values ​​(…). ) and quantified potential value of prior information ( ), this quantified potential value ( ) and quantified potential value of prior information ( The image is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted areas correspond to the components performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream.

[0059] Typically, latent samples are modeled as Gaussian distributions or Gaussian mixture models (not limited to). According to Figure 4 The context model and the super-prior are jointly used to estimate the probability distribution of potential samples. Since the Gaussian distribution can be defined by the mean and variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as ). ).

[0060] 2.3.5. Encoding process using a joint autoregressive superprior model Figure 4 This corresponds to existing compression methods. The encoding and decoding processes will be described in this section and the next section, respectively. Figure 5 The latest encoding process is shown.

[0061] The above Figure 5 The encoding process is described. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transform representation called the latent value, which is... express. It is then fed into the quantizer block, denoted by Q, to obtain the quantization potential value ( ). Then, using an arithmetic coding module (denoted as AE), it is converted into a bitstream (bits1). The arithmetic coding blocks are sequentially... Each sample point is converted into a bit stream (bits1) one by one.

[0062] The module's encoder utilizing prior information, context, decoder utilizing prior information, and entropy parameter subnetwork are used to estimate the quantization potential. The probability distribution of the sample points. Potential The information is input into an encoder that utilizes prior information, and its output is the potential value of the prior information (denoted as...). The latent value of the prior information is then quantified. The second bitstream (bits2) is generated using the arithmetic coding (AE) module. The decompositional entropy module generates a probability distribution used to encode the quantized prior information latent values ​​into the bitstream. The quantized prior information latent values ​​include information about the quantized latent values ​​(…). Information about the probability distribution of ( ).

[0063] Entropy parameter subnetwork generation was used to encode quantized latent values. The probability distribution is estimated. Information generated from the entropy parameter is typically used together to obtain the mean of the Gaussian probability distribution. And variance (scale) (or variance) Parameters. The Gaussian distribution of the random variable x is defined as follows: , where parameters It is the mean or expected value of the distribution (also the median and mode), while the parameter This is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.

[0064] The decoder of the subnetwork, utilizing prior information, generates part of the information used by the entropy parameter subnetwork, while another part is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the quantized latent values. It is typically a matrix composed of many sample points. Sample points can be indicated using indices, for example... [i,j,k] or [i,j], specifically depends on the matrix. Dimensions. Sample points [i,j] are encoded sequentially by the AE, typically using a raster scan order. In a raster scan order, the matrix rows are processed from top to bottom, with samples in a row processed from left to right. In such scenarios (where the AE encodes samples into the bitstream using a raster scan order), the context module uses samples previously encoded in the raster scan order to generate a sequence with the samples. Information related to [i,j]. The information generated by the context module and the decoder utilizing prior information is combined by the entropy parameter module to generate information used to quantize the latent values. The probability distribution encoded into the bit stream (bits1).

[0065] Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder.

[0066] It is worth noting that other names can be used for the modules described above. In the above description, Figure 5 All elements in the algorithm are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).

[0067] 2.3.6. Decoding process using a joint autoregressive hyperprior model Figure 6 The latest decoding process is illustrated. During decoding, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decompositional entropy subnetwork. The decompositional entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values ​​in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This is the quantized potential value of the prior information. The AD process recovers to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.

[0068] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks, context, decoder utilizing prior information, and entropy parameters used in the decoder are the same as the three sub-networks in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder as in the encoder, which is crucial for reconstructing the quantized latent value without any loss. This is essential. As a result, the quantization potential value can be obtained in the decoder as well as in the encoder. Same version.

[0069] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of quantized potential values ​​one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial, and therefore cannot be accelerated using techniques such as parallelization.

[0070] Finally, the fully reconstructed quantized potential value Input into the synthesis transform (in) Figure 6 The module (represented as decoder) is used to obtain the reconstructed image.

[0071] In the above description, Figure 6 All elements in the array are collectively referred to as decoders. The synthetic transformation that converts quantized latent values ​​into a reconstructed image is also called a decoder (or autodecoder).

[0072] 2.4. Neural Networks for Video Compression Similar to traditional video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, it requires more effort to overcome its challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Thus, inter-frame prediction is a key step in these works. Motion estimation and compensation have been widely adopted, but only recently have they been implemented using trained neural networks.

[0073] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the case of random access, decoding can begin at any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, the aim is to reduce decoding time, so usually only the earlier frames in the time domain can be used as reference frames to decode subsequent frames.

[0074] 2.4.1. Low latency Early work first divided the video sequence frames into blocks, with each block choosing from two available modes (intra-frame encoding / decoding or inter-frame encoding / decoding). If intra-frame encoding / decoding was chosen, an associated autoencoder compressed the block. If inter-frame encoding / decoding was chosen, motion estimation and compensation were performed using conventional methods, and a trained neural network was used for residual compression. The autoencoder output was directly quantized and encoded / decoded using the Huffman method.

[0075] An alternative neural network-based video encoding / decoding scheme utilizing PixelMotionCNN is proposed. Frames are compressed sequentially in the temporal domain, and each frame is divided into blocks, which are compressed in raster scan order. Each frame is first inferred using the two preceding reconstructed frames. When compressing a block, the inferred frame, along with the context of the current block, is fed into PixelMotionCNN to derive a latent representation. The residual is then compressed using a variable-rate imaging scheme. The performance of this scheme is comparable to H.264.

[0076] Then, an alternative video compression framework based on end-to-end neural networks is proposed, where all modules are implemented using neural networks. This scheme takes the current frame and the previous reconstructed frame as input and uses a pre-trained neural network to derive optical flow as motion information. The motion information is warped along with the reference frame, and then the neural network generates motion-compensated frames. The residual and motion information are compressed using two separate neural autoencoders. The entire framework is trained with a single rate-distortion loss function. It achieves better performance than H.264.

[0077] An advanced neural network-based video compression scheme is proposed. It inherits and extends traditional video encoding and decoding schemes utilizing neural networks, and has the following main features: 1) it uses only one autoencoder to compress motion information and residuals; 2) it features motion compensation with multiple frames and multiple optical flows; 3) online state is learned and propagated over time through subsequent frames. This scheme achieves better performance than the HEVC reference software in MS-SSIM.

[0078] Subsequently, an extended video compression framework based on end-to-end neural networks is proposed. In this solution, multiple frames are used as references. By using multiple reference frames and associated motion information, a more accurate prediction of the current frame can be provided. Additionally, motion field prediction is deployed to eliminate motion redundancy along the temporal channels. A post-processing network is also introduced to eliminate reconstruction artifacts from previous processes. The performance is significantly better than H.265 in terms of PSNR and MS-SSIM.

[0079] Scale-space flow was then proposed to replace commonly used optical flow by adding a scale parameter. It reportedly achieves better performance than H.264.

[0080] A multi-resolution representation for optical flow is proposed. Specifically, a motion estimation network generates multiple optical flows with different resolutions, and the network learns which one to select under a loss function. Its performance outperforms H.265.

[0081] 2.4.2. Random Access An initial approach based on frame interpolation was designed. Keyframes are first compressed using a neural image compressor, and the remaining frames are compressed hierarchically. Motion compensation is performed in the receptive domain, i.e., feature maps are derived at multiple spatial scales of the original frames, and motion is used to warp these feature maps, which are then used in the image compressor. This method is reportedly comparable to H.264.

[0082] Then, another interpolation-based video compression method was proposed, in which the interpolation model combines motion information compression and image synthesis, and the same autoencoder is used for both the image and the residual.

[0083] Subsequently, a neural network-based video compression method based on a variational autoencoder and employing a deterministic encoder was proposed. Specifically, the model consists of an autoencoder and an autoregressive prior. Unlike previous methods, this method accepts a group of pictures (GOP) as input and incorporates a 3D autoregressive prior by considering temporal correlations when encoding and decoding the latent representation. It delivers performance comparable to H.265.

[0084] 2.5. Prerequisites An end-to-end image codec includes two operation points. The high operation point (highOP) is designed for codec gain, and the base operation point (baseOP) is designed for codec complexity; that is, the base OP can run on computationally limited devices such as mobile phones and tablets.

[0085] Learning-based reconstruction (called synthetic transformation) consists of two pipelines with the same neural network architecture (such as...). Figure 7 It consists of (as shown), except for the input size and the number of channels.

[0086] The inputs to the composition transformation include: - Spatial size of the potential tensor , - Shape Reconstructing the potential space tensor With auxiliary information tensor splicing - Operation point indicator , - Size of the output tensor , - Specified component Index: For the main components, And for secondary components, , - By pair Defined model parameters for the synthetic transformation network.

[0087] The output of the synthesis transformation is of size Tensor reconstruction of color components .

[0088] Figure 7 The synthesis transformation of the high OP (upper branch) and the basic OP (lower branch) is shown.

[0089] Composite transformation from the principal latent tensor and auxiliary ( The concatenation begins with the input. Then, based on the operation point indicator ( The decoder then executes the subsequent sequence of steps.

[0090] For basic operation points ( ) and main components ( The first step of the synthetic transformation (depth of the deep neural network process = 4) is a process with a channel number of It consists of lightweight residual blocks, followed by a core size of [missing information]. The transposed convolution reduces the number of channels to This transposed convolution is achieved by a clipping layer (with a stride of 1). Depth is And residual activation units are used. For minor components ( The first step of the synthetic transformation (depth=4) is simply the potential combinatorial block (LCB), which reduces the number of channels from... Change to The next step (depth=3) for the two components is a kernel size of The transposed convolution changes the number of channels to This transposed convolution is achieved by a clipping layer (with a stride of 1). Depth is The process involves recurrent activation units and residual activation units. The next steps in this process (depth = 2 and 1) involve a kernel size of... The step size is 1 and the number of channels is The invariant regular convolution is combined with residual activation units. Then, a stride of... convolution It will increase the number of channels from Increase to This is done to ensure that the next layer (i.e., with a step size of...) The output of pixel shuffling has a number of channels This process uses trimming layers (step size is...) Depth is )Finish.

[0091] For high operating points ( ) and main components ( The first step of the synthetic transformation (depth of the deep neural network process = 4) is based on the number of channels. It consists of residual blocks, followed by a core with a size of The transposed convolution reduces the number of channels to The transposed convolution and clipping layer (with stride of ) (depth of 4) and residual activation units are combined. For minor components ( The first step in the synthetic transformation is simply the potential combinatorial block (LCB), which reduces the number of channels from... Change to The next step for both components (depth=3) is to have a kernel size of... The transposed convolution changes the number of channels to The convolution-based attention block is placed in the next (in) Figure 7 The Chinese character is represented as Then comes the clipping layer (step size is...). (with a depth of 3) and residual activation units.

[0092] The next step in this process (depth=2) is to have a core size of The step size is 1 and the number of output channels is The regular convolution. This is done to ensure that the next layer (i.e., with a stride of...) is... The output of pixel shuffling has a number of channels The final step (depth=1) involves the attention module based on the transformer (in... Figure 7 The Chinese character is represented as Start with the clipping layer (step). ,depth ) and core size is The residual activation unit. This process is achieved through its kernel size of... Step size is Number of output channels The transposed convolution is followed by a clipping layer (with a stride of 1). Depth is )Finish.

[0093] 2.5.1.CAB The attention block receiving size based on convolution is tensor and in execution Figure 8 The sequence of steps described in the diagram outputs a tensor of the same size. . Figure 8 A convolution-based attention block (CAB) is shown.

[0094] The process is divided into two branches, operating at different spatial resolutions. The first branch consists of two residual blocks. The second branch begins with a downsampled convolution with stride = 2, followed by two residual blocks and a transposed convolution with stride = 2. The second branch ends with a sigmoid function. The two branches are connected by an element-wise multiplication ⊙. The resulting tensor is then multiplied by the parameters. And added To the input tensor. Using =0, all operations in CAB are essentially bypassed.

[0095] 2.5.2.TAM The transformer-based attention module is represented as: ,like Figure 9 As shown. The received size of this sub-network is Tensor and component indicators ( This module is executing... Figure 10 The layer sequence shown is followed by a modified tensor of the same size and shape. The transformer-based attention module consists of two transformer-based attention blocks (…). TAB Composition of secondary components ( ). Attention blocks based on the transformer are executed at the original spatial resolution. For the principal components ( Tensors are first executed A convolution with a stride of 2 is downsampled, then two transformer-based attention blocks are executed, and finally... A transposed convolution with a stride of 2 returns the spatial size of the tensor, which is equal to the input size.

[0096] Figure 9 Transformer-based attention module (TAM) is shown.

[0097] Transformer-based attention blocks are represented as follows: ,like Figure 10 As shown. The block receiving size is tensor and produce tensors of the same size. .

[0098] This process executes two sequential sub-processes.

[0099] The first subprocess begins with a reshaping operation, after which the result has the following shape: Then, the result is normalized by layers. After layer normalization, the result is reshaped, after which the tensor has the following shape: Then, this subprocess proceeds through a series of steps with a step size of 1 and a kernel size of... Convolution (which increases the number of channels to) ) and group size is Group convolution. After group convolution, a TensorChunk with 3 outputs is generated.

[0100] The three outputs of the TensorChunk have shapes They enter the three branches of this subprocess.

[0101] The first branch is from reshaping to sizing. The process involves composition, followed by normalization of the tensor.

[0102] The second branch moves from reshaping to resizing. The process involves several steps: tensor normalization and matrix transpose, followed by tensor normalization and matrix transpose operations. The resulting tensor shape is... .

[0103] Perform matrix multiplication on the results of the first and second branches (represented as...). ), then the tensor shape is This tensor is used in scaling operations (represented as...). The size of ) Multiplier temperature ( Scaling: The resulting tensor is in the same channel. All elements are subjected to the same multiplier temperature Scaling 。 The tensor is then subjected to a soft-max operation.

[0104] The third branch focuses on reshaping to size. start.

[0105] The output of Softmax and the reshaped third branch tensor are multiplied by a matrix, after which the tensor has the following shape: The tensor was reshaped into and through the core size A convolution with a stride of 1 is performed. The output of this convolution is added to the input of the first subprocess.

[0106] Figure 10 Transformer-based attention blocks (TABs) are shown.

[0107] The output of the first sub-procedure becomes the input of the second sub-procedure. The second sub-procedure is reshaped to... Begin with layer normalization and another reshaping. Then, this subprocess uses a kernel size of A convolution with a stride of 1 (which increases the number of channels to...) ) and group size is Step size is 1, core size is Group convolution. Following the group convolution is a TensorChunk, which has two elements of shape... The output of the TensorChunk is processed by a nonlinear exponential linear unit and then merged into the second output of the TensorChunk, i.e., element-wise multiplication. The result is processed with a step size of 1 and a kernel size of . The convolution reduces the number of channels back to... At the end of the process, the result of the convolution is added to the input of the second subprocess.

[0108] 3. Problem 3.1. Core Issues like Figure 7 As shown, the framework includes two synthetic transforms. The basic operation and the high operation require different bitstreams to reconstruct the image; that is, they cannot share a bitstream. This design has several drawbacks. First, the framework is equivalent to having two completely different decoders because they require different bitstreams, which makes the codec redundant and difficult to maintain. Second, the pre-trained models are different, meaning that two sets of pre-trained models must be stored, requiring additional storage. Third, the transform-based attention module (TAM) and the convolution-based attention module (CAM) have the most significant computational complexity in the synthetic transform. However, in the high operation, there is no option to skip these attention modules, making it inflexible and potentially limiting its application in certain scenarios.

[0109] 4. Detailed Solution The detailed solutions below should be considered as examples to illustrate the general concept. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.

[0110] Figure 11 A multi-branch decoder with high-OP synthesis transform is shown, where CAB and TAM can be controlled to be turned on and off independently. For example... Figure 11 As shown, an example is illustrated using a high-op synthesis transform. The invented multi-branch decoder can use two flags. To control TAM and CAB. The same bitstream can be decoded using different configurations, including: – Enable both CAM and TAM, i.e. ; – Disable both CAM and TAB, i.e. ; – Enable CAB and disable TAM, i.e. ; – Enable TAM and disable CAB, i.e. .

[0111] 4.1. Core of the Solution The proposed solution aims to provide a flexible design where the codec includes a multi-branch decoder capable of decoding the same bitstream. One or more modules or layers exist that can be individually controlled to be turned on and off depending on the application scenario.

[0112] General items 1. Whether and / or how the methods disclosed above can be signaled at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0113] 2. Whether and / or how to apply the methods disclosed above may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / image type.

[0114] 3. The methods presented in this document can be used in other codec tools that require chroma blending.

[0115] 4. The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. These can be signed or unsigned.

[0116] 5. The syntax elements disclosed above can be encoded and decoded using at least one context model. Alternatively, they can be encoded and decoded using a bypass method.

[0117] 6. The syntax elements disclosed above can be transmitted via signals in a conditional manner.

[0118] a. SE is transmitted via signal only if the corresponding function is applicable.

[0119] b. SE is transmitted via signal only if the dimensions (width and / or height) of the block meet the requirements.

[0120] 7. The syntax elements disclosed above can be signaled at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0121] 5. Examples 1. In one example, such as Figure 11 As shown, the multi-branch decoder is implemented using a single synthesis transform, but one or more modules / layers can be individually turned on or off. In this example, the CAB attention module and the TAM attention module can respectively utilize flags They are controlled individually. The same bitstream can be decoded using different configurations, including: – ; – ; – ; – .

[0122] 2. In one example, such as Figure 12 As shown, the multi-branch decoder is implemented as multiple synthesis transforms, where each network includes different layers. Figure 12 A multi-branch decoder implemented as multiple synthesis transforms is shown. The syntax `dec_branch` is used to select a branch during the decoding process. Figure 12 In this context, there are four compositional transformations, which are derived from the grammar. control.

[0123] – When dec_branch = 0, the first branch is selected, which has both TAM attention modules and CAB attention modules.

[0124] – When dec_branch = 1, the second branch is selected, which contains only the CAB attention module.

[0125] – When dec_branch = 2, the third branch is selected, which contains only the TAM attention module.

[0126] – When dec_branch = 3, the fourth branch is selected, in which both CAB and TAM are removed.

[0127] 3. According to some embodiments, a decoder is used to reconstruct an image, wherein the decoder consists of neural network layers.

[0128] a. Indicators can be included in the bitstream to indicate whether a module is used in processing or is skipped.

[0129] b. Indicators can be included in the bitstream to indicate whether at least two modules are included in the processing or whether they are skipped.

[0130] i. The two modules can be ordered such that a third module can be placed between the two modules.

[0131] ii. At least two indications may be included in the bitstream, one indication controlling the first of at least two modules, and the second indication controlling the second of at least two modules.

[0132] c. Modules can be attention blocks.

[0133] d. Modules may include neural network-based processing layers, such as convolutional layers, transformer layers, or activation layers.

[0134] e. The module may be part of a synthetic transformation.

[0135] f. Indicators can be flags, syntax elements, or level indicators.

[0136] g. An instruction may indicate any of the following, as described below.

[0137] 4. Indicators (e.g., syntax elements included in a bitstream) can indicate any of the following: a. Processing modules (e.g., attention blocks) are included in the processing performed using the decoder.

[0138] b. The processing module was not included in the processing performed using the decoder.

[0139] c. Both options are possible. The instruction can indicate that both outputs (the first output obtained using the processing module, and the second output obtained without using the processing module) are acceptable.

[0140] Further details of embodiments of this disclosure relating to neural network-based visual data encoding and decoding will now be described. As used herein, the term “visual data” can refer to video, images, pictures in video, or any other visual data suitable for encoding and decoding.

[0141] As discussed above, in existing designs for visual data encoding and decoding based on neural networks (NNs), the synthesis transform for basic operations and the synthesis transform for high operations are configured to process the reconstructed latent representation of visual data derived from different bitstreams. In other words, a bitstream that can be processed by the synthesis transform for basic operations cannot be processed by the synthesis transform for high operations. Therefore, the synthesis transform for basic operations and the synthesis transform for high operations are incompatible with each other, making the codec redundant and difficult to maintain.

[0142] To address the above-mentioned problems and other issues not mentioned, a visual data processing solution is disclosed as described below. The embodiments of this disclosure should be considered as examples illustrating general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.

[0143] Figure 13 A flowchart of a method 1300 for visual data processing according to some embodiments of the present disclosure is shown. Method 1300 may be implemented during the conversion between visual data and a bitstream of visual data using a neural network (NN)-based model. As used herein, the NN-based model can be a model based on neural network techniques. For example, the NN-based model may specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module may include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the NN-based model described herein are illustrative only and should not be construed as limiting the present disclosure in any way.

[0144] like Figure 13 As shown, method 1300 begins at 1302, where a synthesis transform is selected from multiple synthesis transforms in a NN-based model. These multiple synthesis transforms are configured to process the reconstructed latent representation of visual data derived from the same bitstream. In other words, a bitstream that can be processed by one of the synthesis transforms can also be processed by the remaining synthesis transforms. Therefore, the multiple synthesis transforms are compatible with each other.

[0145] In some exemplary embodiments, the synthesized transformation can be selected based on a first indication. For example, the first indication can be indicated in the bitstream. The first indication can be implemented using flags, syntax elements, or level indicators. By way of example, the first indication can be represented as dec_branch. It should be noted that the first indication can also be represented using any other suitable string, such as branchID, decoderID, etc.

[0146] In some other exemplary embodiments, the synthesis transform can be selected based on encoding / decoding information of the visual data (such as quantization information). It should be noted that the synthesis transform can also be selected based on any other suitable information. The scope of this disclosure is not limited in this respect.

[0147] At 1304, a transformation is performed based on the selected synthetic transform. By way of example, and not limitation, the selected synthetic transform can be used to process the reconstructed latent representation of visual data. In some embodiments, the transformation may include encoding the visual data into a bitstream. Additionally or alternatively, the transformation may include decoding the visual data from the bitstream. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0148] In light of the above, multiple synthesis transforms in the NN-based model can process the same bitstream, and one of these synthesis transforms is selected to perform the conversion. Compared to traditional solutions that configure two synthesis transforms to process different bitstreams, the proposed method advantageously improves the compatibility of multiple synthesis transforms, thus providing more options for processing the same bitstream to cater to different applications. In this way, encoding and decoding flexibility is improved, and therefore encoding and decoding efficiency is enhanced.

[0149] In some embodiments, the multiple synthesis transforms may differ from each other. For example, a first synthesis transform among the multiple synthesis transforms may include both a transformer-based attention module (TAM) and a convolution-based attention block (CAB). Furthermore, a second synthesis transform among the multiple synthesis transforms may exclude neither TAM nor CAB. That is, TAM and CAB are absent from the second synthesis transform. Additionally or alternatively, a third synthesis transform among the multiple synthesis transforms may include TAM but exclude CAB, while a fourth synthesis transform among the multiple synthesis transforms may include CAB but exclude TAM. It should be understood that the possible implementations of the multiple synthesis transforms described herein are merely illustrative and should not be construed as limiting this disclosure in any way. The multiple synthesis transforms may also differ from each other in terms of parameter values.

[0150] By way of example, and not limitation, a TAM may include one or more transformer-based attention blocks (TABs). Additionally or alternatively, a TAM may include one or more matrix transpose operations. Conversely, a CAB may include neither TABs nor matrix transpose operations. Exemplary implementations of CABs and TAMs have been described in detail in Sections 2.5.1 and 2.5.2.

[0151] In some embodiments, an indication in the bitstream may indicate whether a processing module in one of a plurality of synthesis transformations is enabled or disabled. Alternatively, at least one indication in the bitstream may indicate whether multiple processing modules in one of a plurality of synthesis transformations are enabled or disabled.

[0152] In some embodiments, a second instruction in at least one instruction may indicate whether a first processing module among a plurality of processing modules is enabled or disabled. Furthermore, a third instruction in at least one instruction may indicate whether a second processing module among a plurality of processing modules is enabled or disabled.

[0153] In some embodiments, the first processing module may immediately follow the second processing module. Alternatively, the first processing module may immediately precede the second processing module. In some other embodiments, another processing module may be arranged between the first and second processing modules.

[0154] In some embodiments, the first processing module may include a CAB, and the second processing module may include a TAM. Additionally or alternatively, the first or second processing module may include attention blocks, convolutional layers, transformer layers, activation layers, and / or the like.

[0155] In some embodiments, the indication in the bitstream may indicate one of the following: a first option, wherein the processing module is included in the synthesis transform; a second option, wherein the processing module may be the synthesis transform; or both the first and second options are acceptable.

[0156] In some embodiments, any of the above indications may be a flag, a syntax element, or a level indicator.

[0157] In some embodiments, first information regarding at least one of the following may be indicated in the bitstream: whether the method is applied, or how the method is applied. For example, the first information may be indicated at one of the following levels: block level, sequence level, picture group level, picture level, stripe level, or slice group level. Additionally or alternatively, the first information may be indicated at one of the following levels: codec structure of codec tree unit (CTU), codec structure of codec unit (CU), codec structure of transform unit (TU), codec structure of prediction unit (PU), codec structure of codec tree block (CTB), codec structure of codec block (CB), codec structure of transform block (TB), codec structure of prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.

[0158] In some embodiments, the first information may depend on encoded information of the visual data. By way of example, encoded information may include block size, color format, single-tree segmentation, dual-tree segmentation, color components, stripe type, image type, and / or the like.

[0159] In some embodiments, the first information may be indicated by a syntax element. For example, a syntax element may be binarized into one of the following: a flag, a fixed-length code, an exponential Golomb code (EG code), a unary code, a rounded unary code, or a rounded binary code. In some embodiments, the syntax element may be encoded or decoded using at least one context model, or the syntax element may be bypassed and encoded / decoded. In some embodiments, the syntax element may be transmitted via signaling based on conditions.

[0160] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding flexibility.

[0161] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by means of visual data processing. The method includes: selecting a synthetic transform from a plurality of synthetic transforms in a neural network (NN)-based model, the plurality of synthetic transforms being configured to process a reconstructed latent representation of visual data derived from the same bitstream; and generating a bitstream using the NN-based model based on the selected synthetic transform.

[0162] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. The method includes: selecting a synthetic transform from a plurality of synthetic transforms in a neural network (NN)-based model, the plurality of synthetic transforms being configured to process a reconstructed latent representation of visual data derived from the same bitstream; generating a bitstream using the NN-based model based on the selected synthetic transform; and storing the bitstream in a non-transitory computer-readable recording medium.

[0163] Implementations of this disclosure can be described according to the following entries, the features of which can be combined in any reasonable manner.

[0164] Item 1. A method for visual data processing, comprising: for a conversion between visual data and a bitstream of the visual data using a neural network (NN)-based model, selecting a synthetic transformation from a plurality of synthetic transformations in the NN-based model, the plurality of synthetic transformations being configured to process a reconstructed latent representation of the visual data derived from the same bitstream; and performing the conversion based on the selected synthetic transformation.

[0165] Item 2. The method according to Item 1, wherein the synthetic transformation is selected based on a first indication.

[0166] Item 3. The method according to Item 2, wherein the first indication is indicated in the bit stream.

[0167] Item 4. The method according to any one of items 1-3, wherein the plurality of synthetic transformations are different from each other.

[0168] Item 5. The method according to any one of items 1-4, wherein the first synthesis transformation of the plurality of synthesis transformations comprises both a transformer-based attention module (TAM) and a convolution-based attention block (CAB), and the second synthesis transformation of the plurality of synthesis transformations does not include either the TAM or the CAB.

[0169] Item 6. The method according to any one of items 1-5, wherein the third synthetic transformation of the plurality of synthetic transformations includes TAM but excludes CAB, or the fourth synthetic transformation of the plurality of synthetic transformations includes CAB but excludes TAM.

[0170] Item 7. The method according to any one of items 1-6, wherein the indication in the bitstream indicates whether a processing module in one of the plurality of synthesis transformations is enabled or disabled.

[0171] Item 8. The method according to any one of Items 1-7, wherein at least one indicator in the bitstream indicates whether to enable or disable multiple processing modules in one of the plurality of synthesis transformations.

[0172] Item 9. The method according to Item 8, wherein the second indication of the at least one indication indicates whether to enable or disable the first processing module among the plurality of processing modules, and the third indication of the at least one indication indicates whether to enable or disable the second processing module among the plurality of processing modules.

[0173] Item 10. The method according to Item 9, wherein another processing module is arranged between the first processing module and the second processing module.

[0174] Item 11. The method according to any one of Items 9-10, wherein the first processing module comprises CAB and the second processing module comprises TAM.

[0175] Item 12. The method according to any one of items 5, 6 and 11, wherein the TAM includes at least one of: a transformer-based attention block (TAB), or a matrix transpose operation.

[0176] Item 13. The method according to any one of Items 9-10, wherein the first processing module or the second processing module comprises at least one of the following: an attention block, a convolutional layer, a transformer layer, or an activation layer.

[0177] Item 14. The method according to any one of Items 1-6, wherein the indication in the bitstream indicates one of the following: a first option, in which the processing module is included in the synthesis transformation; a second option, in which the synthesis transformation does not include the processing module; or both the first option and the second option are acceptable.

[0178] Item 15. The method according to any one of Items 2-14, wherein the indicator is one of the following: a flag, a syntax element, or a level indicator.

[0179] Item 16. The method according to any one of items 1-15, wherein first information regarding at least one of the following is indicated in the bitstream: whether the method is applied, or how the method is applied.

[0180] Item 17. The method according to Item 16, wherein the first information is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0181] Item 18. The method according to Item 16, wherein the first information is indicated in one of the following: the codec structure of a codec tree unit (CTU), the codec structure of a codec unit (CU), the codec structure of a transform unit (TU), the codec structure of a prediction unit (PU), the codec structure of a codec tree block (CTB), the codec structure of a codec block (CB), the codec structure of a transform block (TB), the codec structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a slice parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice group header.

[0182] Item 19. The method according to any one of items 16-18, wherein the first information depends on the encoded / decoded information of the visual data.

[0183] Item 20. The method according to Item 19, wherein the encoded / decoded information includes at least one of the following: block size, color format, single-tree segmentation, dual-tree segmentation, color components, stripe type, or image type.

[0184] Item 21. The method according to any one of items 16-20, wherein the first information is indicated by a syntax element.

[0185] Item 22. According to the method described in Item 21, the syntax element is binarized into one of the following: a flag, a fixed-length code, an exponential Golomb code (EG), a unary code, a rounded unary code, or a rounded binary code.

[0186] Item 23. The method according to any one of items 21-22, wherein the syntax element is encoded or decoded using at least one context model, or wherein the syntax element is encoded or decoded in a bypass manner.

[0187] Item 24. The method according to any one of items 21-23, wherein the syntax element is transmitted via signal based on condition.

[0188] Item 25. The method according to any one of items 1-24, wherein the visual data includes video, a picture of the video, or an image.

[0189] Item 26. The method according to any one of items 1-25, wherein the conversion includes encoding the visual data into the bitstream.

[0190] Item 27. The method according to any one of items 1-25, wherein the conversion includes decoding the visual data from the bitstream.

[0191] Item 28. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-27.

[0192] Item 29. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-27.

[0193] Item 30. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of means for visual data processing, wherein the method comprises: selecting a synthetic transform from a plurality of synthetic transforms in a neural network (NN)-based model, the plurality of synthetic transforms being configured to process a reconstructed latent representation of the visual data derived from the same bitstream; and generating the bitstream using the NN-based model based on the selected synthetic transform.

[0194] Item 31. A method for storing a bitstream of visual data, comprising: selecting a synthetic transformation from a plurality of synthetic transformations in a neural network (NN)-based model, the plurality of synthetic transformations being configured to process a reconstructed latent representation of the visual data derived from the same bitstream; generating the bitstream using the NN-based model based on the selected synthetic transformation; and storing the bitstream in a non-transitory computer-readable recording medium.

[0195] Exemplary device Figure 14 A block diagram of a computing device 1400 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1400 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).

[0196] It should be understood that, Figure 14The computing device 1400 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0197] like Figure 14 As shown, computing device 1400 includes general-purpose computing device 1400. Computing device 1400 may include at least one or more processors or processing units 1410, memory 1420, storage unit 1430, one or more communication units 1440, one or more input devices 1450, and one or more output devices 1460.

[0198] In some embodiments, the computing device 1400 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 1400 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0199] Processing unit 1410 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 1420. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 1400. Processing unit 1410 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0200] Computing device 1400 typically includes various computer storage media. Such media can be any media accessible by computing device 1400, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 1420 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 1430 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or visual data and can be accessed within computing device 1400.

[0201] The computing device 1400 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 14 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more visual data media interfaces.

[0202] Communication unit 1440 communicates with another computing device via a communication medium. Additionally, the functionality of components in computing device 1400 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 1400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0203] Input device 1450 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1460 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 1440, computing device 1400 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 1400 can also communicate with one or more devices that enable a user to interact with computing device 1400, or, if needed, with any device that enables computing device 1400 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).

[0204] In some embodiments, some or all of the components of computing device 1400 may be arranged in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote visual data center locations. Cloud computing infrastructure may provide services through a shared visual data center, although to the user they appear as a single access point. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0205] In embodiments of this disclosure, computing device 1400 can be used to implement visual data encoding / decoding. Memory 1420 may include one or more visual data encoding / decoding modules 1425 having one or more program instructions. These modules are accessible and executable by processing unit 1410 to perform the functions of the various embodiments described herein.

[0206] In an exemplary embodiment of performing visual data encoding, input device 1450 may receive visual data as input 1470 to be encoded. The visual data may be processed, for example, by visual data encoding / decoding module 1425 to generate an encoded bitstream. The encoded bitstream may be provided as output 1480 via output device 1460.

[0207] In an exemplary embodiment of performing visual data decoding, input device 1450 may receive an encoded bitstream as input 1470. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 1425 to generate decoded visual data. The decoded visual data may be provided as output 1480 via output device 1460.

[0208] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for visual data processing, comprising: For the conversion between visual data and the bitstream of the visual data using a neural network (NN)-based model, a synthetic transformation is selected from a plurality of synthetic transformations in the NN-based model, the plurality of synthetic transformations being configured to process the reconstructed latent representation of the visual data derived from the same bitstream; as well as The transformation is performed based on the selected synthetic transformation.

2. The method of claim 1, wherein the synthetic transformation is selected based on a first indication.

3. The method of claim 2, wherein the first indication is indicated in the bit stream.

4. The method according to any one of claims 1-3, wherein the plurality of synthetic transformations are different from each other.

5. The method according to any one of claims 1-4, wherein the first synthesis transformation in the plurality of synthesis transformations comprises both a transformer-based attention module (TAM) and a convolution-based attention block (CAB), and The second synthetic transformation in the plurality of synthetic transformations does not include either the TAM or the CAB.

6. The method according to any one of claims 1-5, wherein the third synthetic transformation in the plurality of synthetic transformations includes TAM but excludes CAB, or The fourth synthetic transformation in the plurality of synthetic transformations includes the CAB but excludes the TAM.

7. The method according to any one of claims 1-6, wherein the indication in the bitstream indicates whether a processing module in one of the plurality of synthesis transformations is enabled or disabled.

8. The method according to any one of claims 1-7, wherein at least one indicator in the bitstream indicates whether to enable or disable multiple processing modules in one of the plurality of synthesis transformations.

9. The method of claim 8, wherein the second indication of the at least one indication indicates whether the first processing module among the plurality of processing modules is enabled or disabled, and The third instruction in the at least one instruction indicates whether the second processing module among the plurality of processing modules is enabled or disabled.

10. The method of claim 9, wherein another processing module is arranged between the first processing module and the second processing module.

11. The method according to any one of claims 9-10, wherein the first processing module comprises CAB and the second processing module comprises TAM.

12. The method according to any one of claims 5, 6, and 11, wherein the TAM comprises at least one of the following: Transformer-based attention blocks (TABs), or Matrix transpose operation.

13. The method according to any one of claims 9-10, wherein the first processing module or the second processing module comprises at least one of the following: Attention block Convolutional layer Converter layer, or Activation layer.

14. The method according to any one of claims 1-6, wherein the indication in the bitstream indicates one of the following: In the first option, the processing module is included in the synthetic transformation. In the second option, the synthesis transformation does not include the processing module, or Both the first option and the second option are acceptable.

15. The method according to any one of claims 2-14, wherein the instruction is one of the following: logo, Syntax elements, or Grade indicator.

16. The method according to any one of claims 1-15, wherein first information regarding at least one of the following is indicated in the bitstream: Whether to apply the method, or How to apply the method.

17. The method of claim 16, wherein the first information is indicated in one of the following: Block level, sequence level, Image group level, Image quality, strip level, or Film series level.

18. The method of claim 16, wherein the first information is indicated in one of the following: The encoding / decoding structure of a codec tree unit (CTU). The encoding / decoding structure of the codec unit (CU), The encoding / decoding structure of the Transform Unit (TU) The encoding / decoding structure of the prediction unit (PU), The encoding / decoding structure of the codec tree block (CTB). The encoding / decoding structure of the codec block (CB). The encoding / decoding structure of the transform block (TB). The encoding and decoding structure of prediction blocks (PB), Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

19. The method according to any one of claims 16-18, wherein the first information depends on the encoded / decoded information of the visual data.

20. The method of claim 19, wherein the encoded / decoded information comprises at least one of the following: Block size, Color format, Single tree segmentation, Two-tree partitioning, Color components, Strip type, or Image type.

21. The method according to any one of claims 16-20, wherein the first information is indicated by a syntax element.

22. The method of claim 21, wherein the syntax element is binarized into one of the following: logo, Fixed-length code Index Columbus (EG) code, One-yuan code, Rounding down a single-digit code, or Rounding binary code.

23. The method according to any one of claims 21-22, wherein the syntax element is encoded or decoded using at least one context model, or The syntax elements therein are bypassed and encoded / decoded.

24. The method according to any one of claims 21-23, wherein the syntax elements are transmitted via signals based on conditions.

25. The method according to any one of claims 1-24, wherein the visual data includes video, a picture of the video, or an image.

26. The method according to any one of claims 1-25, wherein the conversion comprises encoding the visual data into the bitstream.

27. The method according to any one of claims 1-25, wherein the conversion comprises decoding the visual data from the bitstream.

28. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-27.

29. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-27.

30. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by means of visual data processing, wherein the method includes: A synthetic transformation is selected from a plurality of synthetic transformations in a neural network (NN)-based model, the plurality of synthetic transformations being configured to process a reconstructed latent representation of the visual data derived from the same bitstream; as well as The bitstream is generated using the NN-based model based on the selected synthesis transformation.

31. A method for storing a bitstream of visual data, comprising: A synthetic transformation is selected from a plurality of synthetic transformations in a neural network (NN)-based model, the plurality of synthetic transformations being configured to process a reconstructed latent representation of the visual data derived from the same bitstream; Based on the selected synthesis transformation, the bitstream is generated using the NN-based model; as well as The bitstream is stored in a non-transitory computer-readable recording medium.