Visual data processing method, device and medium

By determining the conversion of multiple codewords associated with semantic elements or potential variable information in visual data processing, the problem of insufficient encoding and decoding efficiency and effectiveness in the prior art is solved, and a more efficient encoding and decoding effect is achieved, which is suitable for existing and neural network-based video/image encoding and decoding standards.

CN120513637APending Publication Date: 2025-08-19DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007785.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-13
Filing Date
2024-01-12
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing classic video encoding and codec solutions and neural network-based video compression methods still have room for improvement in encoding and codec efficiency and effectiveness, especially in the use of statistical dependence of potential variables and the processing of unstructured data.

Method used

By converting between the current visual unit of visual data and the bit stream, multiple codewords are determined to be associated with semantic element information or potential variable information, and the conversion is performed based on multiple codewords, and a multi-threaded processing method is adopted to improve the encoding and decoding efficiency and effectiveness.

Benefits of technology

It improves the encoding and decoding efficiency and effectiveness, is suitable for existing video/image encoding and decoding standards and neural network-based video compression technology, achieving higher encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513637A_ABST
    Figure CN120513637A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for visual data processing. In the method, for a conversion between a current visual unit of visual data and a bitstream of visual data, a plurality of threads for encoding and decoding residual information for the current visual unit are determined. The conversion is performed based on a plurality of threads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to visual data processing techniques, and more particularly, to multiple codewords in a bitstream. Background Art

[0002] Image / video compression is an important technology for reducing the cost of image / video transmission and storage in a lossless or lossy manner. Image / video compression technology can be divided into two branches: classical video codec methods and neural network-based video compression methods. Classical video codec schemes use transform-based solutions, in which researchers exploit statistical dependencies in latent variables (e.g., wavelet coefficients) by carefully hand-designing entropy codecs that model the dependencies in the quantization mechanism. Neural network-based video compression comes in two forms: neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded as a codec tool into existing classical video codecs as part of the framework, while the latter is a separate framework developed based on neural networks and does not rely on classical video codecs. The codec efficiency of image / video codecs is generally expected to be further improved. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for visual data processing.

[0004] In a first aspect, a method for visual data processing is provided. The method comprises: determining, for conversion between a current visual unit of visual data and a bitstream of the visual data, a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of the current visual unit; and performing the conversion based on the plurality of codewords. The method according to the first aspect of the present disclosure applies a plurality of codewords in the bitstream rather than a single codeword. In this manner, codec efficiency and / or codec effectiveness can be improved.

[0005] In a second aspect, an apparatus for visual data processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a video bitstream generated by a method performed by a data processing apparatus. The method includes: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of visual data; and generating the bitstream based on the plurality of codewords.

[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of visual data; generating a bitstream based on the plurality of codewords; and storing the bitstream in a non-transitory computer-readable recording medium.

[0009] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 shows a block diagram of an example visual data encoding and decoding system according to some embodiments of the present disclosure;

[0012] Figure 2 A typical transform coding scheme is shown;

[0013] Figure 3 An image from the Kodak dataset and different representations of the image are shown;

[0014] Figure 4 The network architecture of the autoencoder implementing the hyper-prior model is shown;

[0015] Figure 5 shows a block diagram of the combined model;

[0016] Figure 6 shows the encoding process of the combined model;

[0017] Figure 7 The decoding process of the combined model is shown;

[0018] Figure 8 shows the structure of the bitstream in the currently studied image compression framework;

[0019] Figure 9 An example of a bitstream structure according to an embodiment of the present disclosure is shown;

[0020] Figure 10 Another example of a bitstream structure according to an embodiment of the present disclosure is shown;

[0021] Figure 11 Another example of a bitstream structure according to an embodiment of the present disclosure is shown;

[0022] Figure 12 Another example of a bitstream structure according to an embodiment of the present disclosure is shown;

[0023] Figure 13 A flowchart showing a method for visual data processing according to an embodiment of the present disclosure; and

[0024] Figure 14 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0025] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION

[0026] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0027] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0028] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, whether or not explicitly described, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in relation to other embodiments.

[0029] It should be understood that although the terms "first" and "second" and the like can be used to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0030] As used herein, the term "visual data" may refer to image data or video data. The term "visual data processing" may refer to image processing or video processing. The term "visual data encoding and decoding" may refer to image encoding and decoding or video encoding and decoding. The term "encoding and decoding visual data" may refer to "encoding visual data (e.g., encoding visual data into a bitstream)" and / or "decoding visual data (e.g., decoding visual data from a bitstream)."

[0031] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises," "includes," and / or "having," when used herein, indicate the presence of the described features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0032] Sample Environment

[0033] Figure 1 is a block diagram illustrating an example visual data encoding and decoding system 100 that may utilize the techniques of the present disclosure. As shown, the visual data encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 may also be referred to as a data encoding device or a visual data encoding device, and the target device 120 may also be referred to as a data decoding device or a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the target device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a data source 112, a data encoder 114, and an input / output (I / O) interface 116.

[0034] The data source 112 may include a source such as a data capture device. Examples of a data capture device include, but are not limited to, an interface that receives data from a data provider, a computer graphics system for generating data, and / or a combination thereof.

[0035] The data may include one or more pictures or one or more images of a video. The data encoder 114 encodes the data from the data source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded data may be transmitted directly to the target device 120 via the network 130A via the I / O interface 116. The coded data may also be stored on the storage medium / server 130B for access by the target device 120.

[0036] Target device 120 may include an I / O interface 126, a data decoder 124, and a display device 122. I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded data from source device 110 or storage medium / server 130B. Data decoder 124 may decode the encoded data. Display device 122 may display the decoded data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120, configured to interface with an external display device.

[0037] The data encoder 114 and the data decoder 124 may operate according to a data codec standard, such as a video codec standard or a still picture codec standard and other existing and / or future standards.

[0038] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to only that section. In addition, although certain embodiments are described with reference to a multifunctional video codec or other specific data codecs, the disclosed technology is also applicable to other codec technologies. In addition, although some embodiments describe the encoding and decoding steps in detail, it should be understood that the corresponding decoding steps will be implemented by a decoder, which will undo the encoding. In addition, the term "data processing" includes data encoding and decoding or compression, data decoding or decompression, and data transcoding, where data is represented from one compression format to another compression format or at different compression bit rates.

[0039] 1. Brief Overview

[0040] The present disclosure relates to video / image codec technology. Specifically, it relates to network-based image and video compression. The concepts can be applied alone or in various combinations to any existing video / image codec standard or non-standard video codec, such as JPEG-AI and IEEE 1857.11. The proposed concepts may also be applicable to future video / image codec standards or video codecs.

[0041] 2. Introduction

[0042] The past decade has witnessed rapid developments in deep learning across various fields, particularly in computer vision and image processing. Inspired by the remarkable success of deep learning techniques in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression techniques. Neural networks were originally proposed in interdisciplinary research between neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression techniques have made significant progress over the past five years. The latest neural network-based image compression algorithms have reportedly achieved comparable real-time performance to the Versatile Video Codec (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement of neural image compression performance, neural network-based video compression has become an actively developing research field. However, due to the inherent difficulty of the problem, neural network-based video codecs are still in their infancy.

[0043] 2.1 Image / Video Compression

[0044] Image / video compression generally refers to computing techniques that compress images / video into binary codes for easier storage and transmission. The binary codes can support or not support lossless reconstruction of the original image / video, referred to as lossless or lossy compression. Since lossless reconstruction is not necessary in most cases, most efforts have been devoted to lossy compression. The performance of image / video compression algorithms is typically evaluated using two metrics: compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes; a lower ratio is better. Reconstruction quality is measured by comparing the reconstructed image / video with the original; a higher ratio is better.

[0045] Image / video compression technology can be divided into two branches, classical video codec methods and neural network-based video compression methods. Classical video codec schemes use transform-based solutions, where researchers exploit statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully hand-designing entropy codecs that model the dependencies in the quantization mechanism. There are two forms of neural network-based video compression, neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded in the existing classical video codec as a codec tool as part of the framework, while the latter is a separate framework developed based on neural networks and does not rely on classical video codecs.

[0046] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-increasing amount of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups, the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), and the ITU-T also has its own Video Coding Experts Group (VCEG), dedicated to standardizing image and video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard, the Versatile Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC reduces bitrate by an average of 50% while maintaining the same visual quality.

[0047] Neural network-based image / video compression is not new, as many researchers have been working on neural network-based image codecs. However, the network architectures are relatively shallow, and performance has been less than satisfactory. With the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, this technology is still far from mature, and many challenges need to be addressed.

[0048] 2.2 Neural Networks

[0049] Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques that typically consist of multiple processing layers, each composed of multiple simple but nonlinear computational primitives. One advantage of such deep networks is believed to be their ability to process data at multiple levels of abstraction and transform it into different kinds of representations. Note that these representations are not manually designed; rather, the deep network comprising the processing layers is learned from large amounts of data using a common machine learning process. Because deep learning eliminates the need for handcrafted representations, it is considered particularly well-suited for processing raw, unstructured data, such as acoustic and visual signals, which has long been a challenge in the field of artificial intelligence.

[0050] 2.3 Neural Networks for Image Compression

[0051] Existing neural networks for image compression can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding and decoding strategies, while the latter is a transformation-based solution. Sometimes, these two methods are combined in the literature.

[0052] 2.3.1 Pixel Probabilistic Modeling

[0053] According to Shannon's information theory, the optimal lossless codec method achieves a minimum decoding rate of -log2p(x), where p(x) is the probability of symbol x. Numerous lossless codec methods have been developed in the literature, with arithmetic coding being considered one of the best. Given a probability distribution p(x), arithmetic coding ensures that the decoding rate is as close to its theoretical limit of -log2p(x) as possible, without taking into account rounding errors. Therefore, the remaining question is how to determine the probability, which is very challenging for natural images and videos due to the curse of dimensionality.

[0054] Following the predictive coding strategy, one way to model p(x) is to predict pixel probabilities one by one in raster scan order based on previous observations, where x is an image.

[0055] p(x)=p(x1)p(x2|x1)…p(x i |x1,…,x i-1 )…p(x m×n |x1,…,x m×n-1 ) (1) where m and n are the height and width of the image, respectively. The previous observation is also called the context of the current pixel. When the image is large, estimating the conditional probability can be difficult, so a simpler approach is to limit the scope of its context.

[0056] p(x)=p(x1)p(x2|x1)…p(x i |x i-k,…,x i-1 )…p(x m×n |x m×n-k ,…,x m×n-1 ) (2) Where k is a predefined constant that controls the scope of the context.

[0057] It should be noted that this condition can also take into account the sample values of other color components. For example, when encoding and decoding RGB color components, the R sample depends on the previously encoded and decoded pixels (including R / G / B samples), and the current G sample can be encoded and decoded based on the previously encoded and decoded pixels and the current R sample. For encoding and decoding the current B sample, the previously encoded and decoded pixels and the current R and G samples can also be considered.

[0058] Neural networks were originally introduced for computer vision tasks and have been shown to be effective in regression and classification problems. Therefore, it is proposed to use neural networks to estimate the value of a given context x1, x2, ..., x i-1 p(x i ). For binary images, the pixel probability is proposed, that is, x i ∈{-1,+1}. The Neural Autoregressive Distribution Estimator (NADE) is designed for pixel-wise probability modeling, where is a feed-forward network with a single hidden layer. Similar work has been proposed, where the feed-forward network also has skip connections to the hidden layer and the parameters are also shared. These methods are experimented on the binarized MNIST dataset. NADE is extended to a real-valued model RNADE, where the probability p(x i |x1,…,x i-1 ) is derived using a Gaussian mixture. Their feedforward network also has one hidden layer, but the hidden layer is rescaled to avoid saturation and uses rectified linear units (ReLU) instead of sigmoids. NADE and RNADE are improved by reorganizing the order of pixels and using a deeper neural network.

[0059] Most of the above methods directly model the probability distribution in the pixel domain. Some researchers also try to model the probability distribution as a conditional probability distribution based on explicit or latent representations. That is, it can estimate

[0060]

[0061] Where h is an additional condition, and p(x)=p(h)p(x|h), which means that modeling is divided into unconditional modeling and conditional modeling. The additional condition can be image label information or high-level representation.

[0062] 2.3.2 Autoencoder

[0063] Autoencoders originate from the famous work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part converts the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but more channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. The ability of autoencoders to automatically learn representations and eliminate the need for handcrafted features is considered one of the most important advantages of neural networks.

[0064] Figure 2 is a schematic diagram illustrating an example transform coding scheme 200. The original image x is analyzed by the analysis network g a The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the codec rate. Then the synthetic network g s Inverse transform to obtain the reconstructed image Distortion (D) is calculated by using the function g in the perceptual space p Transform x and To calculate, we get z and Compare them to get D.

[0065] Applying autoencoder networks to lossy image compression is intuitive. It simply encodes the learned latent representation from a well-trained neural network. However, adapting autoencoders to image compression is non-trivial, as the original autoencoder is not optimized for compression and is therefore inefficient by directly using a trained autoencoder. Furthermore, other major challenges exist: First, the low-dimensional representation should be quantized before encoding, but quantization is non-differentiable, which is required for backpropagation when training neural networks. Second, the objectives in compression scenarios are different, as both distortion and rate need to be considered. Estimating the rate is challenging. Third, practical image coding and decoding schemes need to support variable rates, scalability, encoding / decoding speed, and interoperability. To address these challenges, many researchers have been actively contributing to this field.

[0066] Prototype autoencoders for image compression such as Figure 2 As shown, it can be regarded as a transform coding strategy. The original image x is analyzed using the network y = g a (x), where y is the latent representation that will be quantized and encoded. The synthesis network will inverse transform the quantized latent representation To obtain the reconstructed image The framework is trained using a rate-distortion loss function, i.e. where D is the value of x and The distortion between them, R is expressed from the quantization The rate calculated or estimated, and λ is the Lagrange multiplier. It should be noted that D can be calculated in the pixel domain or the receptive domain. All existing research works follow this prototype, and the difference may only be the network structure or loss function.

[0067] 2.3.3 Super-prior model

[0068] Figure 3 Example latent representations of images are shown, including an image 300 from the Kodak dataset, a visualization of the latent value 310 representation y of the image 300, the standard deviation σ 320 of the latent values 310, and the latent values y after introducing a super-prior network 330. The super-prior network includes a super-encoder and a decoder.

[0069] In such Figure 2 In the transform codec method for image compression shown in the figure, the encoder subnetwork (Section 2.3.2) uses a parametric analysis transform The image vector x is transformed into a latent representation y, which is then quantized to form because is discrete, so It can be losslessly compressed using entropy coding techniques such as arithmetic coding and transmitted as a bit sequence.

[0070] from Figure 3 It can be clearly seen that the potential value 310 and standard deviation σ320 There is a significant spatial dependence between the elements of . It is worth noting that their scaling (standard deviation σ320) appears to be coupled in the spatial domain. An additional set of random variables can be introduced to capture spatial dependencies and further reduce redundancy. In this case, the image compression network such as Figure 4 shown.

[0071] Figure 4 is a diagram 400 showing an example network architecture of an autoencoder implementing a super-prior model. The upper side shows the image autoencoder network, and the lower side corresponds to the super-prior subnetwork. The analysis and synthesis transforms are denoted as g a and g a Q represents quantization, and AE and AD represent arithmetic encoder and arithmetic decoder respectively. The super prior model consists of two sub-networks, the super encoder (denoted as h a ) and super decoder (denoted as h s ). The hyper-prior model generates quantified hyper-potential values It includes the potential value of Information related to the probability distribution of the sample points. is included in the bitstream and is Transmitted together to the receiver (decoder).

[0072] In the diagram 400, the upper side of the model is the encoder g as discussed above. a and decoder g s The lower side is used to obtain The additional super encoder h a and super decoder h s In this architecture, the encoder subjects the input image x to g a , producing a response y with a spatially varying standard deviation. The response y is fed to h a , summarize the distribution of standard deviations in z. Then, z is quantized is compressed and transmitted as side information. The encoder then uses the quantized vector To estimate the spatial distribution of the standard deviation σ, and use σ to compress and transmit the quantized image representation The decoder first recovers the compressed signal The decoder then uses h s to obtain σ, which provides the correct probability estimate to the decoder to also successfully recover The decoder will then Feed to g s to obtain the reconstructed image.

[0073] When the super encoder and super decoder are added to the image compression network, the quantized latent value The spatial redundancy is reduced. Figure 3 The potential value y 330 in corresponds to the quantized potential value when using a super encoder / decoder. Compared with the standard deviation σ 320, the spatial redundancy is significantly reduced because the correlation of the samples of the quantized potential value is smaller.

[0074] 2.3.4 Context Model

[0075] Although the hyper-prior model improves the quantified potential , but further improvement can be achieved by utilizing an autoregressive model (context model) that predicts the quantized potential value from its causal context.

[0076] The term autoregressive means that the output of a process is later used as its input. For example, the context model subnetwork generates one sample of potential values, which is later used as input to obtain the next sample.

[0077] Figure 5is a diagram 500 illustrating an example combined model configured to jointly optimize a context model with a hyper-prior and an autoencoder. The combined model jointly optimizes an autoregressive component (context model) that estimates a latent probability distribution from a latent causal context with a hyper-prior and an underlying autoencoder. The real-valued latent representation is quantized (Q) to create a quantized latent value and quantified excess potential It is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The dotted area corresponds to the components executed by a receiver (eg, a decoder) to restore the image from the compressed bitstream.

[0078] The example system utilizes a joint architecture where both the super-prior model sub-network (super-encoder and super-decoder) and the context model sub-network are utilized. The super-prior and context models are combined to learn the quantized latent values The probability model on the quantized potential value is then used for entropy coding and decoding. As shown in the schematic diagram 500, the outputs of the context sub-network and the super decoder sub-network are combined by a sub-network called the entropy parameter, which generates the mean μ and scale (or variance) σ parameters for the Gaussian probability model. The Gaussian probability model is then used to encode the quantized potential value samples into a bitstream with the help of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized potential value from the bitstream through the arithmetic decoder (AD) module.

[0079] In one example, the potential sample is modeled as a Gaussian distribution or a Gaussian mixture model (without limitation). In the example according to schematic diagram 500, the context model and the hyper-prior are jointly used to estimate the probability distribution of the potential sample. Since the Gaussian distribution can be defined by a mean and a variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as μ and σ).

[0080] 2.3.5 Gain Variational Autoencoder (G-VAE)

[0081] Typically, neural network-based image / video compression methods require training multiple models to adapt to different rates. Gain Variational Autoencoder (G-VAE) is a variational autoencoder with a pair of gain units, which is designed to achieve continuous variable rate adaptation using a single model. It consists of a pair of gain units that are usually inserted at the output of the encoder and the input of the decoder. The output of the encoder is defined as the potential representation y∈R c*h*w , where c, h, w represent the number of channels, height and width of the potential representation. Each channel of the potential representation is represented as y (i) ∈R h*w , where i = 0, 1, ..., c-1. A pair of gain units includes a gain matrix M∈Rc*n and the inverse gain matrix, where n is the number of gain vectors. The gain vector can be expressed as m s ={α s(0) ,α s(1) ,…,α s(c-1)}、α s(i) ∈R, where s represents the index of the gain vector in the gain matrix.

[0082] The motivation of the gain matrix is similar to the quantization table in JPEG, which controls the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.

[0083]

[0084] where ⊙ is the channel-wise multiplication, i.e. And α s(i) is the gain vector m s The inverse gain matrix used at the decoder side can be expressed as M′∈R c*n , which consists of n inverse gain vectors, namely M′={δ s(0) ,δ s(1) ,…,δ s(c-1)}、δ s(i) ∈R. The inverse gain process is expressed as

[0085]

[0086] in is the decoded quantized potential value representation, and y′ s is the quantized potential value representation of the inverse gain, which will be fed into the synthesis network.

[0087] To achieve continuously variable rate adjustment, interpolation is used between the vectors. Given two pairs of gain vectors {m t ,m′ t} and {m r ,m′ r}, the interpolated gain vector can be obtained by the following equation.

[0088] m v =[(m r ) l ·(m t ) 1-l ]

[0089] m′ v =[(m′ r ) l ·(m′ t ) 1-l ]

[0090] where l∈R is the interpolation coefficient that controls the corresponding bit rate of the generated gain vector pair. Since l is a real number, any bit rate between a given pair of gain vectors can be achieved.

[0091] 2.5.6 Encoding Latent Information Using Joint Autoregressive Hyper-Prior Models

[0092] Figure 5 The design in corresponds to the example combined compression method. In this section and the next, the encoding and decoding processes are described separately.

[0093] Figure 6 An example encoding process 600 is shown. The input image is first processed by the encoder sub-network. The encoder converts the input image into a transformed representation called a potential, denoted by y. Then, y is input to a quantizer block denoted by Q to obtain a quantized potential value Then, The arithmetic coding block (denoted as AE) is used to convert the bit stream (bits1). The arithmetic coding block is sequentially converted into Each sample point is converted into a bit stream (bits1) one by one.

[0094] The modular super-encoder, context, super-decoder, and entropy parameter sub-networks are used to estimate the quantized latent values The potential value y is input to the super encoder, and the super encoder outputs the super potential (denoted as z). Then, the super potential is quantized The arithmetic encoding (AE) module is used to generate a second bit stream (bits2). The decomposition entropy module generates a probability distribution used to encode the quantized super potential value into a bit stream. The quantized super potential value includes the probability distribution of the quantized potential value. information about the probability distribution of .

[0095] The entropy parameter subnetwork generates the potential value used to encode the quantized The information generated by the entropy parameter usually includes the mean μ and scale (or variance) σ parameters that are used together to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as The parameter μ is the mean or expectation of the distribution (also its median and mode), and the parameter σ is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.

[0096] The subnetwork superdecoder generates part of the information used by the entropy parameter subnetwork, and the other part is generated by an autoregressive module called the context module. The context module uses the samples that have been encoded by the arithmetic coding (AE) module to generate information about the probability distribution of the samples of the quantized potential value. It is usually a matrix composed of many sample points. Dimensions, you can use or The sample point is indicated by an index such as Typically, samples are encoded by the AE one by one using raster scan order. In raster scan order, the rows of the matrix are processed from top to bottom, where the samples in the row are processed from left to right. In such a scenario (where the AE encodes the samples into the bitstream using raster scan order), the context module uses the samples previously encoded in raster scan order to generate the same samples as the sample. The information generated by the context module and the super decoder is combined by the entropy parameter module to generate the potential value that is used to convert the quantized Encoded as a probability distribution of the bitstream (bits1).

[0097] Finally, the first bitstream and the second bitstream are transmitted to a decoder as a result of the encoding process.

[0098] It should be noted that the above modules can also use other names.

[0099] In the above description, Figure 6 All elements in are collectively referred to as encoders. The analysis transformation that transforms the input image into the latent representation is also called an encoder (or autoencoder).

[0100] 2.3.7 Decoding Latent Information Using Joint Autoregressive Hyper-Prior Model

[0101] Figure 7 An example decoding process 700 is shown. Figure 7 The decoding process is depicted separately. During the decoding process, the decoder first receives the first bit stream (bits1) and the second bit stream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module by using the probability distribution generated by the decomposition entropy sub-network. The decomposition entropy module typically uses a predetermined template to generate the probability distribution, such as a predetermined mean and variance value in the case of a Gaussian distribution. The output of the arithmetic decoding process of bits2 is It is the quantized super potential value. The AD process is restored to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized super potential value generated by the encoder is Can be reconstructed at the decoder without any changes.

[0102] In obtaining Afterwards, it is processed by the super decoder, and the output of the super decoder is fed into the entropy parameter module. The three sub-networks employed in the decoder, context, super decoder, and entropy parameter, are the same as those in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder (as in the encoder), which is essential for losslessly reconstructing the quantized potential values. As a result, the same quantized potential values as those obtained in the encoder can be obtained in the decoder. Same version.

[0103] After obtaining the probability distribution (such as mean and variance parameters) through the entropy parameter sub-network, the arithmetic decoding module decodes the quantized potential value samples one by one from the bitstream bits1. From a practical point of view, the autoregressive model (context model) is inherently serial and therefore cannot be accelerated using techniques such as parallelization.

[0104] Finally, the fully reconstructed quantized potential value is input to the composite transform (in Figure 7 ) module to obtain the reconstructed image.

[0105] In the above description, Figure 7 All elements in are collectively referred to as decoders. The synthetic transform that converts the quantized latent values into the reconstructed image is also called a decoder (or autodecoder).

[0106] 2.3.8 Encoding and Decoding Processes with Syntax Elements

[0107] In addition to the encoding and decoding potential information mentioned in 2.3.6 and 2.3.7, some syntax elements (e.g., image size, quality parameters, models used in post-processing, etc.) are also required in the decoding stage, which also need to be encoded and decoded in the bitstream. Therefore, the complete encoding process of the currently learned image compression is as follows.

[0108] Table 1. Algorithmic description of encoding with syntax elements

[0109]

[0110] During encoding, if the current syntax element is needed during decoding, it will be encoded by the arithmetic encoder (AE). After encoding and decoding the syntax element, the latent information is also stored in the bitstream by the same AE, as described in 2.3.6.

[0111] The decoding process with syntax elements can be viewed as the inverse operation of the encoding process, which is also given in Table 2.

[0112] Table 2. Algorithm description of encoding using syntax elements

[0113]

[0114]

[0115] During the decoding process, the arithmetic decoder (AD) is used to decode all the information needed for the decoding process. Table 3 shows the syntax elements that can be used during the encoding and decoding process.

[0116] Table 3. Syntax elements that can be used in the encoding and decoding process

[0117] name describe Image Head Image information (shape, bit depth, file format, etc.) res mutator Parameters used for reshaping (reshape dimensions, etc.) ICC grade Parameters for grades beta Scaling factor for variable rates active_tool_idx Model idx piece For the size of the piece y_max_symbol Maximum symbol for arithmetic coding and decoding mask_scale Tools for latent domain manipulation (scaling, block-based skipping, etc.) LDAO Parameters used for latent domain adaptation optimization wavefront Parameters for the wavefront ICCI Parameters used for post-processing

[0118] 2.4 Neural Networks for Video Compression

[0119] Similar to traditional video codecs, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology has lagged behind that of neural network-based image compression technology. However, due to its complexity, it requires more effort to address the challenges. Since 2017, several researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Inter-frame prediction is a key step in these efforts. Motion estimation and compensation are widely used, but only recently have they been implemented using trained neural networks.

[0120] Research on neural network-based video compression can be divided into two categories based on the target scenario: random access and low latency. In the case of random access, it is required that decoding can be started from any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, it aims to reduce decoding time, so typically only the previous frame in the temporal domain can be used as a reference frame to decode subsequent frames.

[0121] 2.5 Preliminary Knowledge

[0122] Almost all natural images / videos are in digital format. Grayscale digital images can be obtained by Indicates that is the set of pixel values, m is the image height, and n is the image width. For example, is a common setting, in which case Therefore a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed ones have certainly fewer bits.

[0123] Color images are usually represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by Representation, where three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24bpp. Digital images / videos can be represented in different color spaces. Neural network-based video compression schemes are mostly developed in the RGB color space, while traditional codecs typically use the YUV color space to represent video sequences. In the YUV color space, the image is decomposed into three channels, namely Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. The advantage is that Cb and Cr are usually downsampled for pre-compression because the human visual system is less sensitive to the chrominance components.

[0124] A color video sequence consists of multiple color images (called frames) to record scenes at different time stamps. For example, in the RGB color space, a color video can be represented by X = {x0, x1, ..., x t ,…,x T-1}, where T is the number of frames in the video sequence, If m=1080, n=1920, And the video has 50 frames per second (fps), then the data rate of the uncompressed video is 1920×1080×8×3×50=2,488,320,000 bits per second (bps), about 2.32 Gbps, which requires a lot of storage, so it definitely needs to be compressed before transmission over the Internet.

[0125] Typically, lossless methods can achieve a natural image compression ratio of approximately 1.5 to 3, which is clearly lower than required. Therefore, lossy compression has been developed to achieve higher compression ratios, but at the expense of distortion. Distortion can be measured by calculating the mean squared difference between the original image and the reconstructed image, also known as the mean squared error (MSE). For grayscale images, the MSE can be calculated using the following equation.

[0126]

[0127] Accordingly, the quality of the reconstructed image compared to the original image can be measured by the Peak Signal-to-Noise Ratio (PSNR):

[0128]

[0129] in yes The maximum value in , for example, for an 8-bit grayscale image is 255. There are other quality assessment metrics such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).

[0130] To compare different lossless compression schemes, it is sufficient to compare the compression ratio for a given rate, or vice versa. However, to compare different lossy compression methods, both the rate and the reconstruction quality must be considered. For example, calculating the relative rates at several different quality levels and then averaging the rates is a common approach; the average relative rate is called the Bjorntgaard delta rate (BD rate). Other important aspects of evaluating image / video codecs include encoding / decoding complexity, scalability, and robustness.

[0131] 3 questions

[0132] As described in 2.3.8, in the existing image compression framework, all semantic element information and latent information are compressed under the same codeword, such as Figure 8 As shown, Figure 8 The structure of the bitstream in the currently studied image compression framework is shown.

[0133] from Figure 8 As can be seen in Figure 2, all information is stored in a single codeword without any explicit bitstream structure. Following this structure, if specific information is needed, due to the characteristics of the arithmetic encoder, it must first decode all the information inside the bitstream.

[0134] Furthermore, in existing frameworks, luma and chroma information are stored in the same codeword. This requires decoding all information in the bitstream when only specific color information is needed. This behavior in existing solutions significantly reduces the efficiency of the decoding process and affects the flexibility and robustness of the bitstream.

[0135] Arithmetic coding differs from other forms of entropy coding, such as Huffman coding, in that instead of separating the input into its constituent symbols and replacing each with a code, arithmetic coding encodes the entire message into a single number (codeword), an arbitrary precision fraction q, where 0.0 ≤ q < 1.0. The codeword here specifies a bitstream containing information about multiple symbols.

[0136] A problem with coding methods such as arithmetic coding and decoding is that multiple symbols are encoded into a single codeword (eg, a single bit stream), so it is not possible to decode a specific symbol individually.

[0137] A codeword herein may correspond to a bitstream that is the output of an entropy encoder or an arithmetic encoder or any other form of variable length encoder. A codeword typically specifies one or more symbols that are encoded as a single bitstream, where in order to decode a second symbol in the bitstream, the first symbol also needs to be decoded. This happens, for example, if the first symbol is encoded and decoded using a variable length codec. In this case, the number of bits used to encode and decode the first symbol can only be determined after the value of the first symbol is decoded. In other words, the starting point of the second symbol in the bitstream can only be determined after the first symbol is decoded. Therefore, the first symbol and the second symbol are considered to be included in a single codeword (or single bitstream) because the decoding of the second symbol requires the decoding of the first symbol. The same is true if the first symbol and the second symbol are encoded and decoded using, for example, an arithmetic codec.

[0138] 4 Detailed solutions

[0139] The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way.

[0140] 4.1 Public Core

[0141] According to the present disclosure, a bitstream is encoded and decoded using multiple codewords, with different codewords being used to represent semantic element information and latent variable information for different color components. Furthermore, additional byte overhead is used in the bitstream to locate the corresponding position of each codeword containing different information.

[0142] 4.2 Details of the Implementation

[0143] 1. It is proposed that a (eg, first) codeword may be used to record common syntax elements used in decoding of luma and chroma components, and contains syntax elements required for decoding luma information.

[0144] • In one example, image size information (eg, width, height, and bit depth) may be stored in the first codeword.

[0145] • In one example, scaling information (eg, scaling width, scaling height, scaling method) may be stored in the first codeword.

[0146] ●In one example, the grade information may be stored in the first codeword.

[0147] • In one example, a parameter for variable rate (eg, beta in GVAE) can be used in the first codeword.

[0148] • In one example, the quality parameter / model index may be stored in the first codeword.

[0149] • In one example, parameters for wavefront parallel encoding and decoding may be stored in the first codeword.

[0150] • In one example, tile information (eg, the number of tiles) may be used in the first codeword.

[0151] • Alternatively, the ymax symbol may be stored in the first codeword.The y_max_symbol syntax element may identify the maximum symbol value of the quantized potential information, which may be used in arithmetic coding of the second codeword.

[0152] In one example, parameters for potential domain operations (e.g., scaling in potential information, block-based skipping in entropy coding) can be stored in the first codeword. Such parameters can include scalers or thresholds that can be used to scale or filter (select) potential samples in the process.

[0153] • In one example, potential domain adaptation optimization parameters may be stored in a first codeword.

[0154] • In one example, slice information (eg, the number of slices) for the luma component may be stored in the first codeword.

[0155] • In one example, parameters for post-processing may be stored in the first codeword.

[0156] • In one example, all syntax elements stored in the first codeword may be encoded and decoded by entropy encoding and decoding.

[0157] o In one example, some syntax elements may be coded via arithmetic coding.

[0158] ■ In one example, uniform distribution can be utilized as the probability distribution of the encoded and decoded symbols.

[0159] ■ In one example, the probability distribution can be calculated in a context-adaptive manner.

[0160] o In one example, some syntax elements may be encoded or decoded via fixed pattern bit strings.

[0161] o In one example, some syntax elements may be encoded or decoded via exp-Golomb encoding or decoding.

[0162] o In one example, some syntax elements can be encoded or decoded directly via unsigned / signed integers.

[0163] 2. It is proposed that a (eg, second) codeword can be used to preserve the latent variable information of the luminance component.

[0164] • In one example, latent variable information may contain hyper-prior information.

[0165] • In one example, the latent variable information may include residual information.

[0166] • In one example, latent variable information can be derived from the output of the synthesis network.

[0167] • In one example, latent variable information can be used as input to a synthesis network.

[0168] • Additionally or alternatively, the second codeword may comprise at least 2 codewords, wherein the first one (the first codeword of the second codeword) may comprise the super-prior information and the second one may comprise the residual latent information.

[0169] 3. A (eg, third) codeword is proposed to preserve the latent variable information of the chrominance component. All codewords will be concatenated to form the final bitstream.

[0170] • In one example, latent variable information may contain hyper-prior information.

[0171] • In one example, the latent variable information may include residual information.

[0172] • In one example, latent variable information can be derived from the output of the synthesis network.

[0173] • In one example, latent variable information can be used as input to a synthesis network.

[0174] • Additionally or alternatively, the third codeword may include at least 2 codewords, wherein the first one (the first codeword of the third codeword) may include super-prior information and the second one may include residual latent information.

[0175] 4. It is proposed that in order to obtain the position of each codeword in the bitstream, one or more bytes will be retained at the beginning of the bitstream.

[0176] • In one example, the relative offset from the initial position will be used to locate the starting position of each part.

[0177] • In one example, the size of each portion can be used to locate the position of each codeword.

[0178] ● In one example, 1 / 2 / 4 / 8 / 16 bytes can be used to locate the position of a codeword.

[0179] • In one example, to record position data, a big endian codec will be utilized.

[0180] 5. It is proposed that a (eg, fourth) codeword may be responsible for storing high-level syntax information required in the decoding of chroma components.

[0181] • In one example, parameters for wavefront parallel encoding and decoding may be stored in codewords.

[0182] • In one example, tile information (eg, number of tiles) may be used in the codeword.

[0183] • In one example, parameters for latent domain operations (eg, scaling in latent information, block-based skipping in entropy coding) can be stored in codewords.

[0184] 6. It is proposed that each part of a codeword (eg, those mentioned above) can be encoded and decoded in a parallel manner.

[0185] 7. It is proposed that luma-specific syntax and / or latent variables can be decoded without decoding chroma information.

[0186] • In one example, to decode only the luma syntax, only the codewords of the luma syntax in the bitstream will be decoded, and parsing of other parts will be skipped.

[0187] • In one example, to decode the luma latent variable, only the relevant codewords will be parsed and decoding of the rest will be skipped.

[0188] 8. It is proposed that additional byte alignment can be performed at the end of at least one of the codewords, i.e., if the codeword is not a multiple of 8 bits, padding bits can be used to make the length of the codeword a multiple of 8 bits. In one example, codeword byte alignment can be applied at the end of all codewords.

[0189] 9. It is proposed that codeword termination may be performed at the end of at least one of the codewords. In one example, codeword termination may be applied at the end of all codewords.

[0190] 10. It is proposed that arithmetic coding in end-to-end video / image coding can be context-based. The probability of the codeword can depend on the information that has been encoded / decoded.

[0191] 11. It is proposed that binary arithmetic coding can be applied in end-to-end video / image coding and decoding.

[0192] • In one example, a Context-Adaptive Binarized Arithmetic Coding (CABAC) codec, such as used in H.264 / AVC / HEVC / VVC, may be applied in the entropy codec of an end-to-end video / image codec.

[0193] Example 1:

[0194] Examples of embodiments are Figure 9 Depicted in.

[0195] According to an embodiment of the present disclosure, a bit stream may be divided into 5 parts.

[0196] The first part uses offsets to record the positions of the remaining 4 parts. For each offset, 4 bytes will be used to store the value, and it will be written through big-endian encoding and decoding.

[0197] The second part contains common information that can be used in the encoding and decoding of luma and chroma information (picture header, res changer, icc profile, y max symbol, mask_scale, ICCI). It also contains basic information used in the encoding and decoding of luma information (LDAO y, slice y, wavefront y). All syntax elements are encoded and decoded using arithmetic coding with uniform distribution.

[0198] ●The third part contains the latent residual information of brightness and the related super prior information.

[0199] ●The fourth part contains syntax elements that can only be used in the encoding and decoding of chroma components (LDAO uv, slice uv, wavefront uv).

[0200] ●The fifth part contains the latent residual information of chrominance and the related super prior information.

[0201] Example 2:

[0202] Another example of an embodiment of the present disclosure is also Figure 10 Depicted in.

[0203] According to an embodiment of the present disclosure, a bit stream may be divided into 6 parts.

[0204] The first part uses offsets to record the positions of the remaining 5 parts. For each offset, 4 bytes will be used to store the value, and it will be written through big-endian encoding and decoding.

[0205] • The second part contains general information that can be used in the coding and decoding of luma and chroma information (picture header, res changer, icc profile, y max symbol, mask_scale, ICCI). All syntax elements are coded by arithmetic coding with uniform distribution.

[0206] • The third part contains basic information used in the coding and decoding of luma information (LDAO y, slice y, wavefront y). All syntax elements are coded and decoded by arithmetic coding with uniform distribution.

[0207] ●The fourth part contains the latent residual information of brightness and the related super prior information.

[0208] ●The fifth part contains syntax elements that can only be used in the encoding and decoding of chroma components (LDAO uv, slice uv, wavefront uv).

[0209] ●The sixth part contains the latent residual information of chrominance and the related super prior information.

[0210] Example 3:

[0211] Another example of an embodiment of the present disclosure is also Figure 11 Depicted in.

[0212] According to an embodiment of the present disclosure, a bit stream may be divided into 6 parts.

[0213] The first part uses offsets to record the positions of the remaining 5 parts. For each offset, 4 bytes will be used to store the value, and it will be written through big-endian encoding and decoding.

[0214] The second part contains common information (picture header, res_modifier, icc_profile, y_max_symbol, mask_scale) that can be used in the encoding and decoding of luma and chroma information (without post-processing). It also contains basic information (LDAO y, slice y, wavefront y) used in the encoding and decoding of luma information. All syntax elements are encoded and decoded using arithmetic coding with uniform distribution.

[0215] ●The third part contains the latent residual information of brightness and the related super prior information.

[0216] ●The fourth part contains syntax elements that can only be used in the encoding and decoding of chroma components (LDAO uv, slice uv, wavefront uv).

[0217] ●The fifth part contains the chromaticity information and related super-prior information.

[0218] ●The sixth part contains information for post-processing.

[0219] Example 4:

[0220] Another example of an embodiment of the present disclosure is also Figure 12 Depicted in.

[0221] According to an embodiment of the present disclosure, a bit stream may be divided into 4 parts.

[0222] The first part contains common information that can be used in the encoding and decoding of luma and chroma information (picture header, res changer, icc profile, y max symbol, mask_scale, ICCI). It also contains basic information used in the encoding and decoding of luma information (LDAO y, slice y, wavefront y). All syntax elements are encoded and decoded using arithmetic coding with uniform distribution.

[0223] ●The second part contains the latent residual information of brightness and the related super prior information.

[0224] The third part contains syntax elements that can only be used in the encoding and decoding of chroma components (LDAO uv, slice uv, wavefront uv).

[0225] ●The fourth part contains the latent residual information of chrominance and the related super prior information.

[0226] ●At the beginning of each part, 4 bytes will be used to record the buffer size of each part.

[0227] 4.3 Benefits of the Implementation

[0228] According to an embodiment of the present disclosure, the structure of the bitstream is redesigned. Based on the functions of syntax elements and potential symbols, the bitstream will be divided into multiple parts, thereby improving the efficiency of decoding specific elements and enhancing the flexibility and robustness of the bitstream.

[0229] 5. Embodiment

[0230] 1. Decoder:

[0231] An image or video decoding method, including a neural sub-network, comprises the following ordered steps:

[0232] - get the position of each codeword in the bitstream,

[0233] - for a codeword containing a syntax element, decoding the syntax element according to the bits,

[0234] -Build the structure of neural network based on grammatical elements

[0235] - For code words containing latent information, decode the latent information according to the bits,

[0236] -Perform decoding of the latent information to obtain a reconstructed image.

[0237] More details will be discussed further below. Figure 13 FIG1 is a flowchart of a method 1300 for visual data processing according to an embodiment of the present disclosure. The method 1300 is implemented for conversion between a current visual unit of visual data and a bit stream of the visual data.

[0238] At block 1310, a plurality of codewords in a bitstream are determined. The codewords are associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit. At block 1320, a conversion is performed based on the plurality of codewords.

[0239] As an example, a bitstream can be encoded and decoded using multiple codewords. Different codewords are used to represent semantic element information and potential variable information for different color components, respectively. The term "codeword" herein may refer to a bitstream unit or bitstream segment that includes or stores at least one syntax element or information. A codeword can be separated from another codeword. As used herein, a "codeword" may be referred to as a "segment in a bitstream," "portion in a bitstream," "bitstream segment," or "unit in a bitstream." Figures 9 to 12 Several examples of bitstreams with multiple codewords are shown.

[0240] Method 1300 enables the application of multiple codewords rather than a single codeword. For example, these codewords can be processed in parallel. In this way, codec efficiency and / or codec effectiveness can be improved.

[0241] In some embodiments, the plurality of codewords includes a first codeword for at least one system element for encoding and decoding luma information and at least one color component. For example, the at least one color component includes at least one of a luma component or a chroma component. In other words, the (e.g., first) codeword may be used to record common syntax elements used in decoding luma and chroma components and include syntax elements required for decoding luma information.

[0242] In some embodiments, the first codeword includes at least one of the following: image size information of visual data, scaling information of visual data, grade information, parameters for variable rate, quality parameters, model index, parameters for wavefront parallel encoding and decoding, slice information, a syntax element indicating the maximum symbol value of quantized potential information, parameters for potential domain operation, potential domain adaptive optimization parameters, slice information for the luminance component, or parameters for post-processing.

[0243] In some embodiments, the image size information includes at least one of: width, height, or bit depth.

[0244] In some embodiments, the image size information includes at least one of: a scaled width, a scaled height, or a scale tool.

[0245] In some embodiments, the tile information includes the number of tiles.

[0246] In some embodiments, a syntax element indicating a maximum symbol value of the quantized potential information is used in arithmetic coding of a second codeword in the plurality of codewords. For example, a y_max_symbol syntax element may be stored in the first codeword. The y_max_symbol syntax element may identify a maximum symbol value of the quantized potential information, which may be used in arithmetic coding of the second codeword.

[0247] In some embodiments, the latent domain operation comprises at least one of scaling in latent information or block-based skipping in entropy coding, and the parameters for the latent domain operation comprise at least one of a scaler or a threshold for scaling or filtering latent samples for processing.

[0248] In some embodiments, the slice information for the luma component includes the number of slices for the luma component.

[0249] In some embodiments, at least one syntax element in the first codeword is coded by at least one of: entropy coding, arithmetic coding, fixed pattern bit string, or Exponential-Golomb coding.

[0250] In some embodiments, at least one syntax element is encoded and decoded by arithmetic coding, and uniform distribution is used as the probability distribution of the encoded at least one syntax element.

[0251] In some embodiments, at least one syntax element is encoded and decoded by arithmetic coding, and a probability distribution of the encoded at least one syntax element is determined in a context-adaptive manner.

[0252] In some embodiments, at least one syntax element in the first codeword is encoded or decoded using an unsigned or signed integer. For example, some syntax elements may be directly encoded or decoded using an unsigned / signed integer.

[0253] In some embodiments, the plurality of codewords includes at least one second codeword for latent variable information of a luma component.For example, the (eg, second) codeword may be used to store latent variable information of a luma component.

[0254] In some embodiments, the latent variable information of the luminance component includes at least one of the following: super-prior information or residual information.

[0255] In some embodiments, latent variable information for the luma component is determined from the output of the synthesis network for the conversion.

[0256] In some embodiments, the latent variable information of the luma component is an input to the synthesis network for the conversion.

[0257] In some embodiments, the at least one second codeword includes a plurality of second codewords including a first second codeword for super-prior information and a second second codeword for residual latent information.

[0258] In some embodiments, the plurality of codewords includes at least one third codeword for latent variable information of a chroma component.For example, the (eg, third) codeword may hold the latent variable information of the chroma component.

[0259] In some embodiments, the latent variable information of the chroma component includes at least one of the following: super-prior information or residual information.

[0260] In some embodiments, latent information for chroma components is determined based on the output of a synthesis network for the transformation.

[0261] In some embodiments, the latent information of the chroma components is an input to the synthesis network for the transformation.

[0262] In some embodiments, the at least one third codeword includes a plurality of third codewords, the plurality of third codewords including a first third codeword for super-prior information and a second third codeword for residual latent information.

[0263] In some embodiments, method 1300 further includes: generating a bitstream by concatenating multiple codewords. For example, all codewords may be concatenated to form a final bitstream.

[0264] In some embodiments, position information of the plurality of codewords is included in the bitstream.

[0265] In some embodiments, the position information is included in at least one byte at the beginning of the bitstream. For example, to obtain the position of each codeword in the bitstream, one or more bytes will be retained at the beginning of the bitstream. Additional byte overhead can be used in the bitstream to locate the corresponding position of each codeword containing different information.

[0266] In some embodiments, the at least one byte includes one of: one byte, two bytes, 4 bytes, 8 bytes, or 16 bytes.

[0267] In some embodiments, the position information includes a relative offset of the codeword from an initial position in the bitstream, and the starting position of the codeword is determined based on the relative offset.

[0268] In some embodiments, the position information includes a size of the codeword, and the position of the codeword is determined based on the size.

[0269] In some embodiments, the location information is encoded and decoded via bit-side encoding and decoding.

[0270] In some embodiments, the plurality of codewords includes a fourth codeword for high-level syntax information used in encoding and decoding of chroma components.For example, the (eg, fourth) codeword may be responsible for storing high-level syntax information required in decoding of chroma components.

[0271] In some embodiments, the fourth codeword includes at least one of the following: parameters for wavefront parallel encoding and decoding, slice information, and parameters for latent domain operations.

[0272] In some embodiments, the tile information includes the number of tiles.

[0273] In some embodiments, the latent domain operation comprises at least one of: scaling in the latent information or block-based skipping in entropy coding.

[0274] In some embodiments, byte alignment is performed at the end of a codeword in a plurality of codewords. As an example, the bit length of the codeword is not a multiple of the predefined bit length, and byte alignment is performed at the end of the codeword by adding padding bits, and the bit length of the aligned codeword is a multiple of the predefined bit length. For example, the predefined bit length can be 8 bits. That is, additional byte alignment can be performed at the end of at least one codeword in the codeword, that is, if the codeword is not a multiple of 8 bits, padding bits can be used to make the length of the codeword a multiple of 8 bits. In one example, codeword byte alignment can be applied at the end of all codewords.

[0275] In some embodiments, codeword termination is performed at at least one end of at least one codeword in the plurality of codewords.

[0276] In some embodiments, codeword termination is performed at the respective end of each of the plurality of codewords.

[0277] In some embodiments, multiple codewords are encoded and decoded in parallel.

[0278] In some embodiments, at least one of the first syntax element for the luma component or the second syntax element for the luma latent variable is encoded without encoding chroma information.

[0279] In some embodiments, a codeword for a first syntax element in a bitstream is encoded and decoded without parsing other codewords in the bitstream.

[0280] In some embodiments, codewords related to the luma latent variable are encoded without encoding or decoding other codewords in the bitstream.

[0281] In some embodiments, the conversion is performed by end-to-end video or image codec, and the arithmetic codec used in the end-to-end video or image codec is context-based.

[0282] In some embodiments, the probability of a codeword in the plurality of codewords is based on the encoded information.

[0283] In some embodiments, the conversion is performed by an end-to-end video or image codec, and the binarization arithmetic codec is applied in the end-to-end video or image codec.

[0284] In some embodiments, the binarization arithmetic codec includes context-adaptive binarization arithmetic codec (CABAC), and CABAC is applied in the entropy codec of the end-to-end video or image codec. For example, the CABAC codec such as that used in H.264 / AVC / HEVC / VVC can be applied in the entropy codec of the end-to-end video / image codec.

[0285] In some embodiments, converting comprises decoding the current visual unit from a bitstream.

[0286] In some embodiments, performing the conversion includes: determining corresponding positions of multiple codewords in a bitstream; for a codeword having at least one syntax element of the multiple codewords, decoding at least one syntax element based on the bitstream; determining a structure of a neural network for the conversion based on the decoded at least one syntax element; for another codeword having potential information of the multiple codewords, decoding the potential information based on the bitstream; and determining a reconstructed image based on the decoded potential information.

[0287] In some embodiments, converting includes encoding the current visual unit into a bitstream.

[0288] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of visual data, the bitstream generated by a method performed by a visual data processing apparatus. In the method, multiple codewords in the bitstream are determined. The codewords are associated with at least one of the following: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data. The bitstream is generated based on the multiple codewords.

[0289] According to further embodiments of the present disclosure, a method for storing a bitstream of visual data is provided. In this method, multiple codewords in a bitstream are determined. The codewords are associated with at least one of the following: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data. A bitstream is generated based on the multiple codewords. The bitstream is stored in a non-transitory computer-readable recording medium.

[0290] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0291] Item 1. A method for visual data processing, comprising: determining, for a conversion between a current visual unit of visual data and a bitstream of the visual data, a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component; and performing the conversion based on the plurality of codewords.

[0292] Item 2. The method of Item 1, wherein the plurality of codewords comprises a first codeword of at least one system element for encoding and decoding luma information and the at least one color component.

[0293] Item 3. The method of Item 2, wherein the at least one color component comprises at least one of a luma component or a chroma component.

[0294] Item 4. A method according to Item 2 or 3, wherein the first codeword includes at least one of the following: image size information of the visual data, scaling information of the visual data, profile information, parameters for variable rate, quality parameters, model index, parameters for wavefront parallel coding and decoding, slice information, a syntax element indicating the maximum symbol value of quantized potential information, parameters for potential domain operation, potential domain adaptive optimization parameters, slice information for the luminance component, or parameters for post-processing.

[0295] Item 5. The method of Item 4, wherein the image size information comprises at least one of: width, height, or bit depth.

[0296] Item 6. The method of Item 4, wherein the image size information comprises at least one of: a scaled width, a scaled height, or a scale tool.

[0297] Clause 7. The method of clause 4, wherein the slice information comprises the number of slices.

[0298] Clause 8. The method of clause 4, wherein the syntax element indicating a maximum symbol value of the quantized potential information is used for arithmetic coding of a second codeword in the plurality of codewords.

[0299] Item 9. A method according to Item 4, wherein the potential domain operation includes at least one of the following: scaling in potential information or block-based skipping in entropy coding and decoding, and the parameters for the potential domain operation include at least one of the following: a scaler or threshold for scaling or filtering potential samples for processing.

[0300] Clause 10. The method of clause 4, wherein the slice information for the luma component comprises the number of slices for the luma component.

[0301] Item 11. A method according to any one of Items 2 to 10, wherein at least one syntax element in the first codeword is encoded by at least one of the following: entropy coding, arithmetic coding, fixed pattern bit string or exponential-Golomb coding.

[0302] Item 12. The method of Item 11, wherein the at least one syntax element is encoded and decoded by the arithmetic coding, and a uniform distribution is used as the probability distribution of the at least one encoded syntax element.

[0303] Item 13. The method of Item 11, wherein the at least one syntax element is encoded and decoded by the arithmetic coding, and a probability distribution of the at least one encoded and decoded syntax element is determined in a context-adaptive manner.

[0304] Clause 14. A method according to any one of clauses 2 to 10, wherein at least one syntax element in the first codeword is encoded by an unsigned or signed integer.

[0305] Item 15. The method of any one of Items 1 to 14, wherein the plurality of codewords includes at least one second codeword for the latent variable information of a luma component.

[0306] Item 16. The method of Item 15, wherein the latent variable information of the luminance component comprises at least one of: super-prior information or residual information.

[0307] Item 17. The method of Item 15 or 16, wherein latent variable information for the luma component is determined from an output of a synthesis network for the transformation.

[0308] Item 18. The method of Item 15 or 16, wherein the latent variable information of the luminance component is an input to a synthesis network for the conversion.

[0309] Item 19. A method according to any one of Items 15 to 18, wherein the at least one second codeword comprises a plurality of second codewords, the plurality of second codewords comprising a first second codeword for super-prior information and a second second codeword for residual latent information.

[0310] Item 20. A method according to any one of Items 1 to 19, wherein the plurality of codewords comprises at least one third codeword for latent variable information of a chroma component.

[0311] Item 21. The method of Item 20, wherein the latent variable information of the chrominance component comprises at least one of: super-prior information or residual information.

[0312] Item 22. The method of Item 20 or 21, wherein the latent information of the chrominance component is determined based on an output of a synthesis network for the transformation.

[0313] Item 23. A method according to Item 20 or 21, wherein the latent information of the chrominance components is an input to a synthesis network for the transformation.

[0314] Item 24. A method according to any one of Items 20 to 23, wherein the at least one third codeword comprises a plurality of third codewords, the plurality of third codewords comprising a first third codeword for super-prior information and a second third codeword for residual latent information.

[0315] Item 25. The method of any one of Items 1 to 24, further comprising: generating the bitstream by concatenating the plurality of codewords.

[0316] Clause 26. A method according to any one of clauses 1 to 25, wherein position information of the plurality of codewords is included in the bitstream.

[0317] Item 27. The method of Item 26, wherein the position information is included in at least one byte at the beginning of the bitstream.

[0318] Item 28. The method of Item 27, wherein the at least one byte comprises one of: one byte, two bytes, 4 bytes, 8 bytes, or 16 bytes.

[0319] Item 29. A method according to any one of Items 26 to 28, wherein the position information comprises a relative offset of a codeword from an initial position in the bitstream, the starting position of the codeword being determined based on the relative offset.

[0320] Clause 30. A method according to any one of clauses 26 to 28, wherein the position information comprises a size of a codeword, the position of the codeword being determined based on the size.

[0321] Item 31. A method according to any one of Items 26 to 30, wherein the position information is encoded by bit-end encoding and decoding.

[0322] Item 32. A method according to any one of Items 1 to 31, wherein the plurality of codewords includes a fourth codeword for high-level syntax information used in encoding and decoding of chroma components.

[0323] Item 33. The method of Item 32, wherein the fourth codeword comprises at least one of: parameters for wavefront parallel coding and decoding, slice information, parameters for latent domain operations.

[0324] Clause 34. The method of clause 33, wherein the slice information comprises the number of slices.

[0325] Item 35. The method of Item 33, wherein the latent domain operation comprises at least one of: scaling in latent information or block-based skipping in entropy coding.

[0326] Item 36. A method according to any one of Items 1 to 35, wherein byte alignment is performed at the end of a codeword in the plurality of codewords.

[0327] Item 37. A method according to Item 36, wherein the bit length of the codeword is not a multiple of a predefined bit length, the byte alignment is performed at the end of the codeword by adding padding bits, and the bit length of the aligned codeword is a multiple of the predefined bit length.

[0328] Clause 38. The method of clause 37, wherein the predefined bit length is 8 bits.

[0329] Item 39. The method of any one of Items 1 to 38, wherein codeword termination is performed at at least one end of at least one codeword of the plurality of codewords.

[0330] Clause 40. The method of any one of clauses 1 to 38, wherein codeword termination is performed at the respective end of each of the plurality of codewords.

[0331] Item 41. A method according to any one of Items 1 to 40, wherein the plurality of codewords are encoded and decoded in parallel.

[0332] Item 42. A method according to any one of Items 1 to 41, wherein at least one of the first syntax element for the luma component or the second syntax element for the luma latent variable is encoded without encoding chroma information.

[0333] Item 43. The method of Item 42, wherein the codeword for the first syntax element in the bitstream is encoded and decoded without parsing other codewords in the bitstream.

[0334] Item 44. The method of Item 42, wherein codewords related to the luma latent variable are encoded without encoding or decoding other codewords in the bitstream.

[0335] Item 45. A method according to any one of Items 1 to 44, wherein the conversion is performed by an end-to-end video or image codec, and the arithmetic codec used in the end-to-end video or image codec is context-based.

[0336] Item 46. The method of Item 45, wherein the probability of a codeword in the plurality of codewords is based on coded information.

[0337] Item 47. A method according to any one of Items 1 to 46, wherein the conversion is performed by an end-to-end video or image codec and binarization arithmetic codec is applied in the end-to-end video or image codec.

[0338] Item 48. The method of Item 47, wherein the binarization arithmetic codec comprises context-adaptive binarization arithmetic codec (CABAC), and wherein the CABAC is applied in entropy coding of the end-to-end video or image codec.

[0339] Item 49. A method according to any one of Items 1 to 48, wherein the converting comprises decoding the current visual unit from the bitstream.

[0340] Item 50. A method according to Item 49, wherein performing the conversion includes: determining corresponding positions of the multiple codewords in the bitstream; for a codeword having at least one syntax element of the multiple codewords, decoding the at least one syntax element based on the bitstream; determining a structure of a neural network for the conversion based on the decoded at least one syntax element; for another codeword having potential information of the multiple codewords, decoding the potential information based on the bitstream; and determining a reconstructed image based on the decoded potential information.

[0341] Item 51. A method according to any one of Items 1 to 48, wherein the converting comprises encoding the current visual unit into the bitstream.

[0342] Item 52. An apparatus for visual data processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 51.

[0343] Item 53. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of Items 1 to 51.

[0344] Item 54. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by a visual data processing device, wherein the method comprises: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data; and generating the bitstream based on the plurality of codewords.

[0345] Item 55. A method for storing a bitstream of visual data, comprising: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data; generating the bitstream based on the plurality of codewords; and storing the bitstream in a non-transitory computer-readable recording medium.

[0346] Example device

[0347] Figure 14 A block diagram of a computing device 1400 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1400 may be implemented as, or included in, the source device 110 (or data encoder 114) or the destination device 120 (or data decoder 124).

[0348] It should be understood that Figure 14 The computing device 1400 shown in FIG. 1 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.

[0349] like Figure 14 As shown, computing device 1400 comprises a general computing device 1400. Computing device 1400 may include at least one or more processors or processing units 1410, memory 1420, storage unit 1430, one or more communication units 1440, one or more input devices 1450, and one or more output devices 1460.

[0350] In some embodiments, computing device 1400 can be implemented as any user terminal or server terminal with computing capability. A server terminal can be a server, a large computing device, etc. provided by a service provider. A user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 1400 can support any type of interface to a user (such as a "wearable" circuit device, etc.).

[0351] Processing unit 1410 may be a physical processor or a virtual processor and may implement various processes based on a program stored in memory 1420. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 1400. Processing unit 1410 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0352] The computing device 1400 typically includes various computer storage media. Such media can be any media accessible by the computing device 1400, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 1420 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 1430 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 1400.

[0353] The computing device 1400 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 14 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0354] The communication unit 1440 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 1400 can be implemented by a single computing cluster or multiple computing machines communicating via a communication connection. Thus, the computing device 1400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0355] Input device 1450 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1460 may be one or more of various output devices, such as a display, speaker, printer, etc. Computing device 1400 may also communicate with one or more external devices (not shown) via communication unit 1440, such as storage devices and display devices. Computing device 1400 may also communicate with one or more devices that enable a user to interact with computing device 1400, or any device that enables computing device 1400 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), if desired. Such communication may occur via an input / output (I / O) interface (not shown).

[0356] In some embodiments, some or all components of computing device 1400 may not be integrated into a single device, but may instead be arranged in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to implement the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to be aware of the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network that can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they appear to be a single access point for users. Thus, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, they may be provided from a conventional server or installed directly or otherwise on a client device.

[0357] The computing device 1400 can be used to implement visual data encoding / decoding in embodiments of the present disclosure. The memory 1420 can include one or more visual data encoding / decoding modules 1425 having one or more program instructions. These modules can be accessed and executed by the processing unit 1410 to perform the functions of the various embodiments described herein.

[0358] In an example embodiment where visual data encoding is performed, input device 1450 may receive visual data as input to be encoded 1470. The visual data may be processed by, for example, visual data encoding / decoding module 1425 to generate an encoded bitstream. The encoded bitstream may be provided as output 1480 via output device 1460.

[0359] In an example embodiment performing visual data decoding, an input device 1450 may receive an encoded bitstream as input 1470. The encoded bitstream may be processed, for example, by a visual data codec module 1425 to generate decoded visual data. The decoded visual data may be provided as output 1480 via an output device 1460.

[0360] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such changes are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for visual data processing, comprising: determining, for conversion between a current visual unit of visual data and a bitstream of the visual data, a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of the current visual unit; and The converting is performed based on the plurality of codewords. 2 . The method of claim 1 , wherein the plurality of codewords comprises a first codeword of at least one system element for encoding and decoding luminance information and the at least one color component. 3 . The method of claim 2 , wherein the at least one color component comprises at least one of a luma component or a chroma component.

4. The method according to claim 2 or 3, wherein the first codeword comprises at least one of the following: Image size information of the visual data, scaling information of the visual data, Grade information, For variable rate parameters, Quality parameters, Model index, For the parameters of wavefront parallel encoding and decoding, Film information, a syntax element indicating the maximum symbol value of the quantized potential information, Parameters for potential domain operations, Latent domain adaptation optimizes parameters, Slice information for the luma component, or Parameters for post-processing.

5. The method according to claim 4, wherein the image size information comprises at least one of the following: width, height, or Bit depth.

6. The method according to claim 4, wherein the image size information comprises at least one of the following: The width of the zoom, the scaled height, or Zoom tool. The method according to claim 4 , wherein the slice information includes the number of the slices.

8. The method of claim 4, wherein the syntax element indicating the maximum symbol value of the quantized potential information is used for arithmetic coding of a second codeword in the plurality of codewords.

9. The method of claim 4, wherein the latent domain operation comprises at least one of: scaling in latent information or block-based skipping in entropy coding, and The parameters for the potential domain operation include at least one of the following: a scaler or a threshold for scaling or filtering potential samples for processing. 10 . The method according to claim 4 , wherein the slice information for the luma component comprises the number of the slices for the luma component.

11. The method according to any one of claims 2 to 10, wherein at least one syntax element in the first codeword is encoded or decoded by at least one of the following: Entropy encoding and decoding, Arithmetic coding and decoding, Fixed pattern bit string, or Exponential-Golomb codec. 12 . The method according to claim 11 , wherein the at least one syntax element is coded by the arithmetic coding, and uniform distribution is used as the probability distribution of the coded at least one syntax element. 13 . The method according to claim 11 , wherein the at least one syntax element is coded by the arithmetic coding, and a probability distribution of the coded at least one syntax element is determined in a context-adaptive manner.

14. The method according to any one of claims 2 to 10, wherein at least one syntax element in the first codeword is encoded by an unsigned or signed integer.

15. The method according to any one of claims 1 to 14, wherein the plurality of codewords comprises at least one second codeword for the latent variable information of a luma component.

16. The method according to claim 15, wherein the latent variable information of the luminance component comprises at least one of the following: Super prior information, or Residual information.

17. The method of claim 15 or 16, wherein the latent variable information of the luma component is determined from an output of a synthesis network for the conversion.

18. The method of claim 15 or 16, wherein the latent variable information of the luma component is an input to a synthesis network for the conversion.

19. The method according to any one of claims 15 to 18, wherein the at least one second codeword comprises a plurality of second codewords, the plurality of second codewords comprising a first second codeword for super-prior information and a second second codeword for residual latent information.

20. The method according to any one of claims 1 to 19, wherein the plurality of codewords comprises at least one third codeword for the latent variable information of a chroma component.

21. The method according to claim 20, wherein the latent variable information of the chrominance component comprises at least one of the following: Super prior information, or Residual information.

22. The method of claim 20 or 21, wherein the latent information for the chroma components is determined based on an output of a synthesis network for the conversion.

23. The method of claim 20 or 21, wherein the latent information of the chroma components is an input to a synthesis network for the conversion.

24. The method according to any one of claims 20 to 23, wherein the at least one third codeword comprises a plurality of third codewords, the plurality of third codewords comprising a first third codeword for super-prior information and a second third codeword for residual latent information.

25. The method according to any one of claims 1 to 24, further comprising: The bit stream is generated by concatenating the multiple codewords.

26. The method according to any one of claims 1 to 25, wherein position information of the plurality of codewords is included in the bitstream.

27. The method of claim 26, wherein the position information is included in at least one byte at the beginning of the bitstream.

28. The method of claim 27, wherein the at least one byte comprises one of: one byte, two bytes, 4 bytes, 8 bytes, or 16 bytes.

29. The method according to any one of claims 26 to 28, wherein the position information comprises a relative offset of a codeword from an initial position in the bitstream, the starting position of the codeword being determined based on the relative offset.

30. The method according to any one of claims 26 to 28, wherein the position information comprises a size of a codeword, and the position of the codeword is determined based on the size.

31. The method according to any one of claims 26 to 30, wherein the position information is encoded and decoded by bit-end encoding and decoding.

32. The method of any one of claims 1 to 31, wherein the plurality of codewords includes a fourth codeword for high-level syntax information used in encoding and decoding of chroma components.

33. The method of claim 32, wherein the fourth codeword comprises at least one of: For the parameters of wavefront parallel encoding and decoding, Film information, or Parameters for potential domain operations. The method according to claim 33 , wherein the slice information includes the number of the slices.

35. The method of claim 33, wherein the potential domain operations include at least one of: zoom in on the underlying information, or Block-based skipping in entropy codecs.

36. The method of any one of claims 1 to 35, wherein byte alignment is performed at the end of a codeword in the plurality of codewords.

37. The method of claim 36, wherein the bit length of the codeword is not a multiple of a predefined bit length, the byte alignment is performed by adding padding bits at the end of the codeword, and the bit length of the aligned codeword is a multiple of the predefined bit length.

38. The method of claim 37, wherein the predefined bit length is 8 bits.

39. The method of any one of claims 1 to 38, wherein codeword termination is performed at at least one end of at least one codeword in the plurality of codewords.

40. The method of any one of claims 1 to 38, wherein codeword termination is performed at the respective end of each of the plurality of codewords.

41. The method according to any one of claims 1 to 40, wherein the plurality of codewords are encoded and decoded in parallel.

42. The method of any one of claims 1 to 41, wherein at least one of the first syntax element for a luma component or the second syntax element for a luma latent variable is encoded without encoding chroma information.

43. The method of claim 42, wherein the codeword for the first syntax element in the bitstream is encoded and decoded without parsing other codewords in the bitstream.

44. The method of claim 42, wherein codewords related to the luma latent variable are encoded and decoded without encoding and decoding other codewords in the bitstream.

45. The method according to any one of claims 1 to 44, wherein the conversion is performed by an end-to-end video or image codec, and the arithmetic codec used in the end-to-end video or image codec is context-based.

46. The method of claim 45, wherein a probability of a codeword in the plurality of codewords is based on coded information.

47. The method according to any one of claims 1 to 46, wherein the conversion is performed by an end-to-end video or image codec, and binarization arithmetic codec is applied in the end-to-end video or image codec.

48. The method of claim 47, wherein the binarization arithmetic codec comprises context-adaptive binarization arithmetic codec (CABAC), and the CABAC is applied in entropy codec of the end-to-end video or image codec.

49. The method of any one of claims 1 to 48, wherein the converting comprises decoding the current visual unit from the bitstream.

50. The method of claim 49, wherein performing the conversion comprises: determining corresponding positions of the plurality of codewords in the bitstream; For a codeword of at least one syntax element having the plurality of codewords, decoding the at least one syntax element based on the bitstream; determining a structure of a neural network for the conversion based on the decoded at least one syntax element; For another codeword having potential information of the plurality of codewords, decoding the potential information based on the bitstream; as well as A reconstructed image is determined based on the decoded latent information.

51. The method of any one of claims 1 to 48, wherein the converting comprises encoding the current visual unit into the bitstream.

52. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 51.

53. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1 to 51.

54. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by a visual data processing apparatus, wherein the method comprises: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data; and The bitstream is generated based on the plurality of codewords.

55. A method for storing a bitstream of visual data, comprising: determining a plurality of codewords in the bitstream, the codewords being associated with at least one of: semantic element information or latent variable information of at least one color component of a current visual unit of the visual data; generating the bitstream based on the plurality of codewords; as well as The bitstream is stored in a non-transitory computer-readable recording medium.