Methods, apparatuses, and media for visual data processing

By dividing the residual sample set into multiple subsets and adding a partitioning indicator to the bitstream, the encoding and decoding efficiency of neural network image/video encoding and decoding is improved, and independent encoding and decoding of different regions is supported.

CN122514780APending Publication Date: 2026-08-04DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2025-01-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing neural network-based image/video encoding and decoding technologies still have room for improvement in encoding and decoding efficiency, especially in supporting independent encoding and decoding of different regions.

Method used

By dividing the residual sample set into multiple residual sample subsets and including an indication in the bitstream of the number of vertical or horizontal divisions of the residual sample set, independent encoding and decoding of different regions can be supported.

Benefits of technology

It improves encoding and decoding efficiency and can better support independent encoding and decoding in different regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122514780A_ABST
    Figure CN122514780A_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a solution for visual data processing. A method for visual data processing is proposed. The method includes: performing a conversion between visual data and a bitstream of visual data using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the residual sample set, or a second indication indicating the number of horizontal divisions of the residual sample set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to visual data processing technology, and more specifically, to visual data encoding and decoding based on neural networks. Background Technology

[0002] Over the past decade, deep learning has made rapid progress across various fields, particularly computer vision and image processing. Neural networks were initially invented as part of interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression technology has made significant strides in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved rate-distortion (RD) performance comparable to that of Multifunctional Video Coding (VVC). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, the encoding and decoding efficiency of neural network-based image / video codecs is generally expected to be further improved. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for visual data processing.

[0004] In a first aspect, a method for visual data processing is proposed. The method includes: performing a conversion between visual data and a bitstream of visual data using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of the following: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0005] Based on the method according to the first aspect of this disclosure, the bitstream includes a first indication indicating the number of vertical divisions of the residual sample set and / or a second indication indicating the number of horizontal divisions of the residual sample set. Compared to conventional solutions lacking both indications, the proposed method can better support independent encoding and decoding of different regions, thereby improving encoding and decoding efficiency.

[0006] In a second aspect, an apparatus for visual data processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream generated by a method performed by a visual data processing device for visual data. The method includes: performing a conversion between visual data and a bitstream using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of the following: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0009] In a fifth aspect, a method for storing a bitstream of visual data is proposed. The method includes: performing a conversion between visual data and a bitstream using a neural network (NN)-based model; and storing the bitstream in a non-transitory computer-readable recording medium, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of the following: a first indication indicating the number of vertical divisions of the residual sample set, or a second indication indicating the number of horizontal divisions of the residual sample set.

[0010] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other objects, features and advantages of exemplary embodiments of the present disclosure will become clearer from the following detailed description with reference to the accompanying drawings, in which the same reference numerals generally refer to the same parts.

[0012] Figure 1A A block diagram of an example visual data encoding / decoding system according to some embodiments of the present disclosure is shown; Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme; Figure 2 An example potential representation of the image is shown; Figure 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model; Figure 4 This is a schematic diagram illustrating an example combined model configured to jointly optimize the context model with the super-prior and the autoencoder; Figure 5 An example encoding process is shown; Figure 6 An example decoding process is shown; Figure 7Example decoding processes according to some embodiments of this disclosure are shown; Figure 8 An example of a learning-based image codec architecture is shown; Figure 9 An example synthetic transform for learning-based image encoding and decoding is shown; Figure 10 An example of a modified linear unit (ReLU) activation function with leakage is shown; Figure 11 An example ReLU activation function is shown; Figure 12 The potential slices in the synthetic transformation are shown; Figure 13 The bitstream layout is shown; Figure 14 An example decoder structure is shown; Figure 15 An example of a variance decoder utilizing prior information is shown; Figure 16 An example decoder utilizing prior information is shown; Figure 17 A diagram illustrating an example multi-level context modeling (MCM) structure is shown. Figure 18 An example implementation of a primary component-guided adaptive upsampling filter is shown; Figure 19 An example bitstream structure is shown; Figure 20 Different types of partitioned regions according to embodiments of the present disclosure are shown; Figure 21 The overlapping according to an embodiment of the present disclosure is shown to be applied to a region boundary; Figure 22 The offset region according to an embodiment of this disclosure is shown; Figure 23 Region-independent encoding and decoding based on embodiments of the present disclosure are illustrated; Figure 24 The symbols for a control region partitioning mechanism according to an embodiment of the present disclosure are shown; Figure 25 Two example grids according to embodiments of this disclosure are shown; Figure 26 A flowchart of a method for visual data processing according to an embodiment of the present disclosure is shown; Figure 27 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0013] In all the accompanying drawings, the same or similar reference numerals generally denote the same or similar elements. Detailed Implementation

[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0015] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0016] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0017] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0019] Example Environment Figure 1AThis is a block diagram illustrating an example visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.

[0020] Visual data source 112 may include sources such as visual data capture devices. Examples of visual data capture devices include, but are not limited to, interfaces for receiving visual data from visual data providers, computer graphics systems for generating visual data, and / or combinations thereof.

[0021] Visual data may include one or more pictures or images from a video. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded / decoded representation of the visual data. The bitstream may include encoded / decoded pictures and associated visual data. An encoded / decoded picture is an encoded / decoded representation of a picture. Associated visual data may include sequence parameter sets, picture parameter sets, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data may be transmitted directly to a destination device 120 via network 130A through I / O interface 116. Encoded visual data may also be stored on storage medium / server 130B for access by the destination device 120.

[0022] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0023] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or further standards.

[0024] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Furthermore, although some embodiments describe encoding steps in detail, it should be understood that the corresponding decoding steps will be implemented by a decoder. Additionally, the term "visual data processing" includes visual data encoding / decoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another or at different compression bitrates. 1. Summary of the Invention This disclosure relates to image and video encoding and decoding based on neural networks (NNs). Specifically, it relates to support for region accessibility, where region accessibility refers to the ability to correctly decode only regional portions of an image (also known as a picture) or video. This idea can be applied alone or in various combinations to image and / or video encoding and decoding methods and specifications.

[0026] 2. Introduction Over the past decade, deep learning has made rapid progress across various fields, particularly computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression. Neural networks were initially invented as part of interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression techniques have made significant progress in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulties of the problem, neural network-based video coding and decoding is still in its early stages.

[0027] 1.1 Image / Video Compression Image / video compression (also known as image / video encoding / decoding) generally refers to the computational techniques used to compress images / videos into binary code to facilitate storage and transmission. The binary code may or may not support lossless reconstruction of the original image / video, a process known as lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is not always necessary. The performance of image / video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. The compression ratio is directly related to the amount of binary code; less is better. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; higher is better.

[0028] Image / video compression techniques can be divided into two branches: classical video codec methods and neural network-based video compression methods. Classical video codec schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) through carefully hand-designed entropy codes that model these dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based codec tools and end-to-end neural network-based video compression. The former is embedded as a codec tool within existing classical video codecs, serving only as part of the framework; the latter is a separate framework developed based on neural networks without relying on classical video codecs.

[0029] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-growing volume of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups, the Joint Picture Experts Group (JPEG) and the Moving Picture Experts Group (MPEG), while the ITU-T also has its own Video Codec Experts Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard, Multi-Functional Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC reduces the bitrate by an average of 50% while maintaining the same visual quality.

[0030] Neural network-based image / video compression is not a new invention, as many researchers have worked on neural network-based image encoding and decoding. However, the network architecture is relatively shallow, resulting in less than ideal performance. Benefiting from abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, this technology is still far from mature and many challenges need to be addressed.

[0031] 1.2 Neural Networks Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform the data into different types of representations. Note that these representations are not manually designed; instead, they are learned from massive amounts of data using general machine learning procedures, including deep networks with processing layers. Deep learning eliminates the need for hand-crafted representations and is therefore considered particularly suitable for processing native unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.

[0032] 1.3 Neural Networks for Image Compression Existing neural networks used for image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature.

[0033] 1.3.1 Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless encoding and decoding can achieve the lowest possible decoding rate. ,in It is a symbol The probability of [the lossy encoding / decoding method]. Many lossless encoding / decoding methods have been developed in the literature, among which arithmetic encoding / decoding is considered one of the best methods. Given a probability distribution... Arithmetic encoding and decoding ensures that the encoding / decoding rate is as close as possible to its theoretical limit without considering rounding errors. Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality.

[0034] Following the predictive encoding / decoding strategy, for One approach to modeling this is to predict pixel probabilities one by one in raster scan order based on previous observations. It's an image.

[0035] (1) in and These are the height and width of the image, respectively. Previous observations are also referred to as the current pixel's... Context When the image is large, estimating the conditional probability can be difficult, so a simplified approach is to limit the scope of its context.

[0036] (2) in It is a predefined constant that controls the scope of the context.

[0037] It should be noted that this condition can also take into account the sample values ​​of other color components. For example, when encoding and decoding RGB color components, the R sample depends on previously encoded and decoded pixels (including R / G / B samples), and the current G sample can be encoded and decoded based on previously encoded and decoded pixels and the current R sample. For the encoding and decoding of the current B sample, previously encoded and decoded pixels as well as the current R and G samples can also be considered.

[0038] Neural networks were initially introduced for computer vision tasks and have proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to estimate values ​​given their context. of The probability of.

[0039] Most methods directly model the probability distribution in the pixel domain. Some researchers have also attempted to model the probability distribution as a conditional probability distribution based on explicit or latent representations. That is, it can be estimated as... (3) in It is an additional condition, and This means that modeling is divided into unconditional modeling and conditional modeling. The additional conditions can be image label information or high-level representations.

[0040] 1.3.2 Automatic Encoder Autoencoders originate from the renowned work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part transforms the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders can automatically learn representations and eliminate the need for handcrafted features, which is considered one of the most significant advantages of neural networks.

[0041] Figure 1B A typical transform encoding / decoding scheme is shown. Original image. Analysis network Transformation to achieve latent representation The latent representation y is quantized and compressed into bits. The number of bits... Used to measure codec rate. Latent representation of quantization. Then by the synthetic network Inverse transform to obtain the reconstructed image By using functions Transformation and Distortion is calculated in the perceptual space.

[0042] Applying autoencoder networks to lossy image compression is intuitive. It simply requires encoding the latent representations learned from a trained neural network. However, adapting autoencoders to image compression is not straightforward, as the original autoencoders are not optimized for compression, making direct use of trained autoencoders inefficient. Furthermore, other major challenges exist: First, low-dimensional representations should be quantized before encoding, but quantization is non-differentiable, necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios, requiring consideration of both distortion and bit rate. Estimating the bit rate is challenging. Third, practical image encoding / decoding schemes need to support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Many researchers have been actively contributing to this field to address these challenges.

[0043] Prototype autoencoders for image compression, such as Figure 1B As shown, it can be regarded as Transform encoding and decoding Strategy. Original image Depend on analyze network Transformation, where It is the potential representation that will be quantized and encoded / decoded. synthesis The network quantizes the latent representation of the inverse transform. To obtain the reconstructed image The framework is trained using a rate-distortion loss function, i.e. ,in yes and Distortion between From quantification The calculated or estimated bit rate, and These are Lagrange multipliers. It should be noted that... It can be computed in the pixel domain or the receptive domain. All existing research follows this prototype, and the differences may only be in the network structure or the loss function.

[0044] 1.3.3 Super-prior model In the transform encoding and decoding method for image compression, the encoder subnetwork (Section 2.3.2) uses parametric analysis of the transform. Transform the image vector x into a latent representation Then quantify it to form .because Since they are discrete values, they can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as bit sequences.

[0045] from Figure 2 The left and right center images clearly show this. Significant spatial dependencies exist among the elements. Notably, their scales (middle right image) appear to be spatially coupled. An additional set of random variables can be introduced. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as Figure 3 As shown.

[0046] exist Figure 3 In the middle, the encoder is on the left side of the model. and decoder (Explained in Section 2.3.2). The right side is used to obtain... Additional encoders utilizing prior information and decoders that utilize prior information Network. In this architecture, the encoder subjectes the input image x to... This produces a response with a standard deviation that varies in the spatial domain. .response fed to In summary The distribution of standard deviations in the data. Then it is quantified ( The data is compressed and transmitted as side information. The encoder then uses the quantization vector... To estimate the spatial distribution of standard deviation And use it to compress and transmit quantized image representations. The decoder first recovers from the compressed signal. Then it uses To obtain This provided it with the correct probability estimate, which also enabled successful recovery. Then it will Feed to To obtain the reconstructed image.

[0047] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. Figure 2 The rightmost image in the diagram corresponds to the quantization latent value when using an encoder / decoder that leverages prior information. Compared to the middle right image, spatial redundancy is significantly reduced because the samples of the quantization latent value have lower correlation.

[0048] exist Figure 2 Middle: Left: Image from the Kodak dataset. Middle left: Visualization of the latent representation y of this image. Middle right: Standard deviation of the latent values. Right side: The latent value y after introducing a super-prior network (an encoder and a decoder that utilize super-prior information).

[0049] Figure 3 The network architecture of an autoencoder implementing a prior model is shown. The left side shows the image autoencoder network, and the right side corresponds to the prior subnetwork. The analytic and synthetic transforms are represented as follows: and Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks, utilizing the hyperprior information of the encoder (denoted as...). ) and decoders that utilize prior information (represented as The advanced prior model generates quantified advanced prior information latent values ​​(). ), which includes information on quantifying potential values Information about the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).

[0050] 1.3.4 Context Model Although the prior model improves the quantification of latent values Modeling the probability distribution of quantified potential values ​​is possible, but additional improvements can be obtained by utilizing an autoregressive model (context model) that predicts quantified potential values ​​from the causal context of quantified potential values.

[0051] The term autoregressive means that the output of a process is later used as its input. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.

[0052] Figure 4 This is a schematic diagram illustrating an example combined model configured to jointly optimize the context model with the super-prior and autoencoder. Table 1 below shows the meaning of the different symbols.

[0053] Table 1 – Symbol Explanation

[0054] A joint architecture can be utilized, which employs both a super-prior model subnetwork (an encoder and a decoder utilizing super-prior information) and a context model subnetwork. The super-prior and context models are combined to learn about the quantized latent value. The probability model is then used for entropy encoding and decoding. For example... Figure 4 As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean for the Gaussian probability model. And variance (scale) (or variance) The parameters are then used. A Gaussian probability model is then used to encode the samples of the quantized latent values ​​into a bitstream with the aid of the arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is used to obtain the quantized latent values ​​from the bitstream via the arithmetic decoder (AD) module. .

[0055] Figure 4 This illustrates a combined model that jointly optimizes the autoregressive component and the superprior, along with the underlying autoencoder. The autoregressive component estimates the probability distribution of the latent values ​​from the causal context of the latent values ​​(context model). The real-valued latent representation is quantized (Q) to create quantized latent values ​​(Q). ) and quantified potential value of prior information ( The image is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted area corresponds to the component performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream.

[0056] Typically, latent samples are modeled as Gaussian distributions or Gaussian mixture models (not limited to). According to Figure 4 The context model and the super-prior are jointly used to estimate the probability distribution of potential samples. Since the Gaussian distribution can be defined by its mean and variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as...). and ).

[0057] 1.3.5 Gain Variational Automatic Encoder (G-VAE) Typically, neural network-based image / video compression methods require training multiple models to adapt to different bit rates. Gain variational autoencoders (G-VAEs) are variational autoencoders with a pair of gain units, designed to achieve continuous variable bit rate adaptation using a single model. It consists of a pair of gain units, typically inserted into the encoder's output and the decoder's input. The encoder's output is defined as the latent representation. ,in This represents the number of channels, height, and width of the latent representation. Each channel of the latent representation is represented as... ,in A pair of gain units includes a gain matrix. and the inverse gain matrix, where This is the number of gain vectors. A gain vector can be represented as... , ,in Indicates the index of the gain vector in the gain matrix.

[0058] The motivation for the gain matrix is ​​similar to the quantization table in JPEG, controlling the quantization loss based on the characteristics of different channels. To apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.

[0059]

[0060] in It is channel multiplication, that is ,and It is the gain vector The first in There are several gain values. The inverse gain matrix used on the decoder side can be represented as... It is made of Composed of an inverse gain vector, i.e. , The inverse gain process is represented as:

[0061] in It is the quantized latent representation of the decoded data, and It is the quantized latent representation of the inverse gain, which will be fed into the synthesis network.

[0062] To achieve continuous variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors... and The interpolation gain vector can be obtained using the following equation.

[0063]

[0064]

[0065] in These are the interpolation coefficients, which control the bit rate of the generated gain vector pairs. Because... Since it is a real number, it is possible to achieve any bit rate between two given gain vector pairs.

[0066] 1.3.6 Encoding process using a joint autoregressive hyperprior model Figure 4 This corresponds to existing compression methods. The encoding and decoding processes will be described in this section and the next section, respectively.

[0067] Figure 5 The encoding process is described. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transform representation called the latent value, which is... express. It is then fed into the quantizer block, denoted by Q, to obtain the quantization potential value ( ). Then, using an arithmetic coding module (denoted as AE), it is converted into a bitstream (bits1). The arithmetic coding blocks are sequentially... Each sample point is converted into a bit stream (bits1) one after another.

[0068] The module utilizes an encoder with prior information, context, a decoder with prior information, and an entropy parameter subnetwork to estimate the quantized latent value. The probability distribution of the sample points. Latent value The information is input into an encoder that utilizes prior information, and its output is the potential value of the prior information (by...). (Representation). Then the potential value of the prior information is quantized ( The arithmetic encoding (AE) module generates a second bitstream (bits2). The decompositional entropy module generates a probability distribution used to encode the quantized prior information latent values ​​into a bitstream. The quantized prior information latent values ​​include information about the quantized latent values ​​(…). Information about the probability distribution of ( ).

[0069] Entropy parameter subnetwork generates latent values ​​for encoding quantization. The probability distribution estimate. Information generated from the entropy parameter typically includes the mean. And variance (scale) (or variance) The parameters, together, are used to obtain the Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as follows: , where parameters It is the mean or expected value of the distribution (also the median and mode), while the parameter This is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. In the existing design, the entropy parameter module is used to estimate the mean and variance values.

[0070] The subnetwork uses a decoder with prior information to generate part of the information used by the entropy parameter subnetwork, while another part is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the quantized latent values. It is usually a matrix composed of many sample points. According to the matrix The dimension, the sample points can use such as [i,j,k] or Indices like [i,j] are used to indicate sample points. [i,j] are typically encoded one after another by the AE using a raster scan order. In the raster scan order, the rows of the matrix are processed from top to bottom, with the samples within each row processed from left to right. In such a scenario (where the AE encodes the samples into a bitstream using the raster scan order), the context module uses the samples previously encoded in the raster scan order to generate a bitstream with the samples. Information related to [i,j]. The information generated by the context module and the decoder utilizing prior information is combined by the entropy parameter module to generate information for quantizing the latent values. The probability distribution encoded as a bit stream (bits1).

[0071] Finally, the first and second bit streams, as the result of the encoding process, are transmitted to the decoder.

[0072] It should be noted that other names can be used for the modules mentioned above.

[0073] In the above description, Figure 5 All elements in the encoder are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).

[0074] 1.3.7 Decoding process using a joint autoregressive superprior model Figure 6 The decoding process is described separately.

[0075] During decoding, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decomposed entropy subnetwork. The decomposed entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values ​​in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This refers to the quantized potential value of the prior information. The AD process reverts to the AE process applied in the encoder. Both the AE and AD processes are lossless, meaning that the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.

[0076] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks used in the decoder—the context, the decoder utilizing prior information, and the entropy parameter—are the same as the three sub-networks in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder as in the encoder, which is crucial for lossless reconstruction of the quantized latent value. This is crucial. As a result, the quantized latent values ​​can be obtained in the decoder as well as those obtained in the encoder. Same version.

[0077] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of quantized potential values ​​one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial, and therefore cannot be accelerated using techniques such as parallelization.

[0078] Finally, the fully reconstructed quantized potential value Input into the synthesis transform (in) Figure 6 The module (represented as decoder) is used to obtain the reconstructed image.

[0079] In the above description, Figure 6 All elements in the array are collectively referred to as the decoder. The synthetic transformation that converts quantized latent values ​​into a reconstructed image is also called a decoder (or automatic decoder).

[0080] 1.4 Neural Networks for Video Compression Similar to traditional video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, more effort is needed to address the challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to remove inter-frame redundancy. Inter-frame prediction is a key step in these works. Motion estimation and compensation have been widely adopted, but only recently have they been implemented using trained neural networks.

[0081] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the case of random access, decoding can begin from any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, the aim is to reduce decoding time; therefore, usually only the earlier frames in the temporal domain can be used as reference frames to decode subsequent frames.

[0082] 1.5 Preliminary Knowledge Almost all natural images / videos are in digital format. Grayscale digital images can be generated by... It means that among them It is a collection of pixel values. It is the image height. It is the image width. For example, This is a common setting, in this case Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while compressed images certainly have fewer bits.

[0083] Color images are typically represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by... This means that three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Most neural network-based video compression schemes are developed in the RGB color space, while traditional codecs typically use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels: Y, Cb, and Cr, where Y is the luminance component and Cb / Cr are the chrominance components. The advantage is that Cb and Cr are often downsampled for pre-compression because the human visual system is less sensitive to the chrominance component.

[0084] A color video sequence consists of multiple color images (called frames) to record a scene at different timestamps. For example, in the RGB color space, a color video can be composed of... It means that among them It is the frame number in the video sequence. .if , , If the video has 50 frames per second (fps), then the data bit rate of the uncompressed video is... Bits per second (bps), approximately 2.32 Gbps, requires a significant amount of storage, so compression is definitely necessary before transmission over the internet.

[0085] Typically, lossless methods can achieve compression ratios of around 1.5 to 3 for natural images, which is clearly below what is required. Therefore, lossy compression has been developed to achieve higher compression ratios, but at the cost of introducing distortion. Distortion can be measured by calculating the mean squared difference between the original and reconstructed images, known as the mean squared error (MSE). For grayscale images, the MSE can be calculated using the following equation.

[0086] (4) Accordingly, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR): (5) in yes The maximum value in the range is 255, for example, for an 8-bit grayscale image. Other quality assessment metrics include structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).

[0087] To compare different lossless compression schemes, it is sufficient to compare the compression ratio at a given bitrate, or vice versa. However, to compare different lossy compression methods, both bitrate and reconstruction quality must be considered. For example, a common method is to calculate the relative bitrates at several different quality levels and then average them; the average relative bitrate is called Bjontegaard's incremental bitrate (BD bitrate). Other important aspects for evaluating image / video encoding / decoding schemes include encoding / decoding complexity, scalability, and robustness.

[0088] 1.6 Separate processing of the luminance and chrominance components of an image Figure 7 The decoding process according to some example embodiments of this disclosure is illustrated. According to one implementation, the luminance and chrominance components of an image can be decoded using separate sub-networks. Figure 7 In this process, the luminance component of the image is processed by sub-networks such as "synthesis", "predictive fusion", "mask convolution (Mask Conv)", "decoder using prior information", and "variance decoder using prior information". The chrominance component is processed by sub-networks such as "synthesis UV", "predictive fusion UV", "mask convolution UV", "decoder UV using prior information", and "variance decoder UV using prior information".

[0089] The advantage of this separate processing is that it reduces the computational complexity of image processing. Typically, in neural network-based image and video decoding, computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is 192, the computational complexity will be proportional to 192 × 192. On the other hand, if the feature maps are divided into 128 for luminance and 64 for chrominance (in the case of separate processing), the computational complexity is proportional to 128 × 128 + 64 × 64, which corresponds to a 45% reduction in complexity. Generally, separate processing of the luminance and chrominance components of an image does not lead to a significant performance degradation because the correlation between the luminance and chrominance components is usually very small.

[0090] Figure 7 The processing (decoding process) in the code can be explained as follows: 1. First, the decompositional entropy model is used to decode the quantization potential values ​​of luminance and chrominance, i.e. Figure 7 In and .

[0091] 2. The probability parameters (e.g., variance) generated by the second network are used to generate quantized residual latent values ​​by performing an arithmetic decoding process.

[0092] 3. For example Figure 7 As shown in orange, the quantized residual latent value is inversely amplified using an inverse gain unit (iGain). The output of the inverse gain unit is expressed for the luminance and chrominance components as follows: and .

[0093] 4. For the luminance component, repeat the following steps in a loop until the desired result is obtained. All elements: a. The first subnetwork is used for... The already obtained samples are used to estimate the quantized potential value ( The mean value parameter of ).

[0094] b. Quantified residual potential value The mean value is used to obtain The next element. 5. In obtaining After all the samples are obtained, a synthetic transformation can be applied to obtain the reconstructed image.

[0095] 6. For the chromaticity components, steps 4 and 5 are the same, but with a separate network set.

[0096] 7. The decoded luminance component is used as additional information to obtain the chrominance component. Specifically, the Inter-Frame Channel Related Information Filter (ICCI) subnetwork is used for chrominance component recovery. Luminance is fed as additional information into the ICCI subnetwork to assist in chrominance component decoding.

[0097] 8. After the luminance and chrominance components are reconstructed, perform adaptive color transformation (ACT).

[0098] The module named ICCI is a neural network-based post-processing module. The exemplary embodiments of this disclosure are not limited to the ICCI sub-network; any other neural network-based post-processing module may also be used.

[0099] Exemplary implementations of some example embodiments of this disclosure are as follows: Figure 7The decoding process is illustrated in the diagram. The framework comprises two branches, one for the luma component and one for the chroma component. Within each branch, the first sub-network includes context, prediction, and optionally a decoder module utilizing prior information. The second network includes a variance decoder module utilizing prior information. The quantized prior information latent value is... and The arithmetic decoding process generates quantized residual latent values, which are then fed into the iGain unit to obtain quantized residual latent values ​​for gain. and .

[0100] After obtaining the residual potential values, a recursive prediction operation is performed to obtain the potential values. and The following steps describe how to obtain potential values. The sample points are processed in the same way as the chromaticity components, but using different networks.

[0101] 1. The autoregressive context module is used to utilize sample points. This is used to generate the first input to the prediction module, where the (m, n) pairs are the indices of the sample points of the obtained potential values.

[0102] 2. Optionally, the second input to the prediction module is obtained by using a decoder that utilizes prior information and quantized potential values ​​of the prior information. And thus obtained.

[0103] 3. Using the first and second inputs, the prediction module generates the mean. .

[0104] 4. Mean And quantified residual potential value Added together to obtain potential value . 5. Repeat steps 1-4 for the next sample point.

[0105] Whether and / or how at least one of the methods disclosed herein can be applied, for example, to transmit signals from the encoder to the decoder in a bitstream.

[0106] Alternatively, whether and / or how to apply at least one of the methods disclosed herein can be determined by the decoder based on encoding / decoding information such as dimensions, color format, etc.

[0107] Alternate or additional sites, named MS1, MS2 or MS3+O (in Figure 7 Modules (in the processing stream) can be included in the processing stream. A module can perform operations on its input by multiplying it by a scalar or adding an additional component to the input to obtain an output. The scalar or additional component used by the module can be indicated in the bitstream.

[0108] Figure 7 The module named RD or AD in the code can be an entropy decoding module. It can be a range decoder or an arithmetic decoder, etc.

[0109] The exemplary embodiments of this disclosure described herein are not limited to those described herein. Figure 7 The example uses a specific combination of units. Some modules may be missing, and some modules may be shifted according to the processing order. Additional modules may also be included. For example: 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined by another module, which can be based on a neural network.

[0110] 2. One or more of the modules named MS1, MS2, or MS3+O can be removed. The core of the proposed solution is unaffected by the removal of one or more of the aforementioned scaling and adding modules.

[0111] exist Figure 7 The bitstream also uses an asterisk (*) to indicate other operations performed during the processing of the luminance and chrominance components. These operations are denoted as MS1, MS2, MS3+0. These operations can be, but are not limited to, adaptive quantization, latent sample scaling, and latent sample offset operations. For example, in adaptive quantization, this might correspond to scaling the samples with a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. Latent scaling might correspond to scaling the samples with a multiplier after the prediction process, where the multiplier value is predefined or indicated in the bitstream. Offset operations might correspond to adding an appended element to the sample, where the value of the appended element can be indicated in the bitstream or estimated or predetermined.

[0112] Another operation can be slicing, where the samples are first sliced ​​(grouped) into overlapping or non-overlapping regions, each of which is processed independently. For example, samples corresponding to the luminance component can be divided into slices with a slice height of 20 samples, while the chrominance component can be divided into slices with a slice height of 10 samples for processing.

[0113] Another operation can be the application of wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the number of samples that can be processed in parallel can be indicated by control parameters. These control parameters can be indicated, estimated, or predetermined in the bitstream. In the case of separate luma and chroma processing, the number of samples that can be processed in parallel can be different, so different indicators can be transmitted in the bitstream to control the operation of luma and chroma processing separately.

[0114] 1.7 Color Separation and Conditional Encoding / Decoding In one example, such as Figure 8 As shown, networks with similar architectures but different numbers of channels are used to separately encode and decode the primary and secondary color components of an image. All boxes with the same name are sub-networks with similar architectures, differing only in input / output tensor sizes and the number of channels. The number of channels for the primary component is... The number of channels for the secondary component is The vertical arrows (pointing downwards) indicate the data flow related to the encoding and decoding of secondary color components. The vertical arrows show the data exchange between the primary and secondary component pipelines.

[0115] The input signal to be encoded is represented as The latent space tensor in the bottleneck of the variational autoencoder is The subscript "Y" indicates the primary component, and the subscript "UV" is used for the spliced ​​secondary components, i.e., the chromaticity components.

[0116] Figure 8 A learning-based image codec architecture is shown.

[0117] First, the input image in RGB color format is converted into primary (Y) and secondary (UV) components. Primary component... and secondary components The primary component is encoded and decoded independently, and the encoded / decoded image size is equal to the input / decoded image size. Secondary components are conditionally encoded and decoded, using data from the primary component. Encoding as auxiliary information And using auxiliary information from the main component. Decoding as a potential tensor Reconstruction. Aside from the number of channels, channel size, and several entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost identical. Therefore, the primary and secondary latent tensors will generate two different bitstreams based on two different entropy models. Before encoding, , The module adjusts the sample point position (labeled "s" in Figure 1) through downsampling, which essentially means that the encoded / decoded image size of the secondary component differs from that of the primary component. The scaling factor s is variable, but the default scaling factor is [value missing]. In conditional encoding and decoding, the size of the auxiliary input tensor is adjusted so that the encoder receives primary and secondary component tensors with the same image size. After reconstruction, the secondary components are rescaled to the original image size using a neural network-based upsampling filter module (“NN color filters” in Figure 1), the output of which is factored... The secondary component that is upsampled.

[0118] Figure 8 The example illustrates an image encoding / decoding system where the input image is first converted into primary (Y) and secondary (UV) components. Output... , This corresponds to the reconstructed output of the primary and secondary components. At the end of the processing, , It is converted back to RGB color format. Typically, this is done before processing using encoding and decoding modules (neural networks). Downsampled (resized). For example, The size can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Therefore, the processing of minor components involves approximately 50% × 50% = 25% fewer samples, resulting in lower computational complexity.

[0119] 1.8 Pruning Operations in Neural Network-Based Encoding and Decoding Figure 9 An example of a synthetic transform for learning-based image encoding / decoding is shown. The example synthetic transform above consists of a series of 4 convolutions with an upsampling stride of 2. The synthetic transform subnetwork is as follows: Figure 9 As shown. The size of the tensors in different parts of the synthesis transform before the clipping layer is Figure 9 The image above.

[0120] Cropping layers will change the tensor size. Change to ,in ;here It is the depth of the convolution performed in the codec architecture. For the principal components, Synthesis Transformation Receive size is The input tensor; The main components Synthesis Transformation The output is ,in .

[0121] For secondary components Synthesis Transformation Receive size is The input tensor; The main components Synthesis Transformation The output is ,in For secondary components, the input size is... ,in It is the scaling factor. For example, the scaling factor could be 2, where the minor components are downsampled by a factor of 2.

[0122] Based on the above explanation, the operation of the clipping layer depends on the output dimensions H and W and the depth of the clipping layer. Figure 9 The leftmost clipping layer has a depth of 0. If the input size of this clipping layer is greater than H or W in the horizontal or vertical dimension, respectively, then the output of this clipping layer must be equal to H or W (output size), and clipping needs to be performed in that dimension. The second clipping layer, counted from left to right, has a depth of 1. The output of the second clipping layer must be equal to... This means that if the input to the second clipping layer is greater than h1 or w1 in any dimension, clipping is applied in that dimension. In summary, the operation of the clipping layers is controlled by the output dimensions H and W. In one example, if both H and W are equal to 16, the clipping layer does not perform any clipping. On the other hand, if both H and W are equal to 17, all four clipping layers will perform clipping.

[0123] 1.9 Displacement Bitwise shift operators can use functions Let be the bitwise AND operator, where n is an integer. If n is greater than 0, it corresponds to the right shift operator (>>), which shifts the bits of the input to the right, and the left shift operator (<<), which shifts the bits to the left. In other words, The operation corresponds to: ,or ,or .

[0124] The output of a shift operation is an integer value. In some implementations, the floor() function can be added to the definition.

[0125] Floor(x) is equal to the largest integer less than or equal to x.

[0126] The " / / " operator, or integer division operator, is an operation that includes division and rounding the result towards zero. For example, 7 / 4 and -7 / -4 are rounded to 1, and -7 / 4 and 7 / -4 are rounded to -1.

[0127] or

[0128] Equation 3: Shift operators as alternative implementations of right or left shift.

[0129] The two's complement representation of x >> yx represents an arithmetic right shift of y bits. This function is defined only for non-negative integer values ​​of y. The bit shifted in by the right shift has the same MSB value as x before the shift operation.

[0130] x << y The two's complement integer representation of x is arithmetically left-shifted by y binary positions. This function is only defined for non-negative integer values of y. The bits shifted into the least significant bit (LSB) due to the left shift have a value equal to 0.

[0131] 1.10 Convolution Operation Convolution is a relatively simple operation in essence: You start with a kernel, which is just a small weight matrix. This kernel "slides" over the input data, performs element-wise multiplication with the part of the input it is currently on, and then sums the results into a single output pixel. In some cases, the convolution operation can include a "bias", which is added to the output of the element-wise multiplication operation.

[0132] The convolution operation can be described by the following mathematical formula. The output out1 can be obtained as:

[0133] where w1 is the multiplication factor, K1 is called the bias (additive term), and is the k-th input, and N is the kernel size in one direction, and P is the kernel size in the other direction. A convolutional layer can consist of convolution operations where more than one output can be generated. Other equivalent depictions of the convolution operation can be found below:

[0134]

[0135] In the above equations, "c" represents the channel number. It is equivalent to the output number, out[1,x,y] is one output, and out[2,x,y] is the second output. Where k is the input number, is one input, is the second input.

[0136] w1 or w describes the weights of the convolution operation.

[0137] 1.11 Leaky ReLU Activation Function The Leaky ReLU activation function is as Figure 10 shown. According to this function, if the input is positive, the output is equal to the input. If the input (y) is negative, the output is equal to a*y. a is typically (but not limited to) a value less than 1 and greater than 0. Since the multiplier a is less than 1, it can be implemented as a multiplication or division operation with a non-integer. The multiplier a can be referred to as the negative slope of the Leaky ReLU function.

[0138] 1.12 ReLU Activation Function The ReLU activation function is as Figure 11As shown in the diagram. According to this function, if the input is positive, the output is equal to the input. If the input (y) is negative, the output is 0.

[0139] 1.13 JPEG AI Image Codec Standard Existing designs utilize some of the aforementioned neural network-based image encoding and decoding methods. Some features are described or summarized below. The numbers in parentheses are reference section numbers.

[0140] 1. (3.5.20) Filler layer The fill layer is represented as ,in These are the height and width of the tensor, which is the input to the analytic transformation. It is the stride of the convolution. This is the depth of the convolutional layer in the deep learning encoder. The fill layer receives data at a size of [size missing]. The tensor and output size is The tensor, in which By default, padding is performed via copying. Different padding models can be specified (e.g., padding with zeros).

[0141] 2. (3.5.22) Cutting layer Clipping layers are represented as ,in These are the height and width of the tensor, which is the output of the composition transformation. It is the stride of the transposed convolution. This refers to the depth of the convolutional layers in the deep learnable reconstruction process. The output size of the clipped layer is... The tensor, in which Defined in Table 2. Note that... Pruning is performed by discarding redundant elements.

[0142] Table 2 Tensor size parameters used for primary component decoding and secondary component decoding.

[0143]

[0144] Table 3 Supported combinations of output image formats and scaling factors

[0145] 3. (8.4) Potential space sheet Latent spatial tile decoding can be enabled independently for each component via a flag in the image header (Section 9.3). If tile_enable_Luma or tile_enable_Chroma is true, then... Figure 12 The potential slice process shown decodes the corresponding components.

[0146] The number and position of the slices are determined by the slice size value transmitted via signal in the image header (Section 9.3). (Equal to the tile_size_Luma of the primary component and the tile_size_Chroma of the secondary component) and sheet overlap (The tile_overlap_Luma of the primary component and the tile_overlap_Chroma of the secondary component) are determined. Furthermore, the ratio between the signal domain size and the latent tensor size is determined. It is known that, for the principal components, it is For secondary components it is .

[0147] The number of vertical slices is The number of horizontal slices is ,in It is the height and width of the output color plane.

[0148] It is the height and width of the output color plane.

[0149] for and

[0150] - Define slice coordinates and dimensions o – Vertical dimension patch start in the signal domain; o –Initiation of vertical dimension patches in latent space; o – Vertical dimension slice size in the signal domain; o – Vertical dimension patch size in the potential space; o – Start of horizontal dimension slices in the signal domain; o – The beginning of a horizontal dimension patch in the latent space; o – Horizontal slice size in the signal domain; o – Horizontal dimension piece size in the latent space.

[0151] - Piece potential space tensor o For , and

[0152]

[0153] o For , and

[0154]

[0155] - Composite transformation for a single piece The composition transformation described in Section 8.3, where the size is and of , As input, and with size and of As output.

[0156] - Merge the reconstructed parts of the tensor into one o - Overlap used on the left edge of the piece; o - Overlap used on the top boundary of the slice; o - Overlap used on the right edge of the piece; o - Overlap used on the lower boundary of the slice; o The vertical dimension of the slice size in the signal domain does not have slice overlap; o The horizontal dimension of the slice size in the signal domain does not have slice overlap; o For , and , . The slice size position and overlap signaling in the image header are described in Appendix C.

[0157] 4. (9.2) Stream Layout The bitstream structure is as follows Figure 13 As shown. The bitstream consists of six parts with byte boundaries, which are: 1. SOC - Start of Stream Marker; 2. PIH (Picture Header Marker), followed by the picture header (Section 9.3); 3. TOH (tool head marker), followed by tool information (Section 9.3.1.1); 4. SOQ (Start of Quality Graph), followed by the bitstream; 5. SOZ (Start of Z-stream marker), followed by the bitstream of the super-prior information tensor z, including ( Figure 14 In , Figure 14 (This illustrates the general JPEG AI decoder structure) and ( Figure 14 In ); 6. SORp (start of the residual stream marked by the major component), followed by the bitstream of the major component residual, which includes... ( Figure 14 In ); 7. SORs (start of residual stream marked by minor components), followed by the bitstream of minor component residuals, which includes... ( Figure 14 In ); 8. EOC - End of stream marker.

[0158] The overall grammatical structure of an image is as follows:

[0159] Each bitstream begins with a 16-bit marker. All markers used in this specification are as follows:

[0160] 5. (9.3.1) Grammar Table

[0161] 6. (9.3.1.1) Model head

[0162]

[0163] 7. (9.3.2) Image header semantics The following service information is transmitted via signal: picture_header_size is the number of bytes in the image header, excluding the first two bytes of the marker; Adding 64 to img_width specifies the width of the input image (from 64 to 65600). img_height plus 64 specifies the height of the input image (from 64 to 65600); picture_format is the data format of the output image (YUV420 = 0, YUV444 = 1, sRGB = 2, YUV422 = 3). bit_depth is the bit depth of the output image ("0" corresponds to 8, "1" corresponds to 10); bit_c_ver is a bit value defined as c_ver = 1 + bit_c_ver. c_ver controls the internal downsampling mode of the secondary components in the vertical direction, as defined in Tables 3 and 2. If bit_c_ver does not exist in the bitstream, then c_ver equals 2.

[0164] bit_c_hor is a bit value defined as c_hor = 1 + bit_c_hor. c_ver controls the internal downsampling mode of the secondary components in the vertical direction, as defined in Tables 3 and 2. If bit_c_hor is not present in the bitstream, then c_hor equals 2.

[0165] `bit_s_ver` is a bit value that defines `s_ver`, which is used to align the encoding / decoding downsampling modes of secondary components in the vertical direction and the downsampling modes in the output image format, as defined in Table 3. The use of `s_ver` is described in Section 7.6. If `bit_c_ver` is not present in the bitstream, then `s_ver` equals 1.

[0166] `bit_s_hor` is a bit value defined as `c_hor = 1 + bit_c_hor`. `c_ver` controls the internal downsampling mode of the secondary components in the vertical direction, as defined in Table 3. The use of `s_ver` is described in Section 7.6. If `bit_c_ver` is not present in the bitstream, then `s_hor` equals 1.

[0167] scale_comp[2] is the size of the primary and secondary components of the encoded and decoded image. and The vertical and horizontal ratios between them; if they do not exist (res_changer_enable = false), then and The allowed values ​​are shown in Table 3.

[0168] independent_beta_uv is a bit rate control parameter that indicates the primary and secondary components. Whether the signs are the same (false / true).

[0169] beta_displacement_log_y – A parameter indicating the ratio between the bit rate control parameter beta chosen by the encoder for the primary component and the bit rate control parameter beta used during model training.

[0170] betaDisplacementLogY = beta_displacement_log_y – 2 11 beta_displacement_log_uv – A parameter indicating the ratio between the bit rate control parameter beta chosen by the encoder for the minor components and the bit rate control parameter beta used during model training.

[0171] betaDisplacementLogUV = beta_displacement_log_uv – 2 11 model_id is an identifier for a pre-stored checkpoint with model weights, where model_id = 0, 1, 2, 3, or 4.

[0172] opIdx is an identifier for the operation point; 0 means "basic" and 1 means "high" operation point.

[0173] tile_enable_Luma and tile_enable_Chroma are enable flags for tiles used for primary and secondary components.

[0174] tile_size_Luma and tile_size_Chroma are the tile sizes for the primary and secondary components.

[0175] tile_overlap_Luma and tile_overlap_Luma are the dimensions of the overlapping area of ​​tiles for the primary and secondary components.

[0176] `cube_group_flag` is a 1-bit unsigned integer. `cube_group_flag=0` indicates that no `cube_flag` is signaled in a group, and the cube flag of a group is set to 1. `cube_group_flag=1` indicates that the cube flag of a group is signaled.

[0177] cube_luma_flag is a size of A 1D array containing cube flags for the principal components. A 1 indicates that a skip mode is applied to a cube of the residual tensor of the principal components. 0 indicates that a cube of the residual tensor of the principal components disables skip mode. .

[0178] cube_chroma_flag is a value of (( A 1D array containing cube flags for minor components. A 1 indicates that a skip mode is applied to a cube of the residual tensor of the minor component. 0 indicates that a cube of the residual tensor for minor components disables skip mode. .

[0179] color_transform_enable is the enable flag for the color transformation module.

[0180] color_transform_matrix[i][j] is the color transformation matrix. If it does not exist (color_transform_enable is false), the default ITU-R BT.709 color transformation is used (see Section 7.8).

[0181] `color_transform_offset[i]` is the offset used for color transformation. If it does not exist (`color_transform_enable` is false), the default ITU-R BT.709 color transformation is used (see Section 7.8).

[0182] 8. (9.5.1.1) Syntax table of the super-prior information tensor

[0183] 9. (9.5.2) Quality Map Information Decoder The input to this process is – The bitstream of quality_map; – Indicators of the Sigma table decoded in the image header: o Quality_map_entropy_index_y for the primary component.

[0184] o Quality map_entropy_index_uv for secondary components.

[0185] The output of this process is – quality_map_delta_Y is a size of , An array containing information for deriving the scaling factor for the residual tensor of the principal components (Section 12.2).

[0186] – quality_map_delta_UV is a size of , , An array containing information for deriving the scaling factor for the residual tensor of minor components (Section 12.2).

[0187] Here, size , , and , , Defined in Table 2.

[0188] q_primary_sigma_Idx is a value of 1000 sigma_idx. A 1D array, where each element of the array is equal to q_sigma_Idx[quality_map_entropy_index_y], where q_sigma_Idx[k] is a 1D array based on the table. 4 It was derived.

[0189] q_secondary_sigma_Idx is a value of size 10 ... A 1D array, where each element in the array is equal to q_sigma_Idx[quality_map_entropy_index_uv].

[0190] q_primary is a size of The 1D array, and the quality_map_delta_Y is determined by the following formula: – quality_map_delta_Y[i, j] = q_primary[ For j = 0... And i = 0.. .

[0191] q_secondary is a size of The 1D array, and the quality_map_delta_UV is determined by the following formula. – quality_map_delta_UV[i, j] = q_secondary[ For j = 0... And i = 0.. .

[0192] Table 4 determines the q_sigma_Idx used for incremental decoding of the quality map.

[0193] 10. (9.5.2.1) Syntax table of quality graph information tensor

[0194] 11. (9.5.3) Main residual stream decoder The input to this process is – Primary component residual bitstream; – For the main components generated by the skip mask (Section 13.3.2), the mask_skip[ , h 4Y , w 4Y ]; – sigma_Idx_primary – Size is The tensor is the output of the sigma quantization of the major components (Section 10.8).

[0195] The output of this process is – r_primary – A one-dimensional array of residual tensor elements It is the input to the decoder skipping process for the main components (Section 13.3.4).

[0196] Here, size , , Defined in Table 2. The one-dimensional array sigma_Idx_primary_1D is determined through the following steps: Set k = 0 for , for , for

[0197] • If mask_skip[c,i,j] equals True, then sigma_Idx_primary_1D[k] is set to equal sigma_Idx_primary[c,i,j], and k is incremented by 1.

[0198] 12. (9.5.3.1) Syntax table of major residual tensors

[0199] 13. (9.5.4) Secondary residual stream decoder The input to this process is – For bitstreams containing minor component residuals; – For minor components, the mask_skip generated by the skip mask (Section 13.3.2)[ , h 4UV , w 4UV ]; – sigma_Idx_secondary – Size is The tensor is the output of the sigma quantization for minor components (Section 10.8).

[0200] The output of this process is – r_secondary – A one-dimensional array of residual tensor elements It is the input to the decoder skipping process for minor components (Section 13.3.4).

[0201] Here, size Defined in Table 2. The one-dimensional array sigma_Idx_secondary_1D is determined through the following steps: Set k = 0 for , for , for

[0202] • If mask_skip[c,i,j] equals True, then sigma_Idx_secondary_1D[k] is set to equal sigma_Idx_secondary[c,i,j], and k is incremented by 1.

[0203] 14. (9.5.4.1) Syntax table for minor residual tensors

[0204] 15. (10.3) Variance decoder using prior information Variance decoder networks utilizing prior information, such as Figure 15 As shown.

[0205] The input to the variance decoder that utilizes prior information is The reconstructed super-prior information tensor Size of input / output tensors , Operation point indicator , From the The model parameters of the variance decoder network that utilizes prior information are defined, and all multiplier parameters in these models are 8-bit integers.

[0206] The output of the variance decoder utilizing prior information is the standard deviation logarithmic tensor. The integer value is in the range of 0. < Within the range, , Defined in Section 0. The dimensions of these tensors for the primary and secondary components are listed in Table 2.

[0207] In the scalable variance decoder utilizing prior information, all operations are integers, all accumulators in the computation are in the 32-bit integer range, and multipliers in the model parameters are quantized to 8-bit integers. This guarantees the bit-accurate behavior of the neural network module. The variance decoder utilizing prior information uses a special type of operation to quantize convolutions. For each quantized convolution in the process, a set of clip values ​​is specified for each channel. and inverse scaling shift parameters ( All limiting values ​​in quantized convolution are... Inverse scaling shift Provided in Section 15.6.

[0208] Note The magnitude of the integer weights in the quantization model does not exceed The combination of shift and limiting values ​​ensures that the register of the quantized convolution is within 32 bits.

[0209] The order of operations in this module does not depend on the operation point indicator. However, the model parameters differ for basic operation points and advanced operation points.

[0210] A variance decoder that utilizes prior information to quantize the kernel size of convolutional layers stride First, then correct the linear unit. Then, the step size is... Quantization kernel size is Quantized convolutions are followed by rectified linear units. The next stride is... Quantized convolution (kernel size) Increase the number of channels to 16 This is followed by pixel shuffle (stride). ), which brings back the number of channels Crop layer (step) ,depth Ensure the size of the output tensor is . The process ends with an absolute value (abs) operation.

[0211] The weighted model is stored in the electronic appendix, in the format specified in Section 4.8, located at...<oper_point> / model_ <mid> / <comp> / hyper_scale_decoder.onnx, where<oper_point> Is it "base" or "high"? <mid>It is an integer from 0 to 5. <comp>It should be either "primary" or "secondary".

[0212] 16. (11.2) Decoder using prior information The learning-based decoder that utilizes prior information consists of two independent pipelines with the same neural network architecture, except for the input size and the number of channels.

[0213] The input to this process is The reconstructed latent tensor of prior information Depend on The defined model parameters of the decoder network utilizing prior information. Operation point indicator .

[0214] The output of this process is It is the explicit prediction (the part of the prediction tensor derived explicitly from information transmitted through the signal), where the channel size is .

[0215] The decoder process utilizing prior information is as follows: Figure 16 As shown.

[0216] A decoder utilizing prior information starts from a kernel size of Start with a stride 1 convolution, followed by a deconvolution (stride 2, kernel size...). ), trimmed layer (depth 6), and linear units with leakage correction. The next step is to have a core size of The first convolution is stride 1, followed by deconvolution (stride 2, kernel size...). ), a clipped layer (depth 5), and linear units with leakage correction. Up to this point, the number of channels remains constant (equal to the number of channels in the input tensor). The decoder utilizing prior information passes through a kernel with a size of To summarize, for higher-level operation points, this convolution increases the number of channels to 1. Furthermore, for basic operation points, the convolution maintains the number of channels unchanged, followed by a linear unit with leakage correction.

[0217] The weighted model is stored in the electronic appendix, in the format specified in Section 4.8, located at...<oper_point> / model_ <mid> / <comp> / hyper_decoder.onnx, where<oper_point> Is it "base" or "high"? <mid>It is an integer from 0 to 5. <comp>It should be either "primary" or "secondary".

[0218] 17. (11.3) Latent tensor reconstruction process The input to this process is Operation point indicator , The reconstructed residual tensor is the output of the SKIP modeling process (13.3.4). Explicit prediction, where the channel size is It utilizes the output of the decoder (11.2) that takes prior information. The output of this process is The potential tensor of the reconstruction.

[0219] The process is as follows.

[0220] if (Basic operation point), then the multi-stage context modeling process is bypassed. Explicit predictions are added to the residuals : if (Advanced operation point), then: Use a multi-stage contextual modeling process (Section 11.4).

[0221] 18. (11.4) Multi-stage context modeling The input to this process is The reconstructed residual tensor is the output of the SKIP modeling process (13.3.4). Explicit prediction, which utilizes the output of the decoder (11.2) with prior information, Eight MCMs k (k=0,…7) model, whose parameters are determined by… definition, The output of this process is The potential tensor of the reconstruction.

[0222] The process includes the following steps: Explicit prediction tensors with Padd layers (depth 5, stride 2) and down-shuffle (11.5.1) M=2 arrive Reconstructed prediction tensor Reconstruction residuals of fill layer (depth 5, stride 2) and downwash (11.5.1) M=1 arrive Reconstructed residual tensor Will Divided into four parts (Each section consists of 2C channels out of 8C channels). Will Divided into eight parts (Each section consists of C / 2 of the 4C channels). For k=0, …,3 oMCM( k ) process, in which As input • The previously reconstructed portion of the reshaped potential space tensor • –Isotopic portion of the reconstructed residual tensor • – Part of the reconstructed explicit prediction tensor Output • Generate .

[0223] exist Channel network procedures on tensors (11.5.4).

[0224] For k=3, ...,7 oMCM( k ) process, in which As input • The previously reconstructed portion of the reshaped potential space tensor • –Isotopic portion of the reconstructed residual tensor • – A portion of the explicitly predicted tensor that has been reshaped.

[0225] Output • Generate .

[0226] Up-Shuffle (11.5.2) M=2 and cut-out layer (depth 5, stride 2) arrive .

[0227] Upward wash (11.5.2) M=2 and cut layer (depth 5, stride 2) arrive (This will be used further in the LSBS procedure 13.4.2).

[0228] Multi-stage context modeling process such as Figure 17 As shown. This is a cyclical process: the later stages use previously obtained elements of the output tensor as input. Figure 17 In the diagram, the data streams used in the previous stage are marked with arrows.

[0229] The weighted model is stored in the electronic attachment, in the format specified in Section 4.8, located in high / model_ <mid> / <comp> / MCM / stage <n>.onnx, where <mid>It is an integer from 0 to 5. <comp>It should be either "primary" or "secondary". <n>It is the number of MCM stages from 0 to 7.

[0230] 19. (13.3.4) Decoder-side skip operation At the decoder, the input to the skip mode process is The code is generated after decoding by me-tANS (Section 9.5.2). 1D array s [num_res_elements], – mask_skip[ C , h 4, w 4]. The output of this process is residual tensor [ C , h 4, w 4]. The output of the lossless decoding process is a 1D array. Its size is equal to mask_skip [ C , h 4, w 4] The total number of "1"s in the tensor.

[0231] In other words, mask_skip [ C , h 4, w 4] Tensor Determination of Residual Tensor Which samples are included in the bitstream? All other samples of the quantized residual tensor are presumed to be equal to 0.

[0232] The residual skip mode process at the decoder is as follows: Dimension Set to equal sigma tensor The number, height, and width of the channels (Table 2).

[0233] tensor [ C , h 4, w 4] was initialized to all zeros.

[0234] counter .

[0235] Apply the following sequential steps: for , for

[0236] for

[0237] If mask_skip [ c,i,j If ] equals True, then = mask_skip [ c,i,j ] ∏( s [ k ]-2 16 -1 +1), and increase k by 1; otherwise, = 0.

[0238] 20. (14.1.1) Inter-frame channel related information filter slices and output selection process The ICCI filter uses the same method as described in Section 8.4 to process tile-based input. The tile size `icci_tile_size` and tile overlap `icci_tile_overlap` are the same for both the primary (luminance) and secondary (chrominance) components and are transmitted via signaling in the image header (Section 9.3). For each component, the ratio between the input domain size and the output tensor size... It is fixed at 1. For the sake of stable calculations in the model selection explained below, the minimum piece size is limited to 176.

[0239] A total of 10 ICCI filters were included in the design. For each color component... comp =0..2("0" –" y ", "1" – " u ", "2"- " v Each piece of ) tileID In this context, the filter is based on the icci_model_idx encoded and decoded in the image header (Section 9.3). comp ][ tileID The value is used to select. If icci_model_idx = 0, then ICCI processing is bypassed.

[0240] 21. (14.2) Adaptive upsampling This section details the primary component-guided adaptive upsampling process. This process utilizes information from the primary component to enhance the secondary components (color information plane) of the image.

[0241] If EFE_upsampler_enabled_flag is true, this process is enabled.

[0242] The input to this process is (Output of the synthesis transformation for the main components). (Output of the synthesis transformation for minor components). The output of this process is Enhanced secondary components It goes to the ICCI filter block (Section 14.1), and It goes to the nonlinear filter block (Section 14.3).

[0243] Figure 18 An example implementation of a primary component-guided adaptive upsampling filter is shown.

[0244] If EFE_upsampler_enabled_flag equals 0, then as described in Section 7.6, Upsampling is performed via bicubic interpolation. Otherwise, the following ordered steps are executed: Call the resolution procedure based on the resolution table in Section 9.4.1 to obtain , , , and .

[0245] The slice procedure, as described in Section 14.2.2, is invoked, with the parsed syntax elements as input and the Tile1 tensor as output.

[0246] The parameter update procedure specified in Section 14.2.1 is invoked, where and As input, and modified and As output.

[0247] vector [2] From [2, Subtract by channel.

[0248] and They were set to equal to and .

[0249] For x = 0.. (W+1) / 2-1, y = 0…(H+1) / 2-1, k = 0..1, i=0… -1, j=0… -1, perform the following operations:

[0250]

[0251]

[0252]

[0253] [k] + [k] and They were set to equal to and .

[0254] 22. (14.4.1) General LEF procedure The LEF process receives the main variance tensor. and the first reconstruction As input, the output of the LEF process is the edge-enhanced modification. .

[0255] Perform the following sequential steps: As described in Section 9.3, the image header parsing process is invoked to obtain... , Set to equal to , for , , The tensor is obtained as follows: , thr[3]=luma_edge_filter_thr_list_table [targetBppIdx]; intensity[4]=luma_edge_filter_intensity_list [targetBppIdx]; , , , , For i = 0... -1, j = 0. -1;

[0256] , For i = 0... -1, j = 0. -1, The following changes have been made: .

[0257] During this process, zero padding is used when the tensor index exceeds the tensor boundary.

[0258] Weight Tensor Defined as follows:

[0259] Defined as follows:

[0260] as follows:

[0261] Defined as follows:

[0262] 3. Problem The existing design has the following limitations: 1) For example Figure 19 As shown, the Q stream, Z stream, and R stream in the bitstream y Flow and R UV The encoded bits of the stream follow a sequence. Therefore, even if the synthesis transform has a slice-based mechanism, the entire bitstream must be decoded to generate a region-decoded image, which introduces additional decoding complexity and potential inconvenience in practical applications.

[0263] 2) Although JPEG AI supports segmenting latent values ​​into multiple segments during the synthesis transform, these segments must be decoded sequentially because each subsequent decoded region uses reference pixels from the previously decoded region. Decoding only a specified region is not supported. Use cases for decoding only a specific region include omnidirectional video-based applications, multi-page document sharing on social media, etc.

[0264] 3) Currently, an image can be segmented into multiple regions, and the correct decoding of each subset of a region can be enabled independently of other regions by making patches subsets of the regions and by patch alignment between luminance and chrominance. However, there is a lack of advanced instructions, such as access to this region in the image header.

[0265] 4) Currently, it is possible to segment an image into multiple regions and enable correct decoding of a subset of each region independently of the others. However, it is not possible to enable correct decoding of the entire region.

[0266] 4. Detailed Solution To address the aforementioned issues, methods outlined below are disclosed. The solutions should be considered as examples for interpreting general concepts, rather than interpreted in a narrow sense. Furthermore, these solutions can be applied individually or combined in any way.

[0267] 1) It is proposed that subsets of samples in an image (such as sub-images, patches, strips, or regions) can be reconstructed independently of samples outside the subset in the image using a NN-based decoder (such as the JPEG-AI decoder).

[0268] 2) To achieve item 1), the following methods can be applied: a. The bitstream was constructed as region-based, replacing the image-based approach.

[0269] b. An image can be divided into multiple regions, either vertically, horizontally, or both. Figure 20 Possible partitioning types are shown.

[0270] i. The number of horizontal (row) regions and vertical (column) regions can be indicated using two flags / signals and transmitted via signals in the bitstream.

[0271] ii. For each partitioned region, a signal / flag can be used to indicate whether the current region is an independently decoded region, that is, the region can be decoded independently of other regions.

[0272] iii. Given the dimensions of an input image, algorithms can be designed to derive the region sizes and the specific start and end positions of each region. Additional constraints can be applied to the algorithm, such as a minimum or maximum allowed region size.

[0273] a. Alternatively, size and / or location information may be explicitly indicated in the bitstream.

[0274] c. The start code / signal can be transmitted into the bitstream via a signal to indicate the start of a region.

[0275] 3) To mitigate regional boundary artifacts, such as Figure 21 As shown, overlap is applied to the dependent region. The overlap mechanism can be used for advanced prior information tensors, residual latent values, and synthetic transformations.

[0276] a. In one example, the number of overlapping pixels can be [number].

[0277] i. Entropy decoding: 1. There is no overlap around the area.

[0278] ii. Decoders utilizing prior information and variance decoders utilizing prior information: 1. There is one sample point overlap on each side of the input.

[0279] 2. Discard four samples on each side of the output.

[0280] iii. MCM: - There are eight overlapping samples on each side of the input.

[0281] - Discard eight samples on each side of the output.

[0282] iv. Composition Transformation: - There are two overlapping samples on each side of the input.

[0283] - Discard 32 or 16 samples on each side of the output.

[0284] 4) As an alternative to item 3), such as Figure 22 As shown, offset regions can be used. These tiles are offset versions of the original tiles used in JPEG AI software for synthesizing transforms. This structure enables pipelined processing, making the decoder faster.

[0285] General items 5) Additional operations can be applied to the proposed method or applied together with the proposed method.

[0286] a) The syntax elements disclosed above (also known as indicators) can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. They can be signed or unsigned.

[0287] b) If a codec tool or codec method is deemed unsuitable or unusable, it means that the syntax elements of the codec tool or codec method may not be transmitted via signal and are implicitly determined to be unused.

[0288] c) The grammatical elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded using a bypass method.

[0289] d) The above-disclosed grammatical elements can be transmitted conditionally via signals.

[0290] e) The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level.

[0291] f) Whether and / or how the methods disclosed above can be applied to signal transmission at the block level / sequence level / picture group level / picture level / strip level / piece group level.

[0292] g) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as color format, color components, and stripe / picture type.

[0293] 6) The proposed method can be applied to other image / video compression solutions involving NN-based encoding and decoding tools.

[0294] 5. Examples The following are some example implementations of the detailed solutions outlined in Section 4 of the previous article.

[0295] 5.1. Example 1 In one example, such as Figure 13 As shown, region-based decoding is implemented using an overlapping region mechanism. A high-level framework is as follows: Figure 23 As shown, the image is divided into multiple regions horizontally, vertically, or both. Each region in the bitstream may have a flag / code associated with it, indicating whether that region is an independently decoded region. To ensure correct decoding of independent regions, overlapping regions are applied only to dependent regions. The number of overlapping pixels can vary for advanced prior information latent values, residual latent values, MCM, and synthetic transforms.

[0296] 5.2. Example 2 Figure 24 The indicator shows that the markers are used to control the region division mechanism, overlapping regions (such as...) Figure 21 (as shown) or the offset area (such as) Figure 22 (As shown).

[0297] In one example, during decoding, both offset region mechanisms and overlapping region mechanisms can exist. Furthermore, flags / code transmitted via signals are used to control the selection of the region partitioning mechanism. An example can be implemented as follows.

[0298] • Composite transformations are implemented using overlapping regions.

[0299] • The entropy decoding part is implemented using overlapping regions.

[0300] Alternatively, the entropy decoding part is implemented using the offset region.

[0301] • Switch control is used to select region segmentation.

[0302] According to some example embodiments of this disclosure: • At least two different grids are used for encoding or decoding the image.

[0303] Based on the decision, either the first grid or the second grid is used.

[0304] The decision can be based on the indications obtained from the bit stream.

[0305] • Instructions can be signs.

[0306] • The indicator can be a grade indicator.

[0307] • Indicators can be derived from image dimensions or model indexes.

[0308] o The first and second grids can be as follows Figure 25 As shown: The first grid and the second grid can include different sets of sample points.

[0309] The size of the blocks in grid 1 and grid 2 can be different.

[0310] Grid 1 and Grid 2 can each contain the same number of blocks.

[0311] At least one block in grid 1 and grid 2 can have different sizes.

[0312] Grids can have overlapping areas. Figure 25 In the original description, grid 1 and grid 2 are depicted as having no overlapping areas. However, according to some example embodiments of this disclosure, the blocks of the grid may have overlapping areas. In other words, block 1 and block 2 may have a small subset of samples shared by both.

[0313] o Grid 1 and Grid 2 can be used in the entropy decoding process.

[0314] The output of the entropy decoding process can correspond to each block of the grid.

[0315] If there are N blocks in the grid, the entropy decoding process can be applied N times. The output of the entropy decoding process can correspond to each block of the grid.

[0316] o Grid 1 and Grid 2 can be used in decoding or variance decoding or sample prediction (latent sample prediction, latent sample reconstruction or multi-level context) processes that utilize prior information.

[0317] The output of the process can correspond to each block of the grid.

[0318] The input to the process can correspond to each block of the grid.

[0319] If there are N blocks in the grid, the process can be applied N times. The output of the process can correspond to each block of the grid.

[0320] If there are N blocks in the grid, the process can be applied N times. The input to the process can correspond to each block of the grid.

[0321] o can have multiple sets with at least two grids.

[0322] For example, there can be two sets of grids, Grid 1 and Grid 2. The first set {Grid 1 and Grid 2} can be used to determine the input of the process (e.g., the process can be an entropy decoding process, a decoding process using prior information, a variance decoding process using prior information, a sample prediction, or a sample reconstruction process). The second set {Grid 1 and Grid 2} can be used to determine the output of the process.

[0323] In another example, the first set of {mesh1 and mesh2} can be used in the first process, and the second set of {mesh1 and mesh2} can be used in the second process.

[0324] The grid can be constructed as follows: The top left block of the second grid can be larger than the top left block of the first grid in both the vertical and horizontal dimensions.

[0325] The bottom right block of the second grid can be smaller than the top left block of the first grid in both the vertical and horizontal dimensions.

[0326] The top row of a block in the second grid can be larger in the vertical dimension than the top row of a block in the first grid.

[0327] The bottom row of a block in the second grid can be smaller in the vertical dimension than the bottom row of a block in the first grid.

[0328] The left column of a block in the second grid can be larger in the vertical dimension than the left column of a block in the first grid.

[0329] The right column of a block in the second grid can be smaller in the vertical dimension than the right column of a block in the first grid.

[0330] In one example implementation, the decoding (or encoding) operation can be divided into two groups.

[0331] The first set of processes may include at least one of the following: • Entropy decoding / encoding process, • The decoding / encoding process utilizing prior information, • Variance decoding / encoding process utilizing prior information • Sample point prediction or sample point reconstruction process.

[0332] The second set of processes may include synthetic / analytical transformations (e.g., performing an inverse transformation operation between latent samples and image samples).

[0333] According to some example embodiments, • At least two grids (grid 1 and grid 2) can be used in the processing of the first group of processes. In other words, the first grid is used based on a determined result, or the second grid is used by a process belonging to the first group of processes.

[0334] As shown in the example above, it can be determined that an action can be taken based on an instruction obtained from the bitstream.

[0335] • Only one grid is used in the processing of the second set of procedures. In other words, no determination is performed, and only one grid is used in the second set of procedures.

[0336] The decoding / encoding operation includes a first set of procedures and a second set of procedures.

[0337] The first grid can be used in parallel processing. In other words, if the first grid is used in a process, that process can be parallelized. Specifically, samples from the first block of the first grid can be obtained in parallel with samples from the second block.

[0338] The second grid can be used in sequential processing. In other words, if the second grid is used in a process, that process cannot be parallelized. Specifically, samples from the second block of the second grid cannot be obtained in parallel with samples from the first block.

[0339] The second grid can be used in pipelined processing or pipelined systems.

[0340] The following describes further details of embodiments of this disclosure related to neural network-based visual data encoding and decoding. As used herein, the term "visual data" can refer to images, pictures in videos, or any other visual data suitable for encoding and decoding. To address the above-mentioned problems and some other issues not mentioned, a visual data processing solution is disclosed as described below. Embodiments of this disclosure should be considered as examples of explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.

[0341] Figure 26 A flowchart of a method 2600 for visual data processing according to some embodiments of the present disclosure is shown. At 2602, a conversion between visual data and a bitstream of visual data is performed using a neural network (NN)-based model. For example, the bitstream may include a sequence of bits. Additionally, the bitstream may also include associated code used as a marker. As used herein, the bitstream may also be referred to as a bitstream.

[0342] In some embodiments, the conversion may include encoding the visual data into a bitstream. Additionally or alternatively, the conversion may include decoding the visual data from the bitstream. This is by way of example, not limitation. Figure 6 The decoding model shown can be used to decode visual data from a bitstream.

[0343] As used herein, a neural network-based model can be a model based on neural network techniques. For example, a neural network-based model can specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module can include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the neural network-based model described herein are illustrative only and should not be construed as limiting this disclosure in any way.

[0344] Additionally, the bitstream includes a first indication of the number of vertical partitions of the residual sample set. Alternatively or additionally, the bitstream includes a second indication of the number of horizontal partitions of the residual sample set. For example, the residual sample set may be a residual tensor, and each of the plurality of residual sample subsets may be a residual subtensor. As used herein, in a vertical partition, a tensor of size W × H may be partitioned into subtensors of size (W / N) × H, where W represents the width of the tensor, H represents the height of the tensor, and N represents the number of vertical partitions. In a horizontal partition, a tensor of size W × H may be partitioned into subtensors of size W × (H / N), where W represents the width of the tensor, H represents the height of the tensor, and N represents the number of horizontal partitions.

[0345] In some example embodiments, the value of the first indicator may be equal to the number of vertical divisions of the residual sample set minus a predetermined number, such as 1. Furthermore, based on the determination that the first indicator does not exist in the bitstream, the value of the first indicator may be presumed to be equal to a predetermined value, such as 0. Similarly, the value of the second indicator may be equal to the number of horizontal divisions of the residual sample set minus a predetermined number, such as 1. Furthermore, based on the determination that the second indicator does not exist in the bitstream, the value of the second indicator may be presumed to be equal to a predetermined value, such as 0.

[0346] In some example embodiments, at least one of the first or second indications may be included in the image header syntax structure in the bitstream.

[0347] In light of the above, the bitstream includes a first indicator indicating the number of vertical divisions of the residual sample set and / or a second indicator indicating the number of horizontal divisions of the residual sample set. Compared to traditional solutions lacking both indicators, the proposed method better supports independent encoding and decoding of different regions, thereby improving encoding and decoding efficiency.

[0348] In some embodiments, the size of a subset of residual samples within a plurality of residual sample subsets can be determined based on the size of the visual data, the number of vertical divisions of the residual sample set, and / or the number of horizontal divisions of the residual sample set. In an example embodiment, the vertical size of the residual sample subset can be determined based on the height of the visual data, the number of horizontal divisions of the residual sample set, a third predetermined value (such as 64, 128, etc.), etc. For example, the vertical size can be equal to the product of the third predetermined value and a value determined based on the height of the visual data and the number of horizontal divisions of the residual sample set. By way of example, and not limitation, the vertical size can be determined as follows: VerSize = floor(floor((img_height + (PA-1)) / PA) / NumHorSplits)*PA, Where VerSize represents the vertical dimension of the residual sample subset, img_height represents the height of the visible data, PA represents the third predefined value, NumHorSplits represents the number of horizontal partitions of the residual sample set, and floor() represents the floor function. For example, the result of floor(x) can be the largest integer less than or equal to x.

[0349] Additionally or alternatively, the horizontal dimension of the residual sample subset can be determined based on the width of the visual data, the number of vertical divisions of the residual sample subset, a fourth predetermined value (such as 64, 128, etc.), etc. For example, the horizontal dimension can be equal to the product of the fourth predetermined value and a value determined based on the width of the visual data and the number of vertical divisions of the residual sample subset. By way of example, and not limitation, the horizontal dimension can be determined as follows: HorSize = floor(floor((img_width + (PB-1)) / PB) / NumVerSplits)*PB, Where HorSize represents the horizontal size of the residual sample subset, img_width represents the width of the visible data, PB represents the fourth predefined value, NumVerSplits represents the number of vertical partitions of the residual sample set, and floor() represents the floor function.

[0350] Furthermore, the size of the residual sample subsets within the multiple residual sample subsets can be constrained using the maximum and / or minimum permissible size. Additionally or alternatively, the bitstream may also include at least one indication indicating the size or location information of the residual sample subsets within the multiple residual sample subsets.

[0351] In some embodiments, the visual data can be segmented into multiple regions, and each subset of residual samples corresponds to one of the multiple regions. For example, each subset of residual samples may correspond to a set of samples covering a region of the visual data. By way of example and not limitation, the shape of the region may be rectangular, square, triangular, etc.

[0352] In some embodiments, the bitstream may also include a third indication indicating whether the residual samples of each of the plurality of regions are in a substream corresponding to that region. By way of example, a third indication equal to a first value (such as 1) may indicate that either the primary or secondary residual data of each region is in the substream. That is, the residual data of different regions are encapsulated in different substreams. In this case, regions can be encoded and decoded independently of other regions. A third indication equal to a second value (such as 0) may indicate that there is only one substream for the primary residual data and only one substream for the secondary residual data. That is, the residual data of components from different regions are encapsulated in the same substream.

[0353] In light of the above, the bitstream includes a third indicator that indicates whether the residual samples of each of the multiple regions are in the substream corresponding to that region. Compared to traditional solutions lacking this indicator, the proposed method can better support independent encoding and decoding of different regions, thereby improving encoding and decoding efficiency.

[0354] In some embodiments, the residual samples of each of the multiple regions can be in the substream corresponding to that region of the bitstream. That is, the bitstream is structured based on regions. In this way, the residual samples of a specific region of the multiple regions can be decoded independently of other regions. In this case, the samples of each of the multiple regions can be reconstructed independently of the samples of the remaining regions of the multiple regions.

[0355] In some embodiments, a substream may begin with an indication used to indicate a marker. For example, PIH (Picture Header Marker) may be followed by a picture header. TOH (Tool Header Marker) may be followed by tool information. SOQ (Quality Graph Start Marker) may be followed by quality graph information. SOZ (Z-Stream Start Marker) may be followed by the bitstream of the super-prior information tensor z. SORp (Major Component Marker Residual Stream Start) may be followed by the major component residual. SORs (Minor Component Marker Residual Stream Start) may be followed by the minor component residual.

[0356] In some embodiments, if a third indication indicates that there is only one substream for the primary residual samples of multiple regions and only one substream for the secondary residual samples of multiple regions, then for processing performed using the first processing module in the NN-based model, the multiple regions overlap. By way of example and not limitation, the first processing module may include a decoding module utilizing prior information, a variance decoding module utilizing prior information, a multi-level context modeling (MCM) module, a synthesis transformation module, etc.

[0357] In some embodiments, at least one sample may be discarded from the output of the first processing module. For example, during the process of merging samples from a set of regions into the output tensor of the first processing module, several samples from each region may be omitted. By way of example, nine regions (each region being 14×14 in size) can be merged into a 30×30 output tensor. In this case, for each region, the left 2 samples, the right 2 samples, the top 2 samples, and the bottom 2 samples can be discarded. It should be understood that the specific values ​​listed herein are intended to be illustrative and not to limit the scope of this disclosure.

[0358] In some embodiments, the amount of overlap between two regions in a plurality of regions may depend on a first processing module. For example, the bitstream may also include at least one of the following: an indication of the amount of overlap used in a decoding module utilizing prior information, or an indication of the amount of overlap used in an MCM module. In some embodiments, the amount of overlap may be different for different processing modules.

[0359] In some embodiments, multiple regions can be offset. Examples are shown in... Figure 22 The diagram is shown. In some embodiments, for processing regions, the tensor used as input to the second processing module may be center-aligned with the tensor used as input to the third processing module. Alternatively, the tensor used as input to the second processing module may be offset relative to the tensor used as input to the third processing module. In one example embodiment, processing of regions utilizing the synthetic transform may be performed using center-aligned tensors, and processing of regions utilizing another processing module may be performed using unaligned tensors.

[0360] In another example embodiment, processing of the region utilizing the synthetic transform can be performed using a center-aligned tensor, and processing of the region utilizing another processing module can be performed using either a center-aligned tensor or an unaligned tensor. Additionally, the bitstream may include an indication indicating whether processing of the region utilizing another processing module is performed using a center-aligned tensor or an unaligned tensor. For example, the other processing module may include at least one of the following: a decoding module utilizing prior information, a variance decoding module utilizing prior information, an entropy decoder, a sample prediction module, an MCM module, or a latent sample reconstruction module.

[0361] In some embodiments, the bitstream may also include an indication that the multiple regions overlap or are offset. In one example embodiment, the synthesis transform is implemented using overlapping regions, and the entropy decoding portion is also implemented using overlapping regions. Alternatively, the entropy decoding portion is implemented using offset regions. Additionally, a switch control is used to select region segmentation. For example, the switch control can be implemented as a flag, syntax element, etc.

[0362] In some embodiments, multiple grids can be used for transformation, each of which is configured to segment the visual data into multiple regions. Additionally, the bitstream may include an indicator that specifies which of the multiple grids is being used. This indicator can be derived from image size, model index, etc.

[0363] In some embodiments, multiple grids may have the same number of regions but different region sizes. Two example grids are shown in... Figure 25 The diagram is shown. In some embodiments, multiple grids may be used in at least one of the following: processing using an entropy decoding module, processing using a decoding module utilizing prior information, processing using a variance decoding module utilizing prior information, processing using a latent sample prediction module, processing using a latent sample reconstruction module, or processing using an MCM module.

[0364] In some embodiments, the multiple grids may include multiple sets of at least two grids. In one example embodiment, the multiple sets of at least two grids may include a first set of at least two grids and a second set of at least two grids, the first set of at least two grids being used to determine the input of a processing module of the NN-based model, and the second set of at least two grids being used to determine the output of the processing module. Alternatively, the first set of at least two grids and the second set of at least two grids may be used in different processing modules of the NN-based model.

[0365] In some embodiments, a neural network-based model may include a first set of processing modules and a second set of processing modules. Multiple grids may be available for the first set of processing modules, and only a single grid may be available for the second set of processing modules. In some embodiments, the multiple grids may include grids for parallel processing, grids for sequential processing, grids for pipelined processing, etc.

[0366] In some embodiments, any of the above indications (such as the first indication, the second indication, the third indication, etc.) can be binarized into one of the following: a flag, a fixed-length code, an exponential Golomb code, a unary code, a rounded unary code, or a rounded binary code.

[0367] In some embodiments, if an indication of an encoding / decoding tool is not present in the bitstream, the encoding / decoding tool may be presumed to be unused. In some embodiments, the indication may be encoded / decoded using at least one context model or by bypass encoding / decoding. In some embodiments, the indication may be conditionally transmitted via signaling.

[0368] In some embodiments, the indication may be given at one of the following levels: block level, sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, whether and / or how the method is applied may be indicated at one of the following levels: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0369] In some embodiments, whether and / or how a method is applied may depend on the encoded / decoded information.

[0370] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding quality.

[0371] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream generated by a method performed by a visual data processing apparatus for visual data. The method includes: performing a conversion between visual data and a bitstream using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into a plurality of subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0372] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. The method includes: performing a conversion between visual data and a bitstream using a neural network (NN)-based model; and storing the bitstream in a non-transitory computer-readable recording medium, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the residual sample set, or a second indication indicating the number of horizontal divisions of the residual sample set.

[0373] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0374] Item 1. A method for visual data processing, comprising: performing a conversion between visual data and a bitstream of the visual data using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into a plurality of subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0375] Item 2. The method according to Item 1, wherein the set of residual samples is a residual tensor, and each of the plurality of residual sample subsets is a residual subtensor.

[0376] Item 3. The method according to any one of items 1-2, wherein the size of the residual sample subset in the plurality of residual sample subsets is determined based on at least one of the following: the size of the visual data, the number of vertical divisions of the residual sample set, or the number of horizontal divisions of the residual sample set.

[0377] Item 4. The method according to any one of items 1-3, wherein the visual data is segmented into multiple regions, and each subset of residual samples in the multiple subsets corresponds to one of the multiple regions.

[0378] Item 5. The method according to Item 4, wherein the bitstream further includes a third indication indicating whether a residual sample of each of the plurality of regions is in a substream of the bitstream corresponding to the region.

[0379] Item 6. The method according to any one of Items 4-5, wherein the residual sample points of each of the plurality of regions are in the substream of the bitstream corresponding to that region.

[0380] Item 7. According to the method described in Item 6, the samples of each of the plurality of regions are reconstructed independently of the samples of the remaining regions of the plurality of regions.

[0381] Item 8. The method according to any one of Items 5-7, wherein the substream begins with an indication for indicating a marker.

[0382] Item 9. The method according to Item 5, wherein if the third indication indicates that there is only one substream for the primary residual samples of the plurality of regions and only one substream for the secondary residual samples of the plurality of regions, then the plurality of regions overlap for processing performed using the first processing module in the NN-based model.

[0383] Item 10. The method according to Item 9, wherein the first processing module comprises at least one of the following: a decoding module utilizing prior information, a variance decoding module utilizing prior information, a multi-level context modeling (MCM) module, or a synthesis transformation module.

[0384] Item 11. The method according to any one of items 9-10, wherein at least one sample is discarded from the output of the first processing module.

[0385] Item 12. The method according to any one of items 9-11, wherein the amount of overlap between two regions in the plurality of regions depends on the first processing module.

[0386] Item 13. The method according to any one of items 1-12, wherein the size of the residual sample subsets in the plurality of residual sample subsets is constrained by at least one of the maximum allowable size or the minimum allowable size.

[0387] Item 14. The method according to any one of items 1-13, wherein the bitstream further includes at least one indication indicating the size or location information of a subset of residual samples in the plurality of residual sample subsets.

[0388] Item 15. The method according to any one of items 4-12, wherein the plurality of regions are offset.

[0389] Item 16. The method according to any one of items 4-5, wherein the bitstream further includes an indication indicating whether the plurality of regions overlap or are offset.

[0390] Item 17. The method according to any one of items 1-16, wherein a plurality of grids are available for the transformation, each of the plurality of grids being configured to divide the visual data into a plurality of regions.

[0391] Item 18. The method according to Item 17, wherein the bitstream further includes an indication indicating one of the plurality of grids being used.

[0392] Item 19. The method according to any one of Items 17-18, wherein the plurality of grids have the same number of regions and different region sizes.

[0393] Item 20. The method according to any one of Items 17-19, wherein the plurality of grids are used in at least one of the following: processing using an entropy decoding module, processing using a decoding module utilizing prior information, processing using a variance decoding module utilizing prior information, processing using a latent sample prediction module, processing using a latent sample reconstruction module, or processing using an MCM module.

[0394] Item 21. The method according to any one of items 17-20, wherein the plurality of grids comprises a plurality of sets of at least two grids.

[0395] Item 22. The method according to Item 21, wherein the plurality of sets of at least two grids includes a first set of at least two grids and a second set of at least two grids, the first set of at least two grids being used to determine the input of the processing module of the NN-based model, and the second set of at least two grids being used to determine the output of the processing module, or the first set of at least two grids and the second set of at least two grids being used in different processing modules of the NN-based model.

[0396] Item 23. The method according to any one of items 1-22, wherein the NN-based model includes a first set of processing modules and a second set of processing modules, multiple grids are available for the first set of processing modules, and only a single grid is available for the second set of processing modules.

[0397] Item 24. The method according to any one of items 17-23, wherein the plurality of grids includes at least one of: a grid for parallel processing, a grid for sequential processing, or a grid for pipelined processing.

[0398] Item 25. The method according to any one of items 1-24, wherein the indication is binarized into one of: a flag, a fixed-length code, an exponential Golomb code, a unary code, a rounded unary code, or a rounded binary code.

[0399] Item 26. The method according to any one of items 1-25, wherein if an indication of an encoding / decoding tool is not present in the bitstream, the encoding / decoding tool is presumed to be unused.

[0400] Item 27. The method according to any one of items 1-26, wherein the instruction is to be encoded or decoded using at least one context model or to be bypassed and decoded.

[0401] Item 28. The method according to any one of items 1-27, wherein the indication is transmitted by signal in a conditional manner.

[0402] Item 29. The method according to any one of items 1-28, wherein the indication is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0403] Item 30. The method according to any one of items 1-29, wherein whether and / or how the method is applied is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0404] Item 31. The method according to any one of items 1-29, wherein whether and / or how the method is applied depends on the encoded / decoded information.

[0405] Item 32. The method according to any one of items 1-31, wherein the visual data is at least a portion of a picture or image of a video.

[0406] Item 33. The method according to any one of items 1-32, wherein the conversion includes encoding the visual data into the bitstream.

[0407] Item 34. The method according to any one of items 1-32, wherein the conversion includes decoding the visual data from the bitstream.

[0408] Item 35. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-34.

[0409] Item 36. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-34.

[0410] Item 37. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a visual data processing apparatus of visual data, wherein the method includes: performing a conversion between the visual data and the bitstream using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into a plurality of subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0411] Item 38. A method for storing a bitstream of visual data, comprising: performing a conversion between the visual data and the bitstream using a neural network (NN)-based model; and storing the bitstream in a non-transitory computer-readable recording medium, wherein a set of residual samples associated with the visual data is segmented into a plurality of subsets of residual samples, and the bitstream includes at least one of: a first indication indicating the number of vertical divisions of the set of residual samples, or a second indication indicating the number of horizontal divisions of the set of residual samples.

[0412] Example device Figure 27 A block diagram of a computing device 2700 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2700 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).

[0413] It should be understood that, Figure 27 The computing device 2700 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0414] like Figure 27 As shown, computing device 2700 includes general-purpose computing device 2700. Computing device 2700 may include at least one or more processors or processing units 2710, memory 2720, storage unit 2730, one or more communication units 2740, one or more input devices 2750, and one or more output devices 2760.

[0415] In some embodiments, the computing device 2700 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2700 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0416] Processing unit 2710 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 2720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 2700. Processing unit 2710 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0417] Computing device 2700 typically includes various computer storage media. Such media can be any media accessible by computing device 2700, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 2720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 2730 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or visual data and can be accessed within computing device 2700.

[0418] The computing device 2700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 27 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to the bus (not shown) via one or more visual data media interfaces.

[0419] Communication unit 2740 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 2700 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 2700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0420] Input device 2750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 2760 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 2740, computing device 2700 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 2700 can also communicate with one or more devices that enable a user to interact with computing device 2700, or any device that enables computing device 2700 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).

[0421] In some embodiments, some or all components of computing device 2700 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider offers applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations within remote visual data centers. Cloud computing infrastructure can be provided through shared visual data centers, although they may appear as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from traditional servers or installed directly or otherwise on client devices.

[0422] The computing device 2700 can be used to implement visual data encoding / decoding in embodiments of this disclosure. The memory 2720 may include one or more visual data encoding / decoding modules 2725 having one or more program instructions. These modules can be accessed and executed by the processing unit 2710 to perform the functions of the various embodiments described herein.

[0423] In an example embodiment of performing visual data encoding, input device 2750 may receive visual data as input 2770 to be encoded. The visual data may be processed, for example, by visual data encoding / decoding module 2725 to generate an encoded bitstream. The encoded bitstream may be provided as output 2780 via output device 2760.

[0424] In an example embodiment of performing visual data decoding, input device 2750 may receive an encoded bitstream as input 2770. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 2725 to generate decoded visual data. The decoded visual data may be provided as output 2780 via output device 2760.

[0425] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.< / n> < / comp> < / mid> < / n> < / comp> < / mid> < / comp> < / mid> < / comp> < / mid> < / comp> < / mid> < / comp> < / mid>

Claims

1. A method for visual data processing, comprising: A conversion between visual data and its bitstream is performed using a neural network (NN)-based model, wherein a set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of the following: A first indication of the number of vertical divisions of the residual sample set, or A second indication indicating the number of horizontal divisions of the residual sample set.

2. The method according to claim 1, wherein the set of residual samples is a residual tensor, and each subset of the plurality of residual sample subsets is a residual subtensor.

3. The method according to any one of claims 1-2, wherein the size of the residual sample subset in the plurality of residual sample subsets is determined based on at least one of the following: The size of the visual data The number of vertical divisions of the residual sample set, or The number of horizontal divisions of the residual sample set.

4. The method according to any one of claims 1-3, wherein the visual data is segmented into a plurality of regions, and each subset of residual sample points in the plurality of residual sample point subsets corresponds to one of the plurality of regions.

5. The method of claim 4, wherein the bitstream further comprises a third indication indicating whether a residual sample of each of the plurality of regions is in a substream of the bitstream corresponding to the region.

6. The method according to any one of claims 4-5, wherein the residual sample points of each of the plurality of regions are in the substream of the bitstream corresponding to that region.

7. The method of claim 6, wherein samples from each of the plurality of regions are reconstructed independently of samples from the remaining regions of the plurality of regions.

8. The method according to any one of claims 5-7, wherein the substream begins with an indication for indicating a marker.

9. The method of claim 5, wherein if the third indication indicates that there is only one subflow for the primary residual samples of the plurality of regions and only one subflow for the secondary residual samples of the plurality of regions, then the plurality of regions overlap for processing performed using the first processing module in the NN-based model.

10. The method of claim 9, wherein the first processing module comprises at least one of the following: Utilizing a decoding module with prior information, Variance decoding module utilizing prior information, Multi-level Context Modeling (MCM) module, or Synthesis and transformation module.

11. The method according to any one of claims 9-10, wherein at least one sample is discarded from the output of the first processing module.

12. The method according to any one of claims 9-11, wherein the amount of overlap between two regions in the plurality of regions depends on the first processing module.

13. The method according to any one of claims 1-12, wherein the size of the residual sample subsets in the plurality of residual sample subsets is constrained by at least one of a maximum permissible size or a minimum permissible size.

14. The method according to any one of claims 1-13, wherein the bitstream further includes at least one indication indicating the size or location information of a subset of residual samples in the plurality of residual sample subsets.

15. The method according to any one of claims 4-12, wherein the plurality of regions are offset.

16. The method according to any one of claims 4-5, wherein the bitstream further includes an indication indicating whether the plurality of regions overlap or are offset.

17. The method according to any one of claims 1-16, wherein a plurality of grids are available for the transformation, each of the plurality of grids being configured to segment the visual data into a plurality of regions.

18. The method of claim 17, wherein the bitstream further includes an indication indicating one of the plurality of grids being used.

19. The method according to any one of claims 17-18, wherein the plurality of grids have the same number of regions and different region sizes.

20. The method according to any one of claims 17-19, wherein the plurality of grids are used in at least one of the following: Using the entropy decoding module for processing, Processing using a decoding module that utilizes prior information, Processing using a variance decoding module that leverages prior information, Processing using the latent sample prediction module, Processing using the latent sample point reconstruction module, or Processing using the MCM module.

21. The method according to any one of claims 17-20, wherein the plurality of grids comprises a plurality of sets of at least two grids.

22. The method of claim 21, wherein the plurality of sets of at least two grids comprises a first set of at least two grids and a second set of at least two grids, the first set of at least two grids being used to determine the input of the processing module of the NN-based model, and the second set of at least two grids being used to determine the output of the processing module, or The first group of at least two grids and the second group of at least two grids are used in different processing modules of the NN-based model.

23. The method according to any one of claims 1-22, wherein the NN-based model comprises a first set of processing modules and a second set of processing modules, a plurality of grids may be used in the first set of processing modules, and only a single grid may be used in the second set of processing modules.

24. The method according to any one of claims 17-23, wherein the plurality of grids comprises at least one of the following: Grid for parallel processing Grids for sequential processing, or Grids for pipeline processing.

25. The method according to any one of claims 1-24, wherein the indication is binarized into one of: a flag, a fixed-length code, an exponential Golomb code, a unary code, a rounded unary code, or a rounded binary code.

26. The method according to any one of claims 1-25, wherein if an indication of an encoding / decoding tool is not present in the bitstream, the encoding / decoding tool is presumed to be unused.

27. The method according to any one of claims 1-26, wherein the instruction is to be encoded or decoded using at least one context model or to be bypassed and encoded / decoded.

28. The method according to any one of claims 1-27, wherein the indication is conditionally transmitted via signal.

29. The method according to any one of claims 1-28, wherein the indication is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

30. The method according to any one of claims 1-29, wherein whether and / or how the method is applied is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

31. The method according to any one of claims 1-29, wherein whether and / or how the method is applied depends on the encoded / decoded information.

32. The method according to any one of claims 1-31, wherein the visual data is at least a portion of a picture or image of a video.

33. The method according to any one of claims 1-32, wherein the conversion comprises encoding the visual data into the bitstream.

34. The method according to any one of claims 1-32, wherein the conversion comprises decoding the visual data from the bitstream.

35. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-34.

36. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-34.

37. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a visual data processing apparatus for visual data, wherein the method includes: The conversion between the visual data and the bitstream is performed using a neural network (NN)-based model, wherein the set of residual samples associated with the visual data is segmented into multiple subsets of residual samples, and the bitstream includes at least one of the following: A first indication of the number of vertical divisions of the residual sample set, or A second indication indicating the number of horizontal divisions of the residual sample set.

38. A method for storing a stream of visual data, comprising: The conversion between the visual data and the bitstream is performed using a neural network (NN) based model; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The set of residual samples associated with the visual data is divided into multiple subsets of residual samples, and the bitstream includes at least one of the following: A first indication of the number of vertical divisions of the residual sample set, or A second indication indicating the number of horizontal divisions of the residual sample set.