Method and device for visual data processing and medium

By adjusting the component samples of visual data and using a neural network model for transformation, the problem of improving encoding and decoding quality in existing technologies has been solved, achieving higher quality reconstruction results.

CN120898421APending Publication Date: 2025-11-04DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480020181.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-29
Filing Date
2024-03-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing neural network-based image/video encoding and decoding technologies still have room for improvement in encoding and decoding quality, especially in fully utilizing cross-component information.

Method used

By adjusting the samples of the first component of the visual data and adjusting the second component based on these samples, a neural network model is used for conversion, thereby improving the encoding and decoding quality.

Benefits of technology

By utilizing cross-component information, the quality of the reconstructed visual data is improved, thereby enhancing the overall performance of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120898421A_ABST
    Figure CN120898421A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for visual data processing. A method for visual data processing is presented. The method comprises, for a conversion between visual data and one or more bitstreams of the visual data using a neural network (NN)-based model, obtaining a set of adjusted first sample points by adjusting first sample points of a first component of the visual data using a set of offsets, each adjusted first sample point in the set of adjusted first sample points corresponds to one offset in the set of offsets; adjusting a second sample point of a second component of the visual data based on the at least one adjusted first sample point, where the at least one adjusted first sample point is determined from the set of adjusted first sample points by comparing each adjusted first sample point in the set of adjusted first sample points to a threshold value; and performing a conversion based on the adjusted second sample point.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to visual data processing techniques, and more particularly, to neural network based visual data coding. BACKGROUND

[0002] In the past decade, deep learning has made rapid progress in various fields, especially in computer vision and image processing. Neural networks were originally invented through interdisciplinary research in neuroscience and mathematics. It shows strong ability in the context of nonlinear transformation and classification. In the past five years, neural network based image / video compression technology has made significant progress. It is reported that the latest neural network based image compression algorithm achieves rate-distortion (R-D) performance comparable to versatile video coding (VVC). With the continuous improvement of neural image compression performance, neural network based video compression has become an actively developing research field. However, the coding quality of neural network based image / video coding is generally expected to be further improved. SUMMARY

[0003] Embodiments of the present disclosure provide a solution for visual data processing.

[0004] In a first aspect, a method for visual data processing is presented. The method comprises: for a conversion between visual data and one or more bitstreams of the visual data utilizing a neural network (NN) based model, obtaining a set of adjusted first samples of a first component of the visual data by adjusting first samples of the first component with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting second samples of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold, and the second component is different from the first component; and performing the conversion based on the adjusted second samples.

[0005] According to the method of the first aspect of the present disclosure, the second component of the visual data is adjusted based on the first component. Compared with the traditional solution that the first component and the second component are processed independently, the proposed method can advantageously utilize cross-component information to improve the quality of the reconstructed visual data, whereby the coding quality can be improved.

[0006] In a second aspect, an apparatus for visual data processing is presented. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is presented. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is presented. The non-transitory computer-readable recording medium stores a bitstream of visual data, the bitstream of visual data being generated by a method performed by an apparatus for visual data processing. The method comprises: obtaining a set of adjusted first samples of a first component of the visual data by adjusting first samples of the first component with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting second samples of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold, and the second component is different from the first component; and generating the bitstream with a neural network (NN) based model based on the adjusted second samples.

[0009] In a fifth aspect, a method for storing a bitstream of visual data is presented. The method comprises: obtaining a set of adjusted first samples of a first component of the visual data by adjusting first samples of the first component with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting second samples of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold, and the second component is different from the first component; generating the bitstream with a neural network (NN) based model based on the adjusted second samples; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] This Summary is intended to introduce some of the concepts discussed below in the of the Invention section. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to the like elements throughout. In the example embodiments of the present disclosure, the same reference numbers in different drawings identify the same or similar components.

[0012] Figure 1A A block diagram of an example visual data coding system according to some embodiments of the present disclosure is shown.

[0013] Figure 1B is a schematic diagram illustrating an example transform coding scheme.

[0014] Figure 2 An example latent representation of an image is shown.

[0015] Figure 3 is a schematic diagram illustrating an example autoencoder implementing a hyperprior model.

[0016] Figure 4 is a schematic diagram illustrating an example combined model configured to jointly optimize a context model with a hyperprior and an autoencoder. The following table 1 shows the meaning of different symbols.

[0017] Table 1 - Symbol explanation

[0018]

[0019] Figure 5 An example encoding process is shown.

[0020] Figure 6 An example decoding process is shown.

[0021] Figure 7 An example decoding process according to the present disclosure is shown.

[0022] Figure 8 An example learning-based image codec architecture is shown.

[0023] Figure 9 An example synthetic transform for learning-based image coding is shown.

[0024] Figure 10 An example Leaky ReLU (LeakyReLU) activation function is shown.

[0025] Figure 11 An example ReLU activation function is shown.

[0026] Figure 12 is a flowchart of an example method of video processing.

[0027] Figure 13 is a flowchart of an example method of video processing.

[0028] Figure 14 is a flowchart of an example method of video processing.

[0029] Figure 15 is a flowchart of an example method of video processing.

[0030] Figure 16 An example neural network is shown.

[0031] Figure 17 An example neural network is shown.

[0032] Figure 18 An example implementation according to embodiments of the disclosure is shown.

[0033] Figure 19 A flowchart of a method for visual data processing according to embodiments of the disclosure is shown.

[0034] Figure 20 A block diagram of a computing device in which various embodiments of the disclosure can be implemented is shown.

[0035] In all of the drawings, like or similar reference numerals are generally employed to refer to like or similar elements throughout the several views. DETAILED DESCRIPTION

[0036] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is merely intended to illustrate and help the person skilled in the art to understand and implement the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.

[0037] In the following description and claims, unless otherwise defined, all scientific and technical terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0038] References in the present disclosure to “one embodiment”, “an embodiment”, “example embodiments”, etc. indicate that the embodiment described can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0039] It should be understood that although the terms “first” and “second” etc. can be used to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be called a second element, and similarly, a second element can be called a first element without departing from the scope of the example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0040] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes" and / or "including," when used herein, specify the presence of stated features, elements and / or components, but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0041] Example Environment

[0042] Figure 1A FIG. 1 is a block diagram illustrating an example visual data coding system 100 that can utilize the techniques of this disclosure. As illustrated, the visual data coding system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a visual data encoding device, and the destination device 120 can also be referred to as a visual data decoding device. In operation, the source device 110 can be configured to generate encoded visual data, and the destination device 120 can be configured to decode the encoded visual data generated by the source device 110. The source device 110 can include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.

[0043] The visual data source 112 can include a source such as a visual data capture device. Examples of the visual data capture device include, but are not limited to, an interface to receive visual data from a visual data provider, a computer graphics system to generate visual data, and / or a combination thereof.

[0044] The visual data can include one or more pictures or one or more images of a video. The visual data encoder 114 encodes the visual data from the visual data source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the visual data. The bitstream can include coded pictures and associated visual data. A coded picture is a coded representation of a picture. The associated visual data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded visual data can be transmitted directly to the destination device 120 via the I / O interface 116 over a network 130A. The encoded visual data can also be stored on a storage medium / server 130B for access by the destination device 120.

[0045] Destination device 120 can include I / O interface 126, visual data decoder 124, and display device 122. I / O interface 126 can include a receiver and / or modem. I / O interface 126 can obtain encoded visual data from source device 110 or storage medium / server 130B. Visual data decoder 124 can decode the encoded visual data. Display device 122 can display the decoded visual data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 configured to interface with an external display device.

[0046] Visual data encoder 114 and visual data decoder 124 can operate according to a visual data coding standard, such as a video coding standard or a still picture coding standard, and other existing and / or further standards.

[0047] Some example embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding only and is not to be construed as limiting the embodiments disclosed in that section to only the section. Also, while specific embodiments are described with reference to multi-function video coding or other specific visual data coders, the disclosed techniques are applicable to other coding techniques as well. Also, while some embodiments describe coding steps in detail, it is understood that corresponding decoding steps would be implemented by a decoder that undo the coding. Also, the term visual data processing encompasses visual data coding or compression, visual data decoding or decompression, and visual data transcoding, where visual data is represented from one compression format to another compression format or at a different compression bit rate.

[0048] 1. Preliminary Discussion

[0049] This document relates to a neural network based image and video compression method that employs modifying components of an image using an adaptive filtering layer. This includes determining whether a value of a sample of a first component is based on a value of a sample of a second component. This document also relates to a neural network based image and video compression method that employs modifying components of an image using an offset. Determining whether to add an offset value to a sample of a second component is based on a value of a sample of a first component.

[0050] 2. Further Discussion

[0051] Deep learning is developing in various fields, such as computer vision and image processing. Inspired by the successful application of deep learning techniques in computer vision, neural image / video compression techniques are being researched and applied to image / video compression techniques. Neural networks are designed based on interdisciplinary research in neuroscience and mathematics. Neural networks show strong capabilities in the context of nonlinear transformation and classification. Image compression algorithms based on example neural networks achieve R-D performance comparable to Versatile Video Coding (VVC), a video coding standard developed by the Joint Video Expert Team (JVET) with experts from the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). Video compression based on neural networks is an actively developing research field, leading to continuous improvement in the performance of neural image compression. However, due to the inherent difficulty of the problems solved by neural networks, video coding based on neural networks is still a largely unexplored discipline.

[0052] 2.1 Image / video compression

[0053] Image / video compression generally refers to a computational technique that compresses video images into binary codes for storage and transmission. Binary codes can or can not support lossless reconstruction of the original image / video. Coding without data loss is called lossless compression, and coding that allows targeted data loss is called lossy compression. Most coding systems use lossy compression because lossless reconstruction is not necessary in most cases. In general, the performance of an image / video compression algorithm is evaluated based on the resulting compression ratio and reconstruction quality. Compression ratio is directly related to the number of binary codes generated by compression, and the fewer binary codes, the better the compression. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video, and the higher the similarity, the better the reconstruction quality.

[0054] Image / video compression techniques can be divided into video coding methods and neural network-based video compression methods. Video coding schemes use transform-based solutions, in which statistical dependencies in latent variables such as discrete cosine transform (DCT) and wavelet coefficients are used to carefully hand-design entropy codes to model dependencies in the quantization domain. Neural network-based video compression can be divided into neural network-based coding tools and end-to-end neural network-based video compression. The former is embedded as a coding tool in existing video codecs and serves only as part of the framework, while the latter is a separate framework developed based on neural networks and does not depend on video codecs.

[0055] A series of video coding standards have been developed to accommodate the growing demand for visual content transmission. The International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC) has two expert groups, namely the Joint Photographic Experts Group (JPEG) and the Moving Picture Experts Group (MPEG). The International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) also has the Video Coding Experts Group (VCEG) for the standardization of image / video coding technologies. Influential video coding standards published by these organizations include Joint Photographic Experts Group (JPEG), JPEG 2000, H.262, H.264 / Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC). The Versatile Video Coding (VVC) standard is developed by the Joint Video Expert Team (JVET) consisting of MPEG and VCEG. VVC reduces the bit rate by an average of 50% compared to HEVC at the same visual quality.

[0056] Neural network-based image / video compression / coding is also under development. Example neural network coding network architectures are relatively shallow, and the performance of such networks is not satisfactory. Neural network-based methods benefit from the support of rich data and powerful computing resources, and thus are better utilized in various applications. Neural network-based image / video compression has shown promising improvements and has been proven to be feasible. However, the technology is far from mature, and many challenges should be addressed.

[0057] 2.2 Neural Networks

[0058] Neural networks, also known as artificial neural networks (ANN), are computational models used in machine learning techniques. Neural networks are typically composed of multiple processing layers, and each layer is composed of multiple simple but non-linear basic computational units. One benefit of such deep networks is the ability to process data with multiple levels of abstraction and transform the data into different kinds of representations. The representations created by neural networks are not manually designed. Instead, using a general-purpose machine learning process, a deep network of processing layers is learned from a large amount of data. Deep learning eliminates the need for handcrafted representations. Therefore, deep learning is considered particularly suitable for processing natural, unstructured data such as acoustic and visual signals. The processing of such data has been a long-standing challenge in the field of artificial intelligence.

[0059] 2.3 Neural Networks for Image Compression

[0060] Neural networks for image compression can be divided into two categories, including pixel probability models and autoencoder models. Pixel probability models employ a predictive coding strategy. Autoencoder models employ a transform-based solution. Sometimes, the two approaches are combined.

[0061] 2.3.1 Pixel probability modeling

[0062] According to Shannon's information theory, the optimal method for lossless coding can achieve the minimum coding rate, which is represented as -log2 p(x), where p(x) is the probability of the symbol x. Arithmetic coding is a lossless coding method that is considered one of the optimal methods. Given a probability distribution p(x), arithmetic coding makes the coding rate as close as possible to the theoretical limit -log2 p(x) without considering the rounding error. Therefore, the remaining problem is to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality. The curse of dimensionality refers to the problem that increasing the dimensionality causes the dataset to become sparse, so a rapidly increasing amount of data is required to effectively analyze and organize data as the number of dimensions increases.

[0063] Following the predictive coding strategy, one way to model p(x) is to predict the pixel probability one by one in the raster scan order based on the previous observations, where x is an image, which can be represented as follows:

[0064] p(x) = p(x1)p(x2|x1)...p(x i |x1,..., x i-1 )...p(x m×n |x1,..., x m×n-1 ) (1)

[0065] where m and n are the height and width of the image, respectively. The previous observations are also referred to as the context of the current pixel. When the image is large, the estimation of the conditional probability can be difficult. Therefore, a simplified approach is to limit the context range of the current pixel as follows:

[0066] p(x) = p(x1)p(x2|x1)...p(x i |x i-k ,..., x i-1 )...p(x m×n |x m×n-k ,..., x m×n-1 ) (2)

[0067] where k is a predefined constant that controls the context range.

[0068] It should be noted that the condition can also consider the sample values of other color components. For example, when coding the red (R), green (G), and blue (B) (RGB) color components, the R sample depends on the previously coded pixels (including R, G, and / or B samples), the current G sample can be coded according to the previously coded pixels and the current R sample. In addition, when coding the current B sample, the previously coded pixels, as well as the current R and G samples, can also be considered.

[0069] Neural networks can be designed for computer vision tasks and can also be effective in regression and classification problems. Thus, neural networks can be used to estimate the probability of p(x i-1 given the context x1, x2,..., x i .

[0070] Most approaches model the probability distribution directly in the pixel domain. Some designs also model the probability distribution as conditional based on explicit or latent representations. Such models can be represented as:

[0071]

[0072] where h is an additional condition and p(x) = p(h)p(x|h) indicates that the modeling is divided into an unconditional model and a conditional model. The additional condition can be image label information or a high-level representation.

[0073] 2.3.2 Autoencoders

[0074] An autoencoder is now described. An autoencoder is trained for dimensionality reduction and includes an encoding component and a decoding component. The encoding component converts a high-dimensional input signal into a low-dimensional representation. The low-dimensional representation can have a reduced spatial dimension but with a larger number of channels. The decoding component recovers the high-dimensional input from the low-dimensional representation. Autoencoders enable automatic learning of representations and eliminate the need for handcrafted features, which is also considered one of the most important advantages of neural networks.

[0075] Figure 1B is a schematic diagram showing an example transform codec scheme. An original image x is transformed by an analysis network g a to achieve a latent representation y. The latent representation y is quantized (q) and compressed into bits. The number of bits R is used to measure the coding rate. The quantized latent representation is then inverse transformed by a synthesis network g s to obtain a reconstructed image The distortion (D) is computed in the perceptual space by transforming x and p with the function g , resulting in z and which are compared to obtain D.

[0076] Autoencoder networks can be applied to lossy image compression. The learned latent representation can be encoded from a well-trained neural network. However, applying autoencoders to image compression is not trivial because the original autoencoder is not optimized for compression, and thus it is inefficient to use as a trained autoencoder directly. Moreover, there are other major challenges. First, the low-dimensional representation should be quantized before being encoded. However, quantization is not differentiable, which is necessary in backpropagation when training a neural network. Second, the goal is different in the compression scenario because both distortion and rate need to be considered. Estimating rate is challenging. Third, a practical image codec scheme should support variable bit rate, scalability, encoding / decoding speed, and interoperability. Various schemes are being developed for these challenges.

[0077] An example autoencoder for image compression using an example transform codec scheme can be viewed as a transform coding strategy. The original image x is transformed with an analysis network y = g a (x) where y is the latent representation to be quantized and coded. The synthesis network inverse transforms the quantized latent representation to obtain the reconstructed image The framework is trained with a rate-distortion loss function where D is the distortion between x and , R is the rate computed or estimated from the quantized representation , and λ is the Lagrange multiplier. D can be computed in the pixel domain or in the perceptual domain. Most example systems follow this prototype, and the differences between these systems can only be in the network structure or the loss function.

[0078] 2.3.3 Hyper-prior model

[0079] Figure 2 An example latent representation of an image is shown. Figure 2 includes an image 201 from the Kodak dataset, a visualization of the latent values 202 representing y for the image 201, a standard deviation σ 203 of the latent values 202, and latent values y 204 after introducing a hyper-prior network. The hyper-prior network includes an encoder and a decoder that utilize hyper-prior information. In the example shown in FIG. 2B, the hyper-prior network is trained on a separate dataset of images 205, and the latent values y 206 are obtained from the separate dataset of images 205. Figure 1B In a transform coding method of image compression as shown in FIG. 3, an encoder subnetwork uses a parametric analysis transform to transform an image vector x into a latent representation y, which is then quantized to form Because is a discrete value, the quantized latent representation can be losslessly compressed using an entropy coding technique such as arithmetic coding and transmitted as a bit sequence.

[0080] From Figure 2The latent values 202 and standard deviations s 203 of the potential values can be seen to be significantly correlated, There is a significant spatial dependency between the elements of the potential values. Notably, their scale (standard deviation s 203) seems to be coupled over the spatial domain. An additional set of random variables can be introduced to capture the spatial dependency and further reduce redundancy. In this case, the image compression network is as shown in Figure 3

[0081] Figure 3 is a schematic diagram showing an example network architecture implementing an autoencoder for a hyper-prior model. The upper side shows the image autoencoder network, and the lower side corresponds to the hyper-prior subnetwork. The analysis and synthesis transformations are denoted as g a and g a . Q denotes quantization, and AE, AD denote the arithmetic encoder and the arithmetic decoder, respectively. The hyper-prior model comprises two subnetworks, an encoder exploiting hyper-prior information (denoted by h a ) and a decoder exploiting hyper-prior information (denoted by h s ). The hyper-prior model generates quantized hyper-prior information latent values which include information about the probability distribution of the samples of the quantized latent values . are included in the bitstream and transmitted to the receiver (decoder) together with

[0082] In Figure 3 , the upper side of the model is the encoder g a and the decoder g s as discussed above. The lower side is an additional encoder exploiting hyper-prior information h a and a decoder exploiting hyper-prior information h s network used to obtain . In this architecture, the encoder subjects the input image x to g a , yielding a response y with a spatially varying standard deviation. The response y is fed into h a , summarizing the distribution of the standard deviation in z. z is then quantized compressed, and transmitted as side information. The encoder then uses the quantized vector to estimate the spatial distribution of the standard deviation s, and uses s to compress and transmit the quantized image representation The decoder first recovers from the compressed signal. The decoder then uses h = h a to obtain s, which provides the decoder with the correct probability estimates to also successfully recover The decoder then feeds into g s to obtain the reconstructed image.

[0083] When the encoder with super-prior information and the decoder with super-prior information are added to the image compression network, the spatial redundancy of the quantized latent values is reduced. Figure 2 The latent values y 204 in the image 202 correspond to the quantized latent values when using the encoder / decoder with super-prior information. Compared to the standard deviation s 203, the spatial redundancy is significantly reduced because the correlation of the samples of the quantized latent values is lower.

[0084] 2.3.4 Context model

[0085] Although the super-prior model improves the modeling of the probability distribution of the quantized latent values additional improvements can be obtained by utilizing an autoregressive model that predicts the quantized latent values from the causal context of the quantized latent values, which can be referred to as a context model.

[0086] The term “autoregressive” indicates that the output of a process is used later as input to that process. For example, the context model subnetwork generates one sample of the latent values, which is later used as input to obtain the next sample.

[0087] Figure 4 is a schematic diagram illustrating an example combined model configured to jointly optimize the context model together with the super-prior and the autoencoder. The combined model jointly optimizes the autoregressive component (which estimates the probability distribution of the latent values from the causal context of the latent values (context model)) together with the super-prior and the base autoencoder. The real-valued latent values are quantized (Q) to create the quantized latent values and the quantized super-prior information latent values are compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The dashed area corresponds to the components performed by the receiver (e.g., decoder) to recover the image from the compressed bitstream.

[0088] The example system utilizes a joint architecture, where both the super-prior model subnetwork (encoder with super-prior information and decoder with super-prior information) and the context model subnetwork are utilized. The super-prior and the context model are combined to learn a probability model over the quantized latent values which is then used for entropy coding. As Figure 4The outputs of the context subnetwork and the decoder subnetwork that utilizes the superprior information are combined by a subnetwork called the entropy parameter, which generates the mean μ and scale (or variance) σ parameters for a Gaussian probability model. The Gaussian probability model is then used to encode the samples of the quantized latent values into a bitstream with the help of an arithmetic encoder (AE) module. In the decoder, the Gaussian probability model is utilized to obtain the quantized latent values from the bitstream through an arithmetic decoder (AD) module

[0089] In an example, the latent samples are modeled as a Gaussian distribution or a Gaussian mixture model (not limited to). In an example according to Figure 4 In an example, the context model and the superprior are jointly used to estimate the probability distribution of the latent samples. Since a Gaussian distribution can be defined by a mean and a variance (also known as sigma or scale), the joint model is used to estimate the mean and the variance (denoted as μ and σ).

[0090] 2.3.5 Gain Variational Autoencoder (G-VAE)

[0091] In an example, a neural network-based image / video compression method needs to train multiple models to adapt to different rates. Gain Variational Autoencoder (G-VAE) is a variational autoencoder with a pair of gain units, which is designed to achieve continuous variable bitrate adaptation using a single model. It includes a pair of gain units, which are usually inserted into the output of the encoder and the input of the decoder. The output of the encoder is defined as a latent representation y e R c*h*w where c, h, w represent the number of channels, height, and width of the latent representation. Each channel of the latent representation is represented as y (i) e R h*w where i = 0, 1,..., c - 1. The pair of gain units includes a gain matrix M e R c*n and an inverse gain matrix where n is the number of gain vectors. A gain vector can be represented as m s = {a s(0) , a s(i) ,..., a s(c-1)}, a s(i) e R, where s represents the index of the gain vector in the gain matrix.

[0092] The motivation of the gain matrix is similar to the quantization table in JPEG, which controls the quantization loss by characteristics of different channels. In order to apply the gain matrix to the latent representation, each channel is multiplied by the corresponding value in the gain vector.

[0093]

[0094] where is the channel multiplication, i.e. and as(i) It is the gain vector m s The i-th gain value in the matrix. The inverse gain matrix used on the decoder side can be represented as M′∈R. c*n It includes n inverse gain vectors, i.e., M′={δ s(0) δ s(1) ,...,δ s(c-1)}, δ s(i) ∈R. The inverse gain process is represented as:

[0095]

[0096] in It is the decoded quantized latent representation, and y′ s It is an inverse gain quantization latent representation that will be fed into the synthesis network.

[0097] To achieve continuous variable bit rate adjustment, interpolation is used between vectors. Given two pairs of gain vectors {m t ,m′ t} and {m r ,m′ r The interpolated gain vector can be obtained through the following equation.

[0098] m v =[(m r ) l ·(m t ) 1-l ]

[0099] m′ v =[(m′ r ) l ·(m′ t )1-l ]

[0100] Where l∈R are interpolation coefficients, which control the corresponding bit rate of the generated gain vector pairs. Since l is a real number, any bit rate between two given gain vector pairs can be achieved.

[0101] 2.3.6 Encoding process using a joint autoregressive hyperprior model

[0102] Figure 4 The design described here corresponds to the example combined compression method. The encoding and decoding processes are described separately in this and the next section.

[0103] Figure 5 An example encoding process is illustrated. The input image is first processed by an encoder subnetwork. The encoder transforms the input image into a transform representation called the latent value, denoted by y. y is then input to a quantizer block, denoted by Q, to obtain the quantized latent value. are then converted into a bitstream (bits1) using an arithmetic encoding module (denoted AE). The arithmetic encoding block converts each sample of in order into a bitstream (bits1).

[0104] The encoder with hyperpriors module, the context module, the decoder with hyperpriors module and the entropy parameter subnetwork are used to estimate the probability distribution of the samples of the quantized latent values The latent values y are input to the encoder with hyperpriors which outputs a hyperpriors latent value (denoted by z). The hyperpriors latent value is then quantized and a second bitstream (bits2) is generated using an arithmetic encoding (AE) module. The factorized entropy module generates the probability distribution used to encode the quantized hyperpriors latent values into a bitstream. The quantized hyperpriors latent values include information about the probability distribution of the quantized latent values .

[0105] The entropy parameter subnetwork generates the probability distribution estimates used to encode the quantized latent values . The information generated by the entropy parameter module usually includes the mean μ and scale (or variance) σ parameters that are used together to obtain a Gaussian probability distribution. The Gaussian distribution of a random variable x is defined as where the parameter μ is the mean or expectation of the distribution (also the median and mode), and the parameter σ is its standard deviation (or scale or variance). To define a Gaussian distribution, the mean and variance values need to be determined. The entropy parameter module is used to estimate the mean and variance values.

[0106] The decoder with hyperpriors module generates the part of the information used by the entropy parameter subnetwork, the other part of the information is generated by an autoregressive module called the context module. The context module uses the samples that have already been encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the samples of the quantized latent values. The quantized latent values are usually matrices composed of many samples. The samples can be indicated using an index, for example or depending on the dimensions of the matrix . The samples are usually encoded by the AE one by one using a raster scan order. In a raster scan order, the rows of the matrix are processed from top to bottom, with the samples in a row being processed from left to right. In such a scenario (where the AE encodes the samples into a bitstream using a raster scan order), the context module uses the samples that have been encoded previously in the raster scan order to generate the information about the probability distribution of the sample Related information. The information generated by the context module and the decoder with super-prior information is combined by the entropy parameter module to generate the probability distribution that is used to convert the quantized latent values into a bitstream (bits1) as a result of the encoding process. It is important to note that other names can be used for the above modules.

[0107] Finally, the first and second bitstreams are transmitted to the decoder as a result of the encoding process. It is important to note that other names can be used for the above modules.

[0108] In the above description, Figure 5 All elements in the encoder are collectively referred to as the encoder. The analytical transform that converts the input image into a latent representation is also referred to as an encoder (or autoencoder).

[0109] 2.3.7 Decoding process using a joint autoregressive super-prior model

[0110] Figure 6 An example decoding process is shown. Figure 6 The decoding process is depicted separately.

[0111] In the decoding process, the decoder first receives the first bitstream (bits1) and the second bitstream (bits2) generated by the corresponding encoder. bits2 is first decoded by the arithmetic decoding (AD) module using the probability distribution generated by the factorized entropy subnetwork. The factorized entropy module typically uses a predetermined template to generate the probability distribution, such as a predetermined mean and variance value in the case of a Gaussian distribution. The output of the arithmetic decoding process of bits2 is which is the quantized super-prior information latent value. The AD process recovers the AE process applied in the encoder. The processes of AE and AD are lossless, which means that the quantized super-prior information latent value can be reconstructed at the decoder without any change.

[0112] After obtaining , it is processed by the decoder with super-prior information, and the output of the decoder with super-prior information is fed to the entropy parameter module. The three subnetworks employed in the decoder (context, decoder with super-prior information, and entropy parameter) are the same as the three subnetworks in the encoder. Therefore, the exact same probability distribution as in the encoder can be obtained in the decoder, which is essential for losslessly reconstructing the quantized latent value . As a result, the exact same version of the quantized latent value

[0113] After obtaining the probability distribution (e.g. mean and variance parameters) by the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of the quantized latent values from the bitstream bits1 one by one. From a practical point of view, autoregressive models (contextual models) are inherently serial, so techniques such as parallelization cannot be used to speed up. Finally, the fully reconstructed quantized latent values are input to the synthesis transform (denoted as decoder in Figure 6 ) module to obtain the reconstructed image.

[0114] In the above description, Figure 6 all elements in are collectively referred to as the decoder. The synthesis transform that converts the quantized latent values to the reconstructed image is also referred to as the decoder (or auto-decoder).

[0115] 2.4 Neural networks for video compression

[0116] Similar to video codec technology, neural image compression is the basis for intra-frame compression in neural video compression. Therefore, the development of neural video compression technology lags behind the development of neural image compression, because neural video compression technology has higher complexity, and therefore requires more effort to address the corresponding challenges. Compared with image compression, video compression requires efficient methods to eliminate inter-frame redundancy. Inter-frame prediction is the main step in these example systems. Motion estimation and compensation are widely used in video codecs, but are usually not implemented by trained neural networks.

[0117] Neural video compression can be divided into two categories according to the target scenario: random access and low latency. In the case of random access, the system allows decoding to start from any point in the sequence, usually divides the entire sequence into multiple separate segments, and allows each segment to be decoded independently. In the case of low latency, the system aims to reduce the decoding time, so that the frames preceding in time can be used as reference frames to decode the subsequent frames.

[0118] 2.5 Preliminaries

[0119] Almost all natural images and / or videos are in digital format. A grayscale digital image can be represented by where is a set of pixel values, m is the image height, and n is the image width. For example, is an example setting, and in this case Therefore, a pixel can be represented by an 8-bit integer. An uncompressed grayscale digital image has 8 bits per pixel (bpp), while a compressed bit is certainly less.

[0120] Color images are typically represented in multiple channels to record color information. For example, in the RGB color space, an image can be represented by where three separate channels store red, green, and blue information. Similar to an 8-bit grayscale image, an uncompressed 8-bit RGB image has 24 bpp. Digital images / videos can be represented in different color spaces. Neural network based video compression schemes are mostly developed in the RGB color space, while video codecs usually use the YUV color space to represent video sequences. In the YUV color space, an image is decomposed into three channels, i.e., luminance (Y), blue-difference chroma (Cb), and red-difference chroma (Cr). Y is the luminance component, and Cb and Cr are the chrominance components. The compression benefit of YUV occurs because Cb and Cr are usually down-sampled to achieve pre-compression, as the human visual system is less sensitive to chrominance components.

[0121] A color video sequence is composed of multiple color images (also called frames) to record the scene at different timestamps. For example, in the RGB color space, a color video can be represented by X = {x0, x1,..., x t ,..., x T-1}, where T is the number of frames in the video sequence, and If m = 1080, n = 1920, and the video has 50 frames per second (fps), the data rate of this uncompressed video is 1920 x 1080 x 8 x 3 x 50 = 2,488,320,000 bits per second (bps). This results in about 2.32 gigabits per second (Gbps), which uses a large amount of storage and should be compressed before transmission over the Internet.

[0122] In general, lossless methods can achieve about 1.5 to 3 compression ratio on natural images, which is significantly lower than the streaming requirement. Therefore, lossy compression is adopted to achieve better compression ratio, but at the cost of the distortion produced. Distortion can be measured by computing the mean squared error between the original image and the reconstructed image, e.g., based on MSE. For a grayscale image, MSE can be computed with the following equation.

[0123]

[0124] Therefore, the quality of the reconstructed image compared to the original image can be measured by the peak signal-to-noise ratio (PSNR):

[0125] where is the maximum value in , e.g., 255 for an 8-bit grayscale image. There are other quality assessment metrics, such as structural similarity (SSIM) and multi-scale SSIM (MS-SSIM).

[0126] To compare different lossless compression schemes, one can compare the compression ratio that gives the resulting rate, and vice versa. However, to compare different lossy compression methods, the comparison must take both the rate and the reconstruction quality into account. This can be done, for example, by computing the relative rate at several different levels of quality, and then averaging the rates. The average relative rate is referred to as the Bjontegaard delta rate (BD rate). There are other aspects of evaluating image and / or video coding schemes, including encoding / decoding complexity, scalability, robustness, etc.

[0127] 2.6 Separate processing of luminance and chrominance components of an image

[0128] Figure 7 An example decoding process according to the present disclosure is shown.

[0129] According to one implementation, the luminance and chrominance components of an image can be decoded using separate subnetworks. In Figure 7 The luminance component of an image is processed by subnetworks “Synthesis”, “Predictive Fusion”, “Masked Convolution”, “Decoder with hyperprior information”, “Variance decoder with hyperprior information”, etc. While the chrominance components are processed by subnetworks: “Synthesis UV”, “Predictive Fusion UV”, “Masked Convolution UV”, “Decoder with hyperprior information UV”, “Variance decoder with hyperprior information UV”, etc.

[0130] The benefit of this separate processing is that by applying separate processing, the computational complexity of the processing of an image is reduced. Typically, in neural network based image and video decoding, the computational complexity is proportional to the square of the number of feature maps. For example, if the total number of feature maps is 192, then the computational complexity will be proportional to 192x192. On the other hand, if the feature maps are divided into 128 for luminance and 64 for chrominance (in case of separate processing), then the computational complexity is proportional to 128x128 + 64x64, which corresponds to a 45% reduction in complexity. Typically, separate processing of the luminance and chrominance components of an image does not result in excessive reduction in performance, since the correlation between the luminance and chrominance components is typically very small.

[0131] Figure 7 The processing (decoding process) in

[0132] 1. First, a factorized entropy model is used to decode the quantized latent values for luminance and chrominance, i.e. Figure 7 The and

[0133] 2. The probability parameters (e.g. variances) generated by the second network are used to generate the quantized residual latent values by performing an arithmetic decoding process.

[0134] 3. The quantized residual latent values are de-gained using an inverse gain unit (iGain) as shown in orange in Figure 7 The output of the inverse gain unit is denoted as and

[0135] 4. For the luma component, the following steps are performed in a loop until all elements of are obtained:

[0136] a. The first sub-network is used to estimate the mean parameters of the quantized latent values using the already obtained samples of .

[0137] b. The quantized residual latent values and the mean are used to obtain the next element of .

[0138] 5. After all samples of are obtained, a synthesis transform can be applied to obtain the reconstructed image.

[0139] 6. For the chroma components, steps 4 and 5 are the same but with separate set of networks.

[0140] 7. The decoded luma component is used as additional information to obtain the chroma components. Specifically, an inter-channel correlation information filter sub-network (ICCI) is used for chroma component recovery. The luma is fed into the ICCI sub-network as additional information to assist the chroma component decoding.

[0141] 8. After the luma and chroma components are reconstructed, an adaptive color transform (ACT) is performed.

[0142] The module named ICCI is a neural network based post-processing module. The example is not limited to the UCCI sub-network. Any other neural network based post-processing module can also be used.

[0143] An exemplary implementation of the disclosure is depicted in Figure 7 (decoding process). The framework includes two branches for luma and chroma components respectively. In each branch, the first sub-network includes a context, a prediction, and an optional decoder module with hyper-prior information. The second network includes a variance decoder module with hyper-prior information. The quantized hyper-prior information latent values are and The arithmetic decoding process generates quantized residual latent values, which are further fed into an iGain unit to obtain the de-gained quantized residual latent values and

[0144] After obtaining the residual latent values, a recursive prediction operation is performed to obtain the latent values and The following steps describe how the latent values are obtained The luma component is processed in the same way but with a different network.

[0145] 1. An autoregressive context module is used to generate a first input of the prediction module using the samples where the pair (m, n) is the index of the sample of the latent values that has already been obtained.

[0146] 2. Optionally, a second input of the prediction module is obtained by using a decoder that makes use of hyper-prior information and the quantized hyper-prior information latent values

[0147] 3. Using the first input and the second input, the prediction module generates a mean value mean[ :, i, j].

[0148] 4. The mean value mean[ :, i, j] and the quantized residual latent values are added together to obtain the latent values

[0149] 5. Steps 1-4 are repeated for the next sample.

[0150] Whether and / or how the at least one method disclosed herein is applied can be signaled from the encoder to the decoder, e.g. in a bitstream.

[0151] Whether and / or how the at least one method disclosed herein is applied can be determined by the decoder based on codec information such as dimensions, color format, etc.

[0152] Furthermore, a module named MS1, MS2 or MS3+O (in Figure 7 ) can be included in the processing flow. This module can perform an operation on its input by multiplying the input with a scalar or adding an additional component to the input to obtain an output. The scalar or the additional component used by this module can be indicated in the bitstream.

[0153] Figure 7 The module named RD in or the module named AD can be an entropy decoding module. It can be a range decoder or an arithmetic decoder, etc.

[0154] The examples described herein are not limited to the specific combination of units of the examples in Figure 7 Some modules can be missing and some modules can be replaced in the processing order. In addition, additional modules can be included. For example: ​

[0155] 1. The ICCI module can be removed. In this case, the outputs of the synthesis module and the synthesis UV module can be combined by way of another module, which can be based on a neural network.

[0156] 2. One or more of the modules named MS1, MS2, or MS3+O can be removed. The core of the disclosure is not affected by the removal of one or more of the scaling and adding modules.

[0157] In Figure 7 , other operations performed during the processing of the luma and chroma components are also indicated using asterisks. These processes are denoted as MS1, MS2, MS3+O. These processes can be, but are not limited to, an adaptive quantization, a potential sample scaling, and a potential sample offset operation. For example, in the adaptive quantization process can correspond to a scaling of the samples with a multiplier before the prediction process, where the multiplier is predefined or its value is indicated in the bitstream. The potential scaling process can correspond to a process where the samples are scaled with a multiplier after the prediction process, where the value of the multiplier is predefined or indicated in the bitstream. The offset operation can correspond to adding an additional element to the samples, again where the value of the additional element can be indicated in the bitstream or be inferred or predetermined.

[0158] Another operation can be a tiling operation, where the samples are first tiled (grouped) into overlapping or non-overlapping regions, where each region is processed independently. For example, the samples corresponding to the luma component can be divided into tiles of height 20 samples, while the chroma components can be divided into tiles of height 10 samples for processing.

[0159] Another operation can be the application of wavefront parallel processing. In wavefront parallel processing, multiple samples can be processed in parallel, and the amount of samples that can be processed in parallel can be indicated by a control parameter. The control parameter can be indicated in the bitstream, can be inferred, or can be predetermined. In the case of separate luma and chroma processing, the number of samples that can be processed in parallel can be different, so different indicators can be signaled in the bitstream to control the operation of the luma and chroma processing separately.

[0160] 2.7 Color separation and conditional coding

[0161] Figure 8 An image codec architecture based on example learning is shown.

[0162] In one example, as Figure 8As shown, using networks with similar architecture but different number of channels, the primary and secondary color components of an image are coded separately. All blocks with the same name are subnetworks with similar architecture, only the input-output tensor dimensions and the number of channels are different. The number of channels for the primary component is C p = 128, and the number of channels for the secondary component is C s = 64. The vertical arrows (arrows pointing downwards) indicate the data flow related to the secondary color component coding. The vertical arrows show the data exchange between the primary and secondary component pipelines.

[0163] The input signal to be encoded is denoted as x, and the latent space tensor in the bottleneck of the variational autoencoder is y. The subscript "Y" indicates the first component, and the subscript "UV" is used for the second component in series, where there is a chroma component.

[0164] First, the input image with RGB color format is converted into a primary (Y) component and a secondary (UV) component. The primary component x Y is coded independently from the secondary component x UV and the coded picture size is equal to the input / decoded picture size. The secondary component is conditionally coded using x Y from the primary component as side information to encode x UV and using x from the primary component with side information to decode the reconstruction. The codec structure for the primary and secondary components is almost identical except for the number of channels, the size of the channels, and the number of entropy models used to convert the latent tensors into bitstreams, so the first and second latent tensors will generate two different bitstreams based on two different entropy models. Before encoding x Y , x UV goes through a module (labeled "s↓" on Figure 8 the top) that adjusts the sample positions by downsampling, which essentially means that the coded picture size of the secondary component is different from the coded picture size of the primary component. The scaling factor s is variable, but the default scaling factor is s = 2. The size of the side input tensor in the conditional coding is adjusted so that the encoder receives the primary component tensor and the second secondary component tensor with the same picture size. After the reconstruction, the secondary component is rescaled to the original picture size with a neural network-based upsampling filter module ("NN color filter s↑" on Figure 8 ), which outputs the secondary component upscaled with a factor of s.

[0165] Figure 8The example in FIG. 1 illustrates an image coding system in which an input image is first converted into a first (Y) component and a second (UV) component. The output is a reconstructed output corresponding to the first and second components. At the end of processing, is converted back to the RGB color format. Typically, x UV is down-sampled (resized) before processing with the encoding and decoding modules (neural networks). For example, x UV can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Thus, the processing of the second component includes about 50% x 50% = 25% less samples and is therefore less computationally complex.

[0166] 2.8 Cropping operation in neural network based coding

[0167] Figure 9 An example synthesis transform for learning based image coding is shown.

[0168] The above example synthesis transform includes a sequence of 4 convolutions with an up-sampling with a stride of 2. The synthesis transform subnetwork is depicted in Figure 9 . The dimensions of the tensors in the different parts of the synthesis transform before the cropping layer are shown in the diagram on Figure 9 .

[0169] The cropping layer changes the tensor dimensions h d x w d to h d-1 x w d-1 , where h d = 2 ceil(H / 2 d ); w d = 2 ceil(W / 2 d ); here d is the depth of the previous convolution in the codec architecture. For the primary component, the synthesis transform receives an input tensor of dimensions h x w, where h = ceil(H / 16); w = ceil(W / 16). The output of the synthesis transform for the primary component is 1 x h0 x w0, where h0 = H; h0 = W.

[0170] For the secondary component, the synthesis transform receives an input tensor of dimensions h UV x w UV ; h UV = ceil(ceil(H / s) / 16); w UV = ceil(ceil(W / s) / 16). The output of the synthesis transform for the primary component is 2 x h UV x w UV0 , where h UV = ceil(H / s); h UV0= ceil(W / s). For the secondary components, the input size is ho = ceil(H / s); wo = ceil(W / s), where s is the scaling factor. For example, the scaling factor can be 2, where the secondary components are down-sampled by a factor of 2.

[0171] Based on the above explanation, the operation of the crop layer depends on the output size H, W and the depth of the crop layer. Figure 9 The depth of the left-most crop layer in the middle is equal to 0. If the size of the input of this crop layer is larger than H or W in the horizontal or vertical dimension, respectively, the output of this crop layer must be equal to H, W (output size) and cropping needs to be performed in that dimension. The second crop layer, counted from left to right, has a depth of 1. The output of the second crop layer must be equal to hi = 2 ceil(H / 2 1 ); wi = 2 ceil(W / 2 1 ), which means that if the input of this second crop layer is larger than hi, wi in any dimension, cropping is applied to that dimension. In summary, the operation of the crop layer is controlled by the output size H, W. In one example, if H and W are both equal to 16, the crop layer does not perform any cropping. On the other hand, if H and W are both equal to 17, all 4 crop layers will perform cropping.

[0172] 2.9 Bitwise shift

[0173] The bitwise shift operator can be represented using the function bitshift(x, n), where n is an integer. If n is larger than 0, it corresponds to the right shift operator (>>), which shifts the bits of the input to the right, and the left shift operator (<<), which shifts the bits to the left. In other words, the bitshift(x, n) operation corresponds to:

[0174] bitshift(x, n) = x * 2 n ,

[0175] or bitshift(x, n) = floor(x * 2 n ),

[0176] or bitshift(x, n) = x / / 2 n .

[0177] The output of the shift operation is an integer value. In some implementations, the floor() function can be added to the definition.

[0178] floor(x) is equal to the largest integer that is smaller than or equal to x.

[0179] The " / / " operator or integer division operator. It is the operation that includes division and truncation of the result to zero. For example, 7 / 4 and -7 / -4 are truncated to 1 and -7 / 4 and 7 / -4 are truncated to -1.

[0180] rightshift(x, n) = x » n or

[0181] leftshift(x, n) = x « n

[0182] Equation 3: Shift operators as alternative implementations of right shift or left shift.

[0183] x » y Arithmetic right shift of the two's complement integer representation of x by y binary digits. The function is defined only for non-negative integer values of y. Bits shifted into the most significant bits (MSB) as a result of the right shift have a value equal to the MSB of x before the shift operation.

[0184] x « y Arithmetic left shift of the two's complement integer representation of x by y binary digits. The function is defined only for non-negative integer values of y. Bits shifted into the least significant bits (LSB) as a result of the left shift have a value equal to 0.

[0185] 2.10 Convolution operation

[0186] The convolution operation starts with a kernel, which is a small matrix of weights. The kernel "slides" over the input data, performing element-wise multiplication with the part of the input it is currently on, and then adding up the results to become a single output pixel. In some cases, the convolution operation can include a "bias", which is added to the output of the element-wise multiplication operation.

[0187] The convolution operation can be described by the following mathematical equation.

[0188] The output out1 can be obtained as:

[0189]

[0190] where w1 is a multiplication factor, K1 is called a bias (additive term), I k is the kth input, N is the kernel size in one direction, and P is the kernel size in the other direction. A convolution layer can include a convolution operation in which more than one output can be generated.

[0191] Other equivalent depictions of the convolution operation can be found below:

[0192]

[0193] In the above equation, "c" indicates the channel number. It is equal to the output number, out[1,x,y] is one output, and out[2,x,y] is the second output. k is the input number, I[1,x,y] is one input, and I[2,x,y] is the second input. wl or w describes the weights of the convolution operation.

[0194] 2.11 LeakyReLU activation function

[0195] Figure 10 An example LeakyReLU activation function is shown. The LeakyReLU activation function is depicted in Figure 10 According to this function, if the input is a positive value, the output is equal to the input. If the input (y) is a negative value, the output is equal to a*y. a is typically (without limitation) a value that is less than 1 and greater than 0. Since the multiplier a is less than 1, it can be implemented as a multiplication or division operation with a non-integer. The multiplier a can be referred to as the negative slope of the LeakyReLU function.

[0196] 2.12 ReLU activation function

[0197] Figure 11 An example ReLU activation function is shown. The ReLU activation function is depicted in Figure 11 According to this function, if the input is a positive value, the output is equal to the input. If the input (y) is a non-positive value, the output is equal to 0.

[0198] 3. Technical problems solved by the disclosed technical solutions

[0199] When components of an image (e.g. luminance component and chrominance components) are processed with different synthesis sub-networks, the correlation between the different components is not fully exploited. In other words, information that can be important for the reconstruction of one component can also be relevant for the reconstruction of a second component. When 2 different synthesis transformations are used to reconstruct 2 different components, this joint information cannot be fully exploited.

[0200] 4. List of solutions and embodiments

[0201] 4.1 Core example

[0202] The objective of the present disclosure is to improve the quality of a component of an image using information from another component. This objective is achieved by:

[0203] • using common processing layers that are used in the neural network implementation.

[0204] • and by including the weights and offset (bias) parameters of said processing layers in the bitstream.

[0205] Decoder operation:

[0206] According to some examples, converting an image to a bitstream using a neural network comprises the following operations:

[0207] - obtaining a weight value from the bitstream,

[0208] - obtaining an offset value,

[0209] - obtaining a result value according to any or all of:

[0210] o applying the offset value to the samples of the first component.

[0211] o applying a threshold function (e.g. a Relu operation) to the samples of the first component.

[0212] o applying the weight value to the samples of the first component.

[0213] - obtaining modified samples of the second component from the result value and the samples of the second component.

[0214] - obtaining the reconstructed image using the samples of the first component and the modified samples of the second component.

[0215] Encoder operations:

[0216] According to some examples, converting an image to a bitstream using a neural network comprises the following operations:

[0217] - obtaining / determining an offset value,

[0218] - obtaining a result value according to any or all of:

[0219] o applying the offset value to the samples of the first component.

[0220] o applying a threshold function (e.g. a Relu operation) to the samples of the first component.

[0221] o computing the weight value.

[0222] - obtaining modified samples of the second component from the result value and the samples of the second component.

[0223] - obtaining the reconstructed image using the samples of the first component and the modified samples of the second component, wherein the weight value is computed (selected) to maximize the quality of the reconstructed image.

[0224] - including the weight value in the bitstream.

[0225] The first component, or the second component, or any of the components described above can be a component of an image.

[0226] - It can be a chroma component or a luma component.

[0227] - The mean value can be subtracted from any component before applying the proposed solution.

[0228] - The mean value can be added to the upsampled component after applying the proposed solution.

[0229] In one example, the first component is Y in YCbCr color format and the second component is the Cb or Cr component.

[0230] In one example, the first component is the G component in RGB color format and the second component is the B / R component.

[0231] In one example, both offsets and / or both weights can be signaled in the bitstream.

[0232] • Alternatively, only one offset and / or one weight can be signaled in the bitstream and the second / third component can share the same value.

[0233] • Alternatively, a prediction coding can be applied to code one of the two weights.

[0234] • Alternatively, a prediction coding can be applied to code one of the two offsets.

[0235] 4.2 Details of the examples

[0236] • The five example implementations of the disclosure can be according to the following equations:

[0237]

[0238] or

[0239]

[0240] or

[0241]

[0242] or

[0243]

[0244] or

[0245]

[0246] In the above equations, the first component is recY (e.g., the luminance component of the image).

[0247] The second component is recU (e.g., the chrominance component of the image).

[0248] The threshold function is a RELU() function.

[0249] The weights are W[n]. In the above equation, M different weight values are used.

[0250] The offsets are b[n]. In the example equation, M different offset values are used.

[0251] The index [1, x, y] indicates the sample at coordinate [1, x, y], which is the coordinate of the sample of the first component or the second component.

[0252] According to the first equation, the multiplication weight values W[n] are first applied to the sample of the first component. Then the addition offset values b[n] are applied to the sample. After that, the threshold function (RELU in the example) is applied. In the example, at most M such weight and offset values are applied to the first sample, and a summation operation The results are added together. Finally, the result of the summation operation is added to the sample of the second component. The third equation is similar to the first equation.

[0253] According to the second equation, the addition offset values b[n] are first applied to the sample. After that, the threshold function (RELU in the example) is applied. Then the multiplication weight values (W[n]) are applied. In the example, at most M such weight and offset values are applied to the first sample, and a summation operation The results are added together. Finally, the result of the summation operation is added to the sample of the second component. The fourth equation is similar to the first equation.

[0254] The mean value can be subtracted from recU or recY before inputting into the process. The mean value can be the mean (average) of the samples of recU or recY.

[0255] The mean value can be added to the modified recU. The mean value can be the mean (average) of the samples of recU or recY.

[0256] Figure 12 is a flowchart of an example method for video processing. Figure 13 is a flowchart of an example method for video processing. Figure 12 and Figure 13 The flowchart in depicts an example implementation of the present disclosure.

[0257] o First, offset subtraction (or addition) is performed on component 1. Then, in Figure 12 The threshold function is performed. Finally, the weight is applied, and the output is added to the second component. At the end of the flowchart, the modified second component is obtained. The reconstructed picture at the end of the decoder or encoder is obtained according to the first component and the modified second component.

[0258] o InFigure 13 In particular, the operations are very similar, except for the fact that the order of the multiplication with the weight and the threshold operation is exchanged. Figure 12

[0259] o The mean value can be subtracted from the input before applying the proposed solution. The mean value can be the mean (average) of the samples of the first component of the second component.

[0260] o The mean value can be added to the output of the method. The mean value can be the mean (average) of the samples of the first component of the second component.

[0261] • The threshold function can be (without limitation) a RELU operation, a leaky Relu operation, a sigmoid operation, a hyperbolic tangent operation, or a MAX(x,y) operation or a MIN(x,y) operation. The MAX(x,y) operation outputs the maximum of the two values x or y, and the MIN(x,y) outputs the minimum of the two values x or y.

[0262] o The sigmoid function can be described as: f(x) = 1 / (1 + e -x ).

[0263] o The hyperbolic tangent can be described as: f(x) = tanh(x) = 2 / (1 + e -2x )-1).

[0264] o In the MAX(x,y) operation or the MIN(x,y) operation, one of the input values can be zero. In other words, the threshold function can be MAX(x,0) or MIN(x,0).

[0265] • The weight value can be implemented as part of the convolution function.

[0266] • The offset value can be implemented as part of the convolution function. More specifically, the offset value can be implemented as a bias value of the convolution function.

[0267] • The first component can be a luminance or luminance component of an image.

[0268] • The second component can be a U chrominance component, or a V chrominance component, or a chrominance component or a chrominance component.

[0269] • Figure 14 is a flowchart of an example method for video processing. Figure 14 The flowchart in depicts another example implementation of the present disclosure. In this example, the first component ​The example in FIG. 1 illustrates the fact that the disclosure can be implemented using the most common neural network processing layers, namely convolutional layers and activation layers such as the relu function. Figure 14 The flowchart in FIG. 1 illustrates another example implementation of the disclosure. This example is similar to the example in FIG. 1, except for the fact that the first component and the second component are both input to a first convolutional layer (e.g., Conv(lxl, 2, 16, bias=l)) and the final addition operation is removed.

[0270] Figure 15 The flowchart in FIG. 1 illustrates another example implementation of the disclosure. This example is similar to the example in FIG. 1, except for the fact that the first component and the second component are both input to a first convolutional layer (e.g., Conv(lxl, 2, 16, bias=l)) and the final addition operation is removed. Figure 15 Figure 14

[0271] In FIG. 1, the following equation can be implemented: Figure 15

[0272]

[0273] • The mean value can be subtracted from recU or recY before being input into the process. The mean value can be the mean (average) of the samples of recU or recY. Figure 16 The flowchart in FIG. 1 illustrates another example implementation of the disclosure. This example is similar to the example in FIG. 1, except for the fact that the first component and the second component are both input to a first convolutional layer (e.g., Conv(lxl, 2, 16, bias=l)) and the final addition operation is removed. Figure 16 The flowchart in FIG. 1 illustrates another example implementation of the disclosure. This example is similar to the example in FIG. 1, except for the fact that the first component and the second component are both input to a first convolutional layer (e.g., Conv(lxl, 2, 16, bias=l)) and the final addition operation is removed.

[0274] • According to the disclosure, the multiplication weight values can be included in the bitstream at the encoder or obtained from the bitstream at the decoder.

[0275] • According to the disclosure, the addition offset values (or bias values) can be obtained from the bitstream.

[0276] • The mean value can be subtracted from recU or recY before being input into the process. The mean value can be the mean (average) of the samples of recU or recY.

[0277] • The mean value can be subtracted from recU or recY before being input into the process. The mean value can be the mean (average) of the samples of recU or recY.

[0278] • The offset value can be obtained according to the maximum value and / or the minimum value.

[0279] o The maximum value can be the maximum value of the samples of the first component.

[0280] o The minimum value can be the minimum value of the samples of the first component. ​​​

[0281] o The maximum or minimum value can be obtained from the bitstream.

[0282] o The offset value can be obtained from a value N that is used to divide the difference between the maximum and minimum values.

[0283] N can be predefined or can be obtained from the bitstream.

[0284] o The offset values 1...N can be obtained as follows:

[0285] where n and N are integer values.

[0286] Figure 17 An example neural network is shown. Figure 17 Another implementation of the disclosure is depicted.

[0287] EFE brightness-aided non-linear filtering process

[0288] The input of the process is and The output of the process is

[0289] The multiplication weight parameter W3

[16] is used.

[0290] The addition bias parameter B2[8] is used.

[0291] For x in 0..W, y in 0..H, and k in 0..1, the following is performed;

[0292]

[0293] 4.3. Explanation and benefits of the example

[0294] The example uses parameters obtained from the bitstream to improve the quality of the reconstructed image. The example is designed such that the following benefits are achieved:

[0295] 1. Some of the parameters used in the equation are obtained from the bitstream. This provides the possibility of content adaptation. In a neural network based image compression network, the network can be pre-trained using a very large dataset. After the training is completed, the network parameters (e.g. weight and / or bias values) cannot be adjusted. However, when the network is used, it is used for brand new images that do not belong to the part of the training dataset. Therefore, there is a difference between the training dataset and real life images. To solve this problem, a small set of parameters that are optimized for the new image are transmitted to the decoder to improve the adaptation to the new content.

[0296] ​A second benefit of including the parameters in the bitstream is that when the parameters are transmitted, much shorter networks can be used to serve the same purpose. In other words, if the parameters are not transmitted as side information, much longer neural networks (including many more convolution and activation layers) can already be necessary to achieve the same purpose.

[0297] 2. The examples can be implemented using the most basic neural network layers. The equations used to explain the examples are designed such that they can be implemented using the most basic processing layers in the neural network literature (i.e. convolution and relu operations). The reason for this intentional choice is that image encoders / decoders are expected to be implemented in a wide range of devices including mobile phones. It is important that an image encoded in one device can be decoded in almost all devices. Although the neural processing chipsets or GPUs in such devices are becoming more and more complex, it is still not possible to implement arbitrary functions on such processing units. As a simple example, the function f(x) = x 2 Although it seems very simple, it cannot be efficiently implemented in a neural processing unit and can only be implemented in a general purpose processing unit such as a CPU. If the function cannot be implemented in a neural processing unit, the processing speed and battery consumption will be greatly increased.

[0298] The examples eliminate the above problem by using the most basic processing layers in the neural network literature. Convolution and relu (and some other activation functions such as leaky relu, sigmoid, etc.) are almost guaranteed to be implemented in a neural processing unit or a GPU. Therefore, a mobile phone with a neural processing unit or a GPU is expected to perform the defined operations efficiently.

[0299] 3. The examples exploit cross-component information to improve the components of an image. According to the examples, the quality of the components is improved, so the reconstructed image is closer to the original image, which is the goal of a good codec. The examples achieve this by exploiting the information included in one component to improve the quality of the second component.

[0300] 5. Further solutions

[0301] 5.1. Technical problem solved by the disclosed technical solution

[0302] When the components of an image (e.g. the luminance component and the chrominance components) are processed with different synthesis subnetworks, the correlation between the different components is not fully exploited. In other words, information that can be important for the reconstruction of one component can also be relevant for the reconstruction of the second component. When two different synthesis transformations are used to reconstruct two different components, this joint information cannot be fully exploited.

[0303] 5.2. List of solutions and embodiments

[0304] In an example, a neural network based image and video compression method is used, which includes modifying components of an image using an offset. It is determined whether to add an offset value to a sample of a second component is based on a value of a sample of a first component.

[0305] 5.2.1 Core example

[0306] Example decoder operation:

[0307] Example 1: According to the present disclosure, converting a bitstream to a reconstructed image using a neural network includes the following operations:

[0308] • obtaining an offset value from the bitstream,

[0309] • determining whether a value of a first sample of a first component is greater than (or less than) a threshold value,

[0310] • if the determination is yes, modifying a value of a second sample of a second component according to the offset value,

[0311] • obtaining the reconstructed image using the first sample and the second sample.

[0312] Example 2: According to the present disclosure, converting a bitstream to a reconstructed image using a neural network includes the following operations:

[0313] • obtaining an offset value from the bitstream,

[0314] • determining whether a value of a first sample of a first component is greater than (or equal to) a first threshold value and less than (or equal to) a second threshold value,

[0315] • if the determination is yes, modifying a value of a second sample of a second component according to the offset value,

[0316] • obtaining the reconstructed image using the first sample and the second sample.

[0317] Example 3: According to the present disclosure, converting a bitstream to a reconstructed image using a neural network includes the following operations:

[0318] • obtaining N offset values, offset[1], offset[2],..., offset[N], from the bitstream,

[0319] • dividing samples of a first component into N groups, group[1], group[2],..., group[N], based on values of the samples,

[0320] • modifying samples of a second component by:

[0321] • if the sample of the first component corresponding to the sample of the second component is included within group[n], add offset[n].

[0322] • obtain the reconstructed image using the first sample and the second sample.

[0323] Example 4: According to the disclosure, converting a bitstream into a reconstructed image using a neural network comprises the following operations:

[0324] • obtaining N offset values, offset[1], offset[2],..., offset[N], from the bitstream,

[0325] • obtaining a first value and a second value,

[0326] • obtaining a gap value using the first value and the second value by gap = (second value - first value) / N.

[0327] • obtaining a first threshold value and a second threshold value by thr n-1 = first value + gap * (n - 1) and thr n = first value + gap * n, respectively.

[0328] • determining whether the value of the first sample of the first component is less than (or equal to) thr n and greater than (or equal to) thr n-1 .

[0329] • if the determination is yes, modifying the second sample of the second component corresponding to the first sample by adding offset[n].

[0330] • obtaining the reconstructed image using the first sample and the second sample.

[0331] According to the example, the first sample and the second sample can have the same coordinates. In other words, the first sample of the first component and the second sample of the second component can be components of a sample of the image.

[0332] According to the example, the first sample and the second sample can have corresponding coordinates. In other words, if the coordinates of the first sample are given by (x, y), the corresponding coordinates of the second sample can be (x / 2, y / 2).

[0333] 5.2.2 Details of the example

[0334] • In example 4:

[0335] • the first value and the second value can be the minimum value and the maximum value, respectively.

[0336] • the first value and the second value can be signaled in the bitstream.

[0337] o The first value and the second value can be computed from values of the samples of the first component.

[0338] o The first value can be a minimum value of the samples of the first component.

[0339] o The second value can be a maximum value of the samples of the first component.

[0340] • According to an example, the first component and the second component can be obtained using a neural network.

[0341] • According to an example, the first component and the second component can be obtained using a synthetic transform.

[0342] o In one example, the first component can be obtained using a first synthetic transform and the second component can be obtained using a second synthetic transform.

[0343] o The first component can be a luminance component.

[0344] o The second component can be a chrominance component.

[0345] ■ The second component can be a chrominance U component,

[0346] ■ The second component can be a chrominance V component,

[0347] ■ The second component can be a chrominance Cb component,

[0348] ■ The second component can be a chrominance Cr component.

[0349] o The first component and the second component can correspond to a rectangular portion of the image. In other words, the first component and the second component can be processed by first being sliced into rectangular portions.

[0350] • The threshold (first threshold or second threshold) can be obtained from a maximum value of the samples of the first component.

[0351] o The maximum value can be a maximum value of all the samples of the first component.

[0352] o The maximum value can be a maximum value of all the samples of a group of samples of the first component.

[0353] o The maximum value can be a maximum value of all the samples within a tile partition of the first component.

[0354] • The threshold (first threshold or second threshold) can be obtained from a minimum value of the samples of the first component.

[0355] o The minimum value can be a minimum value of all the samples of the first component.

[0356] o The minimum value can be a minimum value of all the samples of a group of samples of the first component.

[0357] o The minimum value can be the minimum value of all samples within a tile partition of the first component.

[0358] • The threshold value (first threshold value or second threshold value) can be obtained according to a maximum value and / or a minimum value signaled in the bitstream.

[0359] • The threshold value (first threshold value or second threshold value) can be obtained based on a number signaled in the bitstream. For example, the threshold value can be obtained according to gap = (maximum value - minimum value) / N, where N is signaled in the bitstream.

[0360] o The threshold value can be obtained according to the following equation:

[0361] ■thr n = minimum value + gap x n,

[0362] ■ In the above example, thr n is the n th threshold value.

[0363] o N can correspond to the number of partitions of the first samples.

[0364] • The threshold value can be signaled in the bitstream.

[0365] • The value of the offset can be represented using M bits. For example, a typical value of M can be 16. 16 bits are used to represent each offset value.

[0366] o M can be adjustable and the value of M can be signaled in the bitstream. For example, depending on an indication obtained from the bitstream, the value of M can be 12 or 16.

[0367] Figure 18 An example implementation of the present disclosure is shown. Figure 18 An example implementation of the present disclosure is depicted. The samples of component 1 and the samples of component 2 are obtained using a synthesis transform. They can be obtained using different synthesis transforms. Afterwards, the samples of component 1 are fed as input to a determination unit. The determination unit determines whether the value of a sample is between thr n-1 and thr n . If it is determined to be so, offset n is added to the sample of component 2 to obtain a modified component 2. The sample of component 2 (second sample) and the sample of component 1 (first sample) can have a spatial relationship. For example, the first sample and the second sample can have the same spatial coordinates (x, y). Or there can be a relationship between the coordinates of the first sample and the coordinates of the second sample. For example, the coordinates of the second sample can be (x / 2, y / 2).

[0368] thr n-1 and thr n may be signaled in the bitstream. Or they can be calculated based on the minimum and maximum. The minimum can be the minimum of the samples (all samples or a group of samples) of component 1. Similarly, the maximum can be the maximum of the samples (all samples or a group of samples) of component 1.

[0369] The difference between the consecutive thresholds can be equal to (maximum - minimum) / N, where N is the number of offset values obtained from the bitstream. The number N can be obtained from the bitstream.

[0370] Finally, component 1 and component 2 are used to obtain the reconstructed image.

[0371] 5.2.3. Benefits of the examples

[0372] According to the examples, the correlation between the two components of the image can be more efficiently exploited, especially in case the first component and the second component are obtained using 2 different synthesis transforms. Therefore, the compression efficiency is significantly increased.

[0373] The present disclosure is not limited to the case where the two components are obtained using two different synthesis transforms. The components can be obtained using a single synthesis transform. In addition, the name “synthesis” transform does not limit the present disclosure either. Other names such as inverse transform or just transform generally refer to the same thing. The meaning of synthesis transform is a neural network used to convert a representation of an image from a transform domain to a pixel domain.

[0374] More details of embodiments of the present disclosure related to neural network based visual data coding will be described below. As used herein, the term “visual data” can refer to a video, an image, a picture in a video, or any other visual data suitable to be coded.

[0375] As discussed above, in existing designs of neural network (NN) based visual data coding, components of an image (e.g. luminance component and chrominance components) are processed with different synthesis subnetworks, the correlation between different components is not fully exploited. In other words, the information used to reconstruct a component can also be used to reconstruct another component. However, in existing designs, such cross-component information is not exploited.

[0376] To solve the above problems and some other problems not mentioned, a visual data processing solution as described below is disclosed. Embodiments of the present disclosure should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these embodiments can be applied individually or combined in any way.

[0377] Figure 19A flowchart of a method 1900 for visual data processing according to some embodiments of the present disclosure is shown. The method 1900 can be implemented during a conversion between visual data and a bitstream of the visual data, which is performed with a neural network (NN) based model. As used herein, a NN based model can be a model based on neural network technology. For example, a NN based model can specify a sequence of neural network modules (also referred to as an architecture) and model parameters. A neural network module can include a set of neural network layers. Each neural network layer specifies a tensor operation that receives and outputs a tensor, and each layer has trainable parameters. It should be understood that the possible implementations of the NN based model described herein are merely illustrative, and thus should not be construed as limiting the present disclosure in any way.

[0378] As shown in Figure 19 the method 1900 starts at 1902 by obtaining a set of adjusted first samples by adjusting first samples of a first component of the visual data with a set of offsets, each of the set of adjusted first samples corresponding to one of the set of offsets.

[0379] In some embodiments, each of the set of offsets can be utilized to adjust the first samples to obtain a corresponding adjusted first sample of the set of adjusted first samples. For example, a number of the set of offsets can be predetermined or indicated in the bitstream. In one example, the set of offsets can include only a single offset. Correspondingly, the set of adjusted first samples can include only a single adjusted first sample. Alternatively, the set of offsets can include multiple offsets. Correspondingly, the set of adjusted first samples can include multiple offsets. By way of example and not limitation, the set of offsets can include 8 offsets, and the set of adjusted first samples can include 8 adjusted first samples. It should be understood that the specific values recited herein are intended to be exemplary, and not limit the scope of the present disclosure.

[0380] At 1904, second samples of a second component of the visual data are adjusted based on at least one adjusted first sample. The second component is different from the first component. Further, the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample of the set of adjusted first samples with a threshold value. For example, each adjusted first sample of the at least one adjusted first sample can be greater than the threshold value. For example, the threshold value can be equal to a predetermined value, such as 0, and the like.

[0381] In some embodiments, a threshold function can be utilized to compare each adjusted first sample of the set of adjusted first samples with the threshold value. For example, the threshold function can be a rectified linear unit (ReLU) function, which is defined as follows:

[0382]

[0383] It can be seen that the output of the ReLU function equals the input of the ReLU function if the input of the ReLU function is greater than or equal to 0. In addition, the output of the ReLU function equals 0 if the input of the ReLU function is less than 0. In this case, when the ReLU function is applied to each of the set of adjusted first samples, the output corresponding to the adjusted first sample(s) that are less than or equal to 0 is set to equal 0, and the output corresponding to the adjusted first sample(s) that are greater than 0 is set to equal the adjusted first sample(s) themselves. In this case, the adjusted first sample(s) that are less than or equal to 0 are filtered out and will not affect the subsequent process. Only the adjusted first sample(s) that are greater than 0 will participate in the subsequent process and be considered as the at least one adjusted first sample determined from the set of adjusted first samples.

[0384] It should be appreciated that the threshold function can also be implemented as any other suitable function, such as a leaky ReLU operation, a sigmoid operation, a hyperbolic tangent operation, and the like. The scope of the present disclosure is not limited in this regard.

[0385] In one example, the second component can comprise a secondary component, and the first component can comprise a primary component. Alternatively, the second component can comprise a chroma component, and the first component can comprise a luminance component. In another example, the second component can comprise at least one of a U component or a V component, and the first component can comprise a Y component. It should be appreciated that the above examples are described for purposes of description only. The scope of the present disclosure is not limited in this regard.

[0386] In some embodiments, the second component and the first component can be reconstructed with at least one of the neural network-based models. For example, and without limitation, the synthesis transform can be a neural network for converting a latent representation of visual data from a transform domain to a pixel domain. In one example, the second component and / or the first component can be directly output by the at least one synthesis transform. Alternatively, the second component and / or the first component can be obtained by further processing an output of the at least one synthesis transform. In some embodiments, the at least one synthesis transform can comprise a first synthesis transform and a second synthesis transform different from the first synthesis transform. The second component can be reconstructed with the first synthesis transform, and the first component can be reconstructed with the second synthesis transform. In this case, the second component and the first component are independently reconstructed with the at least one synthesis transform.

[0387] At 1906, a conversion is performed based on the adjusted second samples. For example, and without limitation, the visual data can be reconstructed based on the adjusted second samples. In some embodiments, the conversion can comprise encoding the visual data into a bitstream. Additionally or alternatively, the conversion can comprise decoding the visual data from a bitstream. It should be understood that the above description is merely for the purpose of illustration. The scope of the present disclosure is not limited in this regard.

[0388] In view of the above, the second component of the visual data is adjusted based on the first component. In contrast to conventional solutions where the first and second components are processed independently, the proposed approach can advantageously exploit cross-component information to improve the quality of the reconstructed visual data, whereby the coding quality can be improved.

[0389] In some embodiments, at 1904, the adjustment term can be determined based on a result of weighting the at least one adjusted first sample. For example, if the at least one adjusted first sample comprises only a single adjusted first sample, the adjustment term can be equal to a result of weighting the single adjusted first sample. If the at least one adjusted first sample comprises a plurality of adjusted first samples, the adjustment term can be equal to a weighted sum of the plurality of adjusted first samples.

[0390] For example, and without limitation, the adjustment term can be determined based on the following formula:

[0391]

[0392] where recY(c, x, y) denotes a first sample having a channel index c and coordinates (x, y), b[n] denotes one offset of a set of offsets having an index n, W[n] denotes one weight of a set of weights having an index n, RELU() denotes a ReLU function, the index n ranges from 0 to M, and M can be equal to the number of the set of offsets. In this case, the result of (recY(c, x, y) + b[n] can correspond to a set of adjusted first samples. Based on the above description regarding the ReLU function, the adjusted first sample(s) that are less than or equal to 0 are filtered out and will not affect the subsequent process. The result of the summation function ∑ is equal to the weighted sum of the adjusted first samples that are greater than 0.

[0393] In some embodiments, the at least one weight used to weight the at least one adjusted first sample can be obtained from information indicated in one or more bitstreams. In one example, the at least one weight can be indicated in the bitstream itself. Alternatively, the at least one weight can be determined based on one or more parameters indicated in the bitstream.

[0394] In some embodiments, the adjustment term can be determined utilizing at least one convolutional layer in the neural network based model. This facilitates implementation of the proposed solution utilizing the most basic neural network layer(s).

[0395] Furthermore, the second sample can be adjusted based on the adjustment term. For example, and without limitation, the second sample can be adjusted by adding the adjustment term to the second sample. In some embodiments, the second sample is only adjusted by adding the adjustment term to the second sample when the relevant coding tool is enabled. For example, the second sample can be adjusted in a non-linear filtering process. In this case, the second sample is only adjusted by adding the adjustment term to the second sample when the non-linear filtering process is enabled. This brings greater flexibility to the implementation of the proposed solution.

[0396] In some embodiments, the adjustment of the first sample at 1902 can be performed utilizing one or more convolutional layers in the neural network based model. For example, the set of offsets can be implemented as bias values of the one or more convolutional layers. This facilitates implementation of the proposed solution utilizing the most basic neural network layer(s).

[0397] In some embodiments, the set of offsets can be determined based on at least one of a maximum value or a minimum value. For example, and without limitation, the maximum value can be a maximum value of a set of samples of the first component, and / or the minimum value can be a minimum value of the set of samples of the first component. In one example embodiment, the set of samples can comprise all samples of the first component. That is, the maximum value can be a global maximum value, and / or the minimum value can be a global minimum value.

[0398] In another example embodiment, the set of samples can only comprise partial samples of the first component. In other words, the maximum value can be a local maximum value, and / or the minimum value can be a local minimum value. For example, the first component can be divided into a plurality of tiles, and the set of samples can comprise all samples of one tile of the plurality of tiles. For example, the tile can be a rectangular sub-block of the corresponding component. It should be appreciated that the tile can also be any other suitable shape.

[0399] In some embodiments, a first set of offsets can be used for adjusting at least one sample of a first tile of the plurality of tiles, a second set of offsets can be used for adjusting at least one sample of a second tile of the plurality of tiles, and the first set of offsets can be different from the second set of offsets. That is, different offsets can be used for different tiles. In this way, the coding process can adapt to the content of the visual data, whereby the coding quality can be improved.

[0400] In some embodiments, one or more offsets of the set of offsets can be determined based on a difference between the maximum value and the minimum value. For ease of discussion, a first offset of the set of offsets will be taken as an example. For example, the first offset can be determined based on a division result of dividing the difference by a number of the set of offsets. Additionally, the first offset can be determined based on a product of the division result and an index of the first offset. Furthermore, the first offset can be determined based on a sum of the product and the minimum value.

[0401] For example, but not by way of limitation, the first offset can be determined based on the following equation:

[0402]

[0403] where max denotes the maximum value, min denotes the minimum value, N denotes a number of the set of offsets, and n denotes an index of the first offset. For example, but not by way of limitation, N can be 8, and n can be in a range of 0 to 7.

[0404] In some embodiments, at least one of the maximum value or the minimum value can be indicated in the bitstream. Additionally or alternatively, the set of offsets can be indicated in the bitstream.

[0405] In some embodiments, the second component can comprise two components (such as a U component and a V component), and two sets of weights can be obtained from the information indicated in the bitstream and used to adjust the two components respectively. That is, different components can be processed based on different weights. In this way, the coding process can adapt to the content of the visual data, whereby the coding quality can be improved.

[0406] In view of the above, the solution according to some embodiments of the present disclosure can advantageously utilize the cross-component information to improve the quality of the reconstructed visual data, whereby the coding quality can be improved.

[0407] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data processing. In the method, a set of adjusted first samples is obtained by adjusting first samples of a first component of the visual data with a set of offsets. Each adjusted first sample of the set of adjusted first samples corresponds to an offset of the set of offsets. Additionally, second samples of a second component of the visual data are adjusted based on at least one adjusted first sample. The at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample of the set of adjusted first samples with a threshold value, and the second component is different from the first component. Furthermore, the bitstream is generated based on the adjusted second samples with a neural network (NN) based model.

[0408] According to yet some embodiments of the disclosure, a method for storing a bitstream of visual data is provided. In the method, a set of adjusted first samples is obtained by adjusting first samples of a first component of the visual data with a set of offsets. Each adjusted first sample in the set of adjusted first samples corresponds to an offset in the set of offsets. In addition, second samples of a second component of the visual data are adjusted based on at least one adjusted first sample. The at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold value, and the second component is different from the first component. Furthermore, the bitstream is generated based on the adjusted second samples with a neural network (NN) based model, and is stored in a non-transitory computer readable recording medium.

[0409] Implementations of the disclosure can be described according to the following clauses, which features can be combined in any reasonable manner.

[0410] Clause 1. A method for visual data processing, comprising, for a conversion between visual data with a neural network (NN) based model and one or more bitstreams of the visual data, obtaining a set of adjusted first samples by adjusting first samples of a first component of the visual data with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting second samples of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold value, and the second component is different from the first component; and performing the conversion based on the adjusted second samples.

[0411] Clause 2. The method of clause 1, wherein the threshold value is equal to 0.

[0412] Clause 3. The method of any of clauses 1-2, wherein a threshold function is used to compare each adjusted first sample in the set of adjusted first samples with the threshold value.

[0413] Clause 4. The method of clause 3, wherein the threshold function comprises a rectified linear unit (ReLU) function.

[0414] Clause 5. The method of any of clauses 1-4, wherein each adjusted first sample in the at least one adjusted first sample is greater than the threshold value.

[0415] Item 6. The method of any of items 1-5, wherein adjusting the second sample comprises determining an adjustment term based on a result of weighting the at least one adjusted first sample, and adjusting the second sample based on the adjustment term.

[0416] Item 7. The method of item 6, wherein if the at least one adjusted first sample comprises a single adjusted first sample, the adjustment term is equal to a result of weighting the single adjusted first sample, or if the at least one adjusted first sample comprises a plurality of adjusted first samples, the adjustment term is equal to a weighted sum of the plurality of adjusted first samples.

[0417] Item 8. The method of any of items 6-7, wherein the adjustment term is determined based on:

[0418]

[0419] where recY(c, x, y) represents the first sample with channel index c and coordinates (x, y), b[n] represents one offset of the set of offsets with index n, W[n] represents one weight of the set of weights with index n, RELU() represents a ReLU function, the index n ranges from 0 to M, and M is equal to a number of the set of offsets.

[0420] Item 9. The method of any of items 6-8, wherein the second sample is adjusted by adding the adjustment term to the second sample.

[0421] Item 10. The method of any of items 6-9, wherein at least one weight used for weighting the at least one adjusted first sample is obtained from information indicated in the one or more bitstreams.

[0422] Item 11. The method of any of items 6-10, wherein the adjustment term is determined with at least one convolutional layer in the NN-based model.

[0423] Item 12. The method of any of items 1-11, wherein the adjustment of the first sample is performed with one or more convolutional layers in the NN-based model.

[0424] Item 13. The method of item 12, wherein the set of offsets is implemented as bias values of the one or more convolutional layers.

[0425] Item 14. The method of any of items 1-13, wherein the set of offsets is determined based on at least one of a maximum value or a minimum value.

[0426] Item 15. The method of item 14, wherein the maximum value is a maximum value of a set of samples of the first component or the minimum value is a minimum value of the set of samples.

[0427] Item 16. The method of any of items l-15, wherein a number of the set of offsets is predetermined or indicated in the bitstream.

[0428] Item 17. The method of any of items 14-16, wherein a first offset of the set of offsets is determined based on a difference between the maximum value and the minimum value.

[0429] Item 18. The method of item 17, wherein the first offset is determined based on a division result of dividing the difference by a number of the set of offsets.

[0430] Item 19. The method of item 18, wherein the first offset is determined based on a product of the division result and an index of the first offset.

[0431] Item 20. The method of item 19, wherein the first offset is determined based on a sum of the product and the minimum value.

[0432] Item 21. The method of any of items 15-20, wherein the set of samples includes all samples of the first component.

[0433] Item 22. The method of any of items 15-20, wherein the set of samples includes partial samples of the first component.

[0434] Item 23. The method of any of items 15-20, wherein the first component is divided into a plurality of tiles.

[0435] Item 24. The method of item 23, wherein the set of samples includes all samples of a tile of the plurality of tiles.

[0436] Item 25. The method of any of items 23-24, wherein a first set of offsets is used to adjust at least one sample of a first tile of the plurality of tiles, a second set of offsets is used to adjust at least one sample of a second tile of the plurality of tiles, and the first set of offsets is different from the second set of offsets.

[0437] Item 26. The method of any of items 14-25, wherein at least one of the maximum value or the minimum value is indicated in the bitstream.

[0438] Item 27. The method of any of items 1-13, wherein the set of offsets is indicated in the bitstream.

[0439] Item 28. The method of any of items 1-27, wherein the set of offsets comprises a plurality of offsets.

[0440] Item 29. The method of any of items l-28, wherein the second component comprises two components, and two sets of weights are obtained from information indicated in the bitstream and used to adjust the two components, respectively.

[0441] Item 30. The method of any of items 1-29, wherein the first component and the second component are reconstructed with at least one synthesis transform in the NN-based model.

[0442] Item 31. The method of item 30, wherein the at least one synthesis transform comprises a first synthesis transform and a second synthesis transform different from the first synthesis transform, the first component is reconstructed with the first synthesis transform, and the second component is reconstructed with the second synthesis transform.

[0443] Item 32. The method of any of items 1-31, wherein the first component comprises a primary component, and the second component comprises a secondary component, or wherein the first component comprises a luma component, and the second component comprises a chroma component, or wherein the first component comprises a Y component, and the second component comprises at least one of a U component or a V component.

[0444] Item 33. The method of any of items 1-32, wherein performing the conversion comprises reconstructing the visual data based on the adjusted second samples.

[0445] Item 34. The method of any of items 1-33, wherein obtaining the set of adjusted first samples comprises adjusting the first samples with each offset in the set of offsets to obtain a corresponding adjusted first sample in the set of adjusted first samples.

[0446] Item 35. The method of any of items 1-34, wherein the second samples are adjusted in a non-linear filtering process.

[0447] Item 36. The method of any of items l-35, wherein the visual data comprises a video, a picture of the video, or an image.

[0448] Item 37. The method of any of items 1-36, wherein the conversion comprises encoding the visual data into the one or more bitstreams.

[0449] Item 38. The method of any of items 1-36, wherein the converting comprises decoding the visual data from the one or more bitstreams.

[0450] Item 39. An apparatus for visual data processing comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of items 1-38.

[0451] Item 40. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of items 1-38.

[0452] Item 41. A non-transitory computer-readable recording medium storing a bitstream of visual data, the bitstream of visual data being generated by a method performed by an apparatus for visual data processing, wherein the method comprises: obtaining a set of adjusted first samples of a first component of the visual data by adjusting first samples of the first component of the visual data with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting a second sample of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold, and the second component is different from the first component; and generating the bitstream based on the adjusted second sample with a neural network (NN) based model.

[0453] Item 42. A method for storing a bitstream of visual data, comprising: obtaining a set of adjusted first samples of a first component of the visual data by adjusting first samples of the first component of the visual data with a set of offsets, each adjusted first sample in the set of adjusted first samples corresponding to an offset in the set of offsets; adjusting a second sample of a second component of the visual data based on at least one adjusted first sample, wherein the at least one adjusted first sample is determined from the set of adjusted first samples by comparing each adjusted first sample in the set of adjusted first samples with a threshold, and the second component is different from the first component; generating the bitstream based on the adjusted second sample with a neural network (NN) based model; and storing the bitstream in a non-transitory computer-readable recording medium.

[0454] Example device

[0455] Figure 20A block diagram of a computing device 2000 in which various embodiments of the present disclosure can be implemented is shown. The computing device 2000 can be implemented as the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124), or can be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).

[0456] It should be understood that Figure 20 The computing device 2000 shown in FIG. 13 is for purposes of illustration and explanation only and is not intended to be limiting of the functionality and scope of the embodiments of the present disclosure in any way.

[0457] As Figure 20 shown, the computing device 2000 includes a general-purpose computing device 2000. The computing device 2000 can include at least one or more processors or processing units 2010, a memory 2020, a storage unit 2030, one or more communication units 2040, one or more input devices 2050, and one or more output devices 2060.

[0458] In some embodiments, the computing device 2000 can be implemented as any user terminal or server terminal having computing capability. The server terminal can be a server provided by a service provider, a mainframe computing device, or the like. The user terminal may, for example, be any type of mobile terminal, fixed terminal, or portable terminal including a mobile telephone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, and includes accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 2000 can support any type of interface to the user (such as "wearable" circuitry, etc.).

[0459] The processing unit 2010 can be a physical processor or a virtual processor and can implement various processing based on programs stored in the memory 2020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 2000. The processing unit 2010 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0460] The computing device 2000 typically includes a variety of computer storage media. Such media can be any media that is accessible by the computing device 2000 and can include, without limitation, both volatile and non-volatile media, or removable and non-removable media. The memory 2020 can be volatile (such as register, cache, and / or random access memory (RAM)), non-volatile (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage 2030 can be additional, removable, or non-removable, and can include, but is not limited to, machine-readable media such as memory, flash drives, disks, or any other storage medium that can be used to store information and / or visual data and that can be accessed by the computing device 2000.

[0461] The computing device 2000 can also include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 20 , a disk drive for reading from and / or writing to a removable, non-removable, and / or non-volatile media, such as a disk drive, and an optical disk drive for reading from and / or writing to a removable, non-removable, and / or non-volatile media such as an optical disk, can be provided. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces.

[0462] The communication unit 2040 communicates with another computing device via a communication medium. In addition, the functionality of the components of the computing device 2000 can be implemented by a single computing cluster or multiple computer machines in communication via a communication connection. Thus, the computing device 2000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other common network nodes.

[0463] The input device 2050 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, or the like. The output device 2060 can be one or more of various output devices, such as a display, speakers, a printer, and the like. By means of the communication unit 2040, the computing device 2000 can also communicate with one or more external devices (not shown), such as a storage device or a display device, which enable a user to interact with the computing device 2000, or any device (e.g., a network card, a modem, etc.) that enables the computing device 2000 to communicate with one or more other computing devices. Such communication can occur via an input / output (I / O) interface (not shown).

[0464] In some embodiments, some or all of the components of computing device 2000 can not be integrated in a single device, but can also be arranged in a cloud computing architecture. In a cloud computing architecture, the components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computation, software, data access, and storage services that do not require end-user knowledge of the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing delivers services via the internet using appropriate protocols. For example, cloud computing providers deliver applications via the internet from a remote location that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and corresponding visual data can be stored on servers at remote locations. Computing resources in a cloud computing environment can be consolidated or distributed at locations within a remote visual data center. Cloud computing infrastructure can provide services through shared visual data centers, although they appear as a single point of access for the user. Thus, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, they can be provided from a conventional server or installed directly or otherwise on a client device.

[0465] In embodiments of the disclosure, computing device 2000 can be used to implement visual data encoding / decoding. Memory 2020 can include one or more visual data codec modules 2025 having one or more program instructions. These modules can be accessed and executed by processing unit 2010 to perform the functions of the various embodiments described herein.

[0466] In example embodiments performing visual data encoding, input device 2050 can receive visual data as input 2070 to be encoded. The visual data can be processed by, for example, visual data codec module 2025 to generate an encoded bitstream. The encoded bitstream can be provided as output 2080 via output device 2060.

[0467] In example embodiments performing visual data decoding, input device 2050 can receive an encoded bitstream as input 2070. The encoded bitstream can be processed by, for example, visual data codec module 2025 to generate decoded visual data. The decoded visual data can be provided as output 2080 via output device 2060.

[0468] While the disclosure has been particularly shown and described with reference to the preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such changes are intended to fall within the scope of the application. Accordingly, the foregoing description of embodiments of the application are not intended to be limiting.

Claims

1. A method for visual data processing, comprising: For the conversion between visual data using a neural network (NN)-based model and one or more bitstreams of the visual data, a set of adjusted first points is obtained by adjusting the first point of a first component of the visual data using a set of offsets, each of the adjusted first points in the set of adjusted first points corresponding to one of the offsets in the set of offsets. Based on at least one adjusted first sample point, the second sample point of the second component of the visual data is adjusted, wherein the at least one adjusted first sample point is determined from the set of adjusted first sample points by comparing each adjusted first sample point in the set of adjusted first sample points with a threshold, and the second component is different from the first component; as well as The transformation is performed based on the adjusted second sample points.

2. The method according to claim 1, wherein the threshold is equal to 0.

3. The method according to any one of claims 1-2, wherein a threshold function is used to compare each of the adjusted first points in the set of adjusted first points with the threshold.

4. The method of claim 3, wherein the threshold function comprises a modified linear unit (ReLU) function.

5. The method according to any one of claims 1-4, wherein each of the at least one adjusted first common point is greater than the threshold.

6. The method according to any one of claims 1-5, wherein adjusting the second sample point comprises: The adjustment term is determined based on the weighted result of the at least one adjusted first point; as well as The second sample point is adjusted based on the aforementioned adjustment terms.

7. The method of claim 6, wherein if the at least one adjusted first common point includes a single adjusted first common point, then the adjustment term is equal to the weighted result of the single adjusted first common point, or If the at least one adjusted first identical point includes multiple adjusted first identical points, then the adjustment term is equal to the weighted sum of the multiple adjusted first identical points.

8. The method according to any one of claims 6-7, wherein the adjustment item is determined based on the following: Where recY(c,x,y) represents the first sample point with channel index c and coordinates (x,y), b[n] represents an offset with index n in the set of offsets, W[n] represents a weight with index n in the weights, RELU() represents the ReLU function, the index n ranges from 0 to M, and M is equal to the number of offsets in the set.

9. The method according to any one of claims 6-8, wherein the second sample is adjusted by adding the adjustment term to the second sample.

10. The method according to any one of claims 6-9, wherein information indicating at least one weight used to weight the at least one adjusted first point is obtained from the one or more bitstreams.

11. The method according to any one of claims 6-10, wherein the adjustment term is determined using at least one convolutional layer in the NN-based model.

12. The method according to any one of claims 1-11, wherein the adjustment of the first point is performed using one or more convolutional layers in the NN-based model.

13. The method of claim 12, wherein the set of offsets is implemented as bias values ​​of the one or more convolutional layers.

14. The method according to any one of claims 1-13, wherein the set of offsets is determined based on at least one of a maximum value or a minimum value.

15. The method of claim 14, wherein the maximum value is the maximum value of the sample set of the first component, or the minimum value is the minimum value of the sample set.

16. The method according to any one of claims 1-15, wherein the number of said set of offsets is predetermined or indicated in the bit stream.

17. The method according to any one of claims 14-16, wherein the first offset in the set of offsets is determined based on the difference between the maximum value and the minimum value.

18. The method of claim 17, wherein the first offset is determined based on the result of a division of the difference by the number of the set of offsets.

19. The method of claim 18, wherein the first offset is determined based on the product of the division result and the index of the first offset.

20. The method of claim 19, wherein the first offset is determined based on the sum of the product and the minimum value.

21. The method according to any one of claims 15-20, wherein the sample set includes all samples of the first component.

22. The method according to any one of claims 15-20, wherein the sample set includes a portion of the samples of the first component.

23. The method according to any one of claims 15-20, wherein the first component is divided into a plurality of slices.

24. The method of claim 23, wherein the sample set comprises all samples from one of the plurality of slices.

25. The method according to any one of claims 23-24, wherein a first set of offsets is used to adjust at least one sample point of a first piece of the plurality of pieces, a second set of offsets is used to adjust at least one sample point of a second piece of the plurality of pieces, and the first set of offsets is different from the second set of offsets.

26. The method according to any one of claims 14-25, wherein at least one of the maximum value or the minimum value is indicated in the bit stream.

27. The method according to any one of claims 1-13, wherein the set of offsets is indicated in the bit stream.

28. The method according to any one of claims 1-27, wherein the set of offsets comprises a plurality of offsets.

29. The method according to any one of claims 1-28, wherein the second component comprises two components, and two sets of weights are obtained from information indicated in the bitstream and used to adjust the two components respectively.

30. The method according to any one of claims 1-29, wherein the first component and the second component are reconstructed using at least one synthetic transformation in the NN-based model.

31. The method of claim 30, wherein the at least one synthetic transformation comprises a first synthetic transformation and a second synthetic transformation different from the first synthetic transformation, the first component being reconstructed using the first synthetic transformation, and the second component being reconstructed using the second synthetic transformation.

32. The method according to any one of claims 1-31, wherein the first component comprises a primary component, and the second component comprises a secondary component, or The first component includes a luminance component, and the second component includes a chromaticity component, or The first component includes a Y component, and the second component includes at least one of a U component or a V component.

33. The method according to any one of claims 1-32, wherein performing the conversion comprises: The visual data is reconstructed based on the adjusted second sample points.

34. The method according to any one of claims 1-33, wherein obtaining the set of adjusted first similarities comprises: The first sample point is adjusted using each offset in the set of offsets to obtain the corresponding adjusted first sample point in the set of adjusted first sample points.

35. The method according to any one of claims 1-34, wherein the second sample point is adjusted during the nonlinear filtering process.

36. The method according to any one of claims 1-35, wherein the visual data includes video, a picture of the video, or an image.

37. The method of any one of claims 1-36, wherein the conversion comprises encoding the visual data into the one or more bitstreams.

38. The method according to any one of claims 1-36, wherein the conversion comprises decoding the visual data from the one or more bitstreams.

39. An apparatus for visual data processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-38.

40. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-38.

41. A non-transitory computer-readable recording medium storing a bitstream of visual data, the bitstream of visual data being generated by a method performed by means of means for visual data processing, wherein the method includes: A set of adjusted first identical points is obtained by adjusting the first identical point of the first component of the visual data using a set of offsets, each of the adjusted first identical points in the set of offsets corresponding to one of the offsets in the set of offsets; Based on at least one adjusted first sample point, the second sample point of the second component of the visual data is adjusted, wherein the at least one adjusted first sample point is determined from the set of adjusted first sample points by comparing each adjusted first sample point in the set of adjusted first sample points with a threshold, and the second component is different from the first component; as well as The bitstream is generated using a neural network (NN)-based model based on the adjusted second sample points.

42. A method for storing a bitstream of visual data, comprising: A set of adjusted first identical points is obtained by adjusting the first identical point of the first component of the visual data using a set of offsets, each of the adjusted first identical points in the set of offsets corresponding to one of the offsets in the set of offsets; The second sample point of the second component of the visual data is adjusted based on at least one adjusted first sample point, wherein the at least one adjusted first sample point is determined from the set of adjusted first sample points by comparing each adjusted first sample point in the set of adjusted first sample points with a threshold, and the second component is different from the first component; The bitstream is generated using a neural network (NN)-based model based on the adjusted second sample points; as well as The bitstream is stored in a non-transitory computer-readable recording medium.