Method and device for transcoding visual data and medium

By using a neural network-based visual data transcoding method and apparatus, the compatibility problem between different codecs is solved, enabling a wider range of application scenarios.

CN121925838APending Publication Date: 2026-04-24DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480062657.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-29
Filing Date
2024-09-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing neural network-based image/video transcoding technologies suffer from compatibility issues and cannot effectively support transcoding between different codecs, thus limiting their application scenarios.

Method used

A visual data transcoding method and apparatus based on neural networks is proposed. By reconstructing and encoding the bitstream of visual data, a neural network-based model is used for reconstruction or encoding, supporting compatibility with different encoding and decoding schemes.

Benefits of technology

It improves the compatibility of visual data transcoding and expands the application scenarios of neural network-based encoding and decoding schemes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925838A_ABST
    Figure CN121925838A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a solution for transcoding visual data. A method for transcoding visual data is presented. The method comprises: reconstructing visual data from a first bitstream of visual data; and encoding the reconstructed visual data as a second bitstream of visual data, at least one of the reconstruction or the encoding being performed using a neural network (NN)-based model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to visual data transcoding techniques, and more specifically, to neural network-based visual data transcoding. Background Technology

[0002] Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression technology has made significant strides in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved rate-distortion (RD) performance comparable to Multifunctional Video Coding (VVC). With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, neural network-based image / video transcoding is generally expected to be studied. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for visual data transcoding.

[0004] In a first aspect, a method for visual data transcoding is proposed. The method includes: reconstructing visual data from a first bitstream of visual data; and encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN) based model.

[0005] Based on the method according to the first aspect of this disclosure, the reconstruction and / or encoding steps involved in visual data transcoding are performed using a neural network (NN)-based model. Compared to conventional solutions that do not support NN-based visual data encoding / decoding schemes in transcoding, the proposed method can advantageously support NN-based visual data encoding / decoding schemes, thereby improving the compatibility of visual data transcoding. In this way, the application scenarios of NN-based visual data encoding / decoding schemes are also expanded.

[0006] In a second aspect, an apparatus for visual data transcoding is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data transcoding, wherein the method includes: reconstructing the visual data from a first bitstream of visual data; and encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN)-based model.

[0009] In a fifth aspect, a method for storing bitstreams of visual data is proposed. The method includes: reconstructing visual data from a first bitstream of visual data; encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN)-based model; and storing the second bitstream in a non-transitory computer-readable recording medium.

[0010] This summary is provided to present, in a simplified form, the selected concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1A A block diagram illustrating an exemplary visual data encoding / decoding system according to some embodiments of the present disclosure is shown; Figure 1B This is a schematic diagram illustrating an example transform encoding / decoding scheme; Figure 2 An exemplary potential representation of the image is shown; Figure 3 This is a schematic diagram illustrating an example autoencoder that implements a hyperprior model; Figure 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model together with a super-prior and an autoencoder; Figure 5 An exemplary encoding process is shown; Figure 6 An exemplary decoding process is shown; Figure 7 A learning-based image codec architecture is shown; Figure 8 An example Joint Image Experts Group (JPEG) Artificial Intelligence (AI) decoder structure is shown; Figure 9 An example JPEG AI decoder for a single component is shown; Figure 10 Examples of transcoding for different resolutions, bit rates, or compatible codecs available on playback devices are shown; Figure 11 This illustrates an example scenario where a JPEG-encoded file cannot be decoded on a device with a JPEG AI image codec. Figure 12 Example transcoding schemes for transcoding JPEG-encoded files into JPEGAI-decoded files according to some embodiments of this disclosure are shown; Figure 13 An example transcoding scheme using filtering before a second decoder is shown according to some embodiments of this disclosure; Figure 14 An example transcoding scheme using at least two encoders with a second codec is shown according to some embodiments of the present disclosure; Figure 15 A flowchart of a method for visual data transcoding according to embodiments of the present disclosure is shown; and Figure 16 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0013] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation

[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0015] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0016] The terms "an embodiment," "embodiment," "exemplary embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an exemplary embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0017] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0019] Example Environment Figure 1A This is a block diagram illustrating an exemplary visual data encoding / decoding system 100 from which the techniques of this disclosure can be utilized. As shown, the visual data encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a visual data encoding device, and the destination device 120 may also be referred to as a visual data decoding device. In operation, the source device 110 may be configured to generate encoded visual data, and the destination device 120 may be configured to decode the encoded visual data generated by the source device 110. The source device 110 may include a visual data source 112, a visual data encoder 114, and an input / output (I / O) interface 116.

[0020] Visual data source 112 may include sources such as visual data acquisition devices. Examples of visual data acquisition devices include, but are not limited to, interfaces for receiving visual data from visual data content providers, computer graphics systems for generating visual data, and / or combinations thereof.

[0021] Visual data may include one or more images. A visual data encoder 114 encodes the visual data from a visual data source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the visual data. The bitstream may include encoded images and associated data. The encoded image is an encoded representation of an image. The associated data may include a sequence parameter set, an image parameter set, and other syntax structures. An I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded visual data can be directly transmitted to a destination device 120 via network 130A through the I / O interface 116. The encoded visual data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0022] The destination device 120 may include an I / O interface 126, a visual data decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded visual data from the source device 110 or the storage medium / server 130B. The visual data decoder 124 may decode the encoded visual data. The display device 122 may display the decoded visual data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0023] The visual data encoder 114 and the visual data decoder 124 can operate according to visual data encoding and decoding standards, such as video encoding and decoding standards or still image encoding and decoding standards and other existing and / or additional standards.

[0024] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-functional video codecs or other specific visual data codecs, the disclosed techniques are also applicable to other codec techniques. Additionally, although some embodiments describe encoding steps in detail, it should be understood that the corresponding decoding steps of the inversion encoding will be implemented by the decoder. Furthermore, the term "visual data processing" includes visual data encoding or compression, visual data decoding or decompression, and visual data transcoding, wherein visual data is represented from one compressed format to another or at a different compression bit rate.

[0025] 1. Brief Overview This disclosure describes a neural network-based image and video compression transcoding method. An image or video codec (i.e., a target codec) comprises multiple neural networks, where an encoder encodes the original image / video into a bitstream, and a decoder reconstructs the image / video. However, it is impossible for a codec to decode a bitstream file encoded using another codec. This paper proposes a transcoding scheme to support decoding using a second codec when a bitstream is generated using a first codec.

[0026] 2. Background Over the past decade, deep learning has made rapid progress across various fields, particularly in computer vision and image processing. Inspired by the tremendous success of deep learning in computer vision, many researchers have shifted their focus from traditional image / video compression techniques to neural image / video compression. Neural networks were initially invented through interdisciplinary research in neuroscience and mathematics. They have demonstrated powerful capabilities in the context of nonlinear transformations and classification. Neural network-based image / video compression techniques have made significant progress in the past five years. It has been reported that the latest neural network-based image compression algorithms have achieved RD performance comparable to Multifunctional Video Coding (VVC), the latest video codec standard developed by the Joint Video Experts Group (JVET), comprised of experts from MPEG and VCEG. With the continuous improvement in the performance of neural image compression, neural network-based video compression has become an actively developing research area. However, due to the inherent difficulty of the problem, neural network-based video coding and decoding is still in its early stages.

[0027] 2.1. Image / Video Compression Image / video compression generally refers to the computational technique of compressing images / videos into binary codes for convenient storage and transmission. Binary codes may or may not support lossless reconstruction of the original image / video, known as lossless compression and lossy compression. Most efforts focus on lossy compression because lossless reconstruction is not always necessary. The performance of image / video compression algorithms is typically evaluated from two aspects: compression ratio and reconstruction quality. The compression ratio is directly related to the number of binary codes; fewer is better. Reconstruction quality is measured by comparing the reconstructed image / video with the original image / video; higher is better.

[0028] Image / video compression techniques can be divided into two branches: classical video encoding / decoding methods and neural network-based video compression methods. Classical video encoding / decoding schemes employ transform-based solutions, where researchers utilize statistical dependencies in latent variables (e.g., DCT or wavelet coefficients) by carefully hand-designing entropy encoding / decoding to model dependencies in the quantization domain. Neural network-based video compression takes two forms: neural network-based encoding / decoding tools and end-to-end neural network-based video compression. The former is embedded as an encoding / decoding tool within existing classical video codecs and exists only as part of the framework, while the latter is a separate framework developed based on neural networks, independent of classical video codecs.

[0029] Over the past three decades, a series of classic video codec standards have been developed to accommodate the ever-growing volume of visual content. The International Organization for Standardization (ISO / IEC) has two expert groups: the Joint Group of Picture Experts (JPEG) and the Moving Picture Experts Group (MPEG). The ITU-T also has its own Video Codec Experts Group (VCEG) for standardizing image / video codec technologies. Influential video codec standards released by these organizations include JPEG, JPEG 2000, H.262, H.264 / AVC, and H.265 / HEVC. Following H.265 / HEVC, the Joint Video Experts Group (JVET), comprised of MPEG and VCEG, has been working on a new video codec standard: Multi-Functional Video Codec (VVC). The first version of VVC was released in July 2020. Compared to HEVC, VVC achieves an average bitrate reduction of 50% while maintaining the same visual quality.

[0030] Neural network-based image / video compression is not a new solution, as many researchers have worked on neural network-based image encoding and decoding. However, the network architectures are relatively shallow, and the performance is unsatisfactory. Thanks to the support of abundant data and powerful computing resources, neural network-based methods have been better utilized in various applications. Currently, neural network-based image / video compression has shown promising improvements, confirming its feasibility. However, the technology is still far from mature and many challenges need to be addressed.

[0031] 2.2. Neural Networks Neural networks, also known as artificial neural networks (ANNs), are computational models used in machine learning techniques. They typically consist of multiple processing layers, each composed of several simple but non-linear basic computational units. One advantage of these deep networks is their ability to process data with multiple levels of abstraction and transform it into different kinds of representations. Notably, these representations are not manually designed; instead, deep networks, including processing layers, are learned from large amounts of data using general machine learning procedures. Deep learning eliminates the need for handcrafted representations and is therefore considered particularly suitable for processing native unstructured data, such as acoustic and visual signals, which has been a long-standing challenge in the field of artificial intelligence.

[0032] 2.3. Neural Networks for Image Compression Existing neural networks used for image compression methods can be divided into two categories: pixel probability modeling and autoencoders. The former belongs to predictive encoding / decoding strategies, while the latter is a transform-based solution. Sometimes, these two methods are combined in the literature.

[0033] 2.3.1. Pixel Probability Modeling According to Shannon's information theory, the optimal method for lossless encoding and decoding can achieve the lowest possible decoding rate. ,in It is a symbol The probability of [the lossless encoding / decoding method]. Many lossless encoding / decoding methods have been proposed in the literature, among which arithmetic encoding / decoding is considered one of the best methods. Given a probability distribution... Arithmetic encoding and decoding ensures that the encoding / decoding rate is as close as possible to its theoretical limit without considering rounding errors. Therefore, the remaining problem is how to determine the probability, which is very challenging for natural images / videos due to the curse of dimensionality.

[0034] Following the predictive encoding / decoding strategy, for One approach to modeling this is to predict pixel probabilities one by one in raster scan order based on previous observations. It's an image.

[0035] (1) in These are the height and width of the image, respectively. Previous observations are also referred to as the current pixel's... Context When the image is large, estimating the conditional probability can be difficult, so a simplified approach is to limit the scope of its context.

[0036] (2) in It is a predefined constant that controls the scope of the context.

[0037] It should be noted that this condition can also take into account the sample values ​​of other color components. For example, when encoding and decoding RGB color components, the R sample depends on previously encoded and decoded pixels (including R / G / B samples). The current G sample can be encoded and decoded based on previously encoded and decoded pixels and the current R sample. For encoding and decoding the current B sample, previously encoded and decoded pixels as well as the current R sample and the current G sample can also be considered.

[0038] Neural networks were initially introduced for computer vision tasks and have proven effective in regression and classification problems. Therefore, it has been proposed to use neural networks to adjust their behavior based on their context. Estimated probability A pixel probability was proposed for binary images, namely... The Neural Autoregressive Distribution Estimator (NADE) is designed for pixel probability modeling, where it is a feedforward network with a single hidden layer. A similar work has been proposed in [8], where the feedforward network also has connections that skip hidden layers and the parameters are shared. Experiments have been conducted on the binarized MNIST dataset. NADE is extended to a real-valued model RNADE, where the probability is... It is derived using Gaussian mixture theory. Their feedforward network also has a hidden layer, but the hidden layer is rescaled to avoid saturation and uses a modified linear unit (ReLU) instead of a sigmoid. NADE and RNADE are improved by reorganizing the order of pixels and using a deeper neural network.

[0039] Designing advanced neural networks plays a crucial role in improving pixel probabilistic modeling. Multidimensional Long Short-Term Memory (LSTM) was proposed, which, along with a mixture of conditional Gaussian scaling mixtures, is used for probabilistic modeling. LSTM is a special type of recurrent neural network (RNN) that has proven adept at modeling sequential data. Spatial variants of LSTM have later been applied to images. Several different neural networks, including RNNs and CNNs, namely PixelRNN and PixelCNN, were investigated. In PixelRNN, two variants of LSTM were proposed, called row LSTM and diagonal BiLSTM, the latter specifically designed for images. PixelRNN incorporates residual connections to aid in training deep neural networks with up to 12 layers. In PixelCNN, [the following is a separate section, likely related to image modeling]. maskConvolutions are used to fit the shape of the context. Compared to previous work, PixelRNN and PixelCNN focus more on natural images: they treat pixels as discrete values ​​(e.g., 0, 1, ..., 255) and predict a multinomial distribution over these discrete values; they handle color images in the RGB color space; and they run well on the large-scale image dataset ImageNet. Gated PixelCNN is proposed to improve upon PixelCNN, achieving performance comparable to PixelRNN but with significantly lower complexity. PixelCNN++ is proposed to improve upon PixelCNN by: using a discretized logistic mixture of likelihood and non-256-way multinomial distributions; using downsampling to capture structure across multiple resolutions; introducing additional shortcut connections to accelerate training; employing dropout for regularization; and combining RGB values ​​into a single pixel. PixelSNAIL is proposed, which combines causal convolutions with self-attention.

[0040] Most of the methods described above directly model the probability distribution in the pixel domain. Some researchers have also attempted to model the probability distribution as a conditional probability distribution based on explicit or latent representations. That is, we can estimate... (3) in It is an additional condition, and This means that modeling is divided into unconditional modeling and conditional modeling. The additional conditions can be image label information or high-level representations.

[0041] 2.3.2. Automatic Encoder Autoencoders originate from the renowned work of Hinton and Salakhutdinov. This method is trained for dimensionality reduction and consists of two parts: encoding and decoding. The encoding part transforms the high-dimensional input signal into a low-dimensional representation, typically with a reduced spatial size but a greater number of channels. The decoding part attempts to recover the high-dimensional input from the low-dimensional representation. Autoencoders can automatically learn representations and eliminate the need for hand-crafted features, which is considered one of the most significant advantages of neural networks.

[0042] Figure 1B This is a schematic diagram of a typical transform encoding / decoding scheme. Original image. Analysis network Transformation to achieve latent representation The latent representation y is quantized and compressed into bits. The number of bits... Used to measure codec rate. Quantized latent representation. Then by the synthetic network Inverse transform to obtain the reconstructed image Distortion is achieved through the use of functions. right The transformation is calculated in the perceptual space.

[0043] Applying autoencoder networks to lossy image compression is intuitive. We simply need to encode the learned latent representation from a trained neural network. However, adapting autoencoders to image compression is not straightforward, as the original autoencoders are not optimized for compression, making direct use of the trained autoencoder inefficient. Furthermore, other major challenges exist: First, the low-dimensional representation should be quantized before encoding, but quantization is non-differentiable, which is necessary for backpropagation during neural network training. Second, the objectives differ in compression scenarios because both distortion and bit rate need to be considered. Estimating the bit rate is challenging. Third, practical image encoding / decoding schemes need to support variable bit rates, scalability, encoding / decoding speeds, and interoperability. Many researchers have been actively contributing to this field to address these challenges.

[0044] An autoencoder prototype for image compression, such as Figure 1B As shown, it can be regarded as Transform encoding and decoding Strategy. Original image use analyze network Transformed, where It is the potential representation that will be quantized and encoded / decoded. synthesis The network will quantify the potential representation Perform inverse transform to obtain the reconstructed image Framework utilization distortion loss function (i.e. Training is conducted, among which Distortion between It is a representation based on quantification. The calculated or estimated bit rate, and These are Lagrange multipliers. It should be noted that... It can be computed in the pixel domain or the receptive domain. All existing research follows this prototype, with differences only in network structure or loss function.

[0045] In terms of network architecture, RNNs and CNNs are the most widely used architectures. Within the RNN-related category, a general framework for variable-rate image compression using RNNs has been proposed. They use binary quantization to generate the code and do not consider the rate during training. This framework does indeed provide scalable encoding and decoding capabilities, with RNNs having both convolutional and deconvolutional layers reportedly performing well. An improved version is then proposed by upgrading the encoder to compress binary codes using a neural network similar to PixelRNN. Using the MS-SSIM evaluation metric, its performance on the Kodak image dataset is reported to be superior to JPEG. The RNN-based solution is further improved by introducing hidden-state start-up. Furthermore, an SSIM-weighted loss function is designed, and a spatial adaptive bitrate mechanism is enabled. Using MS-SSIM as the evaluation metric, they achieve better results than BPG on the Kodak image dataset.

[0046] A general framework for rate-distortion optimization of image compression was designed. They used multivariate quantization to generate integer codes and considered rate during training, i.e., the loss was a joint rate-distortion cost, which could be MSE or other costs. They added random uniform noise to simulate quantization during training and used the differential entropy of the noisy codes as a surrogate for rate. They used Generalized Split Normalization (GDN) as the network structure, consisting of a linear mapping and nonlinear parameter normalization. The effectiveness of GDN on image encoding and decoding was verified. An improved version was proposed, where they used three convolutional layers, each followed by a downsampling layer and a GDN layer as the forward transform. Therefore, they used three layers of inverse GDN, each followed by an upsampling layer and a convolutional layer to simulate the inverse transform. Furthermore, an arithmetic encoding / decoding method was designed to compress integer codes. In terms of MSE, performance on the Kodak dataset was reported to outperform JPEG and JPEG 2000. Furthermore, it was further improved by incorporating a scaling super-prior into the autoencoder. They used a subnetwork... potential representation Transform into ,and This will be quantized and transmitted as side information. Therefore, a subnetwork will be used. To achieve the inverse transformation, this subnetwork attempts to extract information from the quantized edge information. Decoding to quantization The standard deviation, which will be in Their method is further utilized during arithmetic encoding and decoding. On the Kodak image set, their method is slightly inferior to BPG in terms of PSNR. These structures are further explored in the residual space by introducing an autoregressive model to estimate the standard deviation and mean. In the latest work, a Gaussian mixture model is used to further remove redundancy in the residuals. Using PSNR as the evaluation metric, the reported performance is comparable to VVC on the Kodak image set.

[0047] 2.3.3. Pre-Prior Model In transform encoding and decoding methods used for image compression, the encoder subnetwork (Section 2.3.2) uses parametric analysis of the transform. Transform the image vector x into a latent representation Then quantify it to form .because Since they are discrete values, they can be losslessly compressed using entropy encoding and decoding techniques such as arithmetic encoding and decoding, and transmitted as a bit sequence.

[0048] from Figure 2 The left and right center images clearly show this. Significant spatial dependencies exist among the elements. Notably, their scales (middle right image) appear to be spatially coupled. An additional set of random variables was introduced. To capture spatial dependencies and further reduce redundancy. In this case, image compression networks such as Figure 3 As shown.

[0049] exist Figure 3 In the middle, the encoder is on the left side of the model. and decoder (Explained in Section 2.3.2). The right-hand side is used to obtain... Additional encoders utilizing prior information and decoders that utilize prior information Network. In this architecture, the encoder subjectes the input image x to... The response with a standard deviation exhibiting spatial variation is obtained. .response fed to In summary The distribution of standard deviations in the data. Then... Quantified ( The data is compressed and transmitted as side information. The encoder then uses the quantized vector data... To estimate the spatial distribution of standard deviation And use it to compress and transmit quantized image representations. The decoder first recovers the signal from the compressed signal. Then the decoder uses To obtain ,Should To provide the decoder with the correct probability estimate and thus successfully recover the value. Then the decoder will Feed to To obtain a reconstructed image.

[0050] When an encoder and a decoder utilizing prior information are added to an image compression network, the quantization latent value is... Spatial redundancy is reduced. Figure 2 The rightmost image in the middle corresponds to the quantization latent value when using an encoder / decoder that leverages prior information. Compared to the middle right image, spatial redundancy is significantly reduced because the samples of the quantization latent value have lower correlation.

[0051] Figure 3 The network architecture of an autoencoder implementing a prior model is shown. The left side shows the image autoencoder network, and the right side corresponds to the prior subnetwork. The analytic transform and the synthetic transform are represented as follows: Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder, respectively. The hyperprior model consists of two sub-networks, utilizing the hyperprior information of the encoder (denoted as...). ) and decoders that utilize prior information (represented as The prior model generates quantified potential values ​​of prior information. The quantified potential value of prior information ( This includes information about quantifying potential values. Information about the probability distribution of the sample points. Included in the bitstream, and with They are transmitted together to the receiver (decoder).

[0052] 2.3.4. Context Model Although the prior model improves the quantification of latent values Modeling the probability distribution of quantified potential values ​​is possible, but further improvements can be achieved by utilizing an autoregressive model (context model) that predicts quantified potential values ​​from the causal context of quantified potential values.

[0053] The term autoregressive means that the output of a process is later used as its input. For example, a contextual model subnetwork generates a sample of latent values, which is later used as input to obtain the next sample.

[0054] Figure 4 This is a schematic diagram illustrating an exemplary combined model configured to jointly optimize a context model with a super-prior and an autoencoder. Table 1 below shows the meaning of the different symbols.

[0055] Table 1 – Symbol Explanation

[0056] A joint architecture is used, in which both a super-prior model subnetwork (an encoder and a decoder utilizing super-prior information) and a context model subnetwork are utilized. The super-prior and context models are combined to learn about the quantized latent value. The probability model is then used for entropy encoding and decoding. For example... Figure 4 As shown, the outputs of the context subnetwork and the decoder subnetwork utilizing prior information are combined by a subnetwork called the entropy parameter, which generates the mean for the Gaussian probability model. And variance (scale) (or variance) The parameters are then used. The Gaussian probability model is then used by the arithmetic encoder (AE) module to encode the samples of the quantized latent values ​​into the bitstream. In the decoder, the Gaussian probability model is used to obtain the quantized latent values ​​from the bitstream via the arithmetic decoder (AD) module. .

[0057] Figure 4 A combined model is shown, which jointly optimizes the following: an autoregressive component (context model) that estimates the probability distribution of latent values ​​from the causal context of the latent values, and a super-prior and a low-level autoencoder. Real-valued latent representations are quantized (Q) to create quantized latent values ​​(…). ) and quantified potential value of prior information ( ), this quantified potential value ( ) and quantified potential value of prior information ( The image is compressed into a bitstream using an arithmetic encoder (AE) and decompressed by an arithmetic decoder (AD). The highlighted areas correspond to the components performed by the receiver (i.e., the decoder) to recover the image from the compressed bitstream.

[0058] Typically, latent samples are modeled as Gaussian distributions or Gaussian mixture models (not limited to). According to Figure 4 The context model and the super-prior are jointly used to estimate the probability distribution of potential samples. Since the Gaussian distribution can be defined by its mean and variance (also called sigma or scale), the joint model is used to estimate the mean and variance (denoted as...). ).

[0059] 2.3.5. Encoding process using a joint autoregressive superprior model Figure 5 Existing compression methods are shown. The encoding and decoding processes will be described in this and the next section, respectively.

[0060] superior Figure 5 The encoding process is described. The input image is first processed by the encoder sub-network. The encoder transforms the input image into a transform representation called the latent value, which is... express. It is then fed into the quantizer block, denoted by Q, to obtain the quantization potential value ( ). Then, using an arithmetic coding module (denoted as AE), it is converted into a bitstream (bits1). The arithmetic coding blocks are sequentially... Each sample point is converted into a bit stream (bits1) one by one.

[0061] The module's encoder utilizing prior information, context, decoder utilizing prior information, and entropy parameter subnetwork are used to estimate the quantization latent value. The probability distribution of the sample points. Potential The information is input into an encoder that utilizes prior information, and its output is the potential value of the prior information (denoted as...). The latent value of the prior information is then quantified. The second bitstream (bits2) is generated using the arithmetic coding (AE) module. The decompositional entropy module generates the probability distribution used to encode the quantized hyperprior information latent values ​​into a bitstream. The quantized hyperprior information latent values ​​include information about the quantized latent values ​​(…). Information about the probability distribution of ( ).

[0062] Entropy parameter subnetwork generation was used to encode quantized latent values. The probability distribution is estimated. Information generated from the entropy parameter is typically used together to obtain the mean of the Gaussian probability distribution. And variance (scale) (or variance) Parameters. The Gaussian distribution of the random variable x is defined as follows: , where parameters It is the mean or expected value of the distribution (also the median and mode), while the parameter This is its standard deviation (or variance or scale). To define a Gaussian distribution, the mean and variance need to be determined. The entropy parameter module is used to estimate the mean and variance values.

[0063] The decoder of the subnetwork, utilizing prior information, generates part of the information used by the entropy parameter subnetwork, while another part is generated by an autoregressive module called the context module. The context module uses samples already encoded by the arithmetic encoding (AE) module to generate information about the probability distribution of the quantized latent values. It is typically a matrix composed of many sample points. Sample points can be indicated using indices, for example... [i,j,k] or [i,j], specifically depends on the matrix. Dimensions. Sample points [i,j] are encoded sequentially by the AE, typically using a raster scan order. In the raster scan order, the matrix rows are processed from top to bottom, with samples in a row processed from left to right. In such a scenario (where the AE encodes the samples into a bitstream using the raster scan order), the context module uses samples previously encoded in the raster scan order to generate a bitstream with the samples. Information related to [i,j]. The information generated by the context module and the decoder utilizing prior information is combined by the entropy parameter module to generate information used to quantize the latent values. The probability distribution encoded as a bit stream (bits1).

[0064] Finally, as a result of the encoding process, the first bitstream and the second bitstream are transmitted to the decoder.

[0065] It is worth noting that other names can be used for the modules mentioned above.

[0066] In the above description, Figure 5 All elements in the algorithm are collectively referred to as encoders. The analytical transformation that converts the input image into a latent representation is also called an encoder (or autoencoder).

[0067] 2.3.6. Decoding process using a joint autoregressive hyperprior model Figure 6 The existing decoding process is described. In the decoding process, the decoder first receives a first bitstream (bits1) and a second bitstream (bits2) generated by the corresponding encoder. Bits2 is first decoded by the arithmetic decoding (AD) module using a probability distribution generated by a decompositional entropy subnetwork. The decompositional entropy module typically uses a predetermined template to generate the probability distribution, for example, using predetermined mean and variance values ​​in the case of a Gaussian distribution. The output of the arithmetic decoding process for bits2 is... This is the quantized potential value of the prior information. The AD process recovers to the AE process applied in the encoder. The AE and AD processes are lossless, which means that the quantized potential value of the prior information generated by the encoder... It can be reconstructed at the decoder without any changes.

[0068] In obtaining Subsequently, it is processed by a decoder utilizing prior information, and the output of the decoder is fed into the entropy parameter module. The three sub-networks, context, decoder utilizing prior information, and entropy parameters used in the decoder are the same as the three sub-networks in the encoder. Therefore, the exact same probability distribution can be obtained in the decoder as in the encoder, which is crucial for reconstructing the quantized latent value without any loss. This is essential. As a result, the quantization potential value can be obtained in the decoder as well as in the encoder. Same version.

[0069] After obtaining the probability distribution (e.g., mean and variance parameters) through the entropy parameter subnetwork, the arithmetic decoding module decodes the samples of quantized potential values ​​one by one from the bitstream bits1. From a practical perspective, the autoregressive model (contextual model) is inherently serial, and therefore cannot be accelerated using techniques such as parallelization.

[0070] Finally, the fully reconstructed quantized potential value Input into the synthesis transform (in) Figure 6 The module (represented as decoder) is used to obtain the reconstructed image.

[0071] In the above description, Figure 6 All elements in the array are collectively referred to as decoders. The synthetic transformation that converts quantized latent values ​​into a reconstructed image is also called a decoder (or autodecoder).

[0072] 2.4. Neural Networks for Video Compression Similar to traditional video encoding and decoding techniques, neural image compression is based on intra-frame compression in neural network-based video compression. Therefore, the development of neural network-based video compression technology lagged behind that of neural network-based image compression, but due to its complexity, it requires more effort to overcome its challenges. Since 2017, some researchers have been working on neural network-based video compression schemes. Compared to image compression, video compression requires effective methods to eliminate inter-frame redundancy. Thus, inter-frame prediction is a key step in these works. Motion estimation and compensation have been widely adopted, but only recently have they been implemented using trained neural networks.

[0073] Depending on the target scenario, research on neural network-based video compression can be divided into two categories: random access and low latency. In the case of random access, decoding can begin at any point in the sequence, typically dividing the entire sequence into multiple separate segments, each of which can be decoded independently. In the case of low latency, the aim is to reduce decoding time, so usually only the earlier frames in the time domain can be used as reference frames to decode subsequent frames.

[0074] 2.4.1. Low latency Early work first divided the video sequence frames into blocks, and each block was selected from two available modes (intra-frame encoding or inter-frame encoding). If intra-frame encoding was chosen, an associated autoencoder was used to compress the block. If inter-frame encoding was chosen, motion estimation and compensation were performed using conventional methods, and a trained neural network was used for residual compression. The output of the autoencoder was directly quantized and encoded / decoded using the Huffman method.

[0075] Another neural network-based video encoding / decoding scheme utilizing PixelMotionCNN is proposed. Frames are compressed sequentially in the temporal domain, and each frame is divided into blocks, which are compressed in raster scan order. Each frame is first inferred using the two preceding reconstructed frames. When compressing a block, the inferred frame and the context of the current block are fed into PixelMotionCNN to derive the latent representation. Then, a variable-rate imaging scheme is used to compress the residuals. The performance of this scheme is comparable to H.264.

[0076] Then, an alternative video compression framework based on end-to-end neural networks is proposed, where all modules are implemented using neural networks. This scheme takes the current frame and a previously reconstructed frame as input and leverages a pre-trained neural network to derive optical flow as motion information. The motion information is warped along with a reference frame, followed by a neural network that generates motion-compensated frames. The residual and motion information are compressed using two separate neural autoencoders. The entire framework is trained using a single rate-distortion loss function. It achieves better performance than H.264.

[0077] An advanced neural network-based video compression scheme is proposed. It inherits and extends traditional video encoding and decoding schemes with neural networks, possessing the following key features: 1) it uses only one autoencoder to compress motion information and residuals; 2) it features motion compensation across multiple frames and multiple optical flows; and 3) it learns online states and propagates them over time through subsequent frames. This scheme achieves better performance than the HEVC reference software in MS-SSIM.

[0078] An extended end-to-end neural network-based video compression framework was then proposed. In this solution, multiple frames are used as references, enabling more accurate predictions of the current frame by using multiple reference frames and associated motion information. Furthermore, motion field prediction is deployed to eliminate motion redundancy along the temporal channels. A post-processing network is also introduced in this work to eliminate reconstruction artifacts from previous processes. The performance significantly outperforms H.265 in terms of both PSNR and MS-SSIM.

[0079] Then, scale-space flow was proposed, replacing the commonly used optical flow by adding a scale parameter. It reportedly achieves better performance than H.264.

[0080] A multi-resolution representation for optical flow is proposed. Specifically, a motion estimation network generates multiple optical flows with different resolutions, and the network learns which one to select under a loss function. Performance is superior to H.265.

[0081] 2.4.2. Random Access An initial approach based on frame interpolation was designed. First, keyframes are compressed using a neural image compressor, and the remaining frames are then compressed hierarchically. Motion compensation is performed in the receptive domain, i.e., feature maps at multiple spatial scales of the original frame are derived and warped using motion, which is then used in the image compressor. This method is reportedly comparable to H.264.

[0082] Then, another interpolation-based video compression method was proposed, in which the interpolation model combines motion information compression and image synthesis, and the same autoencoder is used for both the image and the residual.

[0083] Subsequently, a neural network-based video compression method based on a variational autoencoder with a deterministic encoder is proposed. Specifically, the model consists of an autoencoder and an autoregressive prior. Unlike previous methods, this method accepts a group of pictures (GOP) as input and incorporates a 3D autoregressive prior by considering temporal correlations when encoding and decoding the latent representation. It delivers performance comparable to H.265.

[0084] 2.5. Prerequisites 2.5.1. Color Separation and Conditional Encoding / Decoding In one example, such as Figure 8 As shown, the primary and secondary color components of the image are encoded and decoded separately using networks with similar architectures but different numbers of channels. All boxes with the same name are sub-networks with similar architectures, differing only in input / output tensor sizes and the number of channels. The primary component has the following number of channels: The number of channels for the secondary component is The vertical arrows (pointing downwards) indicate the data flow related to the encoding and decoding of secondary color components. The vertical arrows show the data exchange between the primary and secondary component pipelines.

[0085] The input signal to be encoded is represented as The latent space tensor in the bottleneck of a variational autoencoder is The subscript "Y" indicates the primary component, while the subscript "UV" is used for the secondary components in the stitching, including the chromaticity component.

[0086] First, the input image in RGB color format is converted into primary (Y) and secondary (UV) components. Primary component... Independent of secondary components The image is encoded and decoded, and the size of the encoded / decoded image is equal to the input / decoded image size. Secondary components use data from the primary components. As auxiliary information, it is conditionally encoded and decoded for use in encoding. And using auxiliary information from the main component. As a potential tensor, used for decoding Reconstruction. Aside from the number of channels, channel sizes, and several entropy models used to convert the latent tensor into a bitstream, the codec structures for the primary and secondary components are almost identical. Therefore, the primary and secondary latent tensors will generate two different bitstreams based on two different entropy models. Before encoding, , This module adjusts the sample point position through downsampling (in...) Figure 8 The above is marked as " This essentially means that the encoded image size of the secondary component differs from that of the encoded image size of the primary component. The scaling factor s is variable, but the default scaling factor is 1. In conditional encoding and decoding, the size of the auxiliary input tensor is adjusted so that the encoder receives primary and secondary component tensors with the same image size. After reconstruction, the secondary component utilizes a neural network-based upsampling filter module (…). Figure 8 "NN color filter" The image is rescaled to its original size, and the module outputs a factor. The secondary component that is upsampled.

[0087] Figure 8 The example illustrates an image encoding / decoding system where the input image is first converted into primary (Y) and secondary (UV) components. Output , This corresponds to the reconstructed output of the primary and secondary components. At the end of the processing, , It is converted back to RGB color format. Typically, It is downsampled (resized) before being processed by the encoding and decoding modules (neural network). For example, The size can be reduced by a factor of 50% in each of the vertical and horizontal dimensions. Therefore, the processing of minor components involves approximately 50% × 50% = 25% fewer samples, making it computationally less complex.

[0088] 2.5.2. JPEG AI Image Encoding and Decoding Standard At the time of writing, the JPEG AI image codec standard is an image codec standard standardized by the JPEG Working Group (WG), JPEG WG being WG 1 of ISO / IEC JTC 1 SC 29. The ISO / IEC number for the JPEG AI standard is ISO / IEC 6048. The latest JPEG AI draft specification is included in the JPEG output document WG1N100602.

[0089] The latest JPEG AI draft specification utilizes some of the neural network-based image encoding and decoding methods mentioned above. The following describes or summarizes some features of the latest JPEG AI specification.

[0090] 2.5.2.1. Functional Overview of the JPEG AI Decoding Process Overall decoder architecture as follows Figure 8 As shown. Data (tensors and streams) is displayed in the "white" box, the neural network modules required for decoding are displayed in the gray shaded box, and the toggleable tools are displayed in the purple shaded box.

[0091] For the primary and secondary color components, the code stream can be parsed independently and reconstructed using modules composed of the same sequence from the same neural network layers, the only difference being the size of the input tensor and the number of tensor channels. A single-component decoder, such as... Figure 9 As shown.

[0092] The first stream z must be decoded by a lossless entropy decoder. (Analysis). Assuming it is used for... The probability distribution of the lossless encoding and decoding is a Gaussian distribution with pre-trained parameters (part of the training model), and the cumulative distribution function (in the model) is calculated based on these pre-trained parameters. Figure 9 The above is represented as It was used in the lossless entropy decoder.

[0093] Decoding the super-prior tensor It is used as input to two different processes: a decoder that utilizes prior information and a variance decoder that utilizes prior information.

[0094] Then stream y must be decoded by a lossless decoder ( (Analysis.) Used for parsing. The probability distribution is assumed to be Gaussian, with a mean of zero, and the standard deviation is given as the output of the following steps: using a variance decoder with prior information (Section 10.3), output a tensor of the standard deviation in the logarithmic domain. Then, based on the rate control parameters inside the sigma scaling... Scale it to produce Then, based on the RVS parameter portion inside the adaptive sigma scaling, it is masked and scaled to produce... Finally, tensors The value is quantized (converted to an index in the probability distribution table). According to the rules specified by the SKIP mode, some elements of the residual tensor are skipped (unencoded / decoded) and replaced with zeros in the decoder SKIP module, which receives the parsed set of syntax elements { from the tANS decoder. sThe module receives mask_sigma from the SKIP mask generation module and outputs the reconstructed residual tensor, which is then reshaped into a 3D shape.

[0095] On the decoder side, according to the parameters residual Scaling by inverse gain unit produces Then, the residual tensor is scaled in the invRVS (inverse residual and variance scaling) module to form the residual tensor. This is used to reconstruct the underlying tensor. .

[0096] Decoders that utilize prior information to generate multi-level context models (MCMs) The input is an eight-level neural network process that will also reconstruct the residuals. As input and output, the potential space tensor After latent scaling-LSBS prior to synthesis, the reconstructed latent space tensor Prepare for signal reconstruction. The latent tensor reconstructions of the primary and secondary components are independent of each other.

[0097] Reconstructed potential space tensor This is the input to the composition transformation. The other input to the composition transformation is the auxiliary tensor. For minor component synthesis, auxiliary tensors are generated from the latent tensors reconstructed from the major components. For major components, no auxiliary tensors are used. (Based on the input image height...) and width and main components ( ) and secondary components ( The scaling factor and tensor dimensions are shown in Table 1. For the principal components, the parameters... This means that the synthesis transformation of the primary components does not receive auxiliary information (independent reconstruction). For the secondary components... The auxiliary method for quadratic transformation synthesis is It is the reconstruction of the latent space tensor by the principal components. Integer factors It is formed by resampling (using nearest neighbor downsampling or nearest neighbor upsampling).

[0098] Table 1 Tensor size parameters used for primary and secondary component decoding

[0099] The synthesis transform for both the primary and secondary components consists of the same neural network layers, differing only in the size of the input tensor and the number of tensor channels. The output tensor of the synthesis transform is... (Tensor dimensions are listed in Table 1).

[0100] like Figure 8 As shown, after the synthesis transformation, the primary and secondary components enter the enhancement filter and output format conversion processing module, which includes resampling, inverse color conversion, and filter set.

[0101] 2.5.3. Transcoding Image / video transcoding is the process of converting video / image files (i.e., bitstream files) from one format to another. It involves changing the encoding format, resolution, bitrate, or other parameters of the image / video file to ensure compatibility with different devices, platforms, or internet connectivity. Transcoding is crucial in video production and distribution. Figure 10 The transcoding example shown involves an image being encoded using a first codec encoder, and then decoded using a first codec decoder, which is part of the transcoder. The reconstructed image is then re-encoded using a second codec encoder, and the resulting bitstream file is distributed over the Internet. When a user requests display, the second codec decoder is used to reconstruct the image.

[0102] Figure 10 Examples of transcoding for different resolutions, bit rates, or compatible codecs available on playback devices are shown.

[0103] Image / video transcoding technology has evolved alongside the digitization of user content, with multiple generations of image / video codecs developed to accommodate the ever-increasing volume of visual content. Whenever a new codec is released, transcoding techniques are required to convert bitstreams / files encoded using another codec (e.g., a previous generation codec) into a format decodable by the latest codec. The importance of image / video transcoding stems primarily from the following reasons.

[0104] 1) Compatibility. Different devices and platforms support different resolutions, bitrates, codecs, and container formats. Transcoding enables seamless conversion to desired formats compatible with a wider range of devices and platforms.

[0105] 2) Optimization. Transcoding allows for video optimization for different devices (desktops, smartphones, tablets, or smart TVs), all of which have different definitions of "best." By transcoding video to ensure the smoothest playback on each device with the best possible image quality, you ensure a better viewing experience for your users.

[0106] 3) Streaming. Transcoding is business-critical for video streaming services like YouTube, Netflix, and Amazon Prime Video. These services need to support a wide range of devices, operating systems, and browsers. They use transcoding as a solution – optimizing their videos for multiple formats and resolutions to ensure smooth playback across a wide range of user platforms and bandwidth.

[0107] 4) Reduced file size. Transcoding video can also reduce the file size of the video, making it easier and faster to upload, share, or stream. This is especially important when dealing with large videos.

[0108] 5) Accessibility. Closed captions and subtitles can be added during the transcoding process, allowing users with hearing impairments to understand the video content. Videos can also be transcoded using audio streams in different languages ​​to cater to non-native speakers.

[0109] 3. Problem 3.1. Core Issues At the time of this draft, the JPEG AI image codec standard was a neural network-based image codec standard under development. Prior to this standard, many other image codec standards existed, such as JPEG and JPEG 2000. Figure 11 As shown, problems arise when the terminal device only has the JPEG AI codec, while the image file is encoded using other codecs (such as JPEG). Although JPEG AI is significantly more efficient than its predecessors (JPEG, JPEG 2000), the lack of a transcoder makes the JPEG AI codec standard difficult to deploy widely.

[0110] Figure 11 This illustrates an example where a file encoded in JPEG cannot be decoded on a device with a JPEG AI image codec.

[0111] Besides compatible codec use cases, there are other scenarios where transcoding is necessary.

[0112] • Change the image resolution or bitrate.

[0113] • Enhance or optimize the quality of reconstructed images based on bit rate or application type.

[0114] • Reduce the size of bitstream files on storage devices.

[0115] • Adjust the bit rate according to internet bandwidth.

[0116] 4. Solution The detailed solutions below should be considered as examples to explain general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0117] Figure 12 This example demonstrates how to transcode a JPEG-encoded file into a JPEG AI-decodeable file. The example is in... Figure 12 As shown in the figure, a transcoder is used to convert a bitstream file generated using a first codec into a compatible format so that a second codec can decode it to obtain the correct reconstructed image.

[0118] 4.1. Core of this disclosure The core of this disclosure is to provide a transcoding solution for neural network-based image codecs (such as the JPEG AI image codec standard), in which a encoded file using a first codec can be correctly decoded using a second codec to obtain a correctly reconstructed image. Figure 13 An example is shown demonstrating the use of filtering before the second codec decoder.

[0119] 5. Examples 1. According to the present invention, a transcoding solution is proposed for converting a first bitstream file generated by a first codec into a second bitstream having a compatible format for a second codec decoder, such that the second decoder can reconstruct the image.

[0120] a. In one example, the second codec could be a neural network-based codec, such as the JPEG AI image codec standard.

[0121] i. Alternatively, the second codec may be a conventional / non-neural network-based (NN-based) codec, such as JPEG, JPEG 2000, H.264, H.265 or H.266.

[0122] b. In one example, the first codec can be a traditional / non-neural network-based codec, such as JPEG, JPEG 2000, H.264, H.265, or H.266.

[0123] i. Alternatively, the first codec could be a neural network-based codec, such as the JPEG AI image codec standard.

[0124] c. In one example, the second codec can remain unchanged, and an intermediate processing module is used between the first codec decoder and the second codec encoder.

[0125] i. In one example, the second codec is a NN-based codec, and the NN weights are not changed.

[0126] ii. In such Figure 13 In one example shown, the intermediate processing could be filtering.

[0127] 1. In one example, filtering can be based on neural networks (NNs).

[0128] d. In such Figure 14 In one example shown, the NN-based second codec may include multiple sets of weights. Figure 14 An example transcoding scheme using at least two encoders with a second codec is illustrated according to some embodiments of this disclosure. One encoder is original, and the other encoders are retrained for transcoding purposes.

[0129] i. For example, the first set of NN weights can be used for purposes other than transcoding.

[0130] ii. For example, the second set of NN weights can be used for transcoding purposes.

[0131] 1. In one example, the NN weights may depend on the first codec.

[0132] 2. Indicators can be transmitted via signals in the bitstream to indicate which set of NN weights to use.

[0133] e. For example Figure 14 As shown, the multiple NN weight sets of the second codec can be organized using different architectures and / or numbers of layers and / or types of neural network layers.

[0134] 2. According to the present invention, a transcoder is used to assist a decoder in reconstructing an image, wherein both the transcoder and the decoder consist of neural network layers.

[0135] a. Indicators can be included in the bitstream to indicate whether the transcoder is used in processing or skipped.

[0136] b. Multiple encoders of the second codec can be used, and an indication can be included in the bitstream to indicate which one to use during the decoding process.

[0137] 3. In one example, at least one indication may be transmitted via signaling in the second bitstream to indicate the type of the original image from which it is encoded.

[0138] a. The type can be the original image.

[0139] b. The type can be an image decoded from the first bitstream encoded by a specific codec (such as JPEG, JPEG 2000, H.264, H.265, or H.266).

[0140] i. The type of a particular codec can be transmitted via signals in the second bitstream, represented by reference values ​​such as 0 or 1.

[0141] c. Post-processing can be applied to the image decoded from the second bitstream.

[0142] i. In one example, whether and / or how post-processing is applied can depend on the type of the original image from which it was encoded.

[0143] 1. For example, if the original image is decoded from a first bitstream encoded and decoded by a particular codec, then specific post-processing can be applied.

[0144] ii. In one example, post-processing could be filtering.

[0145] iii. In one example, post-processing can be based on neural networks.

[0146] 4. According to the present invention, the indicator can be included in the bitstream to indicate: a. Selection of model parameters (e.g., weights of convolutional layers) to be used in decoding the transcoded image.

[0147] i. For example, an indicator can indicate the selection of model coefficients (e.g., parameters or weights) for the processing layer.

[0148] b. Selection of entropy encoding / decoding mode.

[0149] i. For example, an indicator can indicate the choice of tables to be used in entropy decoding.

[0150] ii. In another example, the indicator can indicate the choice of the initialization method for entropy decoding.

[0151] c. Selection of the composition transformation.

[0152] i. For example, if the indicator indicates that the bitstream was obtained through transcoding, different synthetic transformations can be applied.

[0153] d. Selection of color transformations.

[0154] i. The indicator can indicate the choice between color transformation or inverse color transformation.

[0155] General items 5. Whether and / or how the methods disclosed above can be applied to signal transmission at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0156] 6. Whether and / or how to apply the methods disclosed above may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / image type.

[0157] 7. The methods presented in this document can be used in other codec tools that require chroma blending.

[0158] 8. The syntax elements disclosed above can be binarized into flags, fixed-length codes, EG(x) codes, unary codes, rounded unary codes, rounded binary codes, etc. These can be signed or unsigned.

[0159] 9. The syntax elements disclosed above can be encoded or decoded using at least one context model. Alternatively, they can be encoded or decoded using a bypass method.

[0160] 10. The syntax elements disclosed above can be transmitted conditionally via signals.

[0161] a. SE is transmitted via signal only if the corresponding function is applicable.

[0162] b. SE is transmitted via signal only if the dimensions of the block (width and / or height) meet the conditions.

[0163] 11. The syntax elements disclosed above can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0164] Further details of embodiments of this disclosure relating to neural network-based visual data encoding and decoding will now be described. As used herein, the term "visual data" may refer to video, images, pictures in video, or any other visual data suitable for encoding and decoding.

[0165] As discussed above, neural network (NN)-based encoders and decoders are not supported in existing designs for visual data transcoding. Therefore, terminal devices with only NN-based encoding / decoding capabilities cannot reconstruct bitstreams encoded or decoded using non-NN-based encoding / decoding schemes. Consequently, the application scenarios for NN-based encoding / decoding schemes are limited.

[0166] To address the above-mentioned problems and other issues not mentioned, a visual data processing solution is disclosed as described below. The embodiments of this disclosure should be considered as examples illustrating general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments can be applied individually or in any combination.

[0167] Figure 15 A flowchart of a method 1500 for visual data transcoding according to some embodiments of the present disclosure is shown. Method 1500 begins at 1502, where visual data is reconstructed from a first bitstream of visual data. Furthermore, at 1504, the reconstructed visual data is encoded into a second bitstream of visual data. At least one of the reconstruction or encoding is performed using a neural network (NN)-based model. As used herein, the NN-based model can be a model based on neural network techniques. For example, the NN-based model can specify a sequence of neural network modules (also called an architecture) and model parameters. A neural network module can include a set of neural network layers. Each neural network layer specifies tensor operations for receiving and outputting tensors, and each layer has trainable parameters. It should be understood that the possible implementations of the NN-based model described herein are merely illustrative and should not be construed as limiting the present disclosure in any way.

[0168] In one example embodiment, reconstruction at 1502 is performed using a first NN-based model (e.g., an NN-based decoder based on the JPEG AI image codec standard), and encoding at 1504 is performed using a non-NN-based encoder (such as an encoder based on the JPEG, JPEG 2000, H.264, H.265, or H.266 codec standards). In another example embodiment, reconstruction at 1502 is performed using a non-NN-based decoder (such as a decoder based on the JPEG, JPEG 2000, H.264, H.265, or H.266 codec standards), and encoding at 1504 is performed using a second NN-based model (e.g., an NN-based encoder based on the JPEG AI image codec standard). In yet another example embodiment, both reconstruction at 1502 and encoding at 1504 are performed using NN-based models. In this case, the NN-based model may include, for example, the first and second NN-based models described above. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0169] In light of the above, the reconstruction and / or encoding steps involved in visual data transcoding are performed using a neural network (NN)-based model. Compared to traditional solutions that do not support NN-based visual data encoding / decoding schemes during transcoding, the proposed method can advantageously support NN-based visual data encoding / decoding schemes, thereby improving the compatibility of visual data transcoding. In this way, the application scenarios of NN-based visual data encoding / decoding schemes are also expanded.

[0170] In some embodiments, at 1504, the reconstructed visual data can be processed using an intermediate processing module. For example, the intermediate processing module may include a filtering process, etc. Additionally, the filtering process may be based on a neural network (NN). The processed reconstructed visual data can then be encoded into a second bitstream using a second NN-based model. In this case, the weights of the second NN-based model may remain unchanged. For example, the weights of the second NN-based model used for visual data encoding can be directly used in this transcoding scenario.

[0171] In some embodiments, the second NN-based model may include multiple sets of weights. Each set of weights can be configured for a specific purpose. For example, a first set of weights in the multiple sets of weights may be used for purposes other than transcoding, such as encoding / decoding, quality enhancement, bitrate adjustment, etc. Furthermore, a second set of weights in the multiple sets of weights may be used for transcoding. For example, the second set of weights may depend on a first codec for a first bitstream. Additionally, the bitstream may include an indication of the weight sets used in the multiple sets of weights. In some embodiments, the multiple sets of weights may be associated with different architectures, different numbers of layers, different types of NN layers, etc.

[0172] In some embodiments, the visual data can be decoded from a second bitstream. For example, encoding at 1504 can be performed using a second NN-based model, and decoding can be performed using a third NN-based model.

[0173] In some embodiments, the first bitstream and / or the second bitstream may include an indication of whether a transcoder is being used. In some embodiments, multiple encoders may be available for encoding at 1504, and the first bitstream and / or the second bitstream may include an indication of the encoder being used among the multiple encoders.

[0174] In some embodiments, post-processing may be applied to the decoded visual data. Alternatively, information regarding at least one of the following depends on the type of reconstructed visual data: whether or how post-processing is applied to the decoded visual data. In some embodiments, post-processing is applied to the decoded visual data if the first bitstream was encoded or decoded using a particular codec.

[0175] In some embodiments, the post-processing process may include filtering, etc. For example, the post-processing process may be based on neural networks (NNs).

[0176] In some embodiments, the second bitstream may include an indication of the type of reconstructed visual data. For example, the type of reconstructed visual data may be allowed to be the original visual data. Additionally or alternatively, the type of reconstructed visual data may be allowed to be visual data decoded from the first bitstream encoded using the first codec. In this case, the second bitstream may also include an indication of the type of the first codec.

[0177] In some embodiments, the second bitstream may include a first set of indicators indicating parameters of a decoding model used to decode visual data from the second bitstream. For example, the first set of indicators may include indicators of coefficients of the processing layer of the decoding model.

[0178] In some embodiments, the first bitstream and / or the second bitstream may include a second set of indicators indicating the entropy encoding / decoding mode to be used. For example, the second set of indicators may include: an indicator indicating a table used in entropy decoding, an indicator indicating an initialization scheme for entropy decoding, etc.

[0179] In some embodiments, the first bitstream and / or the second bitstream may include an indication of the synthetic transformation to be used.

[0180] In some embodiments, if a first indication in the second bitstream indicates that the second bitstream was obtained without transcoding, then a first synthesis transform can be used to process the second bitstream. If the first indication indicates that the second bitstream was obtained through transcoding, then a second synthesis transform, different from the first synthesis transform, can be used to process the second bitstream.

[0181] In some embodiments, the first bitstream and / or the second bitstream may include indications of color transformation. Additionally or alternatively, the first bitstream and / or the second bitstream may include indications of inverse color transformation.

[0182] In some embodiments, first information regarding at least one of the following can be indicated in the bitstream: whether a method is applied, or how the method is applied. For example, the first information can be indicated at the block level, sequence level, picture group level, picture level, strip level, slice group level, etc.

[0183] In some embodiments, the first information may be indicated by one of the following: the codec structure of a codec tree unit (CTU), the codec structure of a codec unit (CU), the codec structure of a transform unit (TU), the codec structure of a prediction unit (PU), the codec structure of a codec tree block (CTB), the codec structure of a codec block (CB), the codec structure of a transform block (TB), the codec structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice header.

[0184] In some embodiments, the first information may depend on encoded or decoded information of the visual data. For example, and not limited to, the encoded or decoded information may include block size, color format, single-tree segmentation, dual-tree segmentation, color components, stripe type, image type, etc.

[0185] In some embodiments, any of the above indications can be a syntax element. For example, a syntax element can be binarized into one of the following: a flag, a fixed-length code, an exponential Golomb code (EG code), a unary code, a rounded unary code, or a rounded binary code. Additionally, a syntax element can be encoded or decoded using at least one context model. Alternatively, a syntax element can be bypassed for encoding or decoding. In some embodiments, a syntax element can be conditionally transmitted via signaling.

[0186] In some embodiments, syntax elements may be indicated at one of the following levels: block level, sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, syntax elements may be indicated at one of the following levels: codec structure of codec tree unit (CTU), codec structure of codec unit (CU), codec structure of transform unit (TU), codec structure of prediction unit (PU), codec structure of codec tree block (CTB), codec structure of codec block (CB), codec structure of transform block (TB), codec structure of prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[0187] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve the compatibility of visual data transcoding and expand the application scenarios of NN-based visual data encoding and decoding schemes.

[0188] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of visual data generated by a method performed by an apparatus for visual data transcoding. The method includes: reconstructing the visual data from a first bitstream of visual data; and encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN)-based model.

[0189] According to further embodiments of this disclosure, a method for storing a bitstream of visual data is provided. The method includes: reconstructing visual data from a first bitstream of visual data; encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN)-based model; and storing the second bitstream in a non-transitory computer-readable recording medium.

[0190] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0191] Item 1. A method for transcoding visual data, comprising: reconstructing visual data from a first bitstream of visual data; and encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN) based model.

[0192] Item 2. According to the method of Item 1, the reconstruction is performed using a first NN-based model.

[0193] Item 3. The method according to either Item 1 or 2, wherein the encoding is performed using a second NN-based model.

[0194] Item 4. According to the method of Item 3, encoding the reconstructed visual data includes: processing the reconstructed visual data using an intermediate processing module; and encoding the processed reconstructed visual data into a second bitstream using a second NN-based model.

[0195] Item 5. According to the method of Item 4, the weights of the second model based on NN are not changed.

[0196] Item 6. The method according to either Item 4 or 5, wherein the intermediate processing module includes a filtering process.

[0197] Item 7. The method of Item 6, wherein the filtering process is based on NN.

[0198] Item 8. According to any of the methods in Items 3-7, the second NN-based model includes multiple sets of weights.

[0199] Item 9. According to the method of Item 8, the first weight set in a plurality of weight sets is used for purposes other than transcoding.

[0200] Item 10. According to the method of either Item 8 or 9, a second weight set in a plurality of weight sets is used for transcoding.

[0201] Item 11. According to the method of Item 10, wherein the second weight set depends on the first codec for the first bitstream.

[0202] Item 12. The method according to any one of Items 8-11, wherein the bit stream includes an indication of a weight set used in a plurality of weight sets.

[0203] Item 13. According to the method of any one of items 8-12, wherein multiple sets of weights are associated with at least one of the following: different architectures, different numbers of layers, or different types of NN layers.

[0204] Item 14. The method according to any one of items 1-13 further includes: decoding visual data from a second bitstream.

[0205] Item 15. According to the method of Item 14, encoding is performed using a second NN-based model, and decoding is performed using a third NN-based model.

[0206] Item 16. The method according to any one of Items 14-15, wherein at least one of the first bitstream or the second bitstream includes an indication of whether a transcoder is used.

[0207] Item 17. The method according to any one of Items 14-16, wherein a plurality of encoders may be used for encoding, and at least one of the first bitstream or the second bitstream includes an indication of an encoder used among the plurality of encoders.

[0208] Item 18. The method of any one of Items 14-17, wherein a post-processing procedure is applied to the decoded visual data.

[0209] Item 19. The method according to any one of Items 14-18, wherein information about at least one of the following depends on the type of reconstructed visual data: whether post-processing is applied to the decoded visual data, or how post-processing is applied to the decoded visual data.

[0210] Item 20. According to the method of Item 19, if the first bitstream is encoded or decoded using a particular codec, then the post-processing is applied to the decoded visual data.

[0211] Item 21. The method according to any one of items 18-20, wherein the post-processing procedure includes a filtering procedure.

[0212] Item 22. The method of any of Items 18-21, wherein the post-processing is based on NN.

[0213] Item 23. The method according to any one of items 1-22, wherein the second bitstream includes an indication of the type of reconstructed visual data.

[0214] Item 24. According to the method of Item 23, the type is allowed to be raw visual data.

[0215] Item 25. The method of any one of Items 23-24, wherein the type is permitted to be visual data decoded from a first bitstream encoded using a first codec.

[0216] Item 26. According to the method of Item 25, the second bitstream further includes an indication of the type of the first codec.

[0217] Item 27. The method according to any one of Items 1-26, wherein the second bitstream includes a first set of indicators indicating parameters for a decoding model used to decode visual data from the second bitstream.

[0218] Item 28. According to the method of Item 27, wherein the first set of indicators includes indicators of the coefficients of the processing layer of the decoding model.

[0219] Item 29. The method according to any one of Items 1-28, wherein at least one of the first bitstream or the second bitstream includes a second set of indicators indicating the entropy encoding / decoding mode to be used.

[0220] Item 30. According to the method of Item 29, wherein the second set of indicators includes at least one of the following: an indicator indicating a table used in entropy decoding, or an indicator indicating an initialization scheme for entropy decoding.

[0221] Item 31. The method according to any one of items 1-30, wherein at least one of the first bitstream or the second bitstream includes an indication of the synthetic transformation to be used.

[0222] Item 32. The method according to any one of items 1-31, wherein if a first indication in the second bitstream indicates that the second bitstream was obtained without transcoding, then a first synthesis transform is used to process the second bitstream, and if the first indication indicates that the second bitstream was obtained by transcoding, then a second synthesis transform, different from the first synthesis transform, is used to process the second bitstream.

[0223] Item 33. The method according to any one of items 1-32, wherein at least one of the first bitstream or the second bitstream includes an indication of color transformation.

[0224] Item 34. The method according to any one of items 1-33, wherein at least one of the first bitstream or the second bitstream includes an indication of an inverse color transformation.

[0225] Item 35. A method according to any one of items 1-34, wherein first information about at least one of the following is indicated in the bitstream: whether the method is applied, or how the method is applied.

[0226] Item 36. According to the method of Item 35, wherein the first information is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0227] Item 37. According to the method of Item 35, wherein the first information is indicated in one of the following: the codec structure of a codec tree unit (CTU), the codec structure of a codec unit (CU), the codec structure of a transform unit (TU), the codec structure of a prediction unit (PU), the codec structure of a codec tree block (CTB), the codec structure of a codec block (CB), the codec structure of a transform block (TB), the codec structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice header.

[0228] Item 38. The method according to any one of items 35-37, wherein the first information depends on the encoded / decoded information of the visual data.

[0229] Item 39. According to the method of Item 38, the encoded or decoded information includes at least one of the following: block size, color format, single-tree segmentation, double-tree segmentation, color components, stripe type, or image type.

[0230] Item 40. The method according to any of Items 12-39, wherein the instruction includes a grammatical element.

[0231] Item 41. According to the method of Item 40, wherein the syntax elements are binarized into one of the following: a flag, a fixed-length code, an exponential Golomb code (EG), a unary code, a rounded unary code, or a rounded binary code.

[0232] Item 42. The method according to any one of items 40-41, wherein the syntax element is encoded or decoded using at least one context model, or wherein the syntax element is encoded or decoded by bypass.

[0233] Item 43. The method of any of Items 40-42, wherein syntax elements are transmitted via signals based on conditions.

[0234] Item 44. According to the method of any of Items 40-43, wherein the syntax element is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0235] Item 45. According to the method of any one of Items 40-44, wherein the syntax element is indicated in one of the following: the codec structure of a codec tree unit (CTU), the codec structure of a codec unit (CU), the codec structure of a transform unit (TU), the codec structure of a prediction unit (PU), the codec structure of a codec tree block (CTB), the codec structure of a codec block (CB), the codec structure of a transform block (TB), the codec structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice header.

[0236] Item 46. The method according to any one of items 1-45, wherein the visual data includes video, a picture of video, or an image.

[0237] Item 47. An apparatus for visual data transcoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-46.

[0238] Item 48. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute a method according to any one of items 1-46.

[0239] Item 49. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of a visual data transcoding apparatus, wherein the method includes: reconstructing the visual data from a first bitstream of visual data; and encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN) based model.

[0240] Item 50. A method for storing a bitstream of visual data, comprising: reconstructing visual data from a first bitstream of visual data; encoding the reconstructed visual data into a second bitstream of visual data, wherein at least one of the reconstruction or encoding is performed using a neural network (NN)-based model; and storing the second bitstream in a non-transitory computer-readable recording medium.

[0241] Example device Figure 16 A block diagram of a computing device 1600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 1600 may be implemented as a source device 110 (or visual data encoder 114) or a destination device 120 (or visual data decoder 124), or may be included in the source device 110 (or visual data encoder 114) or the destination device 120 (or visual data decoder 124).

[0242] It should be understood that, Figure 16 The computing device 1600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0243] like Figure 16 As shown, computing device 1600 includes general-purpose computing device 1600. Computing device 1600 may include at least one or more processors or processing units 1610, memory 1620, storage unit 1630, one or more communication units 1640, one or more input devices 1650, and one or more output devices 1660.

[0244] In some embodiments, the computing device 1600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 1600 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0245] Processing unit 1610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 1620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1600. Processing unit 1610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0246] Computing device 1600 typically includes various computer storage media. Such media can be any media accessible by computing device 1600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 1620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 1630 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 1600.

[0247] The computing device 1600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 16 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to the bus (not shown) via one or more visual data media interfaces.

[0248] Communication unit 1640 communicates with another computing device via a communication medium. Additionally, the functionality of components in computing device 1600 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, computing device 1600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0249] Input device 1650 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 1660 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 1640, computing device 1600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 1600 can also communicate with one or more devices that enable a user to interact with computing device 1600, or, if needed, with any device (e.g., network card, modem, etc.) that enables computing device 1600 to communicate with one or more other computing devices. Such communication can be performed via an input / output (I / O) interface (not shown).

[0250] In some embodiments, some or all of the components of computing device 1600 may be deployed in a cloud computing architecture rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, visual data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding visual data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote visual data center locations. Cloud computing infrastructure may provide services through shared visual data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0251] In embodiments of this disclosure, computing device 1600 may be used to implement visual data encoding / decoding. Memory 1620 may include one or more visual data encoding / decoding modules 1625 having one or more program instructions. These modules are accessible and executable by processing unit 1610 to perform the functions of the various embodiments described herein.

[0252] In an exemplary embodiment of performing visual data encoding, input device 1650 may receive visual data as input 1670 to be encoded. The visual data may be processed, for example, by visual data encoding / decoding module 1625 to generate an encoded bitstream. The encoded bitstream may be provided as output 1680 via output device 1660.

[0253] In an exemplary embodiment of performing visual data decoding, input device 1650 may receive an encoded bitstream as input 1670. The encoded bitstream may be processed, for example, by a visual data encoding / decoding module 1625 to generate decoded visual data. The decoded visual data may be provided as output 1680 via output device 1660.

[0254] While this disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for visual data transcoding, comprising: Reconstruct the visual data from the first bitstream of the visual data; as well as The reconstructed visual data is encoded into a second bitstream of the visual data, wherein at least one of the reconstruction or the encoding is performed using a neural network (NN)-based model.

2. The method of claim 1, wherein the reconstruction is performed using a first NN-based model.

3. The method according to any one of claims 1 and 2, wherein the encoding is performed using a second NN-based model.

4. The method of claim 3, wherein encoding the reconstructed visual data comprises: The reconstructed visual data is processed using an intermediate processing module; as well as The processed and reconstructed visual data is encoded into the second bitstream using the NN-based second model.

5. The method according to claim 4, wherein the weights of the NN-based second model are not changed.

6. The method according to any one of claims 4 and 5, wherein the intermediate processing module includes a filtering process.

7. The method according to claim 6, wherein the filtering process is based on neural networks (NN).

8. The method according to any one of claims 3-7, wherein the NN-based second model comprises a plurality of weight sets.

9. The method of claim 8, wherein the first weight set in the plurality of weight sets is used for a purpose different from transcoding.

10. The method according to any one of claims 8 and 9, wherein the second weight set in the plurality of weight sets is used for transcoding.

11. The method of claim 10, wherein the second weight set depends on a first codec for the first bitstream.

12. The method according to any one of claims 8-11, wherein the bitstream includes an indication of a weight set used in the plurality of weight sets.

13. The method according to any one of claims 8-12, wherein the plurality of weight sets are associated with at least one of the following: Different architectures, Different numbers of layers, or Different types of NN layers.

14. The method according to any one of claims 1-13, further comprising: The visual data is decoded from the second bitstream.

15. The method of claim 14, wherein the encoding is performed using a second NN-based model, and the decoding is performed using a third NN-based model.

16. The method according to any one of claims 14-15, wherein at least one of the first bitstream or the second bitstream includes an indication of whether a transcoder is used.

17. The method according to any one of claims 14-16, wherein a plurality of encoders are available for the encoding, and at least one of the first bitstream or the second bitstream includes an indication of an encoder used among the plurality of encoders.

18. The method according to any one of claims 14-17, wherein a post-processing procedure is applied to the decoded visual data.

19. The method according to any one of claims 14-18, wherein information regarding at least one of the following depends on the type of the reconstructed visual data: Whether to apply post-processing to the decoded visual data, or How to apply the post-processing procedure to the decoded visual data.

20. The method of claim 19, wherein if the first bitstream is encoded or decoded using a specific codec, the post-processing is applied to the decoded visual data.

21. The method according to any one of claims 18-20, wherein the post-processing process includes a filtering process.

22. The method according to any one of claims 18-21, wherein the post-processing procedure is based on neural networks (NNs).

23. The method according to any one of claims 1-22, wherein the second bitstream includes an indication of the type of the reconstructed visual data.

24. The method of claim 23, wherein the type is permitted to be raw visual data.

25. The method according to any one of claims 23-24, wherein the type is permitted to be visual data decoded from the first bitstream encoded using the first codec.

26. The method of claim 25, wherein the second bitstream further includes an indication of the type of the first codec.

27. The method according to any one of claims 1-26, wherein the second bitstream includes a first set of indicators indicating parameters for a decoding model used to decode the visual data from the second bitstream.

28. The method of claim 27, wherein the first set of indications includes indications of the coefficients of the processing layer of the decoding model.

29. The method according to any one of claims 1-28, wherein at least one of the first bitstream or the second bitstream includes a second set of indicators indicating the entropy encoding / decoding mode to be used.

30. The method of claim 29, wherein the second set of indications comprises at least one of the following: Indicator of the table used in entropy decoding, or Instructions for the initialization scheme used for entropy decoding.

31. The method according to any one of claims 1-30, wherein at least one of the first bitstream or the second bitstream includes an indication of the synthetic transformation to be used.

32. The method according to any one of claims 1-31, wherein if a first indication in the second bitstream indicates that the second bitstream was obtained without transcoding, then a first synthesis transform is used to process the second bitstream, and If the first indication indicates that the second bitstream was obtained through transcoding, then a second synthesis transform, different from the first synthesis transform, is used to process the second bitstream.

33. The method according to any one of claims 1-32, wherein at least one of the first bitstream or the second bitstream includes an indication of color transformation.

34. The method according to any one of claims 1-33, wherein at least one of the first bitstream or the second bitstream includes an indication of inverse color transformation.

35. The method according to any one of claims 1-34, wherein first information regarding at least one of the following is indicated in the bitstream: Whether to apply the method, or How to apply the method.

36. The method of claim 35, wherein the first information is indicated in one of the following locations: block level sequence level, Image group level, Image level, strip level, or Film series level.

37. The method of claim 35, wherein the first information is indicated in one of the following: The encoding / decoding structure of a codec tree unit (CTU) The encoding and decoding structure of the codec unit (CU), The encoding and decoding structure of the Transform Unit (TU) The encoding and decoding structure of the prediction unit (PU), The encoding and decoding structure of the codec tree block (CTB), The encoding and decoding structure of the codec block (CB), Transform Block (TB) encoding / decoding structure Encoding and decoding structure of prediction blocks (PB), Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

38. The method according to any one of claims 35-37, wherein the first information depends on the encoded / decoded information of the visual data.

39. The method of claim 38, wherein the encoded / decoded information comprises at least one of the following: Block size, Color format, Single tree segmentation Two-tree segmentation, Color components, Strip type, or Image type.

40. The method according to any one of claims 12-39, wherein the indication includes grammatical elements.

41. The method of claim 40, wherein the syntax element is binarized into one of the following: logo, Fixed length code Index Columbus (EG) code One-yuan code Rounding down a single-element code, or Rounding binary code.

42. The method according to any one of claims 40-41, wherein the syntax element is encoded or decoded using at least one context model, or The syntax elements therein are bypassed and encoded / decoded.

43. The method according to any one of claims 40-42, wherein the syntax elements are transmitted via signals based on conditions.

44. The method according to any one of claims 40-43, wherein the syntax element is indicated in one of the following locations: block level sequence level, Image group level, Image level, strip level, or Film series level.

45. The method according to any one of claims 40-44, wherein the syntax element is indicated in one of the following: The encoding / decoding structure of a codec tree unit (CTU) The encoding and decoding structure of the codec unit (CU), The encoding and decoding structure of the Transform Unit (TU) The encoding and decoding structure of the prediction unit (PU), The encoding and decoding structure of the codec tree block (CTB), The encoding and decoding structure of the codec block (CB), Transform Block (TB) encoding / decoding structure Encoding and decoding structure of prediction blocks (PB), Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

46. ​​The method according to any one of claims 1-45, wherein the visual data includes video, a picture of the video, or an image.

47. An apparatus for visual data transcoding, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-46.

48. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-46.

49. A non-transitory computer-readable recording medium storing a bitstream of visual data generated by a method performed by means of visual data transcoding, wherein the method includes: Reconstruct the visual data from the first bitstream of the visual data; as well as The reconstructed visual data is encoded into a second bitstream of the visual data, wherein at least one of the reconstruction or the encoding is performed using a neural network (NN)-based model.

50. A method for storing a bitstream of visual data, comprising: Reconstruct the visual data from the first bitstream of the visual data; The reconstructed visual data is encoded into a second bitstream of the visual data, wherein at least one of the reconstruction or the encoding is performed using a neural network (NN) based model; as well as The second bitstream is stored in a non-transitory computer-readable recording medium.