Adaptive Selection of Entropy Coding Parameters
Adaptive entropy coding parameter selection addresses the inefficiencies of predefined alphabet sizes by dynamically adjusting entropy coding parameters, improving signal quality and efficiency across varying bitrates.
Patent Information
- Application Number
- JP2024577047
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Conventional entropy coding methods predefined alphabet sizes lead to clipping effects under high bitrates and inefficiencies under low bitrates, degrading reconstructed signal quality and efficiency.
Adaptive selection of entropy coding parameters based on parameters carried in the bitstream, allowing the encoder to adjust alphabet sizes dynamically, avoiding clipping and optimizing operation across varying bitrates.
This approach enhances reconstructed signal quality at high bitrates by eliminating clipping and reduces bitrates at low bitrates, achieving optimal encoding efficiency.
Smart Images

Figure 2025522817000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to entropy encoding and decoding. In particular, the present disclosure relates to the adaptive selection of entropy coding parameters.
Background Art
[0002] Video coding (video encoding and decoding) is used in a wide range of digital video applications such as, for example, broadcast digital TV, video transmission via the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVDs and Blu-ray discs, video content acquisition and editing systems, mobile device video recording, and camcorders for security applications.
[0003] Even for relatively short videos, the amount of video data required to depict them can be quite substantial, and as a result, difficulties can arise when data is streamed or otherwise communicated across a communication network with limited bandwidth capacity. With network resources being limited and the demand for higher video quality increasing ever more, there is a desire for improved compression and decompression techniques that improve the compression rate without sacrificing much or any of the picture quality. Video encoding and decoding can be performed by standard video encoders and decoders that are compatible with, for example, H.264 / AVC, HEVC (H.265), VVC (H.266) or other video coding techniques. Further, video coding or a part thereof may be performed by a neural network.
[0004] Entropy coding is widely used in any encoding or decoding, or other source signals such as still images, pictures, or feature channels of neural networks. The input alphabet of the entropy encoder is finite, and the size of the input alphabet must be known on both the encoder side and the decoder side. A coder with a larger input alphabet size can encode a wider symbol range, but it is less efficient than the same coder with a smaller input alphabet. Due to such an effect, it is optimal to use the smallest possible alphabet. In the conventional method, the entropy coding parameters, especially the input alphabet size, are predefined and used for all possible input signals, which results in a clipping effect under high bitrate conditions and improper bit waste under low bitrate conditions. As a result, the reconstructed quality and coding efficiency deteriorate. Summary of the Invention Problems to be Solved by the Invention
[0005] Embodiments of the present disclosure provide an apparatus and method for entropy encoding data into a bitstream and entropy decoding data from the bitstream.
[0006] Embodiments of the present invention are defined by the features of the independent claims, and further advantageous implementations of the embodiments are defined by the features of the dependent claims. Means for Solving the Problems
[0007] According to a first aspect, an embodiment of the present application provides a decoding method implemented by a decoder, the decoding method including: receiving a bitstream including encoded data of an input signal and a first parameter; analyzing the bitstream to obtain the first parameter; obtaining an entropy coding parameter based on the first parameter; and reconstructing at least a part of the input signal based on the entropy coding parameter.
[0008] In the conventional method, the entropy coding parameter is usually predefined. For example, the alphabet size M is usually selected once based on the expected tensor range (or the potential tensor range), and is predefined by using the predefined alphabet size M for all cases. Since the size of the input alphabet of the entropy encoder is the same as the size of the output alphabet of the entropy decoder, in this specification, the alphabet size M represents the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder. In such a case, when the actual tensor range is wider than the expected tensor range, the input alphabet size determined based on the expected tensor range is not appropriate, and then clipping is required for the coded tensor values. Such clipping damages the signal, especially when the coded tensor range is very different from the alphabet size. In this case, the damage to the coded tensor is a non-linear distortion that causes an unpredictable error in the reconstructed signal, and thus the quality of the reconstructed signal may be greatly degraded. In one implementation, a very large alphabet size can be selected and used for all cases, but increasing the alphabet size disadvantages the compression efficiency under low bitrate conditions, and using a large alphabet size significantly increases the bitrate without improving the reconstruction quality.
[0009] In an embodiment of the present application, the decoder can obtain entropy coding parameters (especially the alphabet size) based on the parameters carried in the bitstream, and since the parameters carried in the bitstream can be changed, the encoder can adaptively adjust the entropy coding parameters by changing the parameters carried in the bitstream. Therefore, the clipping effect can be avoided under high bitrate conditions, and the rate overhead caused by an unduly large alphabet size can also be avoided under low bitrate conditions. In other words, due to the adaptability of the entropy coding parameters, especially the alphabet size, optimal operation of the entropy encoder is possible at low bitrates (corresponding to a narrow range of coded values), resulting in bitrate savings, and at high bitrates (corresponding to a wide range of coded values), the clipping effect is eliminated, resulting in a higher reconstructed signal quality.
[0010] Here, it should be noted that the "entropy coder" can be used as a synonym for the "entropy coding algorithm" that includes both the encoding algorithm and the decoding algorithm. The entropy encoder may be a module that is part of the encoder, and the entropy decoder may be another module that is part of the decoder. The parameters of the entropy encoder and the entropy decoder should be synchronized for correct operation. Therefore, the term "parameters of the entropy encoder" or "entropy coding parameters" means the parameters of both the entropy encoder and the entropy decoder. In other words, the "entropy coding parameters" may be equal to the "parameters of the entropy encoder and the entropy decoder". The entropy encoder encodes the symbols of the alphabet into one or more bits in the bitstream, and the entropy decoder decodes one or more bits in the bitstream into the symbols of the alphabet. On the entropy encoder side, the alphabet means the input alphabet, and on the entropy decoder side, the alphabet means the output alphabet. The size of the input alphabet on the entropy encoder side is equal to the size of the output alphabet on the entropy decoder side.
[0011] In one possible embodiment, the input signal includes video data, image data, point claud data, motion flow, or motion vectors, or any other type of media data.
[0012] In one possible embodiment, the entropy coding parameter includes at least one of: the size of the alphabet of the entropy coder, where the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the renormalization period of the entropy coder. In some embodiments, the renormalization period can be 8 bits, 16 bits, etc.
[0013] Three possible schemes for alphabet size derivation based on the bitstream are conceivable: 1) Explicit signaling using a predefined predictor; 2) Derivation from the quantization parameter (β); 3) Explicit signaling using a predictor according to the quantization parameter.
[0014] The technology can be applied to any type of coder that uses entropy coding in the pipeline.
[0015] In one possible embodiment, the first parameter is the size of the alphabet, and here, based on the first parameter, the step of obtaining the entropy coding parameter includes using the first parameter as the size of the alphabet.
[0016] In this embodiment, the size of the alphabet is directly signaled in the bitstream using, for example, fixed-length coding, exp-Golomb coding, or some other coding algorithm. Typical values of M can be 256, 512, 1024. For example, when signaling 1024 using a fixed-length code, 11 bits are required (1024 10=(100000000002), when signaling log2(1024) - 9 = 1, only 1 bit is required if only the values 512 and 1024 are allowed, or 2 bits are required if four different alphabet sizes such as 512, 1024, 2048, 4096 are allowed. As a result, the direct signaling of M consumes more bits. However, in some exotic cases (e.g., when the alphabet size M is not a power of 2), the direct signaling of M may be useful.
[0017] In one possible embodiment, the first parameter is p, and the entropy coding parameter includes the size of the alphabet M, where M is a function of p.
[0018] In one possible embodiment, based on the first parameter, the step of obtaining the entropy coding parameter includes M = f -1 (p), where f -1 (p) is the inverse function of f(M), and f(M) = p.
[0019] In this embodiment, instead of M itself, the output p of some invertible function f(M) is signaled within the bitstream. Such p can be signaled using fixed - length coding, exp - Golomb coding, or some other coding algorithm. Thus, on the decoder side, M is derived based on p, specifically, M is derived as M = f -1 (p). The advantage of the above embodiment is that any optimal alphabet size selected on the encoder side can be signaled, thus improving the flexibility in signaling the alphabet size. In some embodiments, p is non - negative, but in other embodiments, p can also be negative. For example, the value p can be in the range [0, 5], and 3 bits are used for signaling. The function f(M) can be negotiated in advance between the encoder side and the decoder side.
[0020] In one possible embodiment, M satisfies one of the following: M = k^p, where k is a natural number; or M = k^(p + C), where k is a natural number and C is a constant; or M = k^(a*p + C), where k is a natural number and a and C are constants; or M = a*p + b, where a and b are constants; or M = p^2. In any one of the embodiments, A^B is A B It should be noted that it means.
[0021] In one possible embodiment, p = log2(M) - 9 and M = f -1 (p) = 2^(p + 9), where f -1 (p) is the inverse function of f(M), and f(M) = log2(M) - 9.
[0022] In one possible embodiment, p is signaled using one of a binary code, a unary code, a truncated unary code, or an exp-Golomb code.
[0023] In one possible embodiment, p is signaled using an exp-Golomb code of degree 0.
[0024] In one possible embodiment, the alphabet size is signaled, for example, within the parameter set section of the bitstream, for example, within the picture parameter set section of the bitstream.
[0025] In one possible embodiment, the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, the density of pixels within a 3D object, or a rate-distortion weight coefficient.
[0026] In the above embodiment, the alphabet size can be derived based on several other parameters. In one exemplary implementation, the alphabet size is derived from a quantization parameter or a rate control parameter. Alternatively, the alphabet size can be derived from an image resolution, a video resolution, a frame rate, the density of pixels within a 3D object, etc. In a trainable codec, the alphabet size can be derived from some parameters of the loss function used during training, such as the rate / distortion weighting factor, or some parameters that affect the selection of the gain vector g. The loss function may include rate and distortion components such as peak signal-to-noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM), video multi-method assessment fusion (VMAF), or some other quality metric. For example, the loss function can be loss = beta * distortion + bits, where distortion is measured by PSNR or MS-SSIM or VMAF, bits is the number of bits spent, and beta is a weighting parameter that controls the ratio of the bitrate to the reconstructed quality, and beta may sometimes be referred to as a rate control parameter. Also, it can be a quantization parameter such as the quantization parameter (qp) in normal codecs like JPEG, HEVC, VVC.
[0027] The advantage of the above embodiment is that since the quantization parameter or the rate control parameter already exists in the bitstream and is used for other procedures, such parameters can be used by the decoder side to derive the alphabet size M, and there is no need for additional signaling for the information specifically used to indicate the alphabet size M, thus the bitrate can be saved.
[0028] In one possible embodiment, based on the first parameter, the step of obtaining the entropy coding parameter includes determining a target sub-range where the first parameter is located. The allowable range of the value of the first parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the first parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter. The step of using the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter, or calculating the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0029] In the above embodiment, assuming such a rate control parameter is β, the range of beta (β) is divided into K intervals (K sub-ranges) as follows: [β_0,β_1),[β_1,β_2),..,[β_(K-1),β_K)
[0030] Each of the intervals / sub-ranges corresponds to one alphabet size value Mi. Note that there is a range allowed for a particular codec value β. For example, for some codecs, β is allowed to be in the range [-∞,∞], and for other codecs, β may only be allowed to be in the range [0,∞]. Within the context of this embodiment, the original large range of allowed β values is divided into several sub-ranges, and for all sub-ranges, there exists a specific value of the alphabet size. After obtaining the parameter β, the decoder can select a target interval based on the β value obtained from the bitstream. Specifically, the decoder determines that β_i≦β≦β_(i+1), then selects the interval [β_i,β_(i+1)] as the target interval, and derives the alphabet size value Mi corresponding to this target interval as the alphabet size value M. In some embodiments, each βi within the range {βi} of beta can correspond to one alphabet size value Mi, and the alphabet size value M corresponding to a particular β is calculated based on one or more values Mi corresponding to βi adjacent to β. The value used to calculate M may be the nearest just value Mi corresponding to the target interval, or linear interpolation, bilinear interpolation, or some other interpolation from two or more Mi corresponding to βi adjacent to β, or some other interpolation from two or more Mi corresponding to intervals adjacent to the target interval.
[0031] In one possible embodiment, the first parameter is D, and the entropy coding parameter includes the size of the alphabet M, where M is obtained based on P and D, and P is a predictor that can be derived by the decoder.
[0032] In the above embodiment, the alphabet size can be derived based on the predictor P and the first parameter signaled in the bitstream. Therefore, when receiving the bitstream, the decoder can derive the predictor P based on a predetermined parameter, analyze the first parameter from the bitstream, and then derive the alphabet size M based on the predictor P and the first parameter. The advantage of the above embodiment is that only the difference between P and M is signaled within the bitstream, so that compared with M signaled within the bitstream, the additional bits consumed are reduced. Also, the difference between P and M can be selected based on the content or the bit rate, and the flexibility in signaling the alphabet size is also improved. Therefore, this embodiment provides flexibility in alphabet size selection while minimizing the additional bits spent on signaling. Even in some rare cases where the size of the alphabet predicted from β does not function well, the encoder can still signal the difference value between M and P. This incurs a cost of several bits, but can solve serious problems regarding the clipping effect.
[0033] In one possible embodiment, the step of obtaining the entropy coding parameter based on the first parameter is M = s -1 (D, P), where s -1 (D, P) is the inverse function of s(M, P) and s(M, P)=D.
[0034] In one possible embodiment, s(M, P) includes the following, i.e., s(M, P)=log k (M)-log k (P), where k is a natural number; or s(M, P)=log k (P)-log k (M), where k is a natural number; or s(M, P)=log k (M)-log k (P)-C, where k is a natural number and C is an integer; or s(M, P)=log k(P)-log k (M)-C, where k is a natural number and C is an integer; or s(M,P) = a*log k (P)-b*log k (M)-c, where k is a natural number and a, b, and c are constants; or s(M,P) = a*M + b*P + c, where a, b, and c are constants.
[0035] In one possible embodiment, M = 2^(D + log2(P)), D = s(M,P) = log2(M) - log2(P).
[0036] The invertible function D = s(M,P) can be considered as D = s P (M), and M = s -1 (D,P) can be considered as M = s -1 P (D), where P can be any fixed number, that is, note that P is a constant coefficient.
[0037] In one possible embodiment, D is signaled using one of the following codes: a binary code or a unary code or a truncated unary code or an exp-Golomb code.
[0038] In one possible embodiment, D is signaled using an exp-Golomb code of degree 0.
[0039] In one possible embodiment, P can be derived based on at least one parameter other than the first parameter carried in the bit stream.
[0040] In one possible embodiment, at least one parameter other than the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient.
[0041] In one possible embodiment, P includes: obtaining a rate control parameter beta (β) from a bitstream; determining a target subrange in which the obtained β is located, wherein the allowable range of the value of the rate control parameter β is [β_0, β_K], the allowable range [β_0, β_K] includes a plurality of subranges, the target subrange is one of the plurality of subranges, each of the plurality of subranges includes at least one value of β, and each of the plurality of subranges corresponds to one value of P; selecting, as the value of P, a value corresponding to the target subrange; or calculating the value of P based on one or more values corresponding to one or more subranges adjacent to the target subrange, and is derived based on at least one parameter.
[0042] In one possible embodiment, the decoding method further includes analyzing the bitstream to obtain a flag, where the flag is used to indicate whether the entropy coding parameter is directly carried in the bitstream.
[0043] In the above embodiment, a flag can be introduced into the bitstream to indicate the switching between three embodiments. In this case, 2 bits may be required for this flag. In another possible embodiment, a flag can be used to indicate the switching between two embodiments. In this case, only 1 bit is required.
[0044] In one possible embodiment, when the flag is equal to a first value, it is specified that the entropy coding parameter is carried in the bitstream. In this case, the first parameter is the entropy coding parameter, or the first parameter is the conversion result of the entropy coding parameter. Or when the flag is equal to a second value, it is specified that the entropy coding parameter is not carried in the bitstream, and the entropy coding parameter can be derived by the decoder.
[0045] Such a solution provides a balance between bit savings and flexibility, and in most cases, if the derived entropy parameter is appropriate, only 1 bit is spent on the indication, while in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0046] In one possible embodiment, when the flag is equal to the third value, it is specified that the difference value between M and P, or the conversion result of the difference value between M and P, is carried in the bitstream. In this case, the first parameter is the difference value between M and P, or the conversion result of the difference value between M and P, where M is the size of the input alphabet and P is a predictor that can be derived by the decoder.
[0047] In one possible embodiment, the entropy coder is an arithmetic coder, a range coder, or an ANS (Asymmetric Numerical Systems) coder.
[0048] In one possible embodiment, based on the entropy coding parameters, the step of reconstructing at least a part of the input signal includes the step of obtaining at least one probability model, where a probability model of the output symbol is used to indicate the probability of each possible value of the output symbol, the step of entropy decoding one or more bits in the bitstream by using the at least one probability model and the entropy coding parameters to obtain one or more output symbols, and the step of reconstructing at least a part of the input signal based on the one or more output symbols.
[0049] In one possible embodiment, the method further includes the step of updating the probability model. For example, the probability model is updated after each output symbol, so that each output symbol has a unique probability distribution of possible values. Note that the probability model may sometimes be called a probability distribution.
[0050] In one possible embodiment, the probability model depends on the entropy coding parameters. For example, symbol probabilities are distributed according to a normal distribution N(μ,σ), where N(μ,σ) represents a Gaussian distribution with a mean equal to μ and a variance equal to σ 2 equal. However, an actual probability model, such as a quantized histogram (which also means a mathematical model or a theoretical model), depends on the alphabet size and probability precision within the entropy coding engine or entropy coder. That is, the entropy coding parameters can affect the histogram configuration inside the entropy encoder. Basically, since the alphabet size is the number of possible symbol values, for example, when the alphabet size is equal to 4, a larger value, such as the value "7", cannot be encoded / decoded. The histogram used in the entropy coder is composed of the quantized probabilities of each symbol value. For example, the alphabet is {0,1,2,3}, and the corresponding probabilities are {7 / 16,7 / 16,1 / 16,1 / 16}. Each probability is not zero, the sum of the probabilities is equal to 1, and each of the probabilities is greater than the minimum probability (probability precision) supported by the entropy coding engine (1 / 16 in this example). If the probabilities of some symbols are lower than the minimum probability supported by the entropy coding engine, the probabilities of at least some symbols need to be adjusted to ensure that the probability of each symbol is greater than the minimum probability supported by the entropy coding engine.
[0051] According to a second aspect, embodiments of the present application provide a decoding method for entropy decoding a bitstream, the method comprising: receiving a bitstream including decoded data of an input signal; analyzing the bitstream to obtain a flag, the flag being used to indicate whether entropy coding parameters are directly carried within the bitstream; obtaining entropy coding parameters based on the flag; and reconstructing at least a part of the input signal based on the entropy coding parameters.
[0052] In the above embodiment, a flag can be introduced into the bitstream to indicate switching between three embodiments, in which case two bits may be required for this flag. In another possible embodiment, a flag can be used to indicate switching between two embodiments, in which case only one bit is required. Such a solution provides a balance between bit savings and flexibility, and in most cases, if the derived entropy parameters are appropriate, only one bit is used for the indication, while in some specific cases, it may be possible to explicitly signal the entropy parameters.
[0053] In one possible embodiment, the entropy coding parameters include at least one of the following: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder.
[0054] In one possible embodiment, when the flag is equal to the first value, it is specified that the entropy coding parameter is conveyed in the bitstream, or the conversion result of the entropy coding parameter is conveyed in the bitstream, or when the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed in the bitstream, but the entropy coding parameter can be derived by the decoder.
[0055] In one possible embodiment, when the flag is equal to the third value, it is specified that the difference value between M and P is conveyed in the bitstream, or the conversion result of the difference value between M and P is conveyed in the bitstream, where M is an entropy coding parameter and P is a predictor that can be derived by the decoder.
[0056] In one possible embodiment, based on the flag, the step of obtaining the entropy coding parameter includes, when the flag is equal to the first value, the step of analyzing the bitstream to obtain a first parameter, where the first parameter is an entropy coding parameter, the step of using the first parameter as the entropy coding parameter, or the first parameter is the conversion result of the entropy coding parameter, and the step of obtaining the entropy coding parameter based on the first parameter.
[0057] In one possible embodiment, the conversion result of the entropy coding parameter is p = f(M), where M is the entropy coding parameter and f(M) includes the following, that is: f(M)=log k (M), where k is a natural number; or f(M)=a*log k(M)-C, where k is a natural number, and a and C are predetermined constants; or f(M) = a*M + R, where a and R are predetermined constants; or f(M) = sqrt(M), and the step of obtaining the entropy coding parameter based on the first parameter is M = f -1 (p), where f -1 (p) is the inverse function of f(M).
[0058] In one possible embodiment, the first parameter is p = log2(M) - 9.
[0059] In one possible embodiment, the step of obtaining the entropy coding parameter based on the flag includes, when the flag is equal to the second value, analyzing the bit stream to obtain a second parameter, where the second parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a pixel density within a 3D object, or a rate distortion weight coefficient, and the step of deriving the entropy coding parameter based on the second parameter.
[0060] In one possible embodiment, the step of deriving the entropy coding parameter based on the second parameter includes determining a target sub-range in which the second parameter is located, where the allowable range of the value of the second parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the second parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter, the step of using the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter, or calculating the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0061] In one possible embodiment, the step of obtaining an entropy coding parameter based on a flag includes, when the flag is equal to a third value, analyzing a bit stream to obtain a third parameter, where the third parameter is a difference value between M and P, or the third parameter is a conversion result of the difference value between M and P, M is an entropy coding parameter, and P is a predictor derived by a decoder; a step of deriving P based on at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient; and a step of obtaining an entropy coding parameter based on the third parameter and P.
[0062] In one possible embodiment, the conversion result of the difference value between M and P is D = s(M, P), where s(M, P) is a reversible function, and s(M, P) includes the following, that is: s(M, P) = log k (M) - log k (P), where k is a natural number, or s(M, P) = log k (P) - log k (M), where k is a natural number, or s(M, P) = log k (M) - log k (P) - C, where k is a natural number and C is an integer, or s(M, P) = log k (P) - log k (M) - C, where k is a natural number and C is an integer, or s(M, P) = a * log k (P) - b * log k (M) - c, where k is a natural number and a, b, and c are constants, or s(M, P) = a * M + b * P + c, where a, b, and c are constants, and the step of obtaining an entropy coding parameter based on the third parameter includes M = s-1 (D, P), where s -1 (D, P) is the inverse function of s(M, P).
[0063] According to a third aspect, an embodiment of the present application provides an encoding method implemented by an encoder. The method includes encoding an input signal and a flag into a bitstream with a first parameter, where the first parameter is used to obtain an entropy coding parameter, and transmitting the bitstream to a decoder.
[0064] In an embodiment of the present application, the decoder can obtain an entropy coding parameter (especially the alphabet size) based on the parameter carried in the bitstream. Since the parameter carried in the bitstream is changeable, the encoder can adaptively adjust the entropy coding parameter by changing the parameter carried in the bitstream. Therefore, the clipping effect can be avoided under high bitrate conditions, and the rate overhead caused by an unduly large alphabet size can also be avoided under low bitrate conditions. In other words, due to the adaptability of the entropy coding parameter, especially the alphabet size, the optimal operation of the entropy encoder is possible at low bitrates (corresponding to a narrow range of coded values), resulting in bitrate savings, and there is no clipping effect at high bitrates (corresponding to a wide range of coded values), resulting in high quality of the reconstructed signal.
[0065] In one possible embodiment, the entropy coding parameter includes at least one of the following, namely: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder.
[0066] In one possible embodiment, the first parameter is the size of the alphabet.
[0067] In one possible embodiment, the first parameter is p, where p is the conversion result of M and M is the entropy coding parameter.
[0068] In one possible embodiment, p = f(M), where f(M) is an invertible function
[0069] In one possible embodiment, f(M) includes the following, namely: f(M)=a*log k (M)-C, where k is a natural number, and a and C are predetermined constants, or f(M)=a*M+b, where a and b are constants, or f(M)=sqrt(M).
[0070] In one possible embodiment, p = log2(M)-9.
[0071] In one possible embodiment, p is signaled using one of the following codes, namely: a binary code, or a unary code, or a truncated unary code, or an exp-Golomb code.
[0072] In one possible embodiment, the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, the density of pixels within a 3D object, or a rate distortion weight factor, and the first parameter is used by an entropy decoder to derive entropy coding parameters.
[0073] In one possible embodiment, the first parameter is D obtained based on P and M, where M is an entropy coding parameter and P is a predictor that can be derived by a decoder.
[0074] In one possible embodiment, D = s(M, P), where s(M, P) is a reversible function.
[0075] In one possible embodiment, s(M, P) includes the following, that is: s(M, P)=log k (M)-log k (P), where k is a natural number, or s(M, P)=log k (P)-log k (M), where k is a natural number, or s(M, P)=log k (M)-log k (P)-C, where k is a natural number and C is an integer, or s(M, P)=log k (P)-log k (M)-C, where k is a natural number and C is an integer, or s(M, P)=a*log k (P)-b*log k (M)-c, where k is a natural number and a, b, and c are constants, or s(M, P)=a*M + b*P + c, where a, b, and c are constants. including.
[0076] In one possible embodiment, D = s(M, P) = log2(P) - log2(M).
[0077] In one possible embodiment, D is signaled using one of the following codes: a binary code, or a unary code, or a truncated unary code, or an exp-Golomb code.
[0078] In one possible embodiment, the encoding method further includes the step of encoding a flag into the bit stream, where the flag is used to indicate whether the entropy coding parameter is directly carried within the bit stream.
[0079] In one possible embodiment, when the flag is equal to a first value, it is specified that the entropy coding parameter is carried within the bit stream and that the first parameter is either the entropy coding parameter or a conversion result of the entropy coding parameter; or when the flag is equal to a second value, it is specified that the entropy coding parameter is not carried within the bit stream but can be derived by the decoder.
[0080] In one possible embodiment, when the flag is equal to a third value, it is specified that the difference value between M and P or a conversion result of the difference value between M and P is carried within the bit stream, where M is the entropy coding parameter and P is a predictor that can be derived by the decoder.
[0081] In one possible embodiment, several possible solutions are proposed for alphabet selection on the encoder side.
[0082] In one possible embodiment, the method comprises the step of obtaining the minimum and maximum values of the latent space elements of the entropy encoder, where the latent space elements are the result of the procession of the input signal, and the step of obtaining the size of the alphabet according to M = ceil(max{y} - min{y}) or M = 2^(ceil(log2(max{y} - min{y}))) and further comprising, where ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space elements, min{y} represents the minimum value of the latent space elements, and M represents the size of the alphabet.
[0083] In this embodiment, the alphabet size is selected as the smallest possible number that is larger than the range of the coded values. For example, the minimum and maximum values of the tensor y are first obtained, and the alphabet size is selected as follows: M = ceil(max{y} - min{y})
[0084] Note that in most entropy encoders, the size of the alphabet should be a power of 2, in which case the size of the alphabet can be selected as M = 2^(ceil(log2(max{y} - min{y}))). It should be noted that there are some cases, for example when the modulus of all y values is less than 1, an additional scaling operation can be performed before entropy coding.
[0085] In one possible embodiment, the method further includes the steps of obtaining at least two values centered around M0, where M0 = ceil(max{y} - min{y}) or M0 = 2^(ceil(log2(max{y} - min{y}))), calculating a loss function for at least two values, and selecting the value with the minimum loss function among at least two values as the size of the alphabet, where ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space elements, and min{y} represents the minimum value of the latent space elements.
[0086] The loss function can include a rate component and a distortion component. For example, the loss function can be: Loss = beta * distortion + bits, where distortion is measured by, for example, peak signal-to-noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM), video multi-method assessment fusion (VMAF), or other quality metrics, bits is the number of bits spent, and beta is a weighting parameter that controls the ratio between the bitrate and the reconstructed quality. Beta is also called a rate control parameter. In this approach, clipping may occur, but the bitrate savings due to the use of a smaller alphabet compensate for a slight increase in distortion.
[0087] According to a fourth aspect, an embodiment of the present application provides an encoding method implemented by an encoder. The method includes encoding an input signal and a flag into a bitstream, where the flag is used to indicate whether entropy coding parameters are directly carried in the bitstream, and transmitting the bitstream to a decoder.
[0088] In the above embodiment, a flag can be introduced into the bit stream to indicate the switching between three embodiments. In this case, two bits may be required for this flag. In another possible embodiment, a flag can be used to indicate the switching between two embodiments. In this case, only one bit is required. Such a solution provides a balance between bit savings and flexibility. In most cases, if the derived entropy parameter is appropriate, only one bit is used for the indication. On the other hand, in some specific cases, the entropy parameter can be explicitly signaled.
[0089] In one possible embodiment, the entropy coding parameter includes at least one of the following, namely: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder.
[0090] In one possible embodiment, when the flag is equal to a first value, it is specified that the entropy coding parameter is conveyed in the bit stream, or the conversion result of the entropy coding parameter is conveyed in the bit stream. Or when the flag is equal to a second value, it is specified that the entropy coding parameter is not conveyed in the bit stream, but the entropy coding parameter can be derived by the decoder.
[0091] In one possible embodiment, when the flag is equal to a third value, it is specified that the difference value between M and P is conveyed in the bit stream, or the conversion result of the difference value between M and P is conveyed in the bit stream, where M is an entropy coding parameter and P is a predictor that can be derived by the decoder.
[0092] In one possible embodiment, the method further includes encoding a first parameter into a bit stream when the flag is equal to a first value, where the first parameter is an entropy coding parameter or the first parameter is a conversion result of an entropy coding parameter.
[0093] In one possible embodiment, the conversion result of the entropy coding parameter is p = f(M), where M is the entropy coding parameter and f(M) can be as follows: f(M) = log k (M), where k is a natural number, or f(M) = a * log k (M) - C, where k is a natural number and a and C are predetermined constants, or f(M) = aM + R, where a and R are constants, or f(M) = sqrt(M).
[0094] In one possible embodiment, the first parameter is p = log2(M) - 9.
[0095] In one possible embodiment, p is signaled using one of the following codes: binary code, or unary code, or truncated unary code, or exp-Golomb code.
[0096] In one possible embodiment, p is signaled using an exp-Golomb code of degree 0.
[0097] In one possible embodiment, the method further includes encoding a third parameter into a bit stream when the flag is equal to a third value, where the third parameter is a difference value between M and P or the third parameter is a conversion result of the difference value between M and P, M is an entropy coding parameter, and P is a predictor that can be derived by a decoder.
[0098] In one possible embodiment, the conversion result of the difference value between M and P is D = s(M, P), where s(M, P) is a reversible function, and s(M, P) includes the following, that is: s(M, P) = log k (M) - log k (P), where k is a natural number, or s(M, P) = log k (P) - log k (M), where k is a natural number, or s(M, P) = log k (M) - log k (P) - C, where k is a natural number and C is an integer, or s(M, P) = log k (P) - log k (M) - C, where k is a natural number and C is an integer, or s(M, P) = a * log k (P) - b * log k (M) - c, where k is a natural number and a, b, and c are constants, or s(M, P) = a * M - b * P + c, where a, b, and c are constants, including The step of obtaining the entropy coding parameter based on the third parameter includes M = s -1 (D, P), where s -1 (D, P) is the inverse function of s(M, P).
[0099] In one possible embodiment, D is signaled using one of the following codes, namely: binary code, or unary code, or truncated unary code, or exp - Golomb code.
[0100] In one possible embodiment, D is signaled using an exp - Golomb code of degree 0.
[0101] According to a fifth aspect, an embodiment of the present application provides a decoding apparatus including a receiver configured to receive a bitstream including encoded data of an input signal and a first parameter, an analysis unit configured to analyze the bitstream to obtain the first parameter, an acquisition unit configured to obtain an entropy coding parameter based on the first parameter, and a reconstruction unit configured to reconstruct at least a part of the input signal based on the entropy coding parameter.
[0102] This apparatus provides the advantages of the above-described method.
[0103] In one possible embodiment, the input signal includes video data, image data, dot cloud data, motion flow, or motion vectors, or any other type of media data.
[0104] In one possible embodiment, the entropy coding parameter includes at least one of the following: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc.
[0105] In one possible embodiment, the first parameter is the size of the alphabet, and the acquisition unit is further configured to use the first parameter as the size of the alphabet.
[0106] In one possible embodiment, the first parameter is p, and the entropy coding parameter includes the size of the alphabet M, where M is a function of p.
[0107] In one possible embodiment, the acquisition unit is further configured to acquire M as M = f -1 (p), where f -1 (p) is the inverse function of f(M) and f(M) = p.
[0108] In one possible embodiment, the acquisition unit is configured to: determine a target sub-range where the first parameter is located, where the allowable range of values of the first parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the first parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter; use the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter; or further configured to calculate the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0109] According to a sixth aspect, an embodiment of the present application provides a decoding apparatus including a functional unit that implements the encoding method in any one of the second aspect or possible embodiments of the second aspect.
[0110] This apparatus provides the advantages of the above-described method.
[0111] According to a seventh aspect, an embodiment of the present application provides an encoding apparatus configured to encode an input signal and a first parameter into a bit stream, where the first parameter includes an encoding unit used to acquire an entropy coding parameter and a transmission unit configured to transmit the bit stream to a decoder. The encoding apparatus further includes other functional units for implementing the encoding method in any one of the possible embodiments of the third aspect.
[0112] According to an eighth aspect, an embodiment of the present application provides an encoding apparatus including an encoding unit configured to encode an input signal and a flag into a bit stream, where the flag is used to indicate whether entropy coding parameters are directly carried in the bit stream, and a transmission unit configured to transmit the bit stream to a decoder. The encoding apparatus further includes other functional units for implementing the encoding method in any one of the possible embodiments of the fourth aspect.
[0113] According to a ninth aspect, an embodiment of the present application provides a decoding apparatus including a processing circuit configured to execute the decoding method described in any one of the first aspects or any possible embodiment of the first aspect.
[0114] According to a tenth aspect, an embodiment of the present application provides a decoding apparatus including a processing circuit configured to execute the decoding method described in any one of the second aspects or any possible embodiment of the second aspect.
[0115] According to an eleventh aspect, an embodiment of the present application provides an encoding apparatus including a processing circuit configured to execute the encoding method described in any one of the third aspects or any possible embodiment of the third aspect.
[0116] According to a twelfth aspect, an embodiment of the present application provides an encoding apparatus including a processing circuit configured to execute the encoding method described in any one of the fourth aspects or any possible embodiment of the fourth aspect.
[0117] According to a thirteenth aspect, an embodiment of the present application provides a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors. The storage medium stores programming for execution by the one or more processors, and when the programming is executed by the one or more processors, configures the decoder to execute the decoding method described in any one of the first aspects or any possible embodiment of the first aspect.
[0118] According to a 14th aspect, an embodiment of the present application provides a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors, the storage medium storing programming for execution by the one or more processors, the programming configuring the decoder to perform the method described in any one of the 2nd aspect or a possible embodiment of the 2nd aspect when executed by the one or more processors.
[0119] According to a 15th aspect, an embodiment of the present application provides an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors, the storage medium storing programming for execution by the one or more processors, the programming configuring the decoder to perform the method described in any one of the 3rd aspect or a possible embodiment of the 3rd aspect when executed by the one or more processors.
[0120] According to a 16th aspect, an embodiment of the present application provides an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors, the storage medium storing programming for execution by the one or more processors, the programming configuring the decoder to perform the method described in any one of the 4th aspect or a possible embodiment of the 4th aspect when executed by the one or more processors.
[0121] According to a 17th aspect, an embodiment of the present application provides a non-transitory computer-readable medium that conveys computer instructions that, when executed by a computer device or one or more processors, cause the computer device or the one or more processors to perform the method described in any one of the 1st aspect or a possible embodiment of the 1st aspect.
[0122] According to an 18th aspect, an embodiment of the present application provides a non-transitory computer-readable medium that, when executed by a computer device or one or more processors, conveys computer instructions that cause the computer device or the one or more processors to execute the method described in any one of the 2nd aspect or possible embodiments of the 2nd aspect.
[0123] According to a 19th aspect, an embodiment of the present application provides a non-transitory computer-readable medium that, when executed by a computer device or one or more processors, conveys computer instructions that cause the computer device or the one or more processors to execute the method described in any one of the 3rd aspect or possible embodiments of the 3rd aspect.
[0124] According to a 20th aspect, an embodiment of the present application provides a non-transitory computer-readable medium that, when executed by a computer device or one or more processors, conveys computer instructions that cause the computer device or the one or more processors to execute the method described in any one of the 4th aspect or possible embodiments of the 4th aspect.
[0125] According to a 21st aspect, an embodiment of the present application provides a non-transitory storage medium that includes a bitstream encoded by the method described in any one of the 3rd aspect or possible embodiments of the 3rd aspect.
[0126] According to a 22nd aspect, an embodiment of the present application provides a non-transitory storage medium that includes a bitstream encoded by the method described in any one of the 4th aspect or possible embodiments of the 4th aspect.
[0127] According to a 23rd aspect, an embodiment of the present application provides a computer program stored in a non-transitory medium and including code instructions that, when executed on one or more processors, cause the steps of the method according to any one of the foregoing aspects or any one of the possible embodiments of the foregoing aspects to be executed.
[0128] According to the 24th aspect, an embodiment of the present application is a system for delivering a bitstream, including at least one storage medium configured to store at least one bitstream generated by the encoding method described in any one of the 3rd aspect or possible embodiments of the 3rd aspect, and a video streaming device configured to obtain a bitstream from one of the at least one storage media and transmit the bitstream to a terminal device, wherein the video streaming device includes a content server or a content delivery server.
[0129] In one possible embodiment, it further includes one or more processors configured to perform an encryption process on at least one bitstream to obtain at least one encrypted bitstream, and at least one storage medium configured to store the encrypted bitstream, or one or more processors configured to convert a bitstream in a first format into a bitstream in a second format, and at least one storage medium configured to store the bitstream in the second format.
[0130] In one possible embodiment, it further includes a receiver configured to receive a first operation request, one or more processors configured to determine a target bitstream in at least one storage medium in response to the first operation request, and a transmitter configured to transmit the target bitstream to a terminal-side device.
[0131] In one possible embodiment, the one or more processors are further configured to encapsulate the bitstream to obtain a transport stream in a first format, and the transmitter is further configured to transmit the transport stream in the first format to a terminal-side device for display or transmit the transport stream in the first format to a storage space for storage.
[0132] The present invention can be implemented in hardware (HW) and / or software (SW), or any combination thereof. Furthermore, a HW-based implementation may be combined with a SW-based implementation.
[0133] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0134] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings.
[0135]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Mode for Carrying Out the Invention
[0136] In the following description, reference is made to the accompanying drawings, which form a part of the present disclosure and illustrate specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention can be used. It is understood that embodiments of the present invention may be used in other aspects and may include structural or logical changes not shown in the figures. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0137] Disclosures related to the methods described, for example, are understood to apply also to corresponding devices or systems configured to perform the methods, and vice versa. For example, when one or more specific method steps are described, the corresponding device may include one or more units, such as functional units (e.g., one unit for performing one or more of the described steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly described or illustrated in the drawings. On the other hand, for example, when a specific device is described based on one or more units, such as functional units, the corresponding method may include one step for performing the functionality of one or more units (e.g., one step for performing the functionality of one or more units, or multiple steps each performing the functionality of one or more of the multiple units), even if such one or more steps are not explicitly described or illustrated. Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specified.
[0138] In the following, an overview is provided of some of the technical terms used and the frameworks within which embodiments of the present disclosure may be utilized.
[0139] Video coding typically refers to the processing of a series of pictures that form a video or video sequence. In the field of video coding, the terms "frame" or "image" may be used synonymously with the term "picture". Video coding includes two parts: video encoding and video decoding. Video encoding is performed on the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture for more efficient storage and / or transmission. Video decoding is performed on the destination side and typically involves the reverse process compared to the encoder to reconstruct the video picture. Embodiments that refer to the "coding" of a video picture (or, as will be described later, a general picture) should be understood to relate to both the "encoding" and "decoding" of the video picture. The combination of the encoding part and the decoding part is also called a CODEC (Coding and DECoding).
[0140] Artificial Neural Network An artificial neural network (ANN) or connectionist system is a computing system that is vaguely inspired by the biological neural networks that make up the animal brain. Such systems "learn" to perform tasks by considering examples, generally without being programmed with task-specific rules. For example, in image recognition, these can learn to identify images that contain cats by analyzing example images that have been manually labeled as "cat" or "no cat" and using the results to identify cats in other images. They can do this without any prior knowledge about cats, such as having fur, a tail, whiskers, and a cat-like face. Instead, they automatically generate discriminative features from the examples they process.
[0141] ANN is based on a collection of connected units or nodes called artificial neurons that roughly model the neurons of a biological brain. Each connection can transmit signals to other neurons, similar to the synapses of a biological brain. An artificial neuron that receives a signal can then process it and send a signal to the neurons connected to it.
[0142] In the implementation of an ANN, the "signal" in a connection is a real number, and the output of each neuron is calculated by some non - linear function of the sum of its inputs. This connection is called an edge. Neurons and edges typically have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal in the connection. A neuron may have a threshold, and a signal is sent only if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals move from the first layer (input layer) to the last layer (output layer), possibly passing through layers multiple times in between.
[0143] The original goal of the ANN approach was to solve problems in the same way as the human brain. Over time, the focus has shifted to performing specific tasks, leading to a deviation from biology. ANNs have been used in a variety of tasks, such as computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even activities that have traditionally been considered limited to humans, such as painting.
[0144] The name "Convolutional Neural Network" (CNN) indicates that this network uses a mathematical operation called convolution. Convolution is a special linear operation. A convolutional network is a neural network that uses convolution instead of general matrix multiplication in at least one of its layers.
[0145] Figure 1 schematically shows a general concept of processing by a neural network such as a CNN. A convolutional neural network is composed of an input layer, an output layer, and a plurality of hidden layers. The input layer is a layer to which an input (such as a part of an image as shown in Figure 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved by multiplication or other inner products. The result of a layer is one or more feature maps (f. map in Figure 1), which may also be called channels. Subsampling may be included in part or all of a layer. As a result, as shown in Figure 1, the feature map may become smaller. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers. These layers are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are colloquially called convolutions, but this is just by convention. Mathematically, it is a collective sliding dot product or cross-correlation. This is important for matrix indices in terms of how weights are determined at specific index points.
[0146] When programming a CNN to process an image, as shown in Figure 1, the input is a tensor having a shape (number of images) x (width of the image) x (height of the image) x (depth of the image). It should be noted that the depth of an image may be composed of channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map having a shape (number of images) x (width of the feature map) x (height of the feature map) x (channels of the feature map). The convolutional layer within a neural network should have the following attributes. A convolutional kernel defined by the width and height (hyperparameters). The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) must be equal to the number of channels (depth) of the input feature map.
[0147] The convolutional layer is a core building block of the CNN. The parameters of the layer are composed of a set of learnable filters (the kernels mentioned above), which have a small receptive field but spread across the entire depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, calculating the dot product between the entries of the filter and the input, and generating a 2D activation map for that filter. As a result, the network learns filters that become active when detecting certain types of features at a certain spatial position within the input.
[0148] Another important concept in CNN is pooling, which is a form of non-linear downsampling. There are several non-linear functions for implementing pooling, among which max pooling is the most common. It partitions the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region. The pooling layer gradually reduces the spatial size of the representation, reducing the number of parameters, the memory footprint, and the amount of computation in the network, and thus functioning to control overfitting. In CNN architectures, it is common to periodically insert pooling layers between consecutive convolutional layers. The pooling operation provides another form of translation invariance.
[0149] The pooling layer operates independently on all depth slices of the input and spatially resizes it. The most common form is a pooling layer with a size 2×2 filter that applies two strides for every two depth slices of the input along both width and height, discarding 75% of the activations. Due to the significant reduction in the size of the representation, there is a recent trend to use smaller filters or even discard the pooling layer completely. "Region of Interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the Fast R-CNN architecture.
[0150] The above ReLU is short for rectified linear unit and applies a non-saturating activation function. By setting negative values to zero, it effectively removes negative values from the activation map. This increases the decision function and the non-linear characteristics of the entire network without affecting the receptive field of the convolutional layer.
[0151] After several convolutional layers and max pooling layers, high-level inferences in a neural network are made through fully connected layers. Neurons in a fully connected layer have connections to all activations in the previous layer, as seen in a normal (non-convolutional) artificial neural network. Thus, those activations can be computed as an affine transformation followed by a bias offset (vector addition of learned or fixed bias terms) after matrix multiplication.
[0152] The "loss layer" (which includes the calculation of the loss function) specifies how to penalize the deviation between the predicted (output) label and the true label during training and is usually the final layer of a neural network. Various loss functions suitable for different tasks may be used. Softmax loss is used to predict a single class out of K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values in [0,1]. Euclidean loss is used for regression to real-valued labels.
[0153] To summarize, Figure 1 shows the data flow in a typical convolutional neural network. First, the input image passes through the convolutional layer and is abstracted into a feature map containing several channels corresponding to some of the learnable filters in this layer's set of filters. Next, the feature map is subsampled using, for example, a pooling layer that reduces the dimension of each channel within the feature map. Next, the data comes to another convolutional layer that can have a different number of output channels. As described above, the number of input channels and output channels are hyperparameters of the layer. To establish the connectivity of the network, these parameters need to be synchronized between two connected layers such that the number of input channels of the current layer is equal to the number of output channels of the previous layer. For the first layer that processes input data, such as an image, the number of input channels is usually equal to the number of channels of the data representation, such as 3 channels for the RGB or YUV representation of an image or video, or 1 channel for a grayscale image or video representation.
[0154] Autoencoder and Unsupervised Learning An autoencoder is a type of artificial neural network used to learn efficient data coding in an unsupervised manner. A schematic diagram of it is shown in Figure 2. The purpose of the autoencoder is to learn the representation (encoding) of a set of data, typically for dimensionality reduction, by training the network to ignore the "noise" in the signal. Along with the reduction side, the reconstruction side is learned, and the autoencoder gets its name because it tries to generate a representation as close as possible to its original input from the reduced encoding. In the simplest case, when given one hidden layer, the encoder stage of the autoencoder takes the input x and maps it to h: h = σ(Wx + b)
[0155] This image h is typically referred to as a code, latent variable, or latent representation. Here, σ is an element-wise activation function such as the sigmoid function or rectified linear unit. W is a weight matrix and b is a bias vector. The weights and biases are typically initialized randomly and then iteratively updated during training through backpropagation. Subsequently, the decoder stage of the autoencoder maps h to a reconstruction x' of the same shape as x: x’ = σ’(W’h’ + b’) Here, σ’, W’, and b’ for the decoder need not be related to the corresponding σ, W, and b for the encoder.
[0156] Recent advancements in the field of artificial neural networks, particularly convolutional neural networks, have led researchers to become interested in applying neural network-based techniques to the task of image and video compression. For example, end-to-end optimized image compression using networks based on variational autoencoders has been proposed.
[0157] Therefore, data compression is considered a fundamental and well-studied problem in engineering and is generally formulated with the aim of designing codes for a given discrete data ensemble with minimum entropy. This solution highly depends on the knowledge of the probabilistic structure of the data, and thus, this problem is closely related to probabilistic source modeling. However, since all practical codes must have finite entropy, continuous-valued data (such as vectors of image pixel intensities) must be quantized to a finite set of discrete values, which introduces errors.
[0158] In this context, known as the non-invertible compression problem, one must trade off two competing costs: the entropy (rate) of the discrete representation and the error (distortion) resulting from quantization. Different compression applications, such as data storage or transmission over channels with limited capacity, require different rate-distortion trade-offs.
[0159] For example, JPEG uses the discrete cosine transform for blocks of pixels, and JPEG2000 uses the multi-scale orthogonal wavelet decomposition. Typically, the three components of the transform coding method, namely, the transform, the quantizer, and the entropy code, are optimized separately (often by manual parameter adjustment). The latest video compression standards such as HEVC, VVC, and EVC also use transform representations to code the residual signal after prediction. For this purpose, several transforms such as the discrete cosine and sine transforms (DCT, DST), as well as the low frequency non-separable manually optimized transforms (LFNST), are used.
[0160] Variational Image Compression The variational autoencoder (VAE) framework can be considered as a non-linear transform coding model. This is illustrated in FIG. 3 showing the VAE framework: The encoder 101 maps the input image x to a latent representation (denoted by y) via the function y = f(x). This latent representation may also be called a part or a point in the "latent space" hereinafter. The function f() is a transform function that transforms the input signal x into a more compressible representation y. The quantizer 102 transforms the latent representation y into
Number
Number
[0161] The latent space can be understood as a representation of compressed data in which similar data points are close to each other within the latent space. The latent space is useful for learning data features and finding a simpler representation of the data for analysis. Quantized latent representation T, ŷ, side information of the hyperprior [Number] (or ẑ) is included (quantized) in the bitstream 2 using arithmetic coding (AE). Further, the quantized latent representation is used to reconstruct the image [Number] (or x̂) is provided with a decoder 104 that converts it to, where [Number] . The signal x̂ is an estimate of the input image x. It is desirable for x to be as close as possible to x̂, in other words, the reconstruction quality should be as high as possible. However, the higher the similarity between x̂ and x, the more side information needs to be transmitted. The side information includes the bitstream 1 and bitstream 2 shown in FIG. 3, which are generated by the encoder and transmitted to the decoder. Usually, the more side information there is, the higher the reconstruction quality. However, a large amount of side information means a low compression rate. Therefore, one of the objectives of the system described in FIG. 3 is to balance the reconstruction quality and the amount of side information transmitted within the bitstream.
[0162] In FIG. 3, the component AE105 is an arithmetic coding module that converts samples of the quantized latent representation ŷ and the side information ẑ into a bitstream 1 of binary representation. The samples of ŷ and ẑ may be composed of, for example, integers or floating-point numbers. One objective of the arithmetic coding module is to convert the sample values (through the quantization process) into a string of binary digits, which is then included in a bitstream that may contain further parts corresponding to the encoded image or further side information.
[0163] Arithmetic decoding (AD) 106 is a process of converting a binary number back to a sample value and returning the binary conversion process. Arithmetic decoding is provided by the arithmetic decoding module 106.
[0164] Note that the present disclosure is not limited to this specific framework. Furthermore, the present disclosure is not limited to image or video compression, and can similarly be applied to object detection, image generation, and recognition systems.
[0165] In FIG. 3, there are two sub-networks connected to each other. A sub-network in this context is a logical division between parts of the entire network. For example, in FIG. 3, modules 101, 102, 104, 105, and 106 are called the "encoder / decoder" sub-network. The "encoder / decoder" sub-network is responsible for encoding (generating) and decoding (analyzing) the first bitstream "bitstream 1". The second network in FIG. 3 includes modules 103, 108, 109, 110, and 107 and is called the "hyper encoder / decoder" sub-network. The second sub-network is responsible for generating the second bitstream "bitstream 2".
[0166] The first sub-network is responsible for the following: · Converting the input image x to its latent representation y (which is easier to compress x) 101, · Quantizing the latent representation y to the quantized latent representation ŷ 102, · Using AE by the arithmetic encoding module 105 to compress the quantized latent representation ŷ to obtain the bitstream "bitstream 1", · Analyzing the bitstream 1 via AD using the arithmetic decoding module 106, · Using the analyzed data to reconstruct the reconstructed image (x̂) 104.
[0167] The purpose of the second subnet is to obtain the statistical characteristics of the samples of "bitstream 1" (such as the mean value, variance, and correlation between samples of bitstream 1) so that the compression of bitstream 1 by the first subnet becomes more efficient. The second subnet generates a second bitstream, "bitstream 2", which contains the aforementioned information (such as the mean value, variance, and correlation between samples of bitstream 1).
[0168] The second network includes converting the quantized latent representation y^ to side information z at 103, quantizing the side information z to quantized side information z^, and encoding (e.g., binarizing) the quantized side information z^ into bitstream 2 at 109. In this example, the binarization is performed by arithmetic encoding (AE). The decoder part of the second network includes arithmetic decoding (AD) 110, which decodes the input bitstream 2 into the decoded quantized side information
Number
Number
[0169] Figure 3 shows an example of a VAE (variational auto encoder), and its details may vary in different implementations.
[0170] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of the signal before converting the signal into binary digits (bits). For example, in the VAE framework, the encoder, which is a non-linear transformation, maps the input image x to y, where y has a smaller width and height than x. Since y has a smaller width and height and thus a smaller size, the dimensionality (size) of the signal is reduced, and it is thus easier to compress the signal y. Note that in general, the encoder does not necessarily need to reduce the size of both (or generally all) dimensions. Rather, some exemplary implementations may provide an encoder that reduces the size in only one (or generally a subset thereof) dimension.
[0171] Such an example of the VAE framework is shown in FIG. 4, which utilizes six downsampling layers marked 401 to 406. The network architecture includes a hyperprior model. The left side (g a ,g s ) shows the image autoencoder architecture, and the right side (h a ,h s ) corresponds to the autoencoder implementing the hyperprior. The factorized-prior model uses the same architecture for the analysis and synthesis transforms g a and g s . Q represents quantization, and AE and AD represent an arithmetic encoder and an arithmetic decoder, respectively. The encoder applies the input image x to g a to generate a response y (latent representation) with spatially varying standard deviation. The encoding g a includes a plurality of convolutional layers with subsampling and, as an activation function, generalized divisive normalization (GDN).
[0172] The response is h ais supplied and the distribution of the standard deviation of z is summarized. z is then quantized, compressed, and transmitted as side information. The encoder then uses the quantized vector z^ to
Number
[0173] The layer including downsampling is indicated by a downward arrow in the layer description. The layer description "Conv Nx5x5 / 2↓" means that the layer is a convolutional layer with N channels and the size of the convolutional kernel is 5×5. As described above, 2↓ means that 2-fold downsampling is performed in this layer. As a result of 2-fold downsampling, one of the dimensions of the input signal is reduced by half in the output. In FIG. 4, 2↓ indicates that both the width and height of the input image are reduced by a factor of 2. Since there are six downsampling layers, if the width and height of the input image 414 (denoted by x) are given by w and h, the output signal z^413 has a width and height equal to w / 64 and h / 64, respectively. The modules indicated by AE and AD are an arithmetic encoder and an arithmetic decoder, respectively, and are described with reference to FIG. 3. The arithmetic encoder and the arithmetic decoder are specific implementations of entropy coding. AE and AD can be replaced by other means of entropy coding. In information theory, entropy coding is a lossless data compression scheme that is a reversible process used to convert the values of symbols into binary representations. Also, "Q" in the figure corresponds to a quantization operation, and the quantization operation and the corresponding quantization unit as part of component 413 or 415 do not necessarily have to exist and / or can be replaced by another unit.
[0174] Cloud Solutions for Machine Tasks Video Coding for Machine (VCM) is another popular direction in computer science today. The main idea behind this approach is to transmit a coded representation of image or video information targeted for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to conventional image and video coding targeted at human perception, the quality metric is the performance of computer vision tasks such as object detection accuracy rather than the reconstruction quality. This is shown in FIG. 5.
[0175] Video coding for machines, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks across mobile cloud infrastructure. By splitting the network between mobile and cloud, it is possible to distribute the computational workload so that the overall energy and / or latency of the system is minimized. Generally, collaborative intelligence is a paradigm in which the processing of neural networks is distributed between two or more different computing nodes, such as between devices, but generally between any functionally defined nodes. Here, the term "node" does not mean the neural network nodes described above. Rather, a (computing) node here refers to a (physically or at least logically) separate device / module that implements a part of the neural network. Such devices may be different servers, different end-user devices, different intelligent vehicles or in-vehicle devices, a mixture of servers and / or user devices and / or cloud and / or processors, etc. In other words, computing nodes may be regarded as nodes that belong to the same neural network and communicate with each other to transfer coded data within / for the neural network. For example, to enable complex calculations to be performed, one or more layers may be executed on a first device and one or more layers on another device. However, the distribution may be finer, and a single layer may be executed on multiple devices. In the present disclosure, the term "plurality" refers to two or more. In some existing solutions, part of the neural network functionality is executed on a device (such as a user device or an edge device) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a collection of processing or computing systems outside the device operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training. In this case, the data flows in both directions, i.e., from the cloud to the mobile during backpropagation during training and from the mobile to the cloud and then to inference during the forward pass during training.
[0176] In some research, semantic image compression was presented by encoding deep features and reconstructing the input images from them. Compression based on uniform quantization was shown, and context-based adaptive arithmetic coding (CABAC) from H.264 was shown. In some scenarios, it may be more efficient to send the output of the hidden layer (deep feature map) from the mobile part to the cloud than to send the compressed natural image data to the cloud and perform object detection using the reconstructed images. Efficient compression of the feature map is beneficial for image and video compression and reconstruction for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a common approach to the compression of deep features (i.e., feature maps).
[0177] Currently, video content occupies more than 80% of Internet traffic, and that percentage is expected to increase further. Therefore, it is important to build an efficient video compression system to generate higher-quality frames with a given bandwidth budget. In addition, most video-related computer vision tasks, such as video object detection or video object tracking, are sensitive to the quality of the compressed video, and efficient video compression can benefit other computer vision tasks. On the other hand, video compression technology is also useful for action recognition and model compression.
[0178] End-to-End Image or Video Compression DNN-based image compression methods can utilize large-scale end-to-end training and advanced non-linear transformations that are not used in traditional approaches. However, it is not straightforward to directly apply these techniques to build an end-to-end learning system for video compression. First, the problem of learning how to generate and compress motion information suitable for video compression remains unsolved. Video compression methods rely heavily on motion information to reduce temporal redundancy in video sequences.
[0179] A simple solution is to represent motion information using learning-based optical flow. However, current learning-based optical flow approaches aim to generate the most accurate flow field possible. An accurate optical flow is often not optimal for certain video tasks. Additionally, the data volume of optical flow increases significantly compared to motion information in conventional compression systems, and directly applying existing compression approaches to compress optical flow values greatly increases the number of bits required to store motion information. Second, it is not clear how to build a DNN-based video compression system by minimizing rate-distortion-based objectives for both residual information and motion information. Rate-distortion optimization (RDO) aims to achieve higher quality (i.e., less distortion) of the reconstructed frames given a number of bits (or bitrate) for compression. RDO is important for the compression performance of videos. To utilize the power of end-to-end training for learning-based compression systems, an RDO strategy that optimizes the entire system is needed.
[0180] In "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015, the authors proposed an end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual coding.
[0181] Such an encoder is shown in FIG. 6A. In particular, FIG. 6A shows the overall structure of an end-to-end trainable video compression framework. To compress motion information, a CNN is designated to convert the optical flow into a corresponding representation better suited for compression. Specifically, an autoencoder-style network is used to compress the optical flow. The motion vector (MV) compression network is shown in FIG. 6B. The network architecture is somewhat similar to ga / gs in FIG. 4. In particular, a series of convolutional operations and non-linear transformations, including GDN and IGDN, are involved. The number of output channels for the convolution (deconvolution) is 128 except for the last deconvolution layer which is equal to 2. Given an optical flow of size M×N×2, the MV encoder generates a motion representation of size M / 16×N / 16×128. Next, the motion representation is quantized, entropy-coded, and sent to the bitstream. The MV decoder receives the quantized representation and reconstructs the motion information using the MV encoder.
[0182] FIG. 6C shows the structure of the motion compensation unit. Here, using the previous previous reconstructed frame x t-1 and the reconstructed motion information, the warping unit generates a warped frame (usually using an interpolation filter such as a bilinear interpolation filter). Then, a separate CNN with three inputs generates the predicted picture. The architecture of the motion compensation CNN is also shown in FIG. 6C.
[0183] The residual information between the original frame and the predicted frame is encoded by the residual encoder network. A highly non-linear neural network is used to convert the residual into a corresponding latent representation. Compared to the discrete cosine transform in conventional video compression systems, this approach can better utilize the power of non-linear transformation and achieve higher compression efficiency.
[0184] From the above summary, it can be seen that the CNN-based architecture can be applied to both image and video compression by considering different parts of the video framework including motion estimation, motion compensation, and residual coding. Entropy coding is a common method used for data compression, which is widely adopted in the industry and is also applicable to the compression of feature maps for either human perception or computer vision tasks.
[0185] In the case of reversible video coding, the reconstructed video picture has the same quality as the original video picture (assuming no transmission errors or other data losses during storage or transmission). In the case of irreversible video coding, additional compression is performed, for example, by quantization, to reduce the amount of data representing the video picture, but the video picture cannot be completely reconstructed by the decoder, that is, the quality of the reconstructed video picture is lower or worse compared to the quality of the original video picture.
[0186] Arithmetic Coding Entropy coding is typically used as a reversible coding. Arithmetic coding is a type of entropy coding, which encodes a message as a binary real number within an interval (range) representing the message. Here, the term "message" refers to a sequence of symbols. Symbols are selected from a pre-defined alphabet of symbols. For example, the alphabet can be composed of two values, 0 and 1. A message using such an alphabet is a sequence of bits. Symbols (0 and 1) may appear with different frequencies within the message. In other words, the symbol probabilities may be non-uniform. In fact, the more non-uniform the distribution, the higher the achievable compression by a general entropy code, especially an arithmetic code. Arithmetic coding utilizes a priori known probability models that specify the symbol probabilities of each symbol in the alphabet. The alphabet does not have to be binary. Rather, the alphabet may be composed of, for example, M values from 0 to M - 1. In general, any alphabet of any size may be used. Typically, the alphabet is given by the range of values of the data being coded.
[0187] A variant of the arithmetic coder improved for practical use is called a range coder, which does not use the interval [0, 1), but instead uses a finite range of integers, for example, from 0 to 255. This range is divided according to the probabilities of the alphabet symbols. If the remaining ranges are too small, the ranges may be renormalized in order to describe all the alphabet symbols according to their probabilities.
[0188] One of the main types of entropy coders assigns a unique code to each unique symbol that occurs in the input. These entropy encoders compress data by replacing each fixed-length input symbol with a corresponding variable-length output codeword. For data streams with certain specific entropy characteristics, simple static codes can be useful. These static codes include universal codes (such as Elias gamma coding or Fibonacci coding) and Golomb codes (such as unary coding or Rice coding). For general data streams, codes can be constructed based on the following rule: the length of each codeword is approximately proportional to the negative logarithm of the probability of occurrence of that codeword. Thus, the most common symbols use the shortest codes. Based on the constructed code table, the coder compresses the data by replacing each fixed-length input symbol with a corresponding variable-length output codeword without a prefix. An example of such coding is Huffman coding. The main problem with such coding is that at least one bit is required for each input symbol, even when its probability is close to 1. As a speedup of arithmetic coding, the asymmetric numerical system (ANS) family of entropy coding techniques was invented. Such coders provide a combination of the compression ratio of arithmetic coding and a processing cost similar to Huffman coding.
[0189] Entropy encoders encode symbols of an input alphabet A of size M into symbols of an alphabet B of size R by using an amount of output symbols that is inversely proportional to the probability of the coded symbol. Usually, the probability p i of a symbol a i from alphabet A i means the probability of occurrence of symbol a i in any sequence of symbols from alphabet A. In other words, the probability p iIt means the probability of an event equal to. Unequal non-uniform probabilities of different symbols from the alphabet give the possibility of compression. If all symbols of the alphabet have the same probability \(p_i = 1 / M\), where \(M\) is the size of the alphabet \(A\), compression is impossible.
[0190] A general scheme of an entropy coder is shown in FIG. 7. In most cases, the output alphabet is \(\{0, 1\}\), the size of the output alphabet is usually equal to 2, the symbols of the output binary alphabet are called bits, and a sequence of bits corresponding to a sequence of coded symbols from the input alphabet is called a bit stream. As can be seen in FIG. 7, the output symbols of the entropy encoder called the bit stream are the input to the entropy decoder. The output of the entropy decoder is the same alphabet as the input to the entropy encoder. Also, the output of the entropy decoder can be called the decoded symbol. In addition, an alphabet means a set of symbols, and the alphabet size means the number of symbols in the input alphabet. Here, since the input symbols are shown as \(0, 1, 2, \ldots, M - 1\), there are a total of \(M\) different symbols in the input alphabet. It should be noted that in this application, the alphabet size \(M\) always means the size of the input alphabet.
[0191] In an autoencoder-based coding scheme, an entropy coder is used to compress the latent space symbols. The distribution estimation can be performed a priori (using a pre-trained histogram) or can be carried out using some additional information from the bitstream and / or information from adjacent latents. A general scheme of the entropy coding used in an autoencoder-based coder is shown in FIG. 8: the input signal x is transformed into a feature (or latent) tensor y, where "x" here means the input signal corresponding to image data and "y" here means the latent space tensor, and the transformation process here can be called feature extraction. The latent space tensor contains latent space elements, which are quantized and placed as input to the entropy encoder. In some possible implementations, the latent space elements are processed by a gain unit as shown in FIG. 9, then quantized, and then placed as input to the entropy encoder. The tensor y contains real numbers (not integers) (e.g., floating-point values), and the ranges of these numbers are unknown a priori. Since the entropy coder can operate only with a finite input alphabet, the tensor y is transformed into an integer tensor y^ that contains values from 0 to M - 1, where M is the alphabet size of the entropy coder. Such a transformation from a continuous set of values to a discrete set is called quantization. The transformation can include clamping, rounding, and scaling operations. In one exemplary implementation, first, y is in the range
Number
[0192] As shown in FIG. 9, in order to control the quantization error and the bitstream size (bit rate), a gain unit can be added to the coding scheme. Here, it should be noted that the higher the bit rate, the better the data can be compressed, and the lower the bit rate, the lower the quality of the compressed data. The process of adjusting the compression system parameters to achieve the desired ratio between the bit rate and the quality (or to achieve the desired bit rate) is called rate control. The parameters used during rate control are called the rate control parameter β. The gain unit is used for rate control. The gain unit multiplies the latent space tensor y by the gain vector g:
Number
[0193] Figure 10 shows the rate-distortion curve for one of the autoencoder-based encoders with a gain unit. The horizontal axis represents the bit rate in bits per pixel (bpp), and the vertical axis represents the peak signal-to-noise ratio (PSNR). As can be seen here, when the bit rate (horizontal axis) is increased above 0.5 bpp, the PSNR of the reconstructed signal (vertical axis) does not increase but rather decreases. This is due to very large y g clipping errors occurring.
[0194] One possible solution is to select a very large alphabet size and use it in all cases. However, increasing the alphabet size imposes a penalty on the compression efficiency under some conditions, such as that a large alphabet size is not required at low bit rates. Using a large alphabet size can significantly increase the bit rate, but the reconstructed quality is not improved.
[0195] In the conventional method, the entropy coding parameters are usually predefined. For example, the alphabet size M is usually selected once based on the expected tensor range (or the latent tensor range), and is predefined by using the predefined alphabet size M for all cases. In such a case, when the actual tensor range is wider than the expected tensor range, the input alphabet size determined based on the expected tensor range becomes inappropriate, and clipping is required for the coded tensor values. Such clipping degrades the signal, especially when the coded tensor range is significantly different from the alphabet size. In this case, the corruption of the encoded tensor is a non-linear distortion that causes an unpredictable error in the reconstructed signal, and thus the quality of the reconstructed signal may be significantly degraded. In one implementation, a very large alphabet size can be selected and used for all cases, but increasing the alphabet size incurs a penalty on the compression efficiency under low bitrate conditions, and the use of a large alphabet size can significantly increase the bitrate without improving the reconstruction quality.
[0196] To solve the above problems, embodiments of the present application propose content / bitrate adaptive entropy coding parameter selection. In particular, the entropy coding parameters can be the input alphabet size, and thus the clipping effect can be avoided at high rates without the rate overhead caused by an inappropriately large alphabet size for low rates. The adaptability of the entropy coding parameters, especially the alphabet size, enables the optimal operation of the entropy encoder at low rates (narrow range of coded values), resulting in bitrate savings, and no clipping effect at high rates (wide range of coded values), resulting in high reconstructed signal quality.
[0197] The basic idea of this solution is the entropy coding parameter, especially the bit rate / content adaptability of the alphabet size. For the proper operation of entropy coding, all parameters should be matched between the encoder and the decoder, and thus, basically two problems need to be solved: 1. How to select the appropriate alphabet size on the encoder side? 2. How to derive the alphabet size (selected on the encoder side) on the decoder side?
[0198] For the alphabet selection on the encoder side, several possible solutions are proposed.
[0199] In one possible implementation, the alphabet size can be selected as the smallest possible number greater than the range of the coded values. For example, the minimum and maximum values of the tensor y are first obtained, and the alphabet size is selected as follows: M = ceil(max{y} - min{y}) Here, ceil(x) is the smallest integer greater than x. In most entropy encoders, the alphabet size should be a power of 2, and in this case, the alphabet size can be selected as M = 2^(ceil(log2(max{y} - min{y}))). Here, {y} means the latent space elements in the latent space, and the latent space elements are the result of the progression of the input signal. The process of converting the input signal into the latent space tensor y is sometimes called feature extraction. Generally, an input signal such as an input image is converted into a latent space (feature space), the latent space elements are quantized, and then encoded by an entropy encoder. Also, the latent space can be additionally processed before quantization (e.g., multiplied by a gain vector). Note that in some cases, for example, when the modulus of all y values is less than 1, an additional scaling operation can be performed before entropy coding.
[0200] In another possible embodiment, the alphabet size can be selected based on a rate-distortion optimization process. First, several values of M are tried around M0 = ceil(max{y}-min{y}), and then the loss function is calculated for all these values. The alphabet size Mi that gives the minimum loss function is selected. The loss function may include rate and distortion components such as PSNR, multi-scale structural similarity index (MS-SSIM), video multi-method assessment fusion (VMAF), or other quality metrics. In this approach, clipping may occur, but the bitrate savings due to the use of a smaller alphabet compensate for a slight increase in distortion. For example, the loss function can be loss = beta * distortion + bits, where distortion is measured by PSNR or MS-SSIM or VMAF, bits is the number of bits spent, and beta is a weighting parameter that controls the ratio of bitrate to reconstructed quality, and beta may also be called a rate control parameter.
[0201] For decoder-side alphabet size derivation, several possible solutions are proposed.
[0202] Embodiment 1 In one possible implementation, the alphabet size can be explicitly signaled in the bitstream. In one embodiment, the alphabet size can be directly signaled within the bitstream using, for example, fixed-length coding, exp-Golomb coding, or some other coding algorithm. Typical values of M can be 256, 512, 1024, and for example, when signaling 1024 with a fixed-length code, 11 bits are required (1024 10=(100000000002), when signaling log2(1024) - 9 = 1, only 1 bit is required if only the values 512 and 1024 are allowed, or 2 bits are required if four different alphabet sizes such as 512, 1024, 2048, 4096 are allowed. As a result, the direct signaling of M consumes more bits. However, in some exotic cases (e.g., when the alphabet size M is not a power of 2), the direct signaling of M may be useful.
[0203] In an alternative embodiment, instead of M itself, the output p of a reversible function f(M) can be signaled within the bitstream, and the output p can be referred to as first indication information. Such p can be signaled using fixed-length coding, exp-Golomb coding, or some other coding algorithm. Thus, on the decoder side, M is derived based on the first indication information. Specifically, M is derived as M = f -1 (p). Examples of such reversible functions f(M) can be as follows: 1. f(M) = log k (M), where k is a natural number, for example, k can be equal to 2; 2. f(M) = log k (M) - C, where C is an integer used as a predictor, for example, C can be equal to 9; 3. f(M) = M + R, where R is an integer used as a predictor; 4. f(M) = sqrt(M)
[0204] A preferred method is to signal p = f(M) = log2(M) - 9.
[0205] In some implementations, p is non - negative, but in other implementations, p can also be negative. For example, the value p can be in the range [0, 5], and 3 bits are used for signaling. The function f(M) is negotiated in advance between the encoder side and the decoder side.
[0206] In one possible embodiment, the alphabet size is signaled, for example, in a parameter set section of a bitstream, such as a picture parameter set section of a bitstream.
[0207] The advantage of the above Embodiment 1 is that any optimal alphabet size selected on the encoder side can be signaled, thus improving the flexibility in signaling the alphabet size. The only drawback is that since a few bits are spent on signaling, the bitstream size increases slightly.
[0208] Embodiment 2 In one possible embodiment, the alphabet size can be derived based on several other parameters. In one exemplary implementation, the alphabet size is derived from a quantization parameter or a rate control parameter. Alternatively, the alphabet size can also be derived from an image resolution, a video resolution, a frame rate, the density of pixels in a 3D object, etc. In a trainable codec, the alphabet size can be derived from some parameters of the loss function used during training, such as the rate / distortion weighting factor, or some parameters that affect the selection of the gain vector g. Also, it can be a quantization parameter such as the quantization parameter (qp) in a normal codec like JPEG, HEVC, VVC. For example, the loss function can be loss = beta * distortion + bitrate, where beta is a weighting factor.
[0209] In one exemplary implementation, if such a rate control parameter is β, the range of beta (β) is divided into K intervals (K sub - ranges) as follows: [β_0,β_1),[β_1,β_2),..,[β_(K - 1),β_K)
[0210] Each of the intervals / subranges corresponds to one alphabet size value Mi. Note that β_0 may be equal to -∞ and β_K may be equal to +∞. Note that there is a range allowed for a specific codec value β. For example, some codec βs are allowed to be within the range [-∞,∞], and other codec βs may only be allowed to be within the range [0,∞]. In any case, there is some large range of allowed β (beta) values. Within the context of this embodiment, the original large range of allowed β values is divided into several subranges, and for all subranges, there is a specific value of the alphabet size. One specific division of the β values for the intervals is shown in FIG. 11.
[0211] In this case, after obtaining the parameter β, the decoder can select a target interval based on the β value obtained from the bitstream. Specifically, the decoder determines that β_i ≦ β ≦ β_(i+1), and then selects the interval [β_i,β_(i+1)] as the target interval, and derives the alphabet size Mi corresponding to this target interval as the input alphabet size M on the decoder side.
[0212] In some embodiments, each βi within the range of beta {βi} can correspond to one alphabet size Mi, and the alphabet size M corresponding to a specific β is calculated based on one or more Mi corresponding to the βi adjacent to β. Note that the value used to calculate M can be the nearest just value Mi corresponding to the target interval, and can be linear interpolation, bilinear interpolation or some other interpolation from two or more Mi corresponding to the βi adjacent to β, or some other interpolation from two or more Mi corresponding to the intervals adjacent to the target interval.
[0213] The advantages of the above-described Embodiment 2 are that quantization parameters or rate control parameters already existing in the bitstream used in other procedures can be used on the decoder side to derive the alphabet size M. Therefore, no additional signaling for information specifically used to indicate the alphabet size M is required, and thus, the bit rate can be saved. The disadvantage of Embodiment 2 is that it lacks flexibility. Therefore, if the alphabet size derived for some reason is not optimal, the encoder and decoder have to use it despite the low compression efficiency.
[0214] Embodiment 3 In one possible implementation, the alphabet size can be derived based on a predictor P and second indication information, where the second indication information is signaled in the bitstream and is used to indicate the difference between P and M. The predictor P can be derived based on one of the techniques described in Embodiment 2 above by the decoder, such as quantization parameters, rate control parameters, parameters of a loss function used during training of a trainable codec, or some parameters affecting the gain vector g selection. The parameters used to derive the predictor P can be selected by the encoder or predefined by a standard. Therefore, upon receiving the bitstream, the decoder can derive the predictor P based on predetermined parameters, parse the second indication information from the bitstream, and then derive the alphabet size M based on the predictor P and the second indication information.
[0215] In one embodiment, the difference between P and M can be directly signaled within the bit stream using, for example, fixed-length coding, exp-Golomb coding, or some other coding algorithm. In an alternative embodiment, the output D of some invertible function s(M, P) can be signaled within the bit stream. Such D can be signaled with fixed-length coding, exp-Golomb coding, or some other coding algorithm. In this case, M is derived on the decoder side as M = s -1 (d, P). Examples of such invertible functions s(M, P) can be as follows: 1. s(M, P) = log k (M) - log k (P), where k is a natural number, for example, k may be equal to 2, 2. s(M, P) = log k (P) - log k (M), where k is a natural number, for example, k may be equal to 2, 3. s(M, P) = log k (M) - log k (P) - C, where C is an integer, 4. s(M, P) = log k (P) - log k (M) - C, where C is an integer, 5. s(M, P) = a * log k (P) + b * log k (M) - c, where a, b, and c are constants, 6. s(M, P) = a * M + b * P + c, where a, b, and c are constants.
[0216] Note here that A * B means A times B or A multiplies B.
[0217] A preferred method is to signal D = s(M, P) = log2(P) - log2(M).
[0218] Correspondingly, in this case, M satisfies one of the following, namely
Number
[0219] Since only the difference between P and M is signaled in the bitstream, the additional bits consumed are reduced compared to M signaled in the bitstream. In addition, the difference between P and M can be selected based on the content or bitrate, and the flexibility in signaling the alphabet size is also improved. Therefore, the above Embodiment 3 combines the advantages of Embodiments 1 and 2 to provide flexibility in alphabet size selection while minimizing the additional bits consumed for signaling. Even in some rare cases where the alphabet size predicted from β does not function well, the encoder can still signal the difference value between M and P. This incurs a cost of several bits, but can solve serious problems regarding the clipping effect.
[0220] In one possible implementation, a flag can be introduced into the bitstream to indicate the switching between Embodiment 1, Embodiment 2, and Embodiment 3. In this case, this flag may require 2 bits. In another possible implementation, a flag can be used to indicate the switching between Embodiment 1 and Embodiment 2. In this case, only 1 bit is required. Such a solution provides a balance between bit savings and flexibility. In most cases, if the derived entropy parameter is appropriate, only 1 bit is used for indication. On the other hand, in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0221] In one possible implementation, when the flag is equal to the first value, Embodiment 1 is used, and it is specified that the entropy coding parameter or the conversion result of the entropy coding parameter is conveyed within the bitstream. When the flag is equal to the second value, Embodiment 2 is used, and it is specified that the entropy coding parameter is not conveyed within the bitstream, but the entropy coding parameter can be derived by the decoder. When the flag is equal to the third value, Embodiment 3 is used, and it is specified that the difference value between M and P, or the conversion result of the difference value between M and P, is conveyed within the bitstream, where M is the size of the input alphabet and P is a predictor that can be derived by the decoder.
[0222] In addition to the above embodiments, alternative signaling schemes can also be considered. Also, the alphabet size M can be derived by using an interpolation or extrapolation process from a predetermined value. For example, p is signaled using one of the following codes, namely: binary code, or unary code, or truncated unary code, or exp-Golomb code. In one possible embodiment, p is signaled using an exp-Golomb code of degree 0.
[0223] The above embodiments can be applied to different entropy coders such as arithmetic coder, range coder or ANS (Asymmetric Numerical Systems) coder, etc.
[0224] In some possible implementations, more parameters of the entropy coding can be adaptively selected based at least on the content or the bit rate. For example, the parameters of the entropy coding may include: the minimum symbol probability supported by the entropy coder; the probability accuracy supported by the entropy coder; or the look-ahead period of the entropy encoder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc. Here, it should be noted that the "entropy coder" can be used as a synonym for the "entropy coding algorithm" that includes both the encoding algorithm and the decoding algorithm. The entropy encoder is a module that is part of the encoder, and the entropy decoder is another module that is part of the decoder. The parameters of the entropy encoder and the entropy decoder should be synchronized for correct operation, and thus, the term "parameters of the entropy coder" or "entropy coding parameters" means the parameters of both the entropy encoder and the entropy decoder. In other words, the "entropy coding parameters" can be made equal to the "parameters of the entropy encoder and the entropy decoder". The entropy encoder encodes the symbols of the alphabet into one or more bits in the bit stream, and the entropy decoder decodes one or more bits in the bit stream into the symbols of the alphabet. On the entropy encoder side, the alphabet means the input alphabet, and on the entropy decoder side, the alphabet means the output alphabet. The size of the input alphabet on the entropy encoder side is equal to the size of the output alphabet on the entropy decoder side.
[0225] FIG. 12 is a flowchart showing an exemplary decoding method implemented by a decoding apparatus, and the method includes the following.
[0226] 1201: Receive a bit stream including the encoded data of the input signal and a first parameter.
[0227] 1202: Analyze the bitstream to obtain the first parameter.
[0228] In one possible embodiment, the input signal includes video data, image data, dot cloud data, motion flow, or motion vectors, or any other type of media data, the encoded data means the encoding result of the input signal, the encoded data is composed of a plurality of bits, and the entropy coding parameters are: the size of the alphabet of the entropy coder, where the size of the alphabet is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc.
[0229] 1203: Obtain the entropy coding parameters based on the first parameter.
[0230] 1204: Reconstruct at least a part of the input signal based on the entropy coding parameters and the encoded data.
[0231] In an embodiment of the present application, the decoder can obtain entropy coding parameters (especially the alphabet size) based on the parameters carried in the bitstream, and since the parameters carried in the bitstream can be changed, the encoder can adaptively adjust the entropy coding parameters by changing the parameters carried in the bitstream. Therefore, the clipping effect can be avoided under high bitrate conditions, and the rate overhead caused by an unduly large alphabet size can also be avoided under low bitrate conditions. In other words, due to the adaptability of the entropy coding parameters, especially the alphabet size, optimal operation of the entropy encoder is possible at low bitrates (corresponding to a narrow range of coded values), resulting in bitrate savings, and at high bitrates (corresponding to a wide range of coded values), the clipping effect is eliminated, resulting in a higher reconstructed signal quality.
[0232] In one possible embodiment, the step of reconstructing at least a part of the input signal based on the entropy coding parameters comprises obtaining at least one probability model, wherein a probability model of the output symbol is used to indicate the probability of each possible value of the output symbol; entropy decoding one or more bits in the bitstream using the at least one probability model and the entropy coding parameters to obtain one or more output symbols; and reconstructing at least a part of the input signal based on the one or more output symbols.
[0233] In one possible embodiment, the method further comprises the step of updating the probability model. For example, since the probability model is updated after each output symbol, all output symbols have a unique probability distribution of possible values. Note that the probability model is also called a probability distribution.
[0234] In one possible embodiment, the probability model is selected according to the entropy coding parameter. For example, the symbol probability is distributed according to a normal distribution N(μ,σ), where N(μ,σ) means a Gaussian distribution with an average value equal to μ and a variance equal to σ 2 equal. However, an actual probability model such as a quantized histogram (which also means a mathematical model or a theoretical model) depends on the alphabet size and the probability accuracy within the entropy coding engine or entropy coder. The probability accuracy can be the minimum probability supported by the entropy coding engine. That is, the entropy coding parameter may affect the histogram configuration inside the entropy coder. Basically, since the alphabet size is the number of possible symbol values, for example, when the alphabet size is equal to 4, a larger value, such as the value "7", cannot be encoded / decoded.
[0235] The histogram used in the entropy coder consists of the quantized probabilities of each symbol value. For example, the alphabet is {0, 1, 2, 3}, the corresponding probabilities are {7 / 16, 7 / 16, 1 / 16, 1 / 16}, each probability is not zero, the sum of the probabilities is equal to 1, and each of the probabilities is greater than the minimum probability (probability precision) supported by the entropy coding engine (1 / 16 in this example). If the probabilities of some symbols are lower than the minimum probability supported by the entropy coding engine, at least the probabilities of some symbols need to be adjusted to ensure that the probability of each symbol is greater than the minimum probability supported by the entropy coding engine. Since the probability should be 1 / 16 or more, for example, when the probabilities of two symbol values are equal to 7 / 16, such as {7 / 16, 7 / 16, 1 / 16, 1 / 16, 0 / 16, 0 / 16, 0 / 16, 0 / 16}, the alphabet size is 8: {0, 1, 2, 3, 4, 5, 6, 7}. Therefore, the probabilities of symbols "0" and "1" must be reduced from 7 / 16 to 5 / 16 in this model, for example, the probability of each symbol must be adjusted to {5 / 16, 5 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16, 1 / 16}. Basically, this is one of the reasons why entropy coders with larger alphabets are less efficient. When there are many different possible symbol values, each of them should then have a probability greater than or equal to the minimum probability supported by the entropy encoder. Therefore, even if the probability of one symbol is huge, such as 0.99999, in the quantized histogram, it is simply 1 - (M - 1) * p min where M is the alphabet size and p min is the minimum probability supported by the entropy encoder. Therefore, the maximum probability in the model depends on the size of the alphabet and the minimum probability supported by the entropy encoder, that is: p max = 1 - (M - 1) * p min . p minis related to the calculation accuracy and thus cannot be made extremely small in practice. Therefore, for example, when p min = 1 / 256 and the alphabet size M is equal to 128, the maximum possible probability [Number] is equal to this value, but this is not very large and may not be sufficient in some cases.
[0236] In one possible embodiment, the first parameter is the size of the alphabet, and here, based on the first parameter, the step of obtaining the entropy coding parameter includes using the first parameter as the size of the alphabet.
[0237] In another possible embodiment, the first parameter is the output p of a certain invertible function f(M) instead of M itself. For example, the first parameter is p = f(M). In this case, the entropy coding parameter is obtained as M = f -1 (p), where f -1 (p) is the inverse function of f(M).
[0238] In one possible embodiment, f(M) can be the following, that is: f(M) = log k (M), where k is a natural number, or f(M) = log k (M) - C, where k is a natural number and C is an integer, or f(M) = M + R, where R is an integer, or f(M) = sqrt(M) can be used.
[0239] In one possible embodiment, p = f(M) = log2(M) - 9.
[0240] Correspondingly, in this case, M can be one of the following, that is: M = k^p, where k is a natural number, or M = k^(p + C), where k is a natural number and C is an integer, or M = k^(a*p + b), where k is a natural number, a and b are constants, or M = a*p + b, where a and b are constants, or M = p^2 satisfies one of the following.
[0241] Note that in any one of the embodiments, A^B means A B in this context.
[0242] In one possible embodiment, p = log2(M) - 9, and M = f -1 (p) = 2^(p + 9), where f -1 (p) is the inverse function of f(M), and f(M) = log2(M) - 9.
[0243] In one possible embodiment, p is signaled using one of the following codes: binary code, or unary code, or truncated unary code, or exp - Golomb code.
[0244] In one possible embodiment, p is signaled using the exp - Golomb code of degree 0.
[0245] In one possible embodiment, the alphabet size is signaled, for example, in the parameter set section of the bitstream, such as in the picture parameter set section of the bitstream.
[0246] In one possible embodiment, the first parameter can be some other parameter, such as a rate control parameter, an image resolution, a video resolution, a frame rate, the density of pixels within a 3D object, some parameters of a loss function used during training of a trainable codec, such as a rate / distortion weighting factor, or some parameters that affect the selection of the gain vector g. The loss function may include rate and distortion components such as peak signal-to-noise ratio (PSNR), multi-scale structural similarity index (MS-SSIM), video multi-method assessment fusion (VMAF), or some other quality metric. For example, the loss function can be loss = beta * distortion + bits, where distortion is measured by PSNR or MS-SSIM or VMAF, bits is the number of bits spent, and beta is a weighting parameter that controls the ratio of bitrate to reconstructed quality, and beta may sometimes be referred to as a rate control parameter. Also, it can be a quantization parameter such as the quantization parameter (qp) in a normal codec like JPEG, HEVC, VVC. In this case, the entropy coding parameter can be derived on the decoder side based on the above other parameters.
[0247] The advantage of the above embodiment is that since the quantization parameter or rate control parameter already exists in the bitstream and is used for other procedures, such parameters can be used by the decoder side to derive the alphabet size M, and there is no need for additional signaling for the information specifically used to indicate the alphabet size M, thus saving bitrate.
[0248] In one possible embodiment, the step of obtaining an entropy coding parameter based on a first parameter includes determining a target sub-range in which the first parameter is located, where the allowable range of values of the first parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the first parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter, the step of using the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter, or calculating the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0249] In one possible embodiment, the first parameter is D, and the entropy coding parameter includes the size of alphabet M, where M is obtained based on P and D, and P is a predictor that can be derived by a decoder.
[0250] In one possible embodiment, the first parameter can be the difference value between M and P, where M is the size of the input alphabet and P is a predictor that can be derived by a decoder by using one of the techniques described in Embodiment 2 above.
[0251] The advantage of the above embodiment is that only the difference between P and M is signaled within the bitstream, so the additional bits used are reduced compared to M signaled in the bitstream. Further, the difference between P and M can be selected based on the content or bitrate, and the flexibility in signaling the alphabet size is also improved. Therefore, this embodiment provides flexibility in alphabet size selection while minimizing the additional bits spent on signaling. Even in some rare cases where the size of the alphabet predicted from β does not function well, the encoder can still signal the difference value between M and P. This incurs a cost of a few bits, but it can solve serious problems regarding the clipping effect.
[0252] In one possible embodiment, the first parameter is a value obtained by processing the difference value between M and P. For example, the first parameter is D = s(M, P), where s(M, P) is a reversible function, and s(M, P) can be as follows, that is: s(M, P)=log k (M)-log k (P), where k is a natural number, or, s(M, P)=log k (P)-log k (M), where k is a natural number, or, s(M, P)=log k (M)-log k (P)-C, where k is a natural number and C is an integer, or, s(M, P)=log k (P)-log k (M)-C, where k is a natural number and C is an integer, or, s(M, P)=a*M - b*P + c, where a, b, and c are constants.
[0253] In one possible embodiment, D = s(M, P)=log2(P)-log2(M).
[0254] In one possible embodiment, the entropy coding parameter is obtained as M = s -1 (D, P), where s -1 (D, P) is the inverse function of s(M, P).
[0255] The invertible function D = s(M, P) can be regarded as D = s P (M), where M = s -1 (D, P) can be regarded as M = s -1 P (D), where P is an arbitrary fixed number, that is, it should be noted that P is a constant coefficient.
[0256] In one possible embodiment, D is signaled using one of the following codes: a binary code, or a unary code, or a truncated unary code, or an exp - Golomb code.
[0257] In one possible embodiment, D is signaled using an exp - Golomb code of degree 0.
[0258] In one possible embodiment, P can be derived based on at least one parameter other than the first parameter carried in the bitstream.
[0259] In one possible embodiment, at least one parameter other than the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, the density of pixels in a 3D object, or a rate - distortion weight factor.
[0260] In one possible embodiment, P is: obtaining a rate control parameter beta (β) from a bitstream; determining a target subrange in which the obtained β is located, wherein the allowable range of values of the rate control parameter β is [β_0, β_K], the allowable range [β_0, β_K] is divided into a plurality of subranges, the target subrange is one of the plurality of subranges, each of the plurality of subranges contains at least one value of β, and each of the plurality of subranges corresponds to one value of P; selecting, as the value of P, a value corresponding to the target subrange; or calculating the value of P based on one or more values corresponding to one or more subranges adjacent to the target subrange, and is derived based on at least one parameter.
[0261] In one possible embodiment, the entropy coder is an arithmetic coder, a range coder or an ANS (Asymmetric Numerical Systems) coder.
[0262] Optionally, the method further includes step 1205 of analyzing the bitstream to obtain a flag, where the flag is used to indicate whether the entropy coding parameter is directly carried in the bitstream.
[0263] In the above embodiment, a flag can be introduced into the bitstream to indicate the switching between three embodiments, in which case 2 bits may be required for this flag. In another possible embodiment, a flag can be used to indicate the switching between two embodiments, in which case only 1 bit is required.
[0264] In one possible embodiment, when the flag is equal to the first value, it is specified that the entropy coding parameter is conveyed within the bitstream, in which case the first parameter is either the entropy coding parameter or the first parameter is the conversion result of the entropy coding parameter, or when the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed within the bitstream and the entropy coding parameter can be derived by the decoder.
[0265] Such a solution provides a balance between bit savings and flexibility, and in most cases, if the derived entropy parameter is appropriate, only 1 bit is used for the indication, while in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0266] In one possible embodiment, when the flag is equal to the third value, it is specified that the difference value between M and P, or the conversion result of the difference value between M and P, is conveyed within the bitstream, in which case the first parameter is the difference value between M and P, or the conversion result of the difference value between M and P, where M is the size of the input alphabet and P is a predictor that can be derived by the decoder.
[0267] FIG. 13 is a flowchart showing an exemplary decoding method implemented by a decoding apparatus, and the method includes the following: 1301: Receive a bitstream including the encoded data of the input signal and a flag; 1302: Analyze the bitstream to obtain the flag, where the flag is used to indicate whether the entropy coding parameter is directly conveyed within the bitstream; 1303: Obtain the entropy coding parameter based on the flag; 1304: Reconstruct at least a part of the input signal based on the entropy coding parameter and the encoded data.
[0268] In the above embodiment, a flag can be introduced into the bitstream to indicate the switching between three embodiments. In this case, two bits may be required for this flag. In another possible embodiment, a flag can be used to indicate the switching between two embodiments, in which case only one bit is required. Such a solution provides a balance between bit savings and flexibility. In most cases, if the derived entropy parameter is appropriate, only one bit is used for the indication. On the other hand, in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0269] In one possible embodiment, the entropy coding parameter includes at least one of the following: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc.
[0270] In one possible embodiment, the input signal includes video data, image data, dot-clocked data, motion flow, or motion vectors, or any other type of media data.
[0271] In one possible embodiment, when the flag is equal to the first value, it is specified that the entropy coding parameter is conveyed in the bitstream, or the conversion result of the entropy coding parameter is conveyed in the bitstream. Note that the conversion result of the entropy coding parameter is the result obtained by processing the entropy coding parameter, for example, a value. When the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed in the bitstream, but the entropy coding parameter can be derived by the decoder. In this case, the flag is used to indicate the switching between the above-described Embodiment 1 and Embodiment 2, and only 1 bit is required.
[0272] Such a solution provides a balance between bit savings and flexibility. In most cases, if the derived entropy parameter is appropriate, only 1 bit is used for the indication. On the other hand, in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0273] Optionally, the flag can be used to indicate the switching between the above-described Embodiment 1, Embodiment 2, and Embodiment 3. In this case, the flag has three possible values and 2 bits are required. When the flag is equal to the third value, it is specified that the difference value between M and P is conveyed in the bitstream, or the conversion result of the difference value between M and P is conveyed in the bitstream, where M is the entropy coding parameter and P is a predictor that can be derived by the decoder. Note that the conversion result of the difference value between M and P means the result of processing the difference value between M and P.
[0274] In one possible embodiment, the step of obtaining entropy coding parameters based on a flag includes, when the flag is equal to a first value, analyzing a bit stream to obtain a first parameter, where the first parameter is an entropy coding parameter, and using the first parameter as the entropy coding parameter, or the first parameter is a conversion result of the entropy coding parameter, and obtaining the entropy coding parameter based on the first parameter.
[0275] In one possible embodiment, the conversion result of the entropy coding parameter is p = f(M), where M is the entropy coding parameter and f(M) includes the following, that is: f(M) = log k (M), where k is a natural number; or f(M) = log k (M) - C, where k is a natural number and C is an integer; or f(M) = M + R, where R is an integer; or f(M) = sqrt(M); where the step of obtaining the entropy coding parameter based on the first parameter includes M = f -1 (p), where f -1 (p) is the inverse function of f(M).
[0276] Correspondingly, M satisfies one of the following, that is: M = k p , where k is a natural number, or M = k p+C , where k is a natural number and C is an integer, or M = ap + b, where a and b are constants, or M = p 2 , satisfies one of the following.
[0277] In one possible embodiment, k = 2.
[0278] In one possible embodiment, the first parameter is p = log2(M) - 9.
[0279] In one possible embodiment, the step of obtaining an entropy coding parameter based on a flag includes, when the flag is equal to a second value, analyzing a bit stream to obtain a second parameter, where the second parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient, and deriving an entropy coding parameter based on the second parameter.
[0280] In one possible embodiment, the step of deriving an entropy coding parameter based on the second parameter includes determining a target sub-range in which the second parameter is located, where the allowable range of values of the second parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the second parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter, using the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter, or calculating the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0281] In one possible embodiment, the step of obtaining an entropy coding parameter based on a flag includes, when the flag is equal to a third value, analyzing a bitstream to obtain a third parameter, where the third parameter is a difference value between M and P, or the third parameter is a conversion result of the difference value between M and P, M is an entropy coding parameter, and P is a predictor that can be derived by a decoder; the step of deriving P based on at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient; and the step of obtaining an entropy coding parameter based on the third parameter and P.
[0282] In one possible embodiment, the conversion result of the difference value between M and P is D = s(M, P), where s(M, P) is a reversible function, and s(M, P) includes the following, that is: s(M, P) = log k (M) - log k (P), where k is a natural number, or s(M, P) = log k (P) - log k (M), where k is a natural number, or s(M, P) = log k (M) - log k (P) - C, where k is a natural number and C is an integer, or s(M, P) = log k (P) - log k (M) - C, where k is a natural number and C is an integer, or s(M, P) = a * log k (P) + b * log k (M) - c, where a, b, and c are constants, or s(M, P) = a * M - b * P + c, where a, b, and c are constants. including, the step of obtaining an entropy coding parameter based on the third parameter includes M = s -1(D, P), where s -1 (D, P) is the inverse function of s(M, P). Note that here, A * B means A times B or A multiplies B.
[0283] In one possible embodiment, M satisfies one of the following, that is,
Number
[0284] Figure 14 is a flowchart showing an exemplary encoding method implemented by an encoding device, and the method includes the following: 1401: Encode an input signal and a first parameter into a bitstream, where the first parameter is used to obtain an entropy coding parameter.
[0285] In one possible embodiment, the input signal includes video data, image data, dot cloud data, motion flow, or motion vectors, or any other type of media data, and the entropy coding parameter includes at least one of the following, that is: the size of the alphabet of the entropy coder, where the size of the alphabet is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc.
[0286] 1402: Transmit the bitstream to a decoder.
[0287] In one possible embodiment, the first parameter is the size of the alphabet.
[0288] In one possible embodiment, the first parameter is p, where p is the conversion result of M, and M is an entropy coding parameter.
[0289] In one possible embodiment, p = f(M), where f(M) is an invertible function.
[0290] In one possible embodiment, f(M) includes the following, that is: f(M)=log k (M), where k is a natural number, or f(M)=log k (M)-C, where k is a natural number and C is an integer, or f(M)=a*M + b, where a and b are constants, or f(M)=sqrt(M), including.
[0291] In one possible embodiment, p = log2(M) - 9.
[0292] In one possible embodiment, p is signaled using one of the following codes: binary code, or unary code, or truncated unary code, or exp-Golomb code.
[0293] In one possible embodiment, the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, the density of pixels in a 3D object, or a rate-distortion weight coefficient, and the first parameter is used by an entropy decoder to derive an entropy coding parameter.
[0294] In one possible embodiment, the first parameter is D obtained based on P and M, where M is an entropy coding parameter and P is a predictor that can be derived by a decoder.
[0295] In one possible embodiment, D = s(M, P), where s(M, P) is an invertible function.
[0296] In one possible embodiment, s(M, P) includes the following, namely: s(M, P) = log k (M) - log k (P), where k is a natural number, or s(M, P) = log k (P) - log k (M), where k is a natural number, or s(M, P) = log k (M) - log k (P) - C, where k is a natural number and C is an integer, or s(M, P) = log k (P) - log k (M) - C, where k is a natural number and C is an integer, or s(M, P) = a * M - b * P + c, where a, b, and c are constants, including.
[0297] In one possible embodiment, D = s(M, P) = log2(P) - log2(M).
[0298] In one possible embodiment, D is signaled using one of the following codes, namely: binary code, or unary code, or truncated unary code, or exp - Golomb code.
[0299] In one possible embodiment, the encoding method further includes the step of encoding a flag into the bitstream, where the flag is used to indicate whether the entropy coding parameter is directly carried within the bitstream.
[0300] In one possible embodiment, when the flag is equal to the first value, the entropy coding parameter is conveyed within the bitstream, and it is specified that the first parameter is the entropy coding parameter or the first parameter is the conversion result of the entropy coding parameter, or when the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed within the bitstream but the entropy coding parameter can be derived by the decoder.
[0301] In one possible embodiment, when the flag is equal to the third value, it is specified that the difference value between M and P or the conversion result of the difference value between M and P is conveyed within the bitstream, where M is the entropy coding parameter and P is a predictor that can be derived by the decoder.
[0302] In one possible embodiment, several possible solutions are proposed for alphabet selection on the encoder side.
[0303] In one possible embodiment, before encoding the first parameter into the bitstream, the encoding method further includes: determining the size of the alphabet of the entropy encoder based on at least one of the bit rate of the image data or the coding value.
[0304] FIG. 15 is a flowchart showing an exemplary method for determining the size of the alphabet of the entropy encoder, and the method includes the following: 1501: Obtaining the minimum and maximum values of the latent space elements; 1502: Obtaining the size of the input alphabet as follows: M = ceil(max{y} - min{y}) where ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space elements, min{y} represents the minimum value of the latent space elements, and M represents the size of the alphabet.
[0305] FIG. 16 is a flowchart showing an exemplary method for determining the size of the alphabet of an entropy encoder, the method including the following: 1601: Obtain the minimum and maximum values of the latent space elements; 1602: Obtain the size of the input alphabet as follows: M = 2^(ceil(log2(max{y}-min{y}))) Here, ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space elements, min{y} represents the minimum value of the latent space elements, and M represents the size of the alphabet. In most entropy encoders, the alphabet size should be a power of 2, and in this case, the alphabet size can be selected as M = 2^(ceil(log2(max{y}-min{y}))). Note that there are some cases, for example when the modulus of all y values is less than 1, that an additional scaling operation can be performed before entropy coding.
[0306] FIG. 17 is a flowchart showing an exemplary method for determining the size of the alphabet of an entropy encoder, the method including the following: 1701: Obtain at least two values around M_0, where M_0 = ceil(max{y}-min{y}) or M_0 = 2^(ceil(log2(max{y}-min{y}))); 1702: Calculate the loss function for at least two values; 1703: Select the value with the minimum loss function among at least two values as the size of the input alphabet. Here, ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space elements, and min{y} represents the minimum value of the latent space elements.
[0307] For example, the loss function can be loss = beta * distortion + bits, where the distortion is measured by PSNR or MS-SSIM or VMAF, the bits are the number of bits spent, and beta is a weighting parameter that controls the ratio of the bitrate to the reconstructed quality, and beta may also be referred to as a rate control parameter. In this approach, clipping may occur, but the bitrate savings due to the use of a smaller alphabet compensate for a slight increase in distortion.
[0308] FIG. 18 is a flowchart showing an exemplary encoding method implemented by an encoding apparatus, the method including the following: 1801: Encode an input signal and a flag into a bitstream, where the flag is used to indicate whether entropy coding parameters are directly carried in the bitstream; 1802: Transmit the bitstream to a decoder.
[0309] In one possible embodiment, the input signal includes video data, image data, dot cloud data, motion flow, or motion vectors, or any other type of media data, and the entropy coding parameters include at least one of the following, namely: the size of the alphabet of the entropy encoder, where the size of the alphabet is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy encoder; or the look-ahead period of the entropy encoder.
[0310] In one possible embodiment, a flag is used to indicate the switching between the above-described Embodiment 1 and Embodiment 2, in which case only 1 bit is required. When the flag is equal to the first value, it is specified that the entropy coding parameter is conveyed in the bit stream or the conversion result of the entropy coding parameter is conveyed in the bit stream, or when the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed in the bit stream but the entropy coding parameter can be derived by the decoder.
[0311] Such a solution provides a balance between bit savings and flexibility. In most cases, if the derived entropy parameter is appropriate, only 1 bit is used for the indication. On the other hand, in some specific cases, it may be possible to explicitly signal the entropy parameter.
[0312] Optionally, a flag can be used to indicate the switching between the above-described Embodiment 1, Embodiment 2, and Embodiment 3. In this case, the flag has three possible values and 2 bits are required. When the flag is equal to the third value, it is specified that the difference value between M and P is conveyed in the bit stream or the conversion result of the difference value between M and P is conveyed in the bit stream, where M is the entropy coding parameter and P is a predictor that can be derived by the decoder.
[0313] In one possible embodiment, when the flag is equal to the first value, the first parameter is encoded into the bit stream, where the first parameter is the entropy coding parameter or the first parameter is the conversion result of the entropy coding parameter.
[0314] In one possible embodiment, the conversion result of the entropy coding parameter is p = f(M), where M is the entropy coding parameter and f(M) can be as follows, that is: f(M)=log k(M), where k is a natural number, or f(M) = log k (M) - C, where k is a natural number and C is an integer, or f(M) = aM + R, where a and R are constants, or f(M) = sqrt(M).
[0315] In one possible embodiment, the first parameter is p = log2(M) - 9.
[0316] In one possible embodiment, p is signaled using one of the following codes: binary code, or unary code, or truncated unary code, or exp-Golomb code.
[0317] In one possible embodiment, p is signaled using an exp-Golomb code of degree 0.
[0318] In one possible embodiment, the method further includes encoding a third parameter into a bit stream when the flag is equal to a third value, where the third parameter is a difference value between M and P, or the third parameter is a conversion result of the difference value between M and P.
[0319] In one possible embodiment, the conversion result of the difference value between M and P is D = s(M, P), where s(M, P) is an invertible function, and where s(M, P) includes the following, namely: s(M, P) = log k (M) - log k (P), where k is a natural number, or, s(M, P) = log k (P) - log k (M), where k is a natural number, or, s(M, P) = log k (M) - log k (P) - C, where k is a natural number and C is an integer, or, s(M, P) = log k(P)-log k (M)-C, where k is a natural number and C is an integer, or, s(M,P)=a*log k (P)-b*log k (M)-c, where a, b, and c are constants, or, s(M,P)=a*M - b*P + c, where a, b, and c are constants, including The step of obtaining the entropy coding parameter based on the third parameter is M = s -1 (D,P) including, where s -1 (D,P) is the inverse function of s(M,P).
[0320] In one possible embodiment, D is signaled using one of the following codes: binary code, or unary code, or truncated unary code, or exp-Golomb code.
[0321] In one possible embodiment, D is signaled using an exp-Golomb code of degree 0.
[0322] Embodiments of the present application provide a decoding apparatus including a receiver configured to receive a bitstream including encoded data of an input signal and a first parameter, an analysis unit configured to analyze the bitstream to obtain the first parameter, an acquisition unit configured to obtain an entropy coding parameter based on the first parameter, and a reconstruction unit configured to reconstruct at least a part of the input signal based on the entropy coding parameter.
[0323] This apparatus provides the advantages of the above-described method.
[0324] In one possible embodiment, the input signal includes video data, image data, dot cloud data, motion flow, or motion vectors, or any other type of media data.
[0325] In one possible embodiment, the entropy coding parameter includes at least one of the following: the size of the alphabet of the entropy coder, where the size of the alphabet of the entropy coder is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder; or the minimum symbol probability supported by the entropy coder; or the look-ahead period of the entropy coder. In some embodiments, the look-ahead period can be 8 bits, 16 bits, etc.
[0326] In one possible embodiment, the first parameter is the size of the alphabet, where the acquisition unit is further configured to use the first parameter as the size of the alphabet.
[0327] In one possible embodiment, the first parameter is p, and the entropy coding parameter includes the size of the alphabet M, where M is a function of p.
[0328] In one possible embodiment, the acquisition unit is further configured to acquire M as M = f -1 (p), where f -1 (p) is the inverse function of f(M), and f(M) = p.
[0329] In one possible embodiment, the acquisition unit determines a target sub-range where the first parameter is located. Here, the allowable range of the value of the first parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the first parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter; and further configured to use the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter; or calculate the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range.
[0330] Embodiments of the present application provide a decoding device including a functional unit that implements any one of the above encoding methods.
[0331] Embodiments of the present application provide an encoding device that is configured to encode an input signal and a first parameter into a bit stream. The first parameter includes an encoding unit used to obtain an entropy coding parameter, and a transmission unit configured to transmit the bit stream to a decoder. The encoding device further includes other functional units for implementing any one of the above encoding methods.
[0332] Embodiments of the present application provide a decoding device including a processing circuit configured to execute any one of the above decoding methods.
[0333] Embodiments of the present application provide an encoding device including a processing circuit configured to execute any one of the above encoding methods.
[0334] Embodiments of the present application provide a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors. The storage medium stores programming for execution by the one or more processors, and the programming configures the decoder to execute any one of the above decoding methods when executed by the one or more processors.
[0335] Embodiments of the present application provide a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the one or more processors. The storage medium stores programming for execution by the one or more processors, and the programming configures the decoder to execute any one of the above encoding methods when executed by the one or more processors.
[0336] Embodiments of the present application provide a non-transitory computer-readable medium that, when executed by a computer device or one or more processors, conveys computer instructions that cause the computer device or the one or more processors to execute any one of the above encoding methods.
[0337] Embodiments of the present application provide a non-transitory computer-readable medium that, when executed by a computer device or one or more processors, conveys computer instructions that cause the computer device or the one or more processors to execute any one of the above decoding methods.
[0338] Embodiments of the present application provide a non-transitory storage medium including a bitstream encoded by any one of the above encoding methods.
[0339] Embodiments of the present application provide a computer program stored in a non-transitory medium and including code instructions that, when executed on one or more processors, cause any one of the above encoding method steps to be executed.
[0340] Embodiments of the present application provide a computer program stored in a non-transitory medium and including code instructions, which, when executed on one or more processors, cause any one of the above decoding methods to be performed.
[0341] Embodiments of the present application are a system for delivering a bitstream, including at least one storage medium configured to store at least one bitstream generated by an encoding method described in any one of the third aspect or possible embodiments of the third aspect, and any one of the fourth aspect or possible embodiments of the fourth aspect, and a video streaming device configured to obtain a bitstream from one of the at least one storage media and transmit the bitstream to a terminal device, where the video streaming device includes a content server or a content distribution server.
[0342] In one possible embodiment, it further includes one or more processors configured to perform an encryption process on at least one bitstream to obtain at least one encrypted bitstream, and at least one storage medium configured to store the encrypted bitstream, or one or more processors configured to convert a bitstream in a first format into a bitstream in a second format, and at least one storage medium configured to store the bitstream in the second format.
[0343] In one possible embodiment, it further includes a receiver configured to receive a first operation request, one or more processors configured to determine a target bitstream in at least one storage medium in response to the first operation request, and a transmitter configured to transmit the target bitstream to a terminal-side device.
[0344] In one possible embodiment, one or more processors are further configured to encapsulate a bitstream to obtain a transport stream in a first format, and the transmitter is further configured to transmit the transport stream in the first format to a terminal-side device for display or to transmit the transport stream in the first format to a storage space for storage.
[0345] In one possible embodiment, an exemplary method for storing a bitstream is provided, the method comprising: obtaining a bitstream according to any one of the encoding methods described above; storing the bitstream in a storage medium; and.
[0346] Optionally, the method further comprises performing an encryption process on the bitstream to obtain an encrypted bitstream; storing the encrypted bitstream in a storage medium; and.
[0347] It should be understood that any of the known encryption methods may be used.
[0348] Optionally, the method further comprises performing a segmentation process on the bitstream to obtain a plurality of bitstream segments; storing the plurality of bitstream segments in a storage medium; and further comprising.
[0349] Optionally, the method further comprises obtaining at least one backup of the bitstream and storing the at least one backup in a storage medium. It should be understood that at least one backup of the bitstream can be stored in a storage medium different from the storage medium storing the original bitstream.
[0350] Optionally, the method receives a plurality of bitstreams generated according to any one of the encoding methods described previously; individually assigns address information or identification information to the plurality of bitstreams; stores the plurality of bitstreams at corresponding positions according to the address information or identification information corresponding to the plurality of bitstreams; and further includes.
[0351] Optionally, the method classifies the bitstreams to obtain at least two bitstreams, where the at least two bitstreams include a first bitstream and a second bitstream; stores the first bitstream in a first storage space and the second bitstream in a second storage space; and further includes.
[0352] Optionally, the method further includes transmitting the bitstream to a terminal device by a video streaming device, where the video streaming device can be a content server or a content delivery server.
[0353] In one possible embodiment, an exemplary system for storing bitstreams is provided, the system including a receiver configured to receive a bitstream generated by any one of the previous encoding methods; a processor configured to perform an encryption process on the bitstream to obtain an encrypted bitstream; a computer-readable storage medium configured to store the encrypted bitstream; and including.
[0354] Optionally, the system includes several storage media, and the several storage media can be deployed at different locations. Also, multiple bitstreams may be stored distributed across different storage media. For example, some storage media include a first storage media configured to store a first bitstream and a second storage media configured to store a second bitstream.
[0355] Optionally, the system includes a video streaming device, which can be a content server or a content delivery server, and the video streaming device is configured to obtain a bitstream from one of the storage media and transmit the bitstream to a terminal device.
[0356] In one possible embodiment, an exemplary method for converting the format of a bitstream is provided, and the method includes receiving a bitstream in a first format generated by any one of the encoding methods described previously; converting the bitstream in the first format into a bitstream in a second format; storing the bitstream in the second format in a storage media; and
[0357] Optionally, the method further includes responding to an access request from a terminal-side device and transmitting the stored bitstream in the second format to the terminal-side device. and
[0358] In one possible embodiment, an exemplary system for converting the format of a bitstream is provided, and the system includes a receiver configured to receive a bitstream in a first format generated by any one of the encoding methods described previously; A processor configured to convert a bitstream in a first format into a bitstream in a second format, The processor is further configured to store the bitstream in the second format in a storage medium, The storage medium includes a processor configured to store the bitstream in the second format, A transmitter configured to transmit the stored bitstream in the second format to a terminal-side device in response to an access request from the terminal-side device, and includes.
[0359] In one possible embodiment, an exemplary method for processing a bitstream is provided, the method comprising: Receiving a transport stream including a video stream and an audio stream, the video stream being generated by any one of the encoding methods described previously, Demultiplexing the transport stream to separate the video stream and the audio stream, Decoding the video stream using a video decoder to obtain video data, Decoding the audio stream using an audio decoder to obtain audio data, and includes.
[0360] Optionally, the method further comprises: Synchronizing the audio data and the video data, Outputting the synchronization result to a player for playback, and further includes.
[0361] Optionally, the method further comprises: Decoding the bitstream to obtain video data or image data, performing at least one of luminance mapping, chroma mapping, resolution adjustment, or format conversion on video data or image data; sending the video data or image data to a display; and further comprising.
[0362] In one possible embodiment, an exemplary method for sending a bitstream based on a user operation request is provided, the method comprising: receiving, by an end-side device, a first operation request, the first operation request being used to request playing a target video; in response to the first operation request, determining, in a storage medium, a bitstream corresponding to the target video, the bitstream corresponding to the target video being a bitstream generated according to any one of the encoding methods described above; sending the target bitstream to the end-side device; and comprising.
[0363] Optionally, the method further comprises: encapsulating the bitstream to obtain a transport stream in a first format; sending the transport stream in the first format to a terminal-side device for display or sending the transport stream in the first format to a storage space for storage; and further comprising.
[0364] In one possible embodiment, an exemplary system for sending a bitstream based on a user operation request is provided, the system comprising: a storage medium configured to store a bitstream, the bitstream being a bitstream generated according to any one of the encoding methods described above; A receiver configured to receive a first operation request, A processor configured to determine a target bitstream in a storage medium in response to the first operation request, A transmitter configured to transmit the target bitstream to a device on the terminal side, and
[0365] Optionally, the processor is further configured to encapsulate the bitstream to obtain a transport stream in a first format, and the system is configured to either transmit the transport stream in the first format to a device on the terminal side for display, or transmit the transport stream in the first format to a storage space for storage, and further includes a transmitter configured as such.
[0366] In one possible embodiment, an exemplary method for downloading a bitstream is provided, the method including obtaining the bitstream from a storage medium, where the bitstream is generated according to any one of the encoding methods described previously, decoding the bitstream to obtain a streaming media file, dividing the streaming media file into a plurality of streaming media segments, and downloading the plurality of streaming media segments individually. and
[0367] In one possible embodiment, an exemplary system for downloading a bitstream is provided, the system including an acquisition unit configured to obtain the bitstream from a storage medium, where the bitstream is generated according to any one of the encoding methods described previously, A decoder configured to decode a bitstream to obtain a streaming media file, A processor configured to split the streaming media file into a plurality of streaming media segments, and the processor is further configured to download the plurality of streaming media segments individually. However, the present invention is not limited to any of these exemplary implementations.
[0368] Arithmetic decoding may be performed in parallel, for example, by a multi-core decoder. In addition, only a part of the arithmetic decoding may be performed in parallel. The method of arithmetic decoding may be implemented as range coding.
[0369] The arithmetic coding of the present disclosure can be easily applied to the encoding of the feature maps of a neural network or to the encoding and decoding of classical pictures (still images or videos). The neural network may be used for any purpose, particularly for the encoding and decoding of pictures (still images or videos), or for the encoding and decoding of picture-related data such as motion flow or motion vectors or other parameters. The neural network may also be used in computer vision applications such as image classification, deep detection, segmentation map determination, object recognition for identification, etc.
[0370] Entropy decoding may be performed in parallel, for example, by a multi-core decoder. Additionally, only a part of the entropy decoding may be performed in parallel. FIG. 19 shows an exemplary scheme of a parallel (e.g., multi-core) encoder 620. Each of the input data channels 610 may be encoded into individual sub-streams including coded bits 630-633 and trailing bits 640-643. The length of the sub-stream 650 is signaled. In an implementation of parallel processing, the bit stream is composed of several sub-streams, which are concatenated in a final step. Each of the sub-streams needs to be finalized. This is because the sub-streams are encoded independently of each other, so the encoding (and thus decoding) of one sub-stream does not require the previous encoding (or decoding) of one or more other sub-streams.
[0371] The input data channel may refer to a channel obtained by processing some data by a neural network. For example, the input data may be a feature channel such as an output channel or a latent representation channel of a neural network. In an exemplary implementation, the neural network is a deep neural network and / or a convolutional neural network, etc. The neural network may be trained to process a picture (still image or video). This processing may be for picture encoding and reconstruction, or for computer vision such as object recognition, classification, segmentation, etc. Generally, the present disclosure is not limited to any specific type of task or neural network. Rather, the present disclosure is applicable for encoding any type of data coming from multiple channels, which should generally be understood as any data source. Further, the channels may be provided by preprocessing of the source data.
[0372] Implementation within Picture Coding One possible expansion can be found in FIGS. 20 and 21.
[0373] FIG. 20 shows a schematic block diagram of an exemplary encoder 20 configured to implement the technology of the present application. In the example of FIG. 20, the encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output 272 (or output interface 272). The entropy coding 270 may implement the arithmetic coding method or apparatus described above.
[0374] The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a partitioning unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The encoder 20 as shown in FIG. 20 may also be referred to as a hybrid encoder or an encoder according to a hybrid video / image codec.
[0375] The encoder 20 may be configured to receive, for example, via the input 201, a picture 17 (or picture data 17 or dot cloud data, motion flow, or other types of media data), for example, a picture among a series of pictures forming a video or a video sequence. The received picture or picture data may be a preprocessed picture 19 (or preprocessed picture data 19). In the following description, for simplicity, it is referred to as picture 17. Picture 17 may also be referred to as the current picture or the picture to be coded (especially in video coding, to distinguish the current picture from other pictures, for example, pictures that have been encoded and / or decoded previously in the same video sequence, i.e., the video sequence including the current picture).
[0376] (Digital) pictures are, or can be considered as, two-dimensional arrays or matrices of samples having intensity values. Samples within the array are sometimes also referred to as pixels (short for picture elements) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are utilized, i.e., the picture can be represented by, or can include, three sample arrays. In the RGB format or color space, the picture includes corresponding sample arrays for red, green, and blue. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, such as YCbCr, which includes a luminance component indicated by Y (alternatively, L may also be used in some cases) and two chrominance components indicated by Cb and Cr. The luminance (or short, luma) component Y represents the brightness or gray-level intensity (such as in a grayscale picture), while the two chrominance (or short, chroma) components Cb and Cr represent the chrominance or color information components. Thus, a picture in the YcbCr format includes a luminance sample array of luminance sample values (Y) and two chrominance sample arrays of chrominance values (Cb and Cr). A picture in the RGB format may be converted to the YcbCr format, and vice versa, and this process is also known as color transformation or conversion. If the picture is monochrome, the picture may include only a luminance sample array. Thus, the picture may be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0377] An embodiment of the encoder 20 may include a picture partitioning unit (not shown in FIG. 20) configured to partition picture 17 into a plurality of (typically non-overlapping) picture blocks 203. These blocks may be referred to as root blocks, macroblocks (H.264 / AVC), coding tree blocks (CTB) or coding tree units (CTU) (H.265 / HEVC and VVC). The picture partitioning unit may be configured to use the same block size and corresponding grid defining the block size for all pictures of the video sequence, or to vary the block size between pictures or subsets or groups of pictures, and partition each picture into corresponding blocks. The abbreviation AVC stands for Advanced Video Coding.
[0378] In a further embodiment, the encoder 20 may be configured to directly receive the blocks 203 of picture 17, for example one, some or all of the blocks forming picture 17. The picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.
[0379] Similar to picture 17, the picture blocks 203 are or can be considered as two-dimensional arrays or matrices of samples having intensity values (sample values) but of smaller dimensions than picture 17. In other words, the block 203 may include, for example, one sample array (e.g., a luma array in the case of a monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma and two chroma arrays in the case of a color picture 17), or any other number and / or type of arrays depending on the color format applied. The number of samples in the horizontal and vertical directions (or axes) of the block 203 defines the size of the block 203. Thus, the block may be, for example, an M×N (M columns × N rows) array of samples or an M×N array of transform coefficients.
[0380] An embodiment of the encoder 20 as shown in FIG. 20 may be configured to encode picture 17 block by block. For example, encoding and prediction may be performed for each block 203.
[0381] An embodiment of the encoder 20 as shown in FIG. 20 may be further configured to segment and / or encode a picture using slices (also called video slices), where the picture may be segmented into one or more slices (typically non-overlapping), or encoded using one or more of these slices, and each slice may include one or more blocks (e.g., CTUs).
[0382] An embodiment of the encoder 20 as shown in FIG. 20 may be further configured to segment and / or encode a picture using tile groups (also called video tile groups) and / or tiles (also called video tiles), where the picture may be segmented into one or more tile groups (typically non-overlapping), or encoded using one or more of these tile groups, and each tile group may include one or more blocks (e.g., CTUs) or one or more tiles, where each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), such as complete or partial blocks.
[0383] FIG. 21 shows an example of a decoder 30 configured to implement the technology of the present application. The decoder 30 is configured to receive, for example, encoded picture data 21 (e.g., encoded bitstream 21) encoded by the encoder 20 and obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, such as data representing picture blocks of encoded slices (and / or tile groups or tiles or sub-pictures) and related syntax elements.
[0384] The entropy decoding unit 304 analyzes the bitstream 21 (or generally the encoded picture data 21), and performs, for example, entropy decoding on the encoded picture data 21 to obtain, for example, quantization coefficients 309 and / or decoded coding parameters (not shown in FIG. 21), such as inter prediction parameters (e.g., reference picture index and motion vector), intra prediction parameters (e.g., intra prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or any or all of other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding scheme described for the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360, and provide other parameters to other units of the decoder 30. The decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used. The entropy decoding may implement any of the arithmetic decoding methods or apparatuses described above.
[0385] The reconstruction unit 314 (e.g., adder or summer 314) may be configured to add the reconstruction residual block 313 to the prediction block 365 by adding, for example, the sample values of the reconstruction residual block 313 and the sample values of the prediction block 365 to obtain a reconstruction block 315 within the sample area.
[0386] An embodiment of decoder 30 shown in FIG. 21 may be configured to segment and / or decode a picture using slices (also called video slices), where a picture may be segmented into one or more slices (typically non-overlapping) or decoded using one or more slices, and each slice may include one or more blocks (e.g., CTUs).
[0387] An embodiment of decoder 30 shown in FIG. 21 may be configured to segment and / or decode a picture using tile groups (also called video tile groups) and / or tiles (also called video tiles), where a picture may be segmented or decoded into one or more tile groups (typically non-overlapping), and each tile group may include, for example, one or more blocks (e.g., CTUs) or one or more tiles, where each tile may be, for example, rectangular in shape and may include one or more blocks (e.g., CTUs), e.g., complete or partial blocks.
[0388] Other variations of decoder 30 may be used to decode the encoded picture data 21. For example, decoder 30 may be able to generate an output video stream without loop filtering unit 320. For example, a non-transform-based decoder 30 may be able to directly inverse quantize the residual signal without using inverse transform processing unit 312 for a particular block or frame. In another implementation, decoder 30 may have an inverse quantization unit 310 and an inverse transform processing unit 312 combined in a single unit.
[0389] In encoder 20 and decoder 30, the processing result of the current step may be further processed and output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as Clip or shift may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0390] Implementation in Hardware and Software Some further implementations in hardware and software are described below.
[0391] Any of the encoding devices described above with reference to FIGS. 22 to 25 may provide means for performing the above-described encoding method and decoding method. In particular, the processing circuit in any of these exemplary devices is configured to perform the above-described encoding method and decoding method.
[0392] In the following embodiments of the coding system 10, the encoder 20 and the decoder 30 will be described with reference to FIGS. 22 and 23 in relation to the above-described FIGS. 20 and 21, or other encoders and decoders such as neural network-based encoders and decoders.
[0393] FIG. 22 is a schematic block diagram illustrating an exemplary coding system 10 that may utilize the technology of the present application, such as a video coding system 10 or a picture coding system 10. The encoder 20 and the decoder 30 of the coding system 10 represent examples of devices that may be configured to perform the techniques according to various examples described in the present application.
[0394] As shown in FIG. 22, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to a destination device 14, for example, to decode the encoded picture data 13.
[0395] The source device 12 includes an encoder 20 and, in addition, that is, optionally, may include a picture source 16, a preprocessor (or preprocessing unit) 18, such as an image preprocessor 18, and a communication interface or communication unit 22. The source device 12 can be a cloud server, a content server, or a content delivery server.
[0396] The picture source 16 may include any kind of picture capture device, such as a camera for capturing real-world pictures, and / or any kind of picture generation device, such as a computer graphics processor for generating computer-animated pictures, or any kind of other device for acquiring and / or providing real-world pictures, computer-generated pictures (such as screen content or virtual reality (VR) pictures) and / or any combination thereof (such as augmented reality (AR) pictures). The picture source may be any kind of memory or storage for storing any of the aforementioned pictures.
[0397] Distinguished from the processing executed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.
[0398] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain preprocessed picture 19 or preprocessed picture data 19. The preprocessing executed by the preprocessor 18 may include, for example, trimming, color format conversion (such as from RGB to YcbCr), color correction or noise removal. It can be understood that the preprocessing unit 18 may be any component.
[0399] The encoder 20 is configured to receive the preprocessed picture data 19 and provide encoded picture data 21 (further details have been described above, for example, based on FIG. 20).
[0400] The communication interface 22 of the source device 12 is configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.
[0401] The destination device 14 includes a decoder 30 and, additionally, i.e., optionally, may include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0402] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or any further processed version thereof) directly from, for example, the source device 12 or from any other source, such as a storage device, for example, an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0403] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, or via any kind of network, such as a wired or wireless network or any combination thereof, or any kind of private and public network, or any combination thereof.
[0404] The communication interface 22 may be configured to package, for example, the encoded picture data 21 into a suitable format, such as a packet, and / or to process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network or transmission medium. The communication interface 22 may be configured to encapsulate, for example, the encoded picture data to obtain a transport stream in a first format, and to transmit the transport stream to a terminal-side device for display or to transmit the transport stream in the first format to a storage area for storage.
[0405] The communication interface 28 forms a counterpart of the communication interface 22 and may be configured to, for example, receive the transmitted data, process the transmitted data using any kind of corresponding transmission decoding or processing and / or inverse packetization, and obtain the encoded picture data 21.
[0406] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow of the communication channel 13 in FIG. 22 pointing from the source device 12 to the destination device 14, and may be configured to, for example, send and receive messages, for example, set up a connection, check and exchange any other information related to the communication link and / or data transmission, and may be configured for example for the transmission of encoded picture data.
[0407] The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (further details have been described above, for example, based on FIG. 21).
[0408] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YcbCr to RGB), color correction, trimming or resampling, or any other processing for preparing the decoded picture data 31, for example, for display by the display device 34.
[0409] The display device 34 of the destination device 14 is configured to receive post - processed picture data 33, for example, to display pictures to a user or viewer. The display device 34 may be any kind of display for displaying a reconstructed picture, such as an integrated or external display or monitor, or may include these. The display may include, for example, a liquid crystal display (LCD), an organic light - emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, Lcos (Liquid Crystal on Silicon), a digital light processor (DLP), or any other kind of display.
[0410] FIG. 22 shows the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality, or both. In such embodiments, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or by any combination thereof.
[0411] As will be apparent to those skilled in the art based on the description, the functionality of the different units, or the presence and (exact) partitioning of functionality within the source device 12 and / or destination device 14 as shown in FIG. 22, may vary depending on the actual device and application.
[0412] Encoder 20 or decoder 30, or both encoder 20 and decoder 30, may be implemented via a processing circuit such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof, as shown in FIG. 23. Encoder 20 may be implemented via processing circuit 46 to embody various modules as described with respect to encoder 20 of FIG. 20 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented via processing circuit 46 to embody various modules as described with respect to decoder 30 of FIG. 21 and / or any other decoder system or subsystem described herein. The processing circuit may be configured to perform various operations, as described below. As shown in FIG. 25, if the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either encoder 20 or decoder 30 may be integrated, for example as shown in FIG. 23, as part of a combined encoder / decoder (CODEC) within a single device.
[0413] The source device 12 and the destination device 14 may include any of a wide range of devices, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content delivery server, etc.), a broadcast receiver device, a broadcast transmitter device, etc., including any type of handheld or fixed device, and may or may not use an operating system, or may use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0414] In some cases, the video coding system 10 shown in FIG. 22 is merely an example, and the technology of the present application may be applied to coding settings (such as video / image coding or video / image decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory and streamed, etc. via a network. The encoding device may encode data and store it in memory, and / or the decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other, but simply encode data in memory and / or retrieve data from memory and decode it.
[0415] For the sake of convenience of explanation, embodiments of the present invention are described herein by reference to the reference software of next-generation video coding standards such as HEVC (High-Efficiency Video Coding), or VVC (Versatile Video Coding) developed by ITU-T's VCEG (Video Coding Experts Group) and ISO / IEC's MPEG (Motion Picture Experts Group)'s JCT-VC (Joint Collaboration Team on Video Coding). Those skilled in the art will understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0416] FIG. 24 is a schematic diagram of a coding device (video coding device or image coding device) 400 according to an embodiment of the present invention. The coding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the coding device 400 may be a decoder such as the decoder 30 of FIG. 22 or an encoder such as the encoder 20 of FIG. 22.
[0417] The coding device 400 includes an inlet port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, logic unit or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an outlet port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. The coding device 400 may also include opto-electrical (OE) components and electro-optical (EO) components coupled to the inlet port 410, the receiver unit 420, the transmitter unit 440 and the outlet port 450 for the outlet or inlet of optical or electrical signals.
[0418] Processor 430 is implemented by hardware and software. Processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. Processor 430 communicates with an inlet port 410, a receiver unit 420, a transmitter unit 440, an outlet port 450, and a memory 460. Processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Therefore, including the coding module 470 provides a significant improvement to the functionality of the video coding device 400 and results in the conversion of the video coding device 400 to different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.
[0419] Memory 460 may include one or more disks, tape drives, and solid state drives and may be used to store a program when the program is selected for execution as an overflow data storage device and to store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or non-volatile and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).
[0420] FIG. 25 is a simplified block diagram of an apparatus 500 that may be used as either or both of the source device 12 and the destination device 14 from FIG. 22 according to an exemplary embodiment.
[0421] The processor 502 within the apparatus 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices capable of manipulating or processing information, whether currently existing or developed in the future. The disclosed implementation can be carried out with a single illustrated processor such as, for example, processor 502, and the advantages of speed and efficiency can be achieved using two or more processors.
[0422] The memory 504 within the apparatus 500 can be, in one implementation, a read-only memory (ROM) device or a random access memory (RAM) device in one implementation. Any other suitable type of storage device can be used as the memory 504. The memory 504 can include code and data 506 that are accessed by the processor 502 using the bus 512. The memory 504 can further include an operating system 508 and an application program 510, and the application program 510 includes at least one program that enables the processor 502 to execute the methods described herein. For example, the application program 510 can include applications 1 through N, and these applications further include video coding applications that execute the methods described herein, including encoding and decoding using the arithmetic coding described above.
[0423] The apparatus 500 can also include one or more output devices such as a display 518. The display 518 can be, in one example, a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. The display 518 can be coupled to the processor 502 via the bus 512.
[0424] Although depicted here as a single bus, the bus 512 of the apparatus 500 can be composed of a plurality of buses. Further, the secondary storage 514 can be directly coupled to other components of the apparatus 500 or can be accessed via a network and can include a single integrated unit such as a memory card or a plurality of units such as a plurality of memory cards. Thus, the apparatus 500 can be implemented in a wide variety of configurations.
[0425] Note that embodiments of the coding system 10, the encoder 20, and the decoder 30 (and corresponding system 10), as well as other embodiments described herein, can be configured for video, still image processing, or coding, i.e., for processing or coding individual pictures independent of preceding or subsequent pictures, such as in video encoding. Generally, when picture processing coding is limited to a single picture 17, only the inter prediction units 244 (encoder) and 344 (decoder) may not be available. All other functionality (also referred to as tools or techniques) of the encoder 20 and the decoder 30 may be equally used for still image processing, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, as well as entropy coding 270 and entropy decoding 304.
[0426] For example, the embodiments of the encoder 20 and the decoder 30, and the functions described herein in relation to, for example, the encoder 20 and the decoder 30 may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored on a computer-readable medium or transmitted as one or more instructions or codes via a communication medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another, for example, in accordance with a communication protocol. Thus, the computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0427] FIG. 27 is a block diagram showing a content supply system 3100 for realizing a content delivery service. This content supply system 3100 includes a capture device 3102 and a terminal device 3106, and optionally includes a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination of these types.
[0428] The capture device 3102 may generate data and encode the data by the encoding method as shown in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown), and the server encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 includes, but is not limited to, a camera, a smartphone or a tablet, a computer or a laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 as described above. When the data includes video, the video encoder 20 included in the capture device 3102 may actually perform video encoding processing. When the data includes audio (i.e., voice), the audio encoder included in the capture device 3102 may actually perform audio encoding processing. In some practical scenarios, the capture device 3102 distributes the encoded video data and audio data by multiplexing them together. In other practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106 separately.
[0429] In the content supply system 3100, the terminal device 3106 receives and plays back the encoded data. The terminal device 3106 can be a device having data reception and recovery capabilities such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or any combination thereof, or something similar having the ability to decode the above-described encoded data. For example, the terminal device 3106 may include the destination device 14 as described above. When the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. When the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding processing.
[0430] In the case of a terminal device having its display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can supply the decoded data to its display. In the case of a terminal device without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is connected thereto to receive and display the decoded data.
[0431] When each device in this system performs encoding or decoding, a picture encoding device or a picture decoding device as shown in the above-described embodiments can be used.
[0432] Figure Y shows the structural example of the terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transmission protocol of the stream. The protocol includes, but is not limited to, Real-Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-Time Transport Protocol (RTP), Real-Time Messaging Protocol (RTMP), or any combination of these types.
[0433] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As described above, in some actual scenarios in a video conferencing system, for example, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is transmitted to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.
[0434] Through this demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. The video decoder 3206 includes the video decoder 30 described in the above embodiment, and decodes the video ES by the decoding method shown in the above embodiment to generate video frames, and supplies this data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames, and supplies this data to the synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in Figure Y) before being supplied to the synchronization unit 3212.
[0435] The synchronization unit 3212 synchronizes video frames and audio frames and supplies video / audio to the video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. The information may be coded within the syntax using time stamps related to the presentation of the coded audio and visual data and time stamps related to the delivery of the data stream itself.
[0436] When subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video frames and audio frames, and supplies video / audio / subtitles to the video / audio / subtitle display 3216.
[0437] The present invention is not limited to the above system, and either the picture encoding device or the picture decoding device in the above embodiment can be incorporated into other systems, for example, in-vehicle systems.
[0438] By way of example and not limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can be properly termed a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0439] The commands may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated in a combined codec. Also, the technology may be fully implemented in one or more circuits or logic elements.
[0440] The techniques of the present disclosure may be implemented in a variety of devices or apparatuses, including wireless handsets, cloud servers, application servers, integrated circuits (ICs) or sets of ICs (e.g., chip sets). In the present disclosure, various components, modules or units are described to emphasize the functional aspects of devices configured to execute the disclosed techniques, but do not necessarily require implementation by different hardware units. Rather, as described above, the various units may be combined within codec hardware units, or may provide a set of interoperable hardware units including one or more processors as described above, together with appropriate software and / or firmware.
Claims
Claim 1 A decoding method implemented by a decoder, comprising: receiving a bitstream including encoded data of an input signal and a first parameter; analyzing the bitstream to obtain the first parameter; obtaining an entropy coding parameter based on the first parameter; reconstructing at least a part of the input signal based on the entropy coding parameter and the encoded data; A decoding method comprising the above steps. Claim 2 The entropy coding parameter includes at least one of: the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder, i.e., the size of the alphabet of the entropy encoder, or the minimum symbol probability supported by the entropy encoder, or the look-ahead period of the entropy encoder. The decoding method according to claim 1, including at least one of the above. Claim 3 When the first parameter is the size of the alphabet, the step of obtaining the entropy coding parameter based on the first parameter includes: using the first parameter as the size of the alphabet. The decoding method according to claim 2. Claim 4 The first parameter is p, and the entropy coding parameter includes the size of the alphabet M, where M is a function of p. The decoding method according to claim 1 or 2. Claim 5 The step of obtaining the entropy coding parameter based on the first parameter includes: M = f -1 (p) including, where f -1 (p) is the inverse function of f(M) and f(M) = p, The decoding method according to claim 4. Claim 6 M satisfies one of the following, i.e., M = k a*p+C , where k is a natural number, a and C are constants, or M = a*p + b, where a and b are constants, or M = p 2 The decoding method according to claim 5, satisfying one of the above. Claim 7 p = log 2 (M) - 9, and M = f -1 (p) = 2 p+9 where f -1 (p) is the inverse function of f(M), and f(M) = log 2 (M) - 9 The decoding method according to claim 6. Claim 8 p is signaled using one of the following codes, i.e., binary code, or unary code, or truncated unary code, or exp-Golomb code. The decoding method according to claims 4 to 7, signaled using one of the above. Claim 9 p is signaled using an exp-Golomb code of degree 0. The decoding method according to claim 8. Claim 10 The first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient. The decoding method according to claim 1 or 2.
11. Based on the first parameter, the step of obtaining the entropy coding parameter includes: Determining a target sub-range where the first parameter is located, wherein an allowable range of values of the first parameter includes a plurality of sub-ranges, the target sub-range is one of the plurality of sub-ranges, each of the plurality of sub-ranges includes at least one value of the first parameter, and each of the plurality of sub-ranges corresponds to one value of the entropy coding parameter. Using the value of the entropy coding parameter corresponding to the target sub-range as the value of the entropy coding parameter, or Calculating the value of the entropy coding parameter based on one or more values of the entropy coding parameter corresponding to one or more sub-ranges adjacent to the target sub-range. The decoding method according to claim 10, including the above.
12. The first parameter is D, and the entropy coding parameter includes the size of alphabet M, where M is obtained based on P and D, and P is a predictor derived by the decoder. The decoding method according to claim 1.
13. Based on the first parameter, the step of obtaining the entropy coding parameter includes: M = s -1 (D, P) including, where s -1 (D, P) is the inverse function of s(M, P) and s(M, P) = D The decoding method according to claim 12.
14. M is one of the following, namely: 【Number 1】 where a, b, and C are predetermined constants, or M = a1*D + b1*P + c1, where a1, b1, and c1 are predetermined constants. The decoding method according to claim 13, satisfying one of the above.
15. 【Fig. 2】 The decoding method according to claim 14, which is as described above.
16. D is signaled using one of the following codes, namely: a binary code, or a unary code, or a truncated unary code, or an exp-Golomb code. The decoding method according to claims 12 to 15, signaled using one of the above.
17. P is derived based on at least one parameter other than the first parameter carried in the bitstream, The decoding method according to any one of claims 12 to 16.
18. The at least one parameter other than the first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a density of pixels in a 3D object, or a rate distortion weight coefficient. The decoding method according to claim 17.
19. P is obtaining a rate control parameter beta (β) from the bitstream, determining a target subrange in which the obtained β is located, wherein an allowable range of values of the rate control parameter β is [β_0, β_K], the allowable range [β_0, β_K] is divided into a plurality of subranges, the target subrange is one of the plurality of subranges, each of the plurality of subranges includes at least one value of β, and each of the plurality of subranges corresponds to one value of P. selecting, as the value of P, the value corresponding to the target subrange, or calculating the value of P based on one or more values corresponding to one or more subranges adjacent to the target subrange, The decoding method according to claim 18, including being derived based on the at least one parameter.
20. The decoding method further includes the step of analyzing the bitstream to obtain a flag, wherein the flag is used to indicate whether the entropy coding parameter is directly carried in the bitstream. The decoding method according to any one of claims 1 to 19.
21. When the flag is equal to a first value, it is specified that the entropy coding parameter is carried in the bitstream. In this case, the first parameter is the entropy coding parameter, or the first parameter is a conversion result of the entropy coding parameter, or When the flag is equal to a second value, it is specified that the entropy coding parameter is not carried in the bitstream, and the entropy coding parameter is derived by the decoder. The decoding method according to claim 20.
22. When the flag is equal to the third value, it means that the difference value between M and P, or the conversion result of the difference value between M and P is conveyed within the bitstream, and the first parameter is the difference value between M and P, or the conversion result of the difference value between M and P, where M is the size of the input alphabet and P is the predictor derived by the decoder. The decoding method according to claim 20 or 21.
23. The entropy coder is an arithmetic coder, a range coder or an ANS (Asymmetric Numerical Systems) coder. The decoding method according to any one of claims 2 to 22.
24. Based on the entropy coding parameter and the encoded data, the step of reconstructing at least a part of the input signal is as follows: Obtaining at least one probability model, wherein a probability model of the output symbol is used to indicate the probability of each possible value of the output symbol. By using the at least one probability model and the entropy coding parameter, entropy-decoding one or more bits of the encoded data to obtain one or more output symbols. Based on the one or more output symbols, reconstructing at least a part of the input signal. The decoding method according to any one of claims 1 to 23, including the above steps.
25. The probability model depends on the entropy coding parameter. The decoding method according to claim 24.
26. A decoding method implemented by a decoder, comprising: Receiving a bitstream including encoded data and a flag of an input signal. Analyzing the bitstream to obtain the flag, wherein the flag is used to indicate whether the entropy coding parameter is directly conveyed within the bitstream. Based on the flag, obtaining the entropy coding parameter. Based on the entropy coding parameter, reconstructing at least a part of the input signal. The decoding method including the above steps.
27. The entropy coding parameter is The alphabet size of the entropy encoder, which is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder, or the minimum symbol probability supported by the entropy encoder, or the look-ahead period of the entropy encoder, The decoding method according to claim 26, comprising at least one of the above. **Claim 28** When the flag is equal to the first value, it is specified that the entropy coding parameter is conveyed within the bit stream, or the conversion result of the entropy coding parameter is conveyed within the bit stream, or When the flag is equal to the second value, it is specified that the entropy coding parameter is not conveyed within the bit stream, but the entropy coding parameter is derived by the decoder. The decoding method according to claim 26 or 27. **Claim 29** When the flag is equal to the third value, it is specified that the difference value between M and P is conveyed within the bit stream, or the conversion result of the difference value between M and P is conveyed within the bit stream, where M is the entropy coding parameter and P is a predictor derived by the decoder. The decoding method according to claim 28. **Claim 30** Based on the flag, the step of obtaining the entropy coding parameter is When the flag is equal to the first value, the step of analyzing the bit stream to obtain a first parameter, where the first parameter is the entropy coding parameter, The step of using the first parameter as the entropy coding parameter, Or The first parameter is the conversion result of the entropy coding parameter, The step of obtaining the entropy coding parameter based on the first parameter, The decoding method according to claim 28, including the above. **Claim 31** Based on the flag, the step of obtaining the entropy coding parameter is When the flag is equal to the second value, a step of analyzing the bit stream to obtain a second parameter, wherein the second parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a pixel density in a 3D object, or a rate distortion weight coefficient; A step of deriving the entropy coding parameter based on the second parameter; The decoding method according to claim 28, comprising:
32. The step of obtaining the entropy coding parameter based on the flag is: When the flag is equal to the third value, a step of analyzing the bit stream to obtain a third parameter, wherein the third parameter is the difference value between M and P, or the third parameter is a conversion result of the difference value between M and P, M is the entropy coding parameter, and P is a predictor derived by the decoder; A step of deriving P based on at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, a pixel density in a 3D object, or a rate distortion weight coefficient; A step of obtaining an entropy coding parameter based on the third parameter and P; The decoding method according to claim 29, comprising:
33. An encoding method implemented by an encoder, comprising: A step of encoding an input signal and a first parameter into a bit stream, wherein the first parameter is used to obtain an entropy coding parameter; A step of transmitting the bit stream to a decoder; An encoding method, comprising:
34. The entropy coding parameter is: The size of the alphabet of the entropy encoder, which is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder, or The minimum symbol probability supported by the entropy encoder, or The feedback period of the entropy encoder, The encoding method according to claim 33, comprising at least one of the above.
35. The first parameter is the size of the alphabet; The encoding method according to claim 34.
36. The first parameter is p, where p is the conversion result of M, and M is the entropy coding parameter. The encoding method according to claim 33 or 34.
37. p = f(M), where f(M) is a reversible function. The encoding method according to claim 36.
38. f(M) includes the following, namely: f(M)=a*log k (M)+b, where k is a natural number, a and b are constants, or f(M) = a*M + b, where a and b are constants, or f(M) = sqrt(M). The encoding method according to claim 37.
39. p = log 2 (M) - 9, The encoding method according to claim 38.
40. p is one of the following codes, namely: binary code, or unary code, or truncated unary code, or exp-Golomb code The decoding method according to claims 4 to 7, which is signaled using one of them.
41. The first parameter includes at least one of a rate control parameter, a quantization parameter (qp), an image resolution, a video resolution, a frame rate, the density of pixels in a 3D object, or a rate-distortion weight coefficient. The first parameter is used by the entropy decoder to derive the entropy coding parameter. The encoding method according to claim 33 or 34.
42. The first parameter is D obtained based on P and M, where M is the entropy coding parameter and P is a predictor derived by the decoder. The encoding method according to claim 33 or 34.
43. D = s(M, P), where s(M, P) is a reversible function. The encoding method according to claim 42.
44. s(M, P) is s(M,P) = a * log k (P) + b * log k (M) - c, where k is a natural number, a, b, and c are constants, or s(M, P) = a*M + b*P + c, where a, b, and c are constants. including The encoding method according to claim 43.
45. D = s(M, P) = log 2 (P) - log 2 (M), where The encoding method according to claim 44.
46. D is one of the following codes, namely: binary code, or unary code, or truncated unary code, or exp-Golomb code The encoding method according to claims 42 to 45, which is signaled using one of them.
47. The encoding method further includes the step of encoding a flag into the bitstream, where the flag is used to indicate whether the entropy coding parameter is directly carried in the bitstream. The encoding method according to any one of claims 33 to 46.
48. When the flag is equal to the first value, it is specified that the entropy coding parameter is carried in the bit stream, and the first parameter is the entropy coding parameter, or the first parameter is a conversion result of the entropy coding parameter, or When the flag is equal to the second value, it is specified that the entropy coding parameter is not carried in the bit stream, but the entropy coding parameter is derived by the decoder. The encoding method according to claim 47.
49. When the flag is equal to the third value, it is specified that the difference value between M and P is carried in the bit stream, or a conversion result of the difference value between M and P is carried in the bit stream, where M is the entropy coding parameter and P is a predictor derived by the decoder. The encoding method according to claim 47 or 48.
50. The method comprises: a step of obtaining a minimum value and a maximum value of a latent space element of the entropy encoder, where the latent space element is a result of progress of the input signal; the size of the alphabet M = ceil(max{y} - min{y}) or M = 2^(ceil(log 2 (max{y} - min{y}))) is obtained according to where ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space element, min{y} represents the minimum value of the latent space element, and M represents the size of the alphabet. The encoding method according to any one of claims 33 to 49.
51. The method comprises: M 0 A step of obtaining at least two values around M 0 = ceil(max{y} - min{y}) or M 0 = 2 ^ (ceil(log 2 (max{y} - min{y}))) is, the step and a step of calculating a loss function for the at least two values; a step of selecting, as the size of the alphabet, a value having the minimum loss function among the at least two values; where ceil(x) is the smallest integer greater than x, max{y} represents the maximum value of the latent space element, and min{y} represents the minimum value of the latent space element. The encoding method according to any one of claims 33 to 49.
52. An encoding method implemented by an encoder, Encoding an input signal and a flag into a bit stream, wherein the flag is used to indicate whether an entropy coding parameter is directly carried within the bit stream, and transmitting the bit stream to a decoder, and an encoding method comprising the steps above. **Claim 53** The entropy coding parameter is the size of the alphabet of an entropy encoder, which is the size of the input alphabet of the entropy encoder or the size of the output alphabet of the entropy decoder, or the minimum symbol probability supported by the entropy encoder, or the look-ahead period of the entropy encoder, The encoding method according to claim 33, comprising at least one of the above. **Claim 54** When the flag is equal to a first value, it is specified that the entropy coding parameter is carried within the bit stream, or a conversion result of the entropy coding parameter is carried within the bit stream, or When the flag is equal to a second value, it is specified that the entropy coding parameter is not carried within the bit stream, but the entropy coding parameter is derived by the decoder, The encoding method according to claim 52 or 53. **Claim 55** When the flag is equal to a third value, it is specified that a difference value between M and P is carried within the bit stream, or a conversion result of the difference value between M and P is carried within the bit stream, where M is the entropy coding parameter and P is a predictor derived by the decoder, The encoding method according to any one of claims 52 to 54. **Claim 56** The method further comprises when the flag is equal to the first value, encoding a first parameter into the bit stream, where the first parameter is the entropy coding parameter or a conversion result of the entropy coding parameter, The encoding method according to any one of claims 52 to 55, further comprising the step above. **Claim 57** The method When the flag is equal to a third value, encoding a third parameter into the bitstream, where the third parameter is a difference value between M and P, or the third parameter is a conversion result of the difference value between M and P, M is the entropy coding parameter, and P is a predictor derived by the decoder; The encoding method according to any one of claims 52 to 55, further comprising.
58. A decoding device, A processing circuit configured to execute the steps of the method according to any one of claims 1 to 25 or any one of claims 26 to 32; A decoding device comprising the same.
59. An encoding device, A processing circuit configured to execute the steps of the method according to any one of claims 33 to 51 or any one of claims 52 to 57; An encoding device comprising the same.
60. One or more processors; A non-transitory computer-readable storage medium coupled to the one or more processors; A decoder comprising: the storage medium stores programming for execution by the one or more processors, and when the programming is executed by the one or more processors, configures the decoder to execute the method according to any one of claims 1 to 25 or any one of claims 26 to 32; Decoder.
61. One or more processors; A non-transitory computer-readable storage medium coupled to the one or more processors; An encoder comprising: the storage medium stores programming for execution by the one or more processors, and when the programming is executed by the one or more processors, configures the encoder to execute the method according to any one of claims 33 to 51 or any one of claims 52 to 57; Encoder.
62. A non-transitory computer-readable medium carrying computer instructions that, when executed by a computer device or one or more processors, cause the computer device or the one or more processors to execute the method according to any one of claims 1 to 57.
63. A non-transitory storage medium comprising a bitstream encoded by the method according to any one of claims 33 to 51 or any one of claims 52 to 57.
64. A computer program stored in a non-transitory medium and including code instructions, which, when executed on one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 57.
65. A system for delivering a bitstream, comprising: at least one storage medium configured to store at least one bitstream generated by the method according to any one of claims 33 to 51 or any one of claims 52 to 57; a video streaming device configured to obtain a bitstream from one of the at least one storage media and transmit the bitstream to a terminal device; wherein the video streaming device comprises a content server or a content delivery server. System.
66. further comprising one or more processors configured to perform an encryption process on at least one bitstream to obtain at least one encrypted bitstream, wherein the at least one storage medium is configured to store the encrypted bitstream, or wherein the one or more processors are configured to convert a bitstream in a first format into a bitstream in a second format, wherein the at least one storage medium is configured to store the bitstream in the second format. The system according to claim 65.
67. a receiver configured to receive a first operation request; one or more processors configured to determine a target bitstream in the at least one storage medium in response to the first operation request; a transmitter configured to transmit the target bitstream to a terminal-side device. The system according to claim 65 or 66.
68. The one or more processors are further configured to encapsulate the bitstream to obtain a transport stream in a first format. The transmitter is further configured to transmit the transport stream in the first format to a terminal-side device for display or transmit the transport stream in the first format to a storage area for storage. The system according to claim 67.
Citation Information
Patent Citations
Data encoding method and device, data decoding method and device, and recording medium
JP3339054B2
Adaptive quantization
US10594338B1
Adaptive binarizer selection for image and video coding
US20170180732A1
Genomic information compression by configurable machine learning-based arithmetic coding
WO2022008311A1