Entropy Coding for Signal Enhancement Coding

By applying the run length encoding operation on the residual data of the video signal, the problem of difficulty in realizing low-complexity entropy encoding in the prior art is solved, and efficient compression and rapid decoding of the residual data are achieved.

CN113228668BActive Publication Date: 2025-05-13V NOVA INT LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980064342.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-15
Filing Date
2019-08-01
Publication Date
2025-05-13
Estimated Expiration
2039-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to provide a low-complexity, simple and fast entropy encoding scheme to effectively compress the residual data of the video signal.

Method used

The residual data is encoded using a run-length encoding operation, by generating a run-length encoded byte stream, which includes a set of symbols representing the count of non-zero data values ​​and consecutive zero values ​​of the residual data.

Benefits of technology

Efficient data compression of residual data is realized, and the unique characteristics of residual data are utilized, such as the relative occurrence rate of zero values ​​and the relative diversity of data values, simplifying the encoding and decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113228668B_ABST
    Figure CN113228668B_ABST
Patent Text Reader

Abstract

A method for encoding a video signal is provided, the method comprising: receiving an input frame; processing the input frame to generate at least one set of residual data, the residual data enabling a decoder to reconstruct the input frame from a reconstructed reference frame; and applying a run-length encoding operation to the residual data set, wherein the run-length encoding operation comprises generating a run-length encoded byte stream, the run-length encoded byte stream comprising a set of symbols representing non-zero data values ​​of the residual data set and counts of consecutive zero values ​​of the residual data set. In some embodiments, the method comprises applying a Huffman encoding operation to the set of symbols. A decoding method, apparatus, and computer-readable medium are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Hybrid backward-compatible coding techniques have been previously proposed in, for example, WO 2014 / 170819 and WO 2018 / 046940 (the contents of which are incorporated herein by reference).

[0002] A method is presented which parses a data stream into a first portion of encoded data and a second portion of encoded data; implements a first decoder to decode the first portion of the encoded data into a first rendition of a signal; implements a second decoder to decode the second portion of the encoded data into reconstruction data, the reconstruction data specifying how to modify the first rendition of the signal; and applies the reconstruction data to the first rendition of the signal to produce a second rendition of the signal.

[0003] Another method is also proposed, wherein a set of residual elements can be used to reconstruct a reproduction of a first time sample of a signal. A set of spatiotemporal correlation elements associated with the first time sample is generated. The set of spatiotemporal correlation elements indicates a degree of spatial correlation between a plurality of residual elements and a degree of temporal correlation between first reference data based on the reproduction and second reference data based on the reproduction of a second time sample of the signal. The set of spatiotemporal correlation elements is used to generate output data.

[0004] Typical video encoding techniques involve applying an entropy coding operation to output data. A low complexity, simple and fast entropy coding scheme is needed that applies data compression to the reconstructed data or residual elements of the above proposed techniques or other similar residual data. Summary of the invention

[0005] According to an aspect of the present invention, a method of entropy encoding or decoding residual data is provided, wherein the residual data may be used to correct or enhance a base stream, such as data of a video frame encoded using a legacy video encoding technique.

[0006] According to one aspect of the present invention, a method for encoding a video signal is provided. The method comprises: receiving an input frame; processing the input frame to generate residual data, the residual data enabling a decoder to reconstruct the input frame from a reconstructed reference frame; and applying a run-length encoding operation to the residual data, wherein the run-length encoding operation comprises generating a run-length encoded byte stream, which comprises a set of symbols representing non-zero data values ​​of the residual data and counts of consecutive zero values ​​of the residual data.

[0007] In this way, the encoding method provides a low complexity, simple and fast solution for data compression of residual data. The solution exploits the unique characteristics of the residual data, such as the relative occurrence of zero values, the possible grouping of zero values ​​based on the scan order of the transform process that generated the residual data (if a transform is used), and the relative diversity of data values ​​and their possible relative frequency / rarity in the residual data.

[0008] The set of symbols may be sequential in the encoded byte stream. A count of consecutive zero values ​​may also be referred to as a run of zeros. The residual data on which a run-length encoding operation is performed may represent quantized residual data, preferably a quantized set of transform coefficients. The quantized set of transform coefficients may be ordered by layer, i.e., by sets of coefficients of the same type, by plane, by quality level, or by surface. These terms are further described and defined herein. The symbols may be sequential in a corresponding scan order of a transform operation so that the residual data may be easily matched with a reconstructed reference frame.

[0009] Preferably, the run-length encoding operation includes: encoding non-zero data values ​​of the residual data into at least a first type of symbol; and encoding a count of consecutive zero values ​​into a second type of symbol, so that the residual data is encoded as a sequence of symbols of different types. Therefore, the residual data can be encoded as a sequence of symbols including data value types and runs of zero types. The types promote speed and ease of decoding at the decoder and facilitate another entropy encoding operation, as described in detail below. Structuring the data into a minimized set of fixed-length codes of different types facilitates subsequent steps.

[0010] In some embodiments, the run-length encoding operation includes encoding the data value of the residual data into a first type of symbol and a third type of symbol, each of which includes a portion of the data value so that the portion can be combined at the decoder to reconstruct the data value. Each type of symbol may have a fixed size and may be one byte. Data values ​​greater than a threshold or greater than the size available in one type of symbol may be easily transmitted or stored using a byte stream. Structuring the data into bytes or symbols not only helps with decoding but also helps with entropy encoding of fixed-length symbols, as explained below. Introducing a third type of symbol allows three types of symbols that are easily distinguishable to be formed from each other at the decoder side. In the case of the type of reference symbol, the terms block, byte (where the symbol is a byte), context, or type may be used.

[0011] The run-length encoding operation may include: comparing the size of each data value of the residual data to be encoded with a threshold; and encoding each data value into a first type of symbol when the size is below the threshold and encoding a portion of each data value into the first type of symbol and encoding a portion of each data value into a third type of symbol when the size is above the threshold.

[0012] If the size is above a threshold, the method may include setting a flag in a symbol of the first type of symbol indicating that a portion of the represented data value is encoded into another symbol of a third type of symbol. The flag may be an overflow flag or overflow bit and in some instances may be the least significant bit of a symbol or byte. In the case where the least significant bit of a symbol is a flag, the data value or portion of the data value may be included in the remaining bits of the byte. Setting the overflow bit as the least significant bit facilitates ease of combination with subsequent symbols.

[0013] The method may further include inserting in each symbol a flag indicating the type of symbol to be encoded next in the run-length encoded byte stream. The flag may be an overflow bit as described above or may further be a run flag or run bit indicating whether the next symbol includes a run of zeros or a data value or a portion of a data value. In the case where the flag indicates whether the next symbol includes a run (or a count of consecutive zeros), the flag may be a sign bit, preferably the most significant bit of the symbol. The count of consecutive zeros or data values ​​may be included in the remaining bits of the symbol. Thus, a 'run of zeros' or second type of symbol includes 7 available bits of a byte for a count and a 'data' or first type of symbol includes 6 available bits for data values ​​where the overflow bit is not set and 7 available bits for data values ​​where the overflow bit is set.

[0014] In general, the flag may be different for each type of symbol and may indicate the type of symbol that follows in the stream.

[0015] In a preferred embodiment of the above aspect, the method further includes applying another entropy coding operation to the set of symbols generated by the run-length coding operation. Thus, the fixed-length symbols intentionally generated by the structure of the run-length coding operation can be converted into variable-length codes to reduce the overall size of the data. The run-length coding structure is designed to form high-frequency fixed-length symbols to facilitate the improvement of another entropy coding operation and the overall reduction of data size. Another entropy coding operation utilizes the probability or frequency of occurrence of symbols in the byte stream formed by the run-length coding operation.

[0016] In a preferred embodiment, the other entropy coding operation is a Huffman coding operation or an arithmetic coding operation. Preferably, the method further comprises applying a Huffman coding operation to the symbol set to generate Huffman coded data including a code set representing the run-length coded byte stream. The Huffman coding operation receives the symbols of the run-length coded byte stream as input and outputs a plurality of variable length codes in a bit stream. Huffman coding is an encoding technique that can effectively reduce fixed length codes. The structure of the run-length coded bytes means that these two types of symbols are very likely to be frequently replicated, which means that only a few bits will be needed for each cyclic symbol. Across planes, the same data value is likely to be replicated (especially when the residual is quantized) and therefore the variable length code can be smaller (but repeated).

[0017] Huffman coding is also well optimized for software implementation (as expected here) and is computationally efficient. It uses minimal memory and has faster decoding when compared to other entropy coding techniques because the processing steps only involve walking through the tree. The computational benefits result from the shape and depth of the tree formed by the specifically designed run-length encoding operation.

[0018] The encoding operation may be performed on a set of coefficients, i.e., a layer, a plane of a frame, a quality level, or a complete surface or frame. That is, the statistics or parameters of each operation may be determined based on the set of data values ​​to be encoded and zero data values ​​or symbols representing those values.

[0019] The Huffman encoding operation may be a canonical Huffman encoding operation such that the Huffman encoded data includes a code length for each unique symbol in the set of symbols, the code length representing the length of the code used to encode the corresponding symbol. Canonical Huffman encoding facilitates a reduction in encoding parameters that need to be signaled between an encoder and a decoder in metadata by ensuring that the same code length is attached to symbols in a sequential manner. The codebook of the decoder can be inferred and only the code length needs to be sent. Thus, a canonical Huffman encoder is particularly efficient for shallow trees where there are a large number of different data values.

[0020] In some examples, the method further includes comparing the data size of at least a portion of the run-length encoded byte stream with the data size of at least a portion of the Huffman encoded data; and outputting the run-length encoded byte stream or the Huffman encoded data in an output bitstream based on the comparison. Each block or segment of the input data is thus selectively sent to reduce the overall data size, thereby ensuring that the decoder can easily distinguish the encoding type. The difference may be within a threshold or tolerance so that although one scheme may produce less data, the operational efficiency may mean that one scheme is preferred.

[0021] To distinguish the schemes, the method may further include adding a flag indicating whether the bitstream represents a run-length encoded byte stream or Huffman encoded data to the configuration metadata accompanying the output bitstream. This is a fast and efficient way of signaling. Alternatively, the decoder may recognize the scheme used by the structure of the data, for example, Huffman encoded data may have a header portion and a data portion, while run-length encoded data may only include a data portion and thus the decoder may be able to distinguish between the two schemes.

[0022] The Huffman coding operation may include generating separate frequency tables for symbols of the first type of symbols and symbols of the second type of symbols. In addition, separate Huffman coding operations may be applied for each type of symbol. These concepts contribute to the specific efficiency of variable length coding schemes. For example, in the case of replicating the same data value across planes, such as in video coding where scenes may have similar colors or errors / enhancements, these values ​​may only require a few codes (where the codes are not deformed by symbols of the 'run of zeros' type). Therefore, this embodiment of the present invention is particularly advantageous for use in conjunction with the above-mentioned run-length coding operation. As shown above, in the case of quantized values, the symbols here may be replicated and there may not be too many different values ​​within the same frame or coefficient group.

[0023] The use of different frequency tables for each type of symbol provides specific benefits of a Huffman encoding operation following a specifically designed run-length encoding operation.

[0024] The method may thus include identifying the type of symbol to be encoded next, selecting a frequency table based on the type of the next symbol, decoding the next symbol using the selected frequency table and outputting the encoded symbols in sequence.The frequency table may correspond to a unique codebook.

[0025] The method may include generating a stream header that may include an indication of the multiple code lengths, so that a decoder can derive the code lengths and corresponding corresponding symbols for the canonical Huffman decoding operation. For example, the code length may be the length of each code used to encode a particular symbol. The code length thus has a corresponding code and a corresponding symbol. The stream header may be a stream header of a first type and further include an indication of the symbols associated with the corresponding code lengths in the multiple code lengths. Alternatively, the stream header may be a stream header of a second type and the method may further include sorting the multiple code lengths in the stream header in a predetermined order based on the symbols corresponding to each of the code lengths, so that the code lengths can be associated with the corresponding symbols at the decoder. Each type will provide efficient signaling of encoding parameters depending on the symbol and length to be decoded. The code length may be signaled as the difference between the length and another length, preferably the minimum code length that is signaled.

[0026] The second type of stream header may further include a flag indicating that symbols in a predetermined order of a possible set of symbols are not present in the set of symbols in the run-length encoded byte stream. In this way, only the required length needs to be included in the header.

[0027] The method may further include comparing the number of unique codes in the Huffman encoded data to a threshold value and generating a first type of stream header or a second type of stream header based on the comparison.

[0028] The method may further include comparing the number of non-zero symbols or data symbols to a threshold and generating a first type of stream header or a second type of stream header based on the comparison.

[0029] According to another aspect of the present invention, a method for decoding a video signal may be provided, the method comprising: retrieving an encoded bit stream; decoding the encoded bit stream to generate residual data, and reconstructing an original frame of the video signal from the residual data and a reconstructed reference frame, wherein the step of decoding the encoded bit stream comprises: applying a run-length encoding operation to generate residual data, wherein the run-length encoding operation comprises identifying a set of symbols representing non-zero data values ​​of the residual data and a set of symbols representing counts of consecutive zero values ​​of the residual data; parsing the set of symbols to derive the non-zero data values ​​and counts of consecutive zero values ​​of the residual data; and generating the residual data from the non-zero data values ​​and the counts of consecutive zero values.

[0030] The run-length encoding operation may include: identifying symbols of a first type of symbol representing non-zero data values ​​of residual data; identifying symbols of a second type of symbol representing a count of consecutive zero values, so that a sequence of different types of symbols are decoded to generate residual data; and parsing a set of symbols according to the corresponding type of each symbol.

[0031] The run-length encoding operation may include identifying symbols of a first type of symbol and symbols of a third type of symbol, wherein the first type of symbol and the third type of symbol each represent a portion of a data value; parsing the symbols of the first type of symbol and the symbols of the third type of symbol to derive portions of the data value; and combining the derived portions of the data value into the data value.

[0032] The method may further include retrieving from a symbol of the first type of symbols an overflow flag indicating whether a portion of a data value of the symbol of the first type of symbols is included in a subsequent symbol of a third type.

[0033] The method may further include retrieving from each symbol a flag indicating a subsequent type of symbol expected in the set of symbols.

[0034] The initial symbol may be assumed to be a symbol of a first type such that the initial symbol is parsed to derive at least a portion of a data value.

[0035] The step of decoding the encoded bitstream may include applying a Huffman coding operation to the encoded bitstream to generate the set of symbols representing non-zero data values ​​of the residual data and the set of symbols representing counts of consecutive zero values ​​of the residual data, wherein the step of applying a run-length encoding operation to the set of symbols to generate the residual data is performed.

[0036] The Huffman encoding operation may be a canonical Huffman encoding operation.

[0037] The method may further include retrieving a flag from configuration metadata accompanying the encoded bitstream indicating whether the encoded bitstream includes a run-length encoded byte stream or Huffman encoded data; and selectively applying the Huffman encoding operation based on the flag.

[0038] The method may further include: identifying an initial type of symbol expected to be derived in an encoded bitstream; applying a Huffman encoding operation to the encoded bitstream to derive the initial symbol based on a set of encoding parameters associated with the expected initial type of symbol; retrieving a flag from the initial symbol indicating a subsequent type of symbol expected to be derived from the bitstream; and further applying the Huffman encoding operation to the bitstream to derive subsequent symbols based on a set of encoding parameters associated with the subsequent type of symbol expected to be derived from the bitstream.

[0039] The method may further include repeatedly retrieving from the decoded symbols a flag indicating a subsequent type of symbol expected to be derived from the bitstream and applying the Huffman encoding operation to the bitstream to derive subsequent symbols based on a set of encoding parameters associated with the subsequent type of symbol expected to be derived from the bitstream.

[0040] Thus, the method may include the steps of decoding a symbol using a Huffman encoding operation, decoding a symbol using a run-length encoding operation, identifying the type of symbol expected next based on the run-length encoding operation, and decoding the next symbol using the Huffman encoding operation.

[0041] The method may further include: retrieving a stream header including an indication of a plurality of code lengths to be used for a canonical Huffman encoding operation; associating each code length with a corresponding symbol according to the canonical Huffman encoding operation; and identifying a code associated with each symbol based on the code length according to the canonical Huffman encoding operation. Thus, the code may be associated with sequential symbols of the same code length.

[0042] The stream header may include an indication of symbols associated with respective ones of the plurality of code lengths and associating each code length with a symbol includes associating each code length with a corresponding symbol in the stream header.

[0043] The step of associating each code length with a corresponding symbol includes associating each code length with a symbol from a predetermined set of symbols according to the order in which each code length is retrieved.

[0044] The method may further include not associating a symbol in the predetermined set of symbols with a corresponding code length, wherein a flag of the stream header indicates that the code length is not present in the stream header for the symbol.

[0045] According to one aspect of the present invention, a method for encoding a video signal is provided. The method comprises: receiving an input frame; processing the input frame to generate residual data, the residual data enabling a decoder to reconstruct the input frame from a reconstructed reference frame; and applying a Huffman coding operation to the residual data, wherein the Huffman operation comprises generating an encoded bitstream comprising a set of codes for encoding a set of symbols representing the residual data.

[0046] According to another aspect of the present invention, a method for decoding a video signal may be provided, the method comprising: retrieving an encoded bit stream; decoding the encoded bit stream to generate residual data, and reconstructing the original frame of the video signal from the residual data and a reconstructed reference frame, wherein the step of decoding the encoded bit stream comprises: applying a Huffman coding operation to generate residual data, wherein the Huffman coding operation comprises generating a set of symbols representing residual data by comparing the encoded bit stream with a reference code mapped to corresponding symbols and generating the residual data by counting non-zero data values ​​and consecutive zero values.

[0047] According to another aspect, a device for encoding a data set into an encoded data set may be provided. The device is configured to encode an input video according to the above steps. The device may include a processor configured to perform the method of any one of the above aspects.

[0048] According to another aspect, a device for decoding a data set into a video reconstructed from the data set may be provided. The device is configured to decode the output video according to the above steps. The device may include a processor configured to perform the method of any one of the above aspects.

[0049] Encoders and decoders are also available.

[0050] According to another aspect of the present invention, a computer readable medium may be provided, which, when executed by a processor, causes the processor to perform any one of the methods of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Examples of systems and methods according to the present invention will now be described with reference to the accompanying drawings, in which:

[0052] Figure 1 A high-level diagram showing the encoding process;

[0053] Figure 2 A high-level diagram showing the decoding process;

[0054] Figure 3 A high-level diagram showing the process of forming residual data;

[0055] Figure 4 a high-level schematic diagram showing the process of forming additional residual data at different quality levels;

[0056] Figure 5 A high-level diagram showing the process of reconstructing a frame from residual data;

[0057] Figure 6 Display hierarchical data structures;

[0058] Figure 7 Shows another hierarchical data structure;

[0059] Figure 8 Demonstrate example encoding process;

[0060] Fig. 9 Demonstrate the example decoding process;

[0061] Fig.10 showing a structure of a first data symbol;

[0062] Fig.11 showing a structure of a second data symbol;

[0063] Fig.12 Display the structure of the run symbol;

[0064] Fig.13 Show the structure of the first stream head;

[0065] Fig.14 Show the structure of the second stream head;

[0066] Fig.15 Show the structure of the third stream head;

[0067] Fig.16 Show the structure of the fourth stream head;

[0068] 17A to 17E show a Huffman tree; and

[0069] Fig.18 Shows the run-length encoding state machine. DETAILED DESCRIPTION

[0070] The present invention relates to methods. In particular, the present invention relates to methods for encoding and decoding signals. Processing data may include but is not limited to obtaining, deriving, outputting, receiving and reconstructing data.

[0071] The encoding techniques discussed herein are flexible, adaptable, efficient, and computationally inexpensive encoding formats that combine a video encoding format, a base codec (eg, AVC, HEVC, or any other current or future codec) with enhanced levels of encoded data encoded using different techniques.

[0072] The techniques use a downsampled source signal that is encoded using a base codec to form a base stream. An enhancement stream is formed using a coded set of residuals that correct or enhance the base stream, such as by increasing the resolution or by increasing the frame rate. There may be multiple levels of enhancement data in a hierarchical structure. It is noteworthy that the base stream is generally expected to be decodable by a hardware decoder, while the enhancement stream is expected to be suitable for software processing implementation with suitable power consumption.

[0073] What is needed are methods and systems for efficiently transmitting and storing enhanced coded information.

[0074] It is critical that any entropy coding operation used in the new coding technique is tailored to the specific requirements or constraints of the enhancement stream and be of low complexity. Such requirements or constraints include: possible reductions in computing power resulting from the need for software decoding of the enhancement stream; the need for combinations of decoded sets of residuals and decoded frames; the possible structure of the residual data, i.e., a relatively high proportion of zero values ​​and highly variable data values ​​over a large range; subtle differences in the input quantization blocks of coefficients; and the structure of the enhancement stream as a set of discrete residual frames separated into planes, layers, etc. Entropy coding must also be suitable for multiple enhancement levels within the enhancement stream.

[0075] Research by the inventors has established that modern entropy coding schemes used in video, such as context-based adaptive binary arithmetic coding (CABAC) or context-adaptive variable length coding (CAVLC), are unlikely to be suitable. For example, given the structure of the input data, predictive mechanisms may be unnecessary or may not yield sufficient benefit to compensate for their computational burden. Furthermore, in general, arithmetic coding is computationally expensive and software implementations are often undesirable.

[0076] It should be noted that the constraints placed on the enhancement stream mean that simple and fast entropy coding operations are essential to enable the enhancement stream to effectively correct or enhance individual frames of the base decoded video. It should be noted that in some scenarios, the base streams are also decoded substantially simultaneously before combining, which puts a strain on resources.

[0077] This current document preferably meets the requirements of the following ISO / IEC documents: "Call for Proposals for Low Complexity Video Coding Enhancements", ISO / IEC JTC1 / SC29 / WG11 N17944, Macau, China, October 2018 and "Requirements for Low Complexity Video Coding Enhancements", ISO / IEC JTC1 / SC29 / WG11 N18098, Macau, China, October 2018. In addition, the methods described herein may be incorporated into the V-Nova International Ltd. supplied In product.

[0078] The general structure of the proposed encoding scheme to which the presently described techniques may be applied uses a downsampled source signal encoded by a base codec, adds a level of correction data to the decoded output of the base codec to generate a corrected image, and then adds another level of enhancement data to an upsampled version of the corrected image.

[0079] Thus, the streams are considered to be a base stream and an enhancement stream.It is worth noting that the base stream is generally expected to be decodable by a hardware decoder, while the enhancement stream is expected to be suitable for software processing implementation with suitable power consumption.

[0080] This structure creates multiple degrees of freedom, allowing great flexibility and adaptability to many situations, making the encoding format suitable for many use cases, including Over-the-Top (OTT) delivery, live streaming, live ultra-high-definition (UHD) broadcasting, etc.

[0081] Although the decoded output of the base codec is not intended for viewing, it is a complete decoded video at a lower resolution, making the output compatible with existing decoders and also usable as a lower resolution output if deemed appropriate.

[0082] In general, residual refers to the difference between the values ​​of a reference array or reference frame and the actual array or frame of data. It should be noted that this generalized example is agnostic to the encoding operation performed and the nature of the input signal. Reference to "residual data" as used herein refers to data derived from a residual set, such as the residual set itself or the output of a data set that processes an operation performed on the residual set.

[0083] In some examples described herein, several encoded streams may be generated and sent independently to a decoder. That is, in an example method of encoding a signal, the signal may be encoded using at least two encoding levels. A first encoding level may be performed using a first encoding algorithm and a second level may be encoded using a second encoding algorithm. The method may include: obtaining a first portion of a bitstream by encoding the signal at the first level; obtaining a second portion of the bitstream by encoding the signal at the second level; and sending the first portion of the bitstream and the second portion of the bytestream as two independent bitstreams.

[0084] It should be noted that the entropy coding techniques presented herein are not limited to the multiple LoQs presented in the figures, but provide utility in any residual data used to provide enhancement or correction of video frames reconstructed from an encoded stream encoded using legacy video coding techniques such as HEVC. For example, the utility of the coding techniques to encode residual data of different LoQs is particularly beneficial.

[0085] It should be noted that the Figures 1 to 5 Examples of encoding and decoding schemes are provided in which the entropy encoding techniques provided herein may provide practicality, but it should be understood that the encoding techniques may be used in general to encode residual data.

[0086] Certain examples described herein relate to generalized encoding and decoding processes that provide hierarchical, scalable, flexible encoding techniques. A first portion of a bitstream or a first independent bitstream may be decoded using a first decoding algorithm, and a second portion of the bitstream or a second or independent bitstream may be decoded using a second decoding algorithm. The first decoding algorithm is capable of decoding using an old decoder using old hardware.

[0087] Returning to the initial process described above of providing two levels of enhancement within the base stream and the enhancement stream, Figure 1 An example of a generalized encoding process is depicted in the block diagram of . An input full-resolution video 100 is processed to generate various encoded streams 101, 102, 103. A first encoded stream (encoded base stream) is produced by feeding a downsampled version of the input video to a base codec (e.g., AVC, HEVC, or any other codec). The encoded base stream may be referred to as a base layer or base level. A second encoded stream (encoded level 1 stream) is produced by processing a residual obtained by taking the difference between a reconstructed base codec video and a downsampled version of the input video. A third encoded stream (encoded level 0 stream) is produced by processing a residual obtained by taking the difference between an upsampled version of a corrected version of the reconstructed encoded base video and the input video.

[0088] A downsampling operation may be applied to an input video to produce a downsampled video to be encoded by the base codec. Downsampling may be performed in both vertical and horizontal directions, or alternatively only in the horizontal direction.

[0089] Each enhancement stream encoding process may not necessarily include an upsampling step. Figure 1 In , the first enhancement stream is conceptually the correction stream, while the second enhancement stream is upsampled to provide an enhancement level.

[0090] Referring to the process of generating the enhancement stream in more detail, to generate the encoded level 1 stream, the encoded base stream is decoded (114) (i.e., a decoding operation is applied to the encoded base stream to generate a decoded base stream). The difference between the decoded base stream and the downsampled input video is then formed (110) (i.e., a subtraction operation is applied to the downsampled input video and the decoded base stream to generate a first set of residuals).

[0091] Here, the term residual is used in the same way as known in the art, i.e., the error between a reference frame and a desired frame. Here, the reference frame is the decoded base stream and the desired frame is the downsampled input video. Thus, the residual used in the first enhancement level can be considered as corrected video, since it 'corrects' the decoded base stream to the downsampled input video used in the base encoding operation.

[0092] Likewise, the entropy encoding operations described later are applicable to any residual data, such as any data associated with a residual set.

[0093] The differences are then encoded (115) to generate the encoded level 1 stream (102) (ie, the encoding operation is applied to the first set of residuals to generate the first enhancement stream).

[0094] As indicated above, the enhancement stream may include a first level of enhancement 102 and a second level of enhancement 103. The first level of enhancement 102 may be considered a corrected stream. The second level of enhancement 103 may be considered another level of enhancement that converts the corrected stream into the original input video.

[0095] Another level of enhancement 103 is formed by encoding another set of residuals, which are the differences (119) between the upsampled (117) version of the decoded (118) Level 1 stream and the input video 100.

[0096] As mentioned, the upsampled stream is compared to the input video, which forms another set of residuals (i.e., the difference operation is applied to the re-formed upsampled stream to generate another set of residuals). The other set of residuals is then encoded (121) as an encoded level-0 enhancement stream (i.e., the encoding operation is then applied to the other set of residuals to generate another encoded enhancement stream).

[0097] Therefore, if Figure 1 As illustrated in and described above, the output of the encoding process is a base stream 101 and one or more enhancement streams 102, 103, which preferably include a first level of enhancements and a further level of enhancements.

[0098] Figure 2 The corresponding generalized decoding process is depicted in the block diagram of . The decoder receives the three streams 101, 102, 103 generated by the encoder and a header containing additional decoding information. The encoded elementary stream is decoded by a base decoder corresponding to the base codec used in the encoder, and its output is combined with the decoded residual obtained from the encoded level 1 stream. The combined video is upsampled and further combined with the decoded residual obtained from the encoded level 0 stream.

[0099] During the decoding process, the decoder can parse the headers (global configuration, picture configuration, data block) and configure the decoder based on those headers. To re-form the input video, the decoder can decode each of the base stream, the first enhancement stream, and the other enhancement stream. The frames of the streams can be synchronized and then combined to derive the decoded video.

[0100] exist Figure 1 and 2 In each of the 0 and 1 level encoding operations, the level 0 and 1 level encoding operations may include the steps of transform, quantization and entropy encoding. Similarly, at the decoding stage, the residual may be passed through an entropy decoder, a dequantizer and an inverse transform module. Any suitable encoding and corresponding decoding operations may be used. However, preferably, the 0 and 1 level encoding steps may be performed in software.

[0101] In summary, the methods and apparatus herein are based on an overall algorithm that is built upon existing encoding and / or decoding algorithms (e.g., MPEG standards, such as AVC / H.264, HEVC / H.265, etc., and non-standard algorithms, such as VP9, ​​AV1, etc.) and used as a baseline for enhancement layers used in different encoding and / or decoding algorithms, respectively. The idea behind the overall algorithm is to encode / decode video frames in a hierarchical manner, as opposed to using a block-based approach used in the MPEG family of algorithms. Encoding frames in a hierarchical manner includes generating a residual for a full frame, and then generating a residual for a decimated frame, etc.

[0102] The video compression residual data of a full-size video frame may be referred to as LoQ-0 (e.g., 1920×1080 for an HD video frame), while the video compression residual data of a decimated frame may be referred to as LoQ-x, where x represents the number of hierarchical decimations. Figure 1 and 2 In the described example, the variable x has a maximum value of 1 and therefore there are 2 hierarchical levels that will generate compressed residuals.

[0103] Figure 3 An example of how LoQ-1 can be generated at an encoding device is illustrated. In the current diagram, the AVC / H.264 encoding / decoding algorithm is used as a baseline algorithm to describe the overall algorithm and method, but it should be understood that other encoding / decoding algorithms can be used as a baseline algorithm without any impact on how the overall algorithm works.

[0104] Of course, it will be understood. Figure 3 The blocks of FIG. 5 are merely examples of how the broad concepts may be implemented.

[0105] Figure 3 The figure shows the process of generating entropy coded residual data for the LoQ-1 hierarchy level. In the example, the first step is to decimate the incoming uncompressed video by a factor of 2. This decimated frame is then passed through the base encoding algorithm (in this case, the AVC / H.264 encoding algorithm), where an entropy coded reference frame is then generated and stored. A decoded version of the encoded reference frame is then generated, and the difference between the decoded reference frame and the decimated frame (LoQ-1 residual) will form the input to the transform block.

[0106] A transform (e.g., a Hadamard-based transform in the illustrated example here) converts this difference into 4 components (or planes), namely A (average), H (horizontal), V (vertical), and D (diagonal). These components are then quantized using a variable called a 'step width' (e.g., the residual value can be divided by the step width and the nearest integer value is selected). A suitable entropy encoding process is the subject of this disclosure and the detailed description below. These quantized residuals are then entropy encoded in order to remove any redundant information. The quantized encoded coefficients or components (Ae, He, Ve, and De) are then placed into a serial stream, with definition packets inserted at the beginning of the stream, using a file serialization routine to complete this final stage. The packet data may include such information as the specification of the encoder, the type of upsampling to be employed, whether the A and D planes are discarded, and other information that enables a decoder to decode the stream.

[0107] Both the reference data (entropy encoded half-size baseline frame) and the entropy encoded LoQ-1 residual data may be buffered, transmitted, or stored for use by the decoder during the reconstruction process.

[0108] To generate LoQ-0 residual data, the quantized output is bifurcated and subjected to inverse quantization and transform procedures in order to reconstruct the LoQ-1 residual, which is then added to the decoded reference data (encoded and decoded) in order to obtain a video frame that closely resembles the originally extracted input frame.

[0109] This process ideally mimics the decoding process and therefore does not use the initially extracted frames.

[0110] Figure 4 An example is illustrated of how LoQ-0 may be generated at an encoding device.To derive the LoQ-0 residual, a reconstructed LoQ-1 size frame is derived as described in the previous section.

[0111] The next step is to perform an upsampling of the reconstructed frame to full size (by a factor of 2). At this point, various algorithms can be used to enhance the upsampling process, such as nearest, bilinear, sharp, or cubic algorithms. This reconstructed full size frame, called the 'predicted' frame, is then subtracted from the original uncompressed video input forming some residual (LoQ-1 residual).

[0112] Similar to the LoQ-1 process, the difference is then transformed, quantized, entropy encoded, and file serialized, which then forms the third and final data. This can be buffered, transmitted, or stored for later use by the decoder. As can be seen, a component called the "predicted average" (described below) can be derived using an upsampling process and used in place of the A (average) component to further improve the efficiency of the encoding algorithm.

[0113] Figure 5 Schematically showing how the decoding process may be performed in a specific example. Entropy coded data, LOQ-1 entropy coded residual data, and LOQ-0 entropy coded residual data (e.g., as file serialized coded data). The entropy coded data comprises a downsized (e.g., half-sized, i.e., having dimensions W / 2 and H / 2 relative to a full frame having dimensions W and H) coded basis.

[0114] The entropy coded data is then decoded using a decoding algorithm corresponding to the algorithm used to encode those data (in the example, the AVC / H.264 decoding algorithm). At the end of this step, a decoded video frame with a reduced size (e.g., half size) is produced (indicated as AVC / H.264 video in this example).

[0115] In parallel, the LoQ-1 entropy coded residual data is decoded. As discussed above, the LoQ-1 residual is encoded using four coefficients or components (A, V, H and D), which, as shown in this figure, have a size of one quarter of the full frame size, i.e., W / 4 and H / 4. This is because, as discussed in previous patent applications US 13 / 893,669 and PCT / EP2013 / 059847 (the contents of which are incorporated herein by reference), the four components contain all the information about the residual and are generated by applying a 2×2 transform kernel to the residual (for LoQ-1, its size will be W / 2 and H / 2, i.e., the same size of the downsized entropy coded data). The four components are entropy decoded, then dequantized and finally transformed back to the original LoQ-1 residual by using an inverse transform (in this case, a 2×2 inverse Hadamard transform).

[0116] The decoded LoQ-1 residual is then added to the decoded video frame to produce a downsized (in this case, half-size), reconstructed video frame identified as a half-2D-size reconstruction.

[0117] This reconstructed video frame is then upsampled to bring it to full resolution (so in this example, from half width (W / 2) and half height (H / 2) to full width (W) and full height (H)) using an upsampling filter such as bilinear, bicubic, sharp, etc. The upsampled reconstructed video frame will be the predicted frame (full size, W×H), to which the LoQ-0 decoded residual is then added.

[0118] Specifically, the LoQ-0 encoded residual data is decoded. As discussed above, the LoQ-0 residual is encoded using four coefficients or components (A, V, H and D), which, as shown in this figure, have a size of half the full frame size, namely W / 2 and H / 2. This is because, as discussed in previous patent applications US 13 / 893,669 and PCT / EP2013 / 059847 (the contents of which are incorporated herein by reference), the four components contain all the information about the residual and are generated by applying a 2×2 transform kernel to the residual (for LoQ-0, their sizes will be W and H, i.e., the same size of the full frame). The four components are entropy decoded (see the following process), then dequantized and finally transformed back to the original LoQ-0 residual by using an inverse transform (in this case, a 2×2 inverse Hadamard transform).

[0119] The decoded LoQ-0 residual is then added to the predicted frame to produce a reconstructed full video frame - output frame.

[0120] exist Figure 6The data structures are represented in an exemplary manner in . As discussed, the above description has been made with reference to a specific size and baseline algorithm, but the above approach is applicable to other sizes and or baseline algorithms, and the above description is given only by way of example of the more general concepts described.

[0121] In the encoding / decoding algorithms described above, there are typically 3 planes (e.g., YUV or RGB) with two quality levels LoQ in each plane, described as LoQ-0 (or top level, full resolution) and LoQ-1 (or lower level, downsized resolution, such as half resolution). Each plane may represent a different color component of the video signal.

[0122] Each LoQ contains four components or layers, namely A, H, V, and D. This gives a total of 3×2×4=24 surfaces, of which 12 are full size (e.g., W×H) and 12 are reduced size (e.g., W / 2×H / 2). Each layer may include coefficient values ​​for a particular one of the components, for example, layer A may include the top left A coefficient value for each 2×2 block of the input image or frame.

[0123] Figure 7 An alternative view illustrating the proposed hierarchical data structure. The encoded data may be separated into blocks. Each payload may be ordered into blocks in a hierarchical manner. That is, each payload is grouped into planes, then within each plane, each level is grouped into layers and each layer comprises a set of blocks for that layer. A level represents an enhancement of each level (either the first or another level of enhancement) and a layer represents a set of transform coefficients.

[0124] The method may include retrieving information blocks for two levels of enhancement for each plane. The method may include retrieving 16 layers for each level (e.g., if a 4×4 transform is used). Thus, each payload is sorted into a set of information blocks for all layers in each level and then sorted into a set of information blocks for all layers in the next level of the plane. The payload then includes a set of information blocks for the layers of the first level of the next plane, etc.

[0125] Thus, the method may decode the header and output entropy coded coefficients grouped by plane, level, and layer belonging to the image enhancement being decoded. Thus, the output may be an (nPlanes)×(nLevel)×(nLayer) array surface with element surface[nPlanes][nLevel][nLayer].

[0126] It should be noted that the entropy encoding techniques provided herein may be performed per group, ie, per surface, per plane, per level (LoQ), or per layer.It should be noted that entropy encoding may be performed on any residual data and not necessarily on quantized and / or transformed coefficients.

[0127] like Figure 4 and 5 As shown, the proposed encoding and decoding operations include an entropy coding stage or operation. It is proposed below that the entropy coding operation includes one or more of a run length coding component (RLE) and a Huffman coding component. These two components can be correlated to provide additional benefits.

[0128] As mentioned and as Figure 8 As illustrated in , the input to the entropy encoder is a surface (e.g., residual data derived from a quantized set of residuals, as illustrated in this example) and the output of the process is an entropy encoded version of the residuals (Ae, He, Ve, De). However, as noted above, it should be noted that entropy encoding can be performed on any residual data and not necessarily on quantized and / or transformed coefficients. Fig. 9 A corresponding advanced decoder with inverse input and output is illustrated. That is, the entropy decoder takes the entropy encoded residual (Ae, He, Ve, De) as input and outputs residual data (eg, quantized residual in this illustrated example).

[0129] It should be noted that, generally speaking, Figure 8 RLE decoding is introduced before Huffman decoding, but as will be apparent throughout this application, this order is not limiting and the stages may be interchangeable, interrelated, or in any order.

[0130] Similarly, it should be noted that run-length encoding may be provided without Huffman encoding, and in comparative cases, both run-length encoding and Huffman encoding may not be performed (replaced by an alternative lossless entropy encoding or entropy encoding may not be performed at all). For example, in a production instance with reduced emphasis on data compression, the encoding pipeline may not include entropy encoding, but the benefit of the residual may be in hierarchical storage for distribution.

[0131] A run length encoder (RLE) can compress data by encoding a sequence of the same data value as a single data value and a count of the data value. For example, the sequence 555500022 can be encoded as (4,5)(3,0)(2,2). That is, there are four runs of 5, three runs of 0, followed by two runs of 2.

[0132] To encode the residual data, a modified RLE encoding operation is proposed. To exploit the structure of the residual data, it is proposed to encode only runs of zeros. That is, each value is sent as a value where each zero is sent as a run of zeros. For example, the modified RLE encoding operation described herein may be used to provide entropy encoding and / or decoding in one or more LoQ layers, such as Figures 1 to 5 as shown in .

[0133] Therefore, the entropy encoding operation consists of parsing the residual data and encoding zero values ​​as runs of consecutive zeros.

[0134] As indicated above, to provide an additional level of data compression (which may be lossless), the entropy encoding operation further applies a Huffman encoding operation to the run-length encoded data.

[0135] Huffman coding and run-length coding have been paired together previously, such as in fax coding (e.g., ITU Recommendations T.4 and T.45) and the JPEG file interchange format, but have been superseded over the past 30 years. The present description proposes techniques and methods for implementing Huffman coding and RLE coding, techniques and methods for efficiently exchanging data and metadata between encoders and decoders, and techniques and methods for reducing the overall data size of this combination when used to encode residual data.

[0136] Huffman codes are optimal prefix codes for data compression (e.g., lossless compression). A prefix code is a system of codes where no code works as a prefix of any other code. That is, the Huffman encoding operation takes a set of input symbols and converts each input symbol into a corresponding code. The codes are selected based on the frequency with which each symbol appears in the original data set. In this way, smaller codes can be associated with more frequently occurring symbols in order to reduce the overall size of the data.

[0137] To facilitate Huffman encoding, the input data is preferably structured as symbols.The output of the run-length encoding operation is preferably structured as a byte stream of encoded data.

[0138] In one example, the run-length encoding operation outputs two types of bytes or symbols. The first type of symbol is the value of a non-zero pixel and the second type of symbol is a run of consecutive zero values. That is, the number of zero values ​​that appear consecutively in the original data set.

[0139] In another example, two bytes or symbols may be combined in order to encode certain data values. That is, for example, where a pixel value is greater than a threshold, two bytes or symbols may be used to encode the data value in the symbol. The two symbols may be combined at the decoder to re-form the original data value of the pixel.

[0140] Each symbol or byte may contain 6 or 7 data bits and one or more flags or bits indicating the symbol type.

[0141] In an implementation example, run-length encoded data may be efficiently encoded by inserting one or more flags in each symbol that indicate the next symbol to be expected in the byte stream.

[0142] To facilitate synchronization between the encoder and decoder, the first byte of a stream of run-length encoded data may be a symbol of the type of data value. This byte may then indicate the next type of byte that has been encoded.

[0143] In some embodiments, the flag indicating the next byte may be different depending on the type of symbol. For example, the flag or bit may be located at the beginning or end of a byte, or both. For example, at the beginning of a byte, the bit may indicate that the next symbol may be a run of zeros. At the end of a byte, the bit may be an overflow bit, indicating that the next symbol contains a portion of a data value to be combined with the data value in the current symbol.

[0144] The following describes a specific implementation example of the run length encoding operation. RLE has three contexts: RLC_RESIDUAL_LSB, RLC_RESIDUAL_MSB and RLC_ZERO_RUN. Fig.10 , 11 The structure of these bytes is described in 1 and 12.

[0145] In the first type of symbols, Fig.10 , the 6 least significant bits of the non-zero pixels are encoded. A run bit can be provided to indicate that the next byte is encoding the count of the run of zero. If the pixel value does not fit within 6 data bits, an overflow bit can be set. When the overflow bit is set, the context of the next byte will be an RLC_RESIDUAL_MSB type. That is, the next symbol will include a data byte, which can be combined with the bit of the current symbol to encode the data value. When the overflow bit is set, the next context cannot be a run of zero and therefore the symbol can be used to encode the data.

[0146] Fig.10 Indicates an instance of this type of symbol. If the data value to be encoded is greater than the threshold of 64, or the pixel value does not fit within the 6 data bits, then the overflow bit may be set by the encoding process. If the overflow bit is set to the least significant bit of the byte, then the remaining bits of the byte may encode the data. If the pixel value fits within 6 bits, then the overflow bit may not be set and the run bit may be included, which indicates whether the next symbol is a data value or a run of zeros.

[0147] Fig.11The second type of symbol, RLC_RESIDUAL_MSB, described in , encodes bits 7 to 13 of the pixel value that does not fit within the 6 data bits. Bit 7 of this type of symbol encodes whether the next byte is a run of zeros.

[0148] Fig.12 The third type of symbol RLC_ZERO_RUN described in 7 encodes the 7 bits of the zero run count. That is, the symbol includes the number of consecutive zeros in the residual data. If more bits are needed to encode the count, the run bits of the symbol are higher. The run bits can be the most significant bits of the byte. That is, in the case where a number of consecutive zeros requires more than 7 available bits, such as 178 or 255 available bits, the run bits indicate that the next bit will indicate another run of zeros.

[0149] Optionally, the additional symbol may comprise a second run of zeros that may be combined with the first run, or may comprise a set of bits that may be combined at a decoder with the bits of the first symbol to indicate the value of the count.

[0150] As indicated above, in order for the decoder to start based on a known context, the first symbol in the encoded stream may be a symbol of a residual data value type, ie may be of the RLC_RESIDUAL_LSB type.

[0151] In a specific example, RLE data may be organized into blocks. Each block may have an output capacity of 4096 bytes. RLE may switch to a new block in the following situations:

[0152] - the current block is full;

[0153] - The current RLE data is a run and there are less than 5 bytes remaining in the current block; and / or

[0154] - The current RLE data produces an LSB / MSB pair and there are less than 2 bytes remaining in the current block.

[0155] In summary, a run-length encoding operation can be performed on the residual data of the new encoding technique, which includes encoding a set of data values ​​and a count of consecutive zero values ​​into a stream. In a specific implementation, the output of the run-length encoding operation can be a stream of bytes or symbols, where each byte is one of three types or contexts. The byte indicates the next type of byte expected in the byte stream.

[0156] It should be noted that these structures are provided as examples, and different bit encoding schemes may be applied while following the functional teachings described herein. For example, the encoding of the least and most significant bits may be swapped and / or different bit lengths may be formulated. Also, the flag bit may be located at different predefined positions within a byte.

[0157] As indicated above, a Huffman encoding operation can be applied to a byte stream to further reduce the size of the data. The process can form a frequency table of the relative occurrences of each symbol in the stream. From the frequency table, the process can generate a Huffman tree. A Huffman code for each symbol can be generated by traversing the tree from the root until the symbol is reached, assigning code bits to each branch obtained.

[0158] In a preferred implementation for the particularity of the run-length encoded symbols for the quantized residual (i.e., residual data), a canonical Huffman coding operation may be applied. Canonical Huffman codes may reduce the storage requirements of a set of codes depending on the structure of the input data. The canonical Huffman procedure provides a way to generate codes that implicitly contain information about which symbol a codeword applies to. The Huffman codes may be set under a set of conditions, for example, all codes of a given length may have lexicographically consecutive values ​​in the same order as the symbols they represent and shorter codes lexicographically precede longer codes. In a canonical Huffman coding implementation, only the codeword and the length of each code need to be transmitted for the decoder to be able to copy each symbol.

[0159] After applying the canonical Huffman coding operation, the data size of the output data can be compared with the data size after the run-length coding operation. If the data size is smaller, then the smaller data can be transmitted or stored. The comparison of the data size can be performed based on the size of the data block, the size of the layer, plane or surface, or the overall size of the frame or video. In order to signal how the data is encoded to the decoder, a flag that the data is encoded using Huffman coding, RLE coding, or both can be transmitted in the header of the data. In an alternative implementation, the decoder may be able to identify that the data has been encoded using a specific coding operation based on the characteristics of the data. For example, as indicated below, the data can be split into a header and a data portion, wherein the canonical Huffman coding is used to signal the code length of the encoded symbol.

[0160] The normalized Huffman coded data includes a portion that signals the code length for each symbol to the decoder and a portion that includes a bit stream representing the coded data. In the specific implementation presented herein, the code length may be signaled in a header portion and the data in a data portion. Preferably, the header portion will signal for each surface but may signal for each block or another subdivision depending on the configuration.

[0161] In the proposed example, the header may be different depending on the amount of non-zero codes to be encoded. The amount of non-zero codes to be encoded (for example, each surface or block) may be compared to a threshold and the header used based on the comparison. For example, when there are more than 31 non-zero codes, the header may indicate all symbols in sequence, starting with a predetermined signal such as 0. In sequence, the length of each symbol may be signaled. When there are less than 31 non-zero codes, each symbol value and the corresponding code length of the symbol may be signaled in the header.

[0162] Figures 13 to 16 The description header format and code length may depend on the specific implementation of the amount of non-zero codes written to the stream header.

[0163] Fig.13 Explains the situation where there can be more than 31 non-zero values ​​in the data. The header contains the minimum and maximum code lengths. The code length of each symbol is sent in sequence. A flag indicates that the length of the symbol is non-zero. The code length bits are then sent as the difference between the code length and the minimum signaled length. This reduces the overall size of the header.

[0164] Fig.14 Description similar to Fig.13 But a header used when there are less than 31 non-zero codes. The header further contains the number of symbols in the data, followed by the symbol value and the length of the codeword for that symbol, also sent as a difference.

[0165] Fig.15 and 16 Indicates additional headers to be sent in peripheral situations. For example, when the frequency is all zeros, the stream header may be Fig.14 The two values ​​31 in the minimum and maximum length fields indicate a special case. When there is only one code in the Huffman tree, the stream header can be as follows Fig.16 As indicated in , a value of 0 in the minimum and maximum length fields is used to indicate a special case and the symbol value is used next.

[0166] In this latter example, when only one symbol value is present, this may indicate that only one data value is present in the residual data.

[0167] The encoding process can thus be summarized as follows −

[0168] Parse the residual data to identify consecutive counts of data values ​​and zero values.

[0169] A set of symbols is generated, where each symbol is a byte and each byte includes a data value or a run of zeros and an indication of the next symbol in the set of symbols. The symbol may include an overflow bit indicating that the next symbol includes a portion of the data value included in the current symbol or a run bit indicating that the next symbol is a run of zeros.

[0170] Each fixed-length symbol is converted into a variable code using Canonical Huffman Coding. Canonical Huffman Coding parses the symbols to identify the frequency with which each symbol appears in the set of symbols and assigns a code to the symbol based on the frequency and the symbol value (for the same code length).

[0171] A set of code lengths is generated and output, each code length being associated with a symbol. The code length is the length of each code assigned to each symbol.

[0172] Combine variable length codes into a bit stream.

[0173] Outputs an encoded bitstream.

[0174] In the particular implementation described, it should be noted that the maximum code length depends on the number of encoded symbols and the number of samples used to derive the Huffman frequency table. For N symbols, the maximum code length is N-1. However, in order for the symbol to have a k-bit Huffman code, the number of samples used to derive the frequency table needs to be at least:

[0175]

[0176] Among them, F i is the i-th Fibonacci number.

[0177] For example, for 8-bit symbols, the theoretical maximum code length is 255. However, for HD 1080p video, the maximum number of samples used to derive the frequency table for the RLE context of the residual is 1920*1080 / 4=518 400, which is between S26 and S27. Therefore, for this example, a symbol cannot have a Huffman code larger than 26 bits. For 4K, this number increases to 29 bits.

[0178] For completeness, using Figure 17, an example of Huffman encoding of a set of symbols is provided, in this example the symbols are alphanumeric. As shown above, Huffman codes are optimal prefix codes that can be used for lossless data compression. A prefix code is a system of codes in which there is no codeword that is a prefix of any other code.

[0179] In order to find the Huffman code for a given set of symbols, a Huffman tree needs to be formed. First, the symbols are classified by frequency, for example:

[0180] symbol frequency A 3 B 8 C 10 D 15 E 20 F 43

[0181] The two lowest elements are removed from the list and become leaf elements where the frequency of the parent node is the sum of the frequencies of the two lower elements. Fig.17a A partial tree is shown in FIG.

[0182] The new classified frequency list is:

[0183] symbol frequency C 10 * 11 D 15 E 20 F 43

[0184] Then, repeat the loop, such as Fig.17b Combine the two lowest elements as described in .

[0185] The new list is:

[0186] symbol frequency D 15 E 20 * 21 F 43

[0187] Repeat this process until Fig.17c , 17d , the list described in 17e and the corresponding table below only holds one element.

[0188] symbol frequency * 21 * 35 F 43

[0189] symbol frequency F 43 * 56

[0190] symbol frequency * 99

[0191] Once the tree is built, to generate the Huffman code for a symbol, these three traverse from the root to this symbol, outputting 0 each time the left branch is taken and 1 each time the right branch is taken. In the above example, this gives the following code:

[0192] symbol Code Code length A 1010 3 B 1011 3 C 100 2 D 110 2 E 111 2 F 0 0

[0193] The code length of a symbol is the length of its corresponding code.

[0194] To decode a Huffman code, the tree is traversed starting at the root, taking the left path when a 0 is read and the right path when a 1 is read. When a leaf is hit, the symbol is found.

[0195] As described above, the RLE symbols may be encoded using Huffman coding. In a preferred embodiment, canonical Huffman coding is used. In an implementation of canonical Huffman coding, a Huffman coding procedure may be used and the resulting code may be transformed into a canonical Huffman code. That is, the table is rearranged so that symbols of the same length are in the same order as the symbols they represent and the value of the subsequent code will always be higher than the previous code. When the symbols are letters, they may be rearranged in alphabetical order. When the symbols are values, they may be in sequential value order and the corresponding codewords may be changed to conform to the above constraints. Canonical Huffman coding and / or other Huffman coding methods described herein may differ from Figures 17a to 17e Basic example of .

[0196] In the examples described above, the Huffman encoding operation may be applied to data symbols encoded using a modified run length encoding operation. It is further proposed to form a frequency table for each RLE context or state. For each type of symbol encoded by the RLE encoding operation, there may be a different set of Huffman codes. In one implementation, there may be a different Huffman encoder for each RLE context or state.

[0197] It is contemplated that one stream header may be sent for each RLE context or symbol type, ie, multiple stream headers may indicate the code length for each code set, ie, each Huffman encoder or each frequency table.

[0198] An example of a specific implementation step of the encoder is described below. In one implementation, encoding of data output by RLE is performed in three main steps:

[0199] 1. Initialize the encoder (for example, through the InitialiseEncode function):

[0200] a. Initialize each encoder (initialize one encoder per RLE state), where the corresponding frequency table is generated by RLE;

[0201] b. Form a Huffman tree;

[0202] c. Calculate the code length of each symbol; and

[0203] d. Determine the minimum and maximum code lengths.

[0204] 2. Write the code length and code table to the stream header for each encoder (e.g., via the WriteCodeLengths function):

[0205] a. Assign codes to symbols (e.g., using the AssignCodes function). It should be noted that the function used to assign codes to symbols is slightly different from the examples of Figures 19a to 19e above, as it ensures that all codes of a given length are sequential, i.e., using canonical Huffman coding;

[0206] b. Write the minimum and maximum code lengths to the stream header; and

[0207] c. Write the code length of each symbol to the stream.

[0208] 3. Encode the RLE data:

[0209] a. Set the RLE current context to RLC_RESIDUAL_LSB;

[0210] b. Encode the current symbol in the input stream through the encoder corresponding to the current context;

[0211] c. Use the RLE state machine to get the next context (see Figure 20). If this is not the last symbol of the stream, go to b; and

[0212] d. If the RLE data is only one block and the encoded stream is larger than RLE, then store the RLE encoded stream instead of the Huffman encoded stream.

[0213] For each RLE block, the Huffman encoder forms a corresponding Huffman-encoded portion.

[0214] The output of the encoding process described herein is therefore a Huffman encoded bit stream or a run-length encoded byte stream. The corresponding decoder therefore employs a corresponding Huffman decoding operation in a manner opposite to the encoding operation. Nevertheless, there are slight differences from the decoding operation described below to improve efficiency and accommodate the specificity of the encoding operation described above.

[0215] When the encoder signals in the header or other configuration metadata that RLE or Huffman encoding operations have been used, the decoder can first identify that Huffman decoding operations should be applied to the stream. Similarly, in the manner described above, the decoder can identify whether Huffman encoding has been used based on the data stream by identifying the number of portions of the stream. If there are header portions and data portions, then Huffman decoding operations should be applied. If there is only a data portion for the stream, then only the RLE decoding portion can be applied.

[0216] The entropy decoder performs the inverse transform of the decoder. In this example, Huffman decoding is followed by run-length decoding.

[0217] The following example describes the decoding process, where in the example of an encoded stream, the bitstream has been encoded using the canonical Huffman encoding operation and the run-length encoding operation as described above. However, it should be understood that the Huffman encoding operation and the run-length encoding operation can be applied separately to decode streams encoded in different ways.

[0218] The input to the process is an encoded bitstream. The bitstream can be retrieved from a storage file or streamed locally or remotely. The encoded bitstream is a series of sequential bits with no immediately discernible structure. Only by applying the Huffman encoding operation to the bitstream can a series of symbols be derived.

[0219] Therefore, the process first applies a Huffman encoding operation to the bitstream.

[0220] Assume that the first encoded symbol is a symbol of a data value type. In this way, when the Huffman encoding operation uses a different frequency table or parameter, the Huffman encoding operation can use the correct codebook for the symbol type.

[0221] A canonical Huffman encoding operation must first retrieve the appropriate encoding parameters or metadata to facilitate decoding the bitstream. In the example here, the parameters are a set of code lengths retrieved from the stream header. For simplicity, one set of code lengths will be described as being exchanged between the encoder and decoder, but it should be noted that, as described above, multiple sets of code lengths may be signaled so that the encoding operation can use different parameters for each type of symbol it wishes to decode.

[0222] The encoding operation assigns the code lengths received in the stream header to the corresponding symbols. As described above, the stream header can be sent in different types. When the stream header contains a set of code lengths and corresponding symbols, each code length can be associated with a corresponding symbol value. When the code lengths are sent sequentially, the decoder can associate each received code length with a symbol in a predetermined order. For example, the code lengths can be retrieved in the order 3, 4, 6, 2, etc. The encoding process can then associate each length with a corresponding symbol. Here, the order is sequential - (symbol, length) - (0,3) (1,4) (3,6) (4,2).

[0223] like Fig.15 As illustrated in , some symbols may not be sent but the flag may indicate that the symbols have a corresponding zero code length, ie, the symbols are not present in the bitstream to be decoded (or are not used).

[0224] The canonical Huffman encoding operation can then assign a code to each symbol based on the code length. For example, when the code length is 2, the code for that symbol would be '1x'. Where x indicates that the value is not significant. In a specific implementation, the code would be '10'.

[0225] When the lengths are the same, the canonical Huffman encoding operation assigns codes based on the sequential order of the symbols. Each code is assigned so that when parsing the bit stream, the decoder can keep checking the next bit until there is only one possible code. For example, the first sequential symbol of length 4 may be assigned code 1110 and the second sequential symbol of length 4 may be assigned code 1111. Therefore, if the first bit of the bit stream is 1110x, then parsing the bit stream, the bit stream may only have encoded the first sequential symbol of length 4.

[0226] Therefore, once the encoding operation has established the code set associated with each symbol, the encoding operation can move to parsing the bit stream. In some embodiments, the encoding operation will build a tree to establish the symbols associated with the codes of the bit stream. That is, the encoding operation will obtain each bit of the bit stream and traverse the tree according to the value of the bit in the bit stream until a leaf is found. Once a leaf is found, the symbol associated with the leaf is output. The process continues the next bit of the bit stream at the root of the tree. In this way, the encoding operation can output a set of symbols derived from the code set stored in the encoded bit stream. In a preferred embodiment, the symbols are each byte, as described elsewhere herein.

[0227] The symbol set may then be passed to a run-length encoding operation. The run-length encoding operation identifies the type of symbol and parses the symbol to extract the data value or the run of zeros. From the data value and the run of zeros, the operation may re-form the encoded original residual data.

[0228] As shown above, the output of Huffman coding is likely to be a set of symbols. The challenge of the next decoding step is to extract the relevant information from those symbols, noting that each symbol includes data in a different format and each symbol will represent different information, even if that information is not immediately discernible from the symbol or byte.

[0229] In a preferred embodiment, the encoding operation consultation is as follows Fig.18 In summary, the encoding operation first assumes that the first symbol is of type data value (the data value can of course be 0). From this, the encoding operation can identify whether the byte contains an overflow flag or a run flag. The overflow and run flags are described above and described in Fig.10 middle.

[0230] The encoding operation may first check the least significant bit of the symbol (or byte). Here, this is the overflow flag (or bit). If the least significant bit is set, the encoding operation recognizes that the next symbol will also be a data value and is of type RLC_RESIDUAL_MSB. The remaining bits of the symbol will be the least significant bits of the data value.

[0231] The encoding process proceeds to parse the next symbol and extract the least significant bits as the remainder of the data value. The bits of the first symbol can be combined with the bits of the second symbol to reconstruct the data value.

[0232] exist Fig.11 In this current symbol, there will also be a run flag, which is the most significant bit here. If this flag is set, the next symbol will be a run symbol instead of a data symbol.

[0233] exist Fig.12 In the run symbol illustrated in , the encoding operation checks the most significant bit to indicate whether the next symbol is a run symbol or a data symbol and extracts the remaining bits as a run of zeros. That is, the maximum run is 127 zeros.

[0234] If the data symbol indicates that there is no overflow (in the overflow flag, here the least significant bit), then the symbol will also contain a run bit in the most significant bit and the encoding operation will be able to identify from this bit whether the next symbol in the set is a run symbol or a data symbol. The encoding operation can therefore extract the data value from bits 6 to 1. Here, there are 31 available data values ​​without overflow.

[0235] As mentioned, Fig.18The state machine only illustrates the process.

[0236] Starting at the RLC_RESIDUAL_LSB symbol, if the overflow bit = 0 and the run bit = 0, then the next symbol will be the RLC_RESIDUAL_LSB symbol. If the overflow bit = 1, then the next symbol will be the RLC_RESIDUAL_MSB. If the overflow bit = 1, then there will be no run bit. If the run bit = 1 and the overflow bit = 0, then the next bit will be the RLC_ZERO_RUN symbol.

[0237] In RLC RLC_RESIDUAL_MSB, if the run bit = 0, then the next symbol will be a RLC_RESIDUAL_LSB symbol. If the run bit = 1, then the next symbol will be a RLC_ZERO_RUN symbol.

[0238] In the RLC_ZERO_RUN symbol, if the run bit = 0, then the next bit will be the RLC_RESIDUAL_LSB symbol. If the run bit = 1, then the next bit will be the RLC_ZERO_RUN symbol.

[0239] The bits can of course be inverted (0 / 1, 1 / 0, etc.) without loss of functionality. Similarly, the sign or position within a byte of a flag is illustrative only.

[0240] The run-length encoding operation may identify the next symbol in the set of symbols and extract the run of data values ​​or zeros. The encoding operation may then combine these values ​​with zeros to re-form the residual data. The order may be in the order of extraction or alternatively in some predetermined order.

[0241] The encoding operation may therefore output residual data that has been encoded into a byte stream.

[0242] The above describes how multiple metadata or encoding parameters can be used in the encoding process. At the decoding side, the encoding operation can include feedback between operations. That is, the expected next symbol can be extracted from the current symbol (e.g., using overflow and run bits) and the expected next symbol is given to the Huffman encoding operation.

[0243] The Huffman encoding operation may optionally generate a separate codebook for each type of symbol (and retrieve multiple stream headers and build a table of multiple codes and symbols).

[0244] Assuming that the first symbol will be a data value, the Huffman encoding operation will decode the first code to match a symbol of that type. Based on the matched symbol, the encoding operation can identify overflow bits and / or run bits and identify the next type of symbol. The Huffman encoding operation can then use the corresponding codebook for the type of symbol to extract the next code in the bitstream. The process uses the current decoded symbol in this iterative manner to extract an indication of the next symbol and changes the codebook accordingly.

[0245] It has been described above how a decoding operation can combine a Huffman encoding operation, preferably a canonical Huffman encoding operation, with a run-length encoding operation to re-form residual data from an encoded bit stream. It should be understood that the technique of the run-length encoding operation can be applied to a byte stream that has not been encoded using a Huffman encoding operation. Similarly, a Huffman encoding operation can be applied without the latter step of the run-length encoding operation.

[0246] An example of a specific implementation step of the decoder is described below. In an example implementation, if the stream has more than one part, the following steps are performed:

[0247] 1. Read the code length from the stream header (for example, using the ReadCodeLengths function):

[0248] a. Set the code length for each symbol;

[0249] b. Assign codes from code lengths to symbols (e.g., using the AssignCodes function); and

[0250] c. Generate a table for searching subsets of codes with the same length. Each element of the table records the first index and the corresponding code (firstIdx, firstCode) of a given length.

[0251] 2. Decode the RLE data:

[0252] a. Set the RLE context to RLC_RESIDUAL_LSB.

[0253] b. Decode the current code, which searches for the correct code length in the generated table and indexes into the code array by: firstIdx - (current_code - firstCode). This is because all codes of a given length are sequential by the construction of the Huffman tree.

[0254] c. Use the RLE state machine to obtain the next context (see Fig.18 ). If the stream is not empty, go to b.

[0255] The run length decoder reads the run length encoded data byte by byte. By construction, the context of the first byte of the data is guaranteed to be RLC_RESIDUAL_LSB. The decoder uses Fig.18 The state machine shown in determines the context of the next byte of data. The context tells the decoder how to interpret the current byte of data as described above.

[0256] It should be noted that in this particular implementation example, the run-length state machine is also used by the Huffman encoding and decoding process to know which Huffman code to use for the current symbol or codeword.

[0257] The methods and processes described herein may be embodied as code (e.g., software code) and / or data at both encoders and decoders implemented, for example, in a streaming media server or client device or a client device decoded from a data storage device. The encoders and decoders may be implemented in hardware or software, as is well known in the art to which data compression belongs. For example, hardware acceleration using a specially programmed graphics processing unit (GPU) or a specially designed field programmable gate array (FPGA) may provide certain efficiencies. For completeness, such code and data may be stored on one or more computer-readable media, which may include any device or medium that can store code and / or data for use by a computer system. When a computer system reads and executes the code and / or data stored on a computer-readable medium, the computer system executes the methods and processes embodied as data structures and codes stored in a computer-readable storage medium. In certain embodiments, one or more of the steps of the methods and processes described herein may be executed by a processor (e.g., a processor of a computer system or a data storage system).

[0258] In general, any of the functions described in this text or illustrated in the figures may be implemented using software, firmware (e.g., fixed logic circuitry), programmable or non-programmable hardware, or a combination of these implementations. In general, the terms "component" or "function" as used herein refer to software, firmware, hardware, or a combination of these. For example, in the case of a software implementation, the term "component" or "function" may refer to program code that performs a specified task when executed on one or more processing devices. The illustrated separation of components and functions into distinct units may reflect any actual or conceptual physical grouping and allocation of such software and / or hardware and tasks.

[0259] In the present application, methods for encoding and decoding signals, in particular video signals and / or image signals, are described.

[0260] In particular, a method of encoding a signal is described, the method comprising receiving an input frame and processing the input frame to generate at least one first set of residual data enabling a decoder to reconstruct an original frame from a reconstructed reference frame.

[0261] In one embodiment, the method comprises obtaining a reconstructed frame from a decoded frame obtained from a decoding module, wherein the decoding module is configured to generate the decoded frame by decoding a first encoded frame that has been encoded according to a first encoding method. The method further comprises downsampling the input frame to obtain a downsampled frame, and passing the downsampled frame to an encoding module configured to encode the downsampled frame according to the first encoding method so as to generate the first encoded frame. Obtaining the reconstructed frame may further comprise upsampling the decoded frame to generate the reconstructed frame.

[0262] In another embodiment, the method includes obtaining a reconstructed frame from a combination of a second set of residual data and a decoded frame obtained from a decoding module, wherein the decoding module is configured to generate the decoded frame by decoding a first encoded frame that has been encoded according to a first encoding method. The method further includes downsampling the input frame to obtain a downsampled frame, and passing the downsampled frame to an encoding module configured to encode the downsampled frame according to the first encoding method so as to generate the first encoded frame. The method further includes generating the second set of residual data by taking the difference between the decoded frame and the downsampled frame. The method further includes encoding the second set of residual data to generate a first set of encoded residual data. Encoding the second set of residual data may be performed according to a second encoding method. The second encoding method includes transforming the second set of residual data into a transformed second set of residual data. Transforming the second set of residual data includes selecting a subset of the second set of residual data and applying a transform to the subset to generate a corresponding subset of the transformed second set of residual data. One of the subsets of the transformed second set of residual data may be obtained by averaging the subsets of the second set of residual data.Obtaining the reconstructed frame may further include up-sampling a combination of the second set of sampled residual data and the decoded frame to generate the reconstructed frame.

[0263] In one embodiment, generating at least one set of residual data comprises taking the difference between a reconstructed reference frame and an input frame. The method further comprises encoding the first set of residual data to generate a first set of encoded residual data. Encoding the first set of residual data may be performed according to a third encoding method. The third encoding method comprises transforming the first set of residual data into a transformed first set of residual data. Transforming the first set of residual data comprises selecting a subset of the first set of residual data and applying a transform to the subset to generate a corresponding subset of the transformed first set of residual data. One of the subsets of the transformed first set of residual data may be obtained by the difference between the average of the subset of the input frame and the corresponding element of the combination of the second set of residual data and the decoded frame.

[0264] In particular, a method of decoding a signal is described, the method comprising receiving an encoded frame and at least one set of encoded residual data. A first encoded frame may be encoded using a first encoding method.

[0265] At least one set of residual data may be encoded using the second and / or third encoding method.

[0266] The method further comprises passing the first encoded frame to a decoding module, wherein the decoding module is configured to generate a decoded frame by decoding an encoded frame that has been encoded according to the first encoding method.

[0267] The method may further comprise decoding at least one set of encoded residual data according to a respective encoding method used to encode it.

[0268] In one embodiment, the first set of encoded residual data is decoded by applying a second decoding method corresponding to the second encoding method to obtain a first set of decoded residual data. The method further comprises combining the first set of residual data with the decoded frame to obtain a combined frame. The method further comprises upsampling the combined frame to obtain a decoded reference frame.

[0269] The method further comprises decoding a second set of encoded residual data by applying a third decoding method corresponding to the third encoding method to obtain a second set of decoded residual data. The method further comprises combining the second set of decoded residual data with a decoded reference frame to obtain a reconstructed frame.

[0270] In another embodiment, the method includes upsampling a decoded frame to obtain a decoded reference frame.

[0271] The method further comprises decoding the set of encoded residual data by applying a second or third decoding method corresponding to the second or third encoding method to obtain a set of decoded residual data. The method further comprises combining the set of decoded residual data with a decoded reference frame to obtain a reconstructed frame.

[0272] The following statements describe preferred or exemplary aspects described and illustrated herein.

[0273] A method of encoding an input video into a plurality of encoded streams such that the encoded streams can be combined to reconstruct the input video, the method comprising:

[0274] Receive full-resolution input video;

[0275] downsampling a full resolution input video to form a downsampled video;

[0276] encoding the downsampled video using a first codec to form an encoded elementary stream;

[0277] reconstructing a video from the encoded video to generate a reconstructed video;

[0278] comparing the reconstructed video to the input video; and

[0279] One or more further encoded streams are formed based on the comparison.

[0280] The input video may be a downsampled video compared to the reconstructed video.

[0281] According to an example method, comparing the reconstructed video to the input video includes:

[0282] The reconstructed video is compared to the downsampled video to form a first set of residuals, and wherein forming the one or more further encoded streams comprises encoding the first set of residuals to form a primary encoded stream.

[0283] The input video may be a full resolution input video and the reconstructed video may be upsampled compared to the reconstructed video.

[0284] According to an example method, comparing the reconstructed video to the input video includes:

[0285] upsampling the reconstructed video to generate an upsampled reconstructed video; and

[0286] The upsampled reconstructed video is compared to the full resolution input video to form a second set of residuals, and wherein forming the one or more further encoded streams comprises encoding the second differences to form a second level encoded stream.

[0287] Thus, in an example, the method may generate an encoded base stream, a primary encoded stream and a secondary encoded stream according to the example method defined above. Each of the primary encoded stream and the secondary encoded stream may contain enhancement data used by the decoder to enhance the encoded base stream.

[0288] The residual may be the difference between two videos or frames.

[0289] The encoded stream may be accompanied by one or more headers, which include parameters indicating various aspects of the encoding process to facilitate decoding. For example, the header may include the codec used, the transform applied, the quantization applied, and / or other decoding parameters.

[0290] The above describes how the steps of sorting and selecting can be applied to the residual data, how the step of subtracting the temporal coefficients can be performed, and the quantization can be adapted. Each of these steps can be predetermined and selectively applied or can be applied based on analysis of the input video, downsampled video, reconstructed video, upsampled video, or any combination of the above videos to improve the overall performance of the encoder. The steps can be selectively applied based on a predetermined set of rules or deterministically applied based on analysis or feedback of performance.

[0291] An example method further comprising:

[0292] The encoded elementary stream is sent.

[0293] An example method further comprising:

[0294] Send the first level coded stream.

[0295] An example method further comprising:

[0296] Send secondary encoded stream.

[0297] According to another aspect of the present disclosure, a decoding method is provided.

[0298] A method of decoding a plurality of encoded streams into a reconstructed output video, the method comprising:

[0299] receiving a first coded elementary stream;

[0300] decoding the first encoded elementary stream according to the first codec to generate a first output video;

[0301] receiving one or more additional encoded streams;

[0302] decoding the one or more additional encoded streams to generate a residual set; and

[0303] The residual set is combined with the first video to generate a decoded video.

[0304] In an example, the method includes retrieving a plurality of decoding parameters from a header. The decoding parameters may indicate which procedural steps are included in an encoding process.

[0305] In an example, the method may include receiving a primary coded stream and receiving a secondary coded stream. In this example, the step of decoding the one or more additional coded streams to generate a residual set includes:

[0306] Decoding the primary coded stream to derive a first set of residuals;

[0307] The step of combining the residual set with the first video to generate a decoded video comprises:

[0308] combining the first set of residuals with the first output video to generate a second output video;

[0309] upsampling the second output video to generate an upsampled second output video;

[0310] decoding the secondary coded stream to derive a second set of residuals; and

[0311] The second set of residuals is combined with the second output video to generate a reconstructed output video.

[0312] The method may further include displaying or outputting the reconstructed output.

Claims

1. A method for encoding a video signal, the method comprising: Receive input frame (100); processing the input frame (100) to generate residual data (110, 119) which forms part of an enhancement stream and enables a decoder to reconstruct the input frame from a reconstructed reference frame; applying a run length encoding operation to the residual data (110, 119), wherein the run-length encoding operation comprises generating a run-length encoded byte stream, the run-length encoded byte stream comprising a set of symbols representing non-zero data values ​​of the residual data and a count of consecutive zero values ​​of the residual data; and Wherein the method is characterized by the following steps: applying a canonical Huffman encoding operation to the set of symbols to generate Huffman-encoded data comprising a set of codes representing the run-length-encoded byte stream; comparing a data size of at least a portion of the run-length encoded byte stream with a data size of at least a portion of the Huffman encoded data; and Based on which has a smaller data size, either the run-length encoded byte stream or the Huffman encoded data is output in an output bitstream.

2. The method of claim 1, wherein the run-length encoding operation comprises: encoding non-zero data values ​​of the residual data into at least a first type of symbols; and The count of consecutive zero values ​​is encoded into symbols of a second type such that the residual data is encoded as a sequence of symbols of a different type.

3. The method according to claim 1 or 2, wherein the run-length encoding operation comprises: encoding a data value of the residual data into a symbol of a first type and a symbol of a third type, the first type of symbol and the third type of symbol each comprising a portion of the data value such that the portions can be combined at a decoder to reconstruct the data value; The encoding of the data value further comprises: Comparing the magnitude of each data value of the residual data to be encoded with a threshold value; encoding each data value into a symbol of the first type when the magnitude is below the threshold and encoding a portion of each data value into a symbol of the first type and a portion of each data value into a symbol of the third type when the magnitude is above the threshold; and If the size is above the threshold, a flag is set in the symbol of the first type of symbols indicating that a portion of the represented data value is encoded into another symbol of the third type of symbols.

4. The method according to claim 2, further comprising: A flag is inserted in each symbol indicating the type of symbol encoded next in the run-length encoded byte stream.

5. The method according to claim 1 or 2, further comprising: A flag indicating whether the bitstream represents the run-length encoded byte stream or the Huffman encoded data is added to configuration metadata accompanying the output bitstream.

6. The method of claim 2, wherein the Huffman encoding operation comprises: Separate frequency tables are generated for symbols of the first type of symbols and symbols of the second type of symbols.

7. The method according to claim 1 or 2, wherein the Huffman coded data comprises a code length for each unique symbol of the symbol set, the code length representing the length of the code used to encode the corresponding symbol, And wherein the method further comprises: Generate a stream header including an indication of the plurality of code lengths so that a decoder can derive the code lengths and corresponding corresponding symbols for a canonical Huffman decoding operation, wherein the stream header is one of: a stream header of a first type, the first type of stream header comprising an indication of the symbol associated with a respective code length of the plurality of code lengths; or a second type of stream header, wherein for the second type of stream header, the method further comprises ordering the plurality of code lengths in the stream header based on a predetermined order of symbols corresponding to each of the code lengths so that the code lengths can be associated with corresponding symbols at the decoder; wherein the second type of stream header further comprises a flag indicating that the predetermined order of symbols in a possible set of symbols is not present in the set of symbols in the run-length encoded byte stream; Wherein the method further comprises: The number of unique codes in the Huffman encoded data is compared to a threshold value and the first type of stream header or the second type of stream header is generated based on the comparison.

8. A method for decoding a video signal, the method comprising: retrieving an encoded bitstream; decoding the encoded bitstream to generate residual data for an enhancement stream, and reconstructing an original frame of the video signal from the residual data and the reconstructed reference frame, The step of decoding the encoded bit stream comprises: applying a run-length encoding operation to generate said residual data, The run length encoding operation comprises: identifying a set of symbols representing non-zero data values ​​of the residual data and a set of symbols representing counts of consecutive zero values ​​of the residual data; parsing the set of symbols to derive the non-zero data values ​​and the count of consecutive zero values ​​of the residual data; and generating said residual data from said counts of said non-zero data values ​​and consecutive zero values; The method is characterized in that the step of decoding the encoded bit stream comprises: A canonical Huffman coding operation is selectively applied to the encoded bitstream based on the metadata to generate a set of symbols representing non-zero data values ​​of residual data and a set of symbols representing counts of consecutive zero values ​​of the residual data, wherein a step of applying a run-length encoding operation to the set of symbols to generate the residual data is performed.

9. The method of claim 8, wherein the run-length encoding operation comprises: identifying a sign of a first type of sign representing a non-zero data value of the residual data; identifying a symbol of a second type of symbol representing a count of consecutive zero values ​​such that a sequence of symbols of different types are decoded to generate the residual data; as well as The set of symbols is parsed according to each symbol's corresponding type.

10. The method according to claim 8 or 9, wherein the run-length encoding operation comprises: identifying a symbol of a first type of symbol and a symbol of a third type of symbol, the first type of symbol and the third type of symbol each representing a portion of a data value; parsing symbols of the first type of symbols and symbols of the third type of symbols to derive portions of a data value; as well as The derived portions of the data value are combined into a data value.

11. The method according to claim 8 or 9, further comprising: A flag is retrieved from each symbol indicating the subsequent type of symbol expected in the set of symbols.

12. The method according to claim 8 or 9, further comprising: retrieving from configuration metadata accompanying the encoded bitstream a flag indicating whether the encoded bitstream comprises a run-length encoded byte stream or Huffman encoded data; and The Huffman encoding operation is selectively applied to the decoded bitstream based on the flag.

13. The method according to claim 8 or 9, further comprising: retrieving a stream header including an indication of a plurality of code lengths to be used for said canonical Huffman encoding operation; Associating each code length with a corresponding symbol according to the canonical Huffman encoding operation; as well as A Huffman encoding operation according to the specification identifies a code associated with each symbol based on the code length.

14. The method of claim 13, wherein the stream header includes an indication of symbols associated with respective ones of the plurality of code lengths and associating each code length with a symbol includes associating each code length with a corresponding symbol in the stream header.

15. The method of claim 13, wherein the step of associating each code length with a corresponding symbol comprises associating each code length with a symbol in a predetermined set of symbols according to an order in which each code length is retrieved.

16. The method of claim 15, further comprising not associating a symbol in the predetermined set of symbols with a corresponding code length, wherein a flag of the stream header indicates that a code length is not present in the stream header for the symbol.

17. An encoding device comprising a processor configured to perform the method according to any one of claims 1 to 7.

18. A decoding device comprising a processor configured to perform the method according to any one of claims 8 to 16.

19. A computer readable medium comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Decomposition of residual data during signal encoding, decoding and reconstruction in a tiered hierarchy

    US9509990B2

  • Hybrid backward-compatible signal encoding and decoding

    WO2014170819A1

  • Video compression using differences between a higher and a lower layer

    WO2018046940A1

  • Methods and Systems for Low-Complexity Data Compression

    US20080101464A1