Exchanging information in hierarchical video coding
By exchanging information in a hierarchical video coding scheme and coordinating coding operations to adjust parameters, the inefficiencies and quality deficiencies in existing technologies are resolved, resulting in more efficient video coding and decoding, and improved data quality and processing efficiency.
Patent Information
- Application Number
- CN202080028272.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-16
- Filing Date
- 2020-04-16
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2040-04-16
AI Technical Summary
Existing hierarchical video coding technologies are insufficient in terms of efficiency and optimization, making it difficult to effectively reduce data size and processing requirements while improving the image quality of the final reconstructed image.
By exchanging information in a hierarchical coding scheme, coordinating the first and second coding operations to adjust coding parameters, and improving overall efficiency or quality, this includes exchanging encoder and decoder information between coding operations, leveraging the specific benefits provided by each level of the hierarchical structure.
It improves data quality and processing efficiency, reduces data size and processing requirements, and enables a more efficient video encoding and decoding process.
Smart Images

Figure CN113994685B_ABST
Abstract
Description
BACKGROUND
[0001] Recent improvements in video coding technology have included the concept of hierarchical video coding. Examples include VC6, which is undergoing standardization at SMPTE as ST 2117, and LCEVC, which is undergoing standardization at MPEG as MPEG-5 Part II. Typically, these hierarchical coding schemes use multiple resolution levels and an encoder (or encoding module) associated with each resolution level.
[0002] Examples of hierarchical coding technology include patent publications WO 2013 / 171173, WO 2014 / 170819, WO 2018 / 046940, and WO 2019 / 111004, the contents of which are incorporated herein by reference.
[0003] In these new coding schemes, efficiency and optimization are sought to reduce data size and / or processing requirements, while improving the picture quality of the final reconstructed image. SUMMARY
[0004] According to a first aspect of the application, a method of encoding a signal can be provided, the method comprising: receiving an input signal; applying a first encoding operation to the input signal using a first codec to produce a first encoded stream; and, applying a second encoding operation to the input signal to produce a second encoded stream, wherein the first and second encoded streams are used to be combined at a decoder; and wherein the method further comprises exchanging information between the first encoding operation and the second encoding operation. Depending on the exchanged information, the exchange of useful information between the encoding operations allows for an improvement in data quality after reconstruction, data simplification, and / or processing efficiency. The signal can be a data stream.
[0005] Preferably, the method can further comprise adapting the first or second encoding operation, or both, based on the information. With such exchange of information, the adaptation can be coordinated to improve overall efficiency or quality or provide a balance throughout the encoding operations. More preferably, the first and second encoding operations can be encoding operations of a hierarchical coding scheme. As each level of the hierarchical structure provides a specific benefit or impact, by adapting the parameters at each level, an overall improvement can be made.
[0006] In certain examples, the hierarchical coding scheme can include generating a base encoded signal by feeding an input signal to an encoder with a down-sampled version of the input signal; generating a first residual signal by obtaining a decoded version of the base encoded signal and using a difference between the decoded version of the base encoded signal and the down-sampled version of the input signal to generate the first residual signal; and encoding the first residual signal to generate a first encoded residual signal. The method can further include generating a second residual signal by decoding the first encoded residual signal to generate a first decoded residual signal, using the first decoded residual signal to correct the decoded version of the base encoded signal to generate a corrected decoded version, up-sampling the corrected decoded version, and using a difference between the up-sampled version and the input signal to generate the second residual signal; wherein the method further includes encoding the second residual signal to generate a second encoded residual signal, wherein the base encoded signal, the first encoded residual signal, and the second encoded residual signal comprise an encoding of the input signal. The first encoding operation can be the step of encoding the first residual signal or the step of encoding the first encoded residual signal. The second encoding operation can be the step of encoding the first residual signal or the step of encoding the second residual signal.
[0007] In certain other examples, the hierarchical coding scheme can include generating a base encoded signal by feeding an input signal to an encoder with a down-sampled version of the input signal, the down-sampled version having undergone one or more down-sampling operations; generating one or more residual signals by up-sampling an output of each down-sampling operation to generate one or more up-sampled signals and using a difference between each up-sampled signal and an input to a corresponding down-sampling operation to generate the one or more residual signals; and encoding the one or more residual signals to generate one or more encoded residual signals. The first and second encoding operations can correspond to the step of encoding either of the one or more residual signals.
[0008] Alternatively, the hierarchical coding scheme can comprise generating a base encoded signal by feeding an input signal to an encoder in a down-sampled version which has undergone a plurality of sequential down-sampling operations; and generating an encoded first residual signal by: up-sampling the down-sampled version of the input signal; using a difference between the up-sampled version of the down-sampled version and an input to a last down-sampling operation of the plurality of sequential down-sampling operations to generate the first residual signal; and, encoding the first residual signal; generating a second residual signal by: up-sampling a sum of the first residual signal and an output of a preceding down-sampling operation to the last down-sampling operation of the plurality of sequential down-sampling operations; using a difference between the up-sampled sum and the input to the preceding down-sampling operation to generate the second residual signal; and, encoding the second residual signal. The first encoding operation can be the step of encoding the first residual signal or the step of encoding the second residual signal.
[0009] According to a second aspect of the application, there can be provided a method of decoding a signal, the method comprising: receiving a first encoded signal and a second encoded signal; applying a first decoding operation to the first encoded signal to generate a first output signal; applying a second decoding operation to the second encoded signal to generate a second output signal; and, combining the first output signal and the second output signal to reconstruct an input signal, wherein the method further comprises exchanging information between the first decoding operation and the second decoding operation.
[0010] Preferably, the method can further comprise adapting the first or second decoding operation, or both, based on the information. More preferably, the first and second decoding operations can be decoding operations of a hierarchical coding scheme.
[0011] In certain examples, the hierarchical coding scheme can comprise: receiving a base encoded signal and indicating a decoding of the base encoded signal to generate a base decoded signal; receiving a first encoded residual signal and decoding the first encoded residual signal to generate a first decoded residual signal; correcting the base decoded signal using the first decoded residual signal to generate a corrected version of the base decoded signal; up-sampling the corrected version of the base decoded signal to generate an up-sampled signal; receiving a second encoded residual signal and decoding the second encoded residual signal to generate a second decoded residual signal; and, combining the up-sampled signal and the second decoded residual signal to generate a reconstructed version of the input signal.
[0012] In certain instances, the first and second encoded signals can comprise first and second sets of components, respectively, the first set of components corresponding to a lower image resolution than the second set of components, the method comprising, for each of the first and second sets of components: decoding the set of components to obtain a decoded set, the method further comprising upscaling the decoded first set of components to increase the corresponding image resolution of the decoded first set of components to equal the corresponding image resolution of the decoded second set of components, and combining the decoded first and second sets of components together to produce a reconstructed set. The method can further comprise receiving one or more further sets of components, wherein each of the one or more further sets of components corresponds to a higher image resolution than the second set of components, and wherein each of the one or more further sets of components corresponds to progressively higher image resolutions, the method comprising, for each of the one or more further sets of components, decoding the set of components to obtain a decoded set, the method further comprising, for each of the one or more further sets of components, in increasing order of corresponding image resolution: upsampling the reconstructed set having the highest corresponding image resolution to increase the corresponding image resolution of the reconstructed set to equal the corresponding image resolution of the further set of components, and combining the reconstructed set with the further set of components together to produce a further reconstructed set.
[0013] Optionally, the step of exchanging information can comprise sending the information with metadata in a stream. Alternatively, the step of exchanging information can comprise embedding the information in a stream. Further alternatively, the step of exchanging information can comprise sending the information using an application programming interface (API). Further alternatively, the step of exchanging information can comprise sending a pointer to a shared memory space. The step of exchanging information can further comprise exchanging information using supplemental enhancement information (SEI).
[0014] The exchanged information can comprise encoding or decoding parameters for modifying encoding or decoding operations, respectively.
[0015] The exchanged information can comprise one or more selected from the group comprising: user input information; metadata describing content of an input video; host device information; content analysis information; perceptual information describing a region of an image, wherein encoding artefacts are less noticeable taking into account the human visual system (HVS); motion information; complexity of a frame to be encoded; frame entropy; estimated number of bits required; frame information; decisions taken during encoding or decoding; type of prediction for a frame to be decoded / encoded; quantization level for a frame to be decoded / encoded, decisions made per group of pixels; statistics information; target video quality metric; and, rate control information.
[0016] The signal can be video and the step of exchanging information can be performed per video, per group of pictures, per segment, per picture, per group of pixels and per pixel
[0017] According to another aspect of the application, a method of encoding a data stream can be provided, the method can comprise receiving an input video, applying a base encoding operation to the input video using a base codec to produce a base encoded stream, applying another encoding operation to the input video to produce an enhancement stream, exchanging information between the base encoding operation and the other encoding operation.
[0018] According to another aspect of the application, a method of decoding a data stream can be provided, the method can comprise receiving a base encoded data stream and enhancement stream data, applying a base decoding operation to the base encoded stream to produce a first output video, applying another decoding operation to the enhancement data stream to produce a set of residuals, and combining the first output video with the set of residuals to reconstruct an input video, wherein the method further comprises exchanging information between the base decoding operation and the other decoding operation.
[0019] According to another aspect, an apparatus for encoding a data set into an encoded data set comprising a header and a payload can be provided. The apparatus is configured to encode an input video according to the steps above. The apparatus can comprise a processor configured to carry out the method of any of the above aspects.
[0020] According to another aspect, an apparatus for decoding a data set from a data set comprising a header and a payload into a reconstructed video can be provided. The apparatus is configured to decode an output video according to the steps above. The apparatus can comprise a processor configured to carry out the method of any of the above aspects.
[0021] An encoding apparatus and a decoding apparatus can also be provided.
[0022] According to further aspects of the application, a computer readable medium which, when executed by a processor, causes the processor to perform any of the methods of the above aspects can be provided. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 High level schematic showing a hierarchical encoding and decoding process;
[0024] Figure 2 High level schematic showing a hierarchical deconstruction process;
[0025] Figure 3 Alternative high level schematic showing a hierarchical deconstruction process;
[0026] Figure 4a high level schematic showing an encoding process suitable for encoding residuals for layered output;
[0027] Figure 5 a high level schematic showing a hierarchical decoding process suitable for decoding each output level from Figure 4
[0028] Figure 6 a high level schematic showing an encoding process of a hierarchical encoding technique; and,
[0029] Figure 7 a high level schematic showing a decoding process suitable for decoding output from Figure 6 DETAILED DESCRIPTION
[0030] The present invention relates to methods. In particular, the present invention relates to methods for encoding and decoding signals. Processing data can include, but is not limited to, obtaining, deriving, outputting, receiving, and reconstructing data. The present invention relates to the exchange of useful information between two (or more encoders) encoding the same content (or portions thereof or representations thereof), in the form of a joint module or a module implemented in one of the encoders but producing information that can be useful to the other. Similarly, each of the concepts described herein also apply to decoder stages where multiple decoding stages can exchange information with each other.
[0031] In preferred examples, the encoder or decoder is part of a hierarchical encoding scheme. In more preferred examples, although the encoder or decoder utilizes techniques incorporated in the VC-6 or LCEVC encoding schemes, the concepts shown herein are not limited to these particular hierarchical encoding schemes.
[0032] Figure 1 A hierarchical encoding scheme is shown in broad outline. Data to be encoded 101 is retrieved by a hierarchical encoder 102 that outputs encoded data 103. The encoded data 103 is then received by a hierarchical decoder 104 that decodes the data and outputs decoded data 105.
[0033] In general, the hierarchical encoding schemes used in the examples herein produce a base or core level that represents the original data of a lower quality level and one or more residual levels that can be used to reproduce the original data of a higher quality level using a decoded version of the base level data. In general, the term “residual” as used herein refers to the difference between the values of a reference array or frame and the actual array or frame of data. The array can be a one or two dimensional array representing an encoding unit. For example, an encoding unit can be a 2x2 or 4x4 collection of residual values corresponding to a similarly sized region of an input video frame.
[0034] It should be noted that the generalized instance is unknowable regarding the properties of the input signal. The reference to "residual data" as used herein refers to data derived from a set of residuals, such as the set of residuals itself or the output of a dataset on which operations are performed on the set of residuals. Throughout this specification, generally, a set of residuals contains multiple residuals or residual elements, each of which corresponds to a signal element, i.e., an element of the signal or original data.
[0035] In certain instances, the data may be an image or a video. In these instances, the residual set corresponds to an image or frame of the video, where each residual is associated with a pixel of the signal, the pixel being a signal element.
[0036] The method described herein can be applied to so-called data planes that reflect different color components of a video signal. For example, the method can be applied to different planes reflecting YUV or RGB data of different color channels. Different color channels can be processed in parallel. The components of each stream can be compared in any logical order.
[0037] A hierarchical coding scheme in which the concepts of the present invention can be deployed will now be described. The scheme is conceptually illustrated in… Figures 2 to 5 In this type of coding technique, residual data is used for progressively higher quality levels. In the proposed technique, the core layer represents the image at a first resolution, and subsequent layers in the hierarchical structure are residual data or adjustment layers necessary for the decoding side to reconstruct the image at higher resolutions. Each layer or level can be referred to as a ladder index, such that the residual data is the data needed to correct low-quality information present in lower ladder indices. In this hierarchical technique, specifically each residual layer, each layer or ladder index is typically a relatively sparse dataset with many zero-valued elements.
[0038] When referring to a tier index, it collectively refers to all tiers or component sets at that tier, such as all subsets produced by the transformation steps performed at that quality tier.
[0039] By employing this specific hierarchical approach, the described data structure eliminates any requirements or dependencies on the aforementioned or preceding quality levels. Quality levels can be encoded and decoded independently without reference to any other layers. Therefore, compared to many other known hierarchical coding schemes that require decoding of the lowest quality level in order to decode any higher quality level, the described method does not require decoding of any other layers. However, the principles of information exchange described below can also be applied to other hierarchical coding schemes.
[0040] like Figure 2As shown, the encoded data represents a set of layers or levels, which are broadly referred to here as a ladder index. The base or core layer represents the original data frame 210, but at the lowest quality level or resolution, and subsequent residual data ladders can be combined with the data at the core ladder index to regenerate the original image at progressively higher resolutions.
[0041] To generate the core hierarchy index, the input data frame 210 can be downsampled using several downsampling operations 201, corresponding to the number of levels or hierarchy indices to be used in the hierarchical coding operation. One less downsampling operation 201 is required than the number of levels in the hierarchical structure. In all the examples shown herein, although there are four levels or hierarchy indices in the output encoded data and therefore three downsampling operations, it will be understood that these are for illustrative purposes only. With n indicating the number of levels, the number of downsamplers is n-1. Core level R 1-n This is the output of the third downsampling operation. As mentioned above, the core level R... 1-n The representation of the input data frame corresponding to the lowest quality level.
[0042] To distinguish downsampling operation 201, each operation will be referred to in the order in which it is performed on the input data 210 or the data represented by its output. For example, in this instance, the third downsampling operation 201... 1-n It can also be called a core downsampler because its output produces a core tier index or tier. 1-n That is, the indices of all tiers at this level are 1-n. Therefore, in this example, the first downsampling operation 201 -1 Corresponding to the R-1 downsampler, the second downsampling operation 201 -2 Corresponding to the R-2 downsampler, and the third downsampling operation 201 1-n Corresponding to the core or R-3 downsampler.
[0043] like Figure 2 As shown in the figure, it represents the core quality level R. 1-n The data underwent an upsampling operation 202 1-n This is referred to here as the core upsampler. In the second downsampling operation 201 -2 The output (output of the R-2 downsampler, i.e., the input to the core downsampler) and the core upsampler 202 1-n The difference 203 between the outputs is output as the first residual data R. -2 This first residual data R -2 Correspondingly, this represents the core level R. -3 The error between the signal used to create the level. Since the signal itself undergoes two downsampling operations in this example, the first residual data R...-2 To adjust the layers, the adjustment layers can be used to reproduce the original signal at a higher quality level than the core quality level but at a lower level than the input data frame 210.
[0044] In Figure 2 and 3 the way of generating the residual data representing the higher quality level is conceptually shown to vary.
[0045] In Figure 2 the output of the second down-sampling operation 201 -2 (or R-2 down-sampler, i.e. to generate the first residual data R -2 ) is up-sampled 202 -2 and the difference 203 -2 between the input to the second down-sampling operation 201 -2 (or R-2 down-sampler, i.e. the output of the R-1 down-sampler) is calculated in the same way as for the first residual data R -1 . This difference is thus the second residual data R -1 and represents an adjustment layer that can be used to reproduce the original signal at a higher quality level using data from a lower level.
[0046] However, in a variation of Figure 3 the output of the second down-sampling operation 201 -2 (or R-2 down-sampler) is combined or summed 304 -2 with the first residual data R -2 to reproduce the output of the core up-sampler 202 1-n . In this variation, the up-sampling 202 -2 is performed on this reproduced data rather than the down-sampled data. The up-sampled data is compared 203 -1 in a similar way to the input to the second down-sampling operation (or R-2 down-sampler, i.e. the output of the R-1 down-sampler) to generate the second residual data R -1 .
[0047] Figure 2 and 3 the variations between the embodiments of Figure 2 yield a minor variation in the residual data between the two embodiments.
[0048] The process or loop repeats to generate a third residual R0. In the examples of Figure 2 and 3 the output residual data R0(i.e. third residual data) corresponds to the highest level and is used at the decoder to reproduce the input data frame. At this level, the difference operation is based on the same input data frame as the input to the first down-sampling operation.
[0049] Figure 4 Example encoding process 401 is shown for encoding each of the hierarchical or tiered indices of data to produce a tiered set of encoded data with tiered indices. This encoding process is only an example of a suitable encoding process for encoding each of the hierarchies, but it should be understood that any suitable encoding process can be used. The input to the process is from... Figure 2 Or the corresponding level of the residual data output by 3, and the output is a set of tiers of encoded residual data, the tiers of encoded residual data together hierarchically representing the encoded data.
[0050] In the first step, transformation 402 is performed. This transformation can be a directional decomposition transformation, wavelet transform, or discrete cosine transform as described in WO2013 / 171173. If a directional decomposition transformation is used, a set of four components can be output. When referring to the ladder index, it collectively refers to all directions (A, H, V, D), i.e., the four ladders. The component set is then quantized 403 before entropy encoding. In this example, the entropy encoding operation 404 is coupled to a sparsification step 405, which utilizes the sparsity of the residual data to reduce the total data size and involves mapping data elements to a sorted quadtree. This coupling of entropy encoding and sparsification is further described in WO2019 / 111004, but the precise details of this process are not relevant to the understanding of this invention. Each array of residuals e can be considered a ladder.
[0051] The process described above corresponds to the encoding process suitable for encoding data used for reconstruction according to the SMPTE ST 2117, VC-6 multiplane image format. VC-6 is a flexible, multi-resolution, intrinsic bitstream-only format capable of compressing any ordered set of integer-element grids, each grid having an independent size and designed for image compression. It employs data-agnostic techniques for compression and can compress low- or high-bit-depth images. The bitstream header can contain various metadata about the image.
[0052] As will be understood, each step or step index can be implemented using a separate encoder or encoding operation. Similarly, the encoding module can be divided into downsampling and comparison steps to generate residual data, and the residuals can then be encoded, or alternatively, each of the steps in the step can be implemented in a combined encoding module. Thus, the process can be implemented, for example, using four encoders: one encoder for each step index, one encoder and multiple encoding modules operating in parallel or serially, or one encoder repeatedly operating on different datasets.
[0053] An example of reconstructing an original data frame that has been encoded using the exemplary process above is set forth below. This reconstruction process can be referred to as pyramid reconstruction. Advantageously, the method provides an efficient technique for reconstructing an image encoded in a received data set that can be received by means of a data stream, for example by individually decoding different sets of components corresponding to different image size or resolution levels and combining image details from one decoded set of components with upscaled decoded image data from a lower resolution set of components. Thus, by performing this process for two or more sets of components, a digital image at progressively higher resolution or greater number of pixels can be reconstructed at the structure or details in the sets of components without requiring the complete or full image details of the highest resolution set of components to be received. In particular, the method facilitates progressively adding higher and higher resolution details while reconstructing an image from lower resolution sets of components in a hierarchical manner.
[0054] Furthermore, decoding each set of components separately facilitates parallel processing of the received sets of components, thus improving reconstruction speed and efficiency in embodiments where multiple processes are available.
[0055] Each resolution level corresponds to a quality level or echelon index. This is a collective term associated with describing a plane (in this example, a representation of a grid of integer valued elements) of all new input or received sets of components and the output reconstructed image for a loop of index-m. For example, the reconstructed image in echelon index zero is the output of the final loop of pyramid reconstruction.
[0056] Pyramid reconstruction can be a process that reconstructs an inverse pyramid starting from an initial echelon index and using a loop of new residuals to derive higher echelon indices up to a maximum quality, quality zero, at echelon index zero. A loop can be considered a step in such pyramid reconstruction, identified by index-m. A step typically includes upsampling data output from a possible previous step, for example upsizing a decoded first set of components, and new residual data as an additional input in order to obtain output data to be upsampled in a possible subsequent step. In the case where only a first and second set of components are received, the number of echelon indices will be two, and there is no possible subsequent step. However, in instances where the number of sets of components or echelon indices is three or greater, output data can be progressively upsam ped in subsequent steps.
[0057] A first set of components typically corresponds to an initial echelon index, which can be represented by echelon index 1-N, where N is the number of echelon indices in the plane.
[0058] Generally, the upgrading of a decoded first component set includes applying an up-sampler to the output of the decoding procedure for the initial tier index. In an example, this involves bringing the resolution of the reconstructed picture output from the decoding of the initial tier index component set into line with the resolution of the second component set corresponding to 2-N. Generally, the upgraded output from a lower tier index component set corresponds to a predicted image at a higher tier index resolution. Due to the low resolution initial tier index image and the up-sampling process, the predicted image generally corresponds to a smooth or blurred picture.
[0059] Adding higher resolution detail from the above tier index to this predicted picture provides a combined reconstructed image set. Advantageously, where the received component sets for one or more higher tier index component sets include residual image data or data indicative of the difference in pixel values between the upgraded predicted picture and the original, uncompressed or pre-encoded image, the amount of received data required to reconstruct an image or data set of a given resolution or quality can be significantly less than the amount or rate of data required to receive the same quality image using other techniques. Thus, by combining the low detail image data received at a lower resolution with progressively greater detail image data received at increasingly higher resolutions according to the method, data rate requirements are reduced.
[0060] Generally, the encoded data set includes one or more further component sets, wherein each of the one or more further component sets corresponds to a higher image resolution than the second component set, and wherein each of the one or more further component sets corresponds to a progressively higher image resolution, the method comprising, for each of the one or more further component sets, decoding the component set so as to obtain a decoded set, the method further comprising, for each of the one or more further component sets, in increasing order of corresponding image resolution: upgrading the reconstructed set having the highest corresponding image resolution so as to increase the corresponding image resolution of the reconstructed set to equal the corresponding image resolution of another component set, and combining the reconstructed set with the another component set so as to produce another reconstructed set.
[0061] In this way, the method can involve taking a reconstructed image output of a given component set tier or tier index, upgrading the reconstructed set, and combining it with a decoded output of an above component set or tier index to produce a new, higher resolution reconstructed picture. It will be appreciated that this can be performed repeatedly for progressively higher tier indices, depending on the total number of component sets in the received set.
[0062] In a typical example, each of the sets of components corresponds to progressively higher image resolutions, with each progressively higher image resolution corresponding to a four-fold increase in the number of pixels in the corresponding image. Typically, thus, the image size corresponding to a given set of components is four times the size or number of pixels of the image corresponding to the set of components below it, or twice the height and twice the width of the image, the set of components below it having a tier index one less than the tier index in question. The set of sets of components received may, for example, facilitate simpler upscaling operations, in which the linear size of each corresponding image is twice the size of the image below it.
[0063] In the example shown, the number of further sets of components is two. Thus, the total number of sets of components in the received set is four. This corresponds to an initial tier index of tier-3.
[0064] The first set of components can correspond to image data, and the second and any further sets of components correspond to residual image data. As noted above, in the case that the lowest tier index, i.e. the first set of components, contains a low resolution or down-sampled version of the image being transmitted, the method provides a particularly advantageous reduction in data rate requirement for a given image size. In this way, at each cycle of the reconstruction starting with a low resolution image, the image is upscaled in order to produce a high resolution, but a smooth version, and then the image is improved by adding the difference between the upscaled predicted picture and the actual image to be transmitted at the resolution, and this addition improvement can be repeated for each cycle. Thus, each set of components above the initial tier index requirement contains only residual data, in order to reintroduce information that can have been lost in down-sampling the original image to the lowest tier index.
[0065] The method provides a way of obtaining image data, for example upon receipt of a set containing data that has been compressed, for example by means of decomposition, quantization, entropy encoding and sparsification, which image data can be residual data.
[0066] The sparsification step is particularly advantageous when used in conjunction with a sparse set of original or pre-transmission data, which can typically correspond to residual image data. The residual can be the difference between elements of a first image and elements of a second image typically at the same location. Such residual image data can typically have a high degree of sparsity. This can be considered to correspond to images in which areas of detail are sparsely distributed amongst areas in which the detail is minimal, negligible or non-existent. Such sparse data can be described as a data array in which the data is organised in at least a two-dimensional structure (e.g. a grid) and in which a significant proportion of the data so organised is zero (logically or numerically) or considered to be below a certain threshold. Residual data is just one example. Additionally, metadata can be sparse and so reduced in size to a considerable extent by this process. Transmitting data that has been sparsified allows for a significant reduction in the required data rate by virtue of not transmitting such sparse areas, and instead reintroducing them at the appropriate location within the received byte set at the decoder.
[0067] Typically, the entropy decoding, dequantisation and directional synthesis transform steps are performed in accordance with parameters defined by the encoder or node from which the received encoded data set is transmitted. For each set of component or tier index, the steps serve to decode the image data in order to arrive at a set that can be combined with different tier indices in accordance with the techniques disclosed above, whilst allowing for the transmission of each level of the set in a data efficient manner.
[0068] A method of reconstructing an encoded data set in accordance with the methods disclosed above can also be provided, in which the decoding of each of the first and second sets of components is performed in accordance with the methods disclosed above. Thus, the advantageous decoding methods of the present disclosure can be used for each set of components or tier index in a received image data set and reconstructed accordingly.
[0069] Reference Figure 5 A decoding example is now described. A set of encoded data 501 is received, in which the set comprises four tier indices, each comprising four tiers: from tier 0, the highest resolution or quality level, to tier -3, the initial tier. The image data carried in the tier -3 component set corresponds to the image data, and the other component set contains the residual data for the transmitted image. Although each of the levels can output data that can be considered to be residual, the residual in the initial tier level, tier -3, actually corresponds to the actual reconstructed image. At stage 503, each of the component sets is processed in parallel in order to decode the encoded set.
[0070] With reference to the initial tier index or core tier index, the following decoding steps are performed for each component set tier -3 to tier 0.
[0071] At step 507, the set of components is de-sparse. In this way, de-sparseness causes the sparse two-dimensional array to be reproduced from the set of encoded bytes received at each tier. By this process, the zero values grouped at positions within the two-dimensional array that were not received (due to being omitted from the transmitted set of bytes, in order to reduce the amount of data transmitted) are repopulated. The non-zero values in the array retain their correct values and positions within the reproduced two-dimensional array, with the de-sparseness step repopulating the transmitted zero values at their appropriate positions or groups of positions in between.
[0072] At step 509, a range decoder is applied to the de-sparse set at each tier in order to replace the encoded symbols within the array with pixel values, the configuration parameters of the range decoder corresponding to those used to encode the transmitted data prior to transmission. The encoded symbols in the received set are replaced with pixel values according to an approximation of the pixel value distribution of the image. Using an approximate distribution, i.e. the relative frequency of each value across all pixel values in the image, rather than the true distribution, permits a reduction in the amount of data required to decode the set, since the range decoder requires distribution information in order to perform this step. As described in this disclosure, the steps of de-sparseness and range decoding are interdependent rather than sequential. This is indicated by the loop formed by the arrows in the flowchart.
[0073] At step 511, the array of values is de-quantized. This process is performed again according to the parameters by which the decomposed image was quantized prior to transmission.
[0074] Following de-quantization, at step 513 the set is transformed by a synthesis transform comprising applying an inverse directional decomposition operation to the de-quantized array. This causes the directional filtering of the 2x2 operators comprising averaging, horizontal, vertical and diagonal operators to be reversed, such that the resulting array is the image data for tier-3 and the residual data for tiers-2 to tier-0.
[0075] Stage 505 shows several loops involved in the reconstruction with the output of the synthesis transform for each of the sets of components 501.
[0076] Stage 515 indicates the reconstructed image data output from the decoder 503 for the initial tier. In the example, the reconstructed picture 515 has a resolution of 64x64. At 516, this reconstructed picture is upsampled to increase the constituent number of its pixels by a factor of four, thereby producing a predicted picture 517 having a resolution of 128x128. At stage 520, the predicted picture 517 is added to the decoded residual 518 from the output of the decoder at tier-2. The addition of these two 128x128 sized images results in a 128x128 sized reconstructed image that contains the smooth image details of the initial tier augmented by the higher resolution details from the residual at tier-2. If the desired output resolution corresponds to tier-2, then this resulting reconstructed picture 519 can be output or displayed. In this example, the reconstructed picture 519 is used for another cycle.
[0077] At step 512, the reconstructed image 519 is upsampled in the same manner as at step 516 to produce a predicted picture 524 of size 256x256. This is then combined at step 528 with the decoded tier-1 output 526, thereby producing a reconstructed picture 527 of size 256x256 that is an upscaled version of the prediction 519 augmented by the higher resolution details of the residual 526. At 530, this process is repeated at the last time and the reconstructed picture 527 is upscaled to a resolution of 512x512 for combination with the tier 0 residual at stage 532. Thereby, a 512x512 reconstructed picture 531 is obtained.
[0078] Another hierarchical coding technique is shown in Figure 6 and 7 by which the principles of the present invention can be utilized. This technique is a flexible, adaptable, efficient and computationally light coding format that combines different video coding formats, base codec (e.g. AVC, HEVC, or any other current or future codec) with at least two levels of encoded data augmentation.
[0079] The general structure of the coding scheme uses a downsampled source signal encoded with a base codec, adds correction data of a first level to the decoded output of the base codec to produce a corrected picture, and then adds augmentation data of another level to an upsampled version of the corrected picture.
[0080] Thus, the stream is considered as a base stream and an enhancement stream. Notably, it is generally expected that the base stream can be decoded by a hardware decoder, while the enhancement stream is suitable for a software processing implementation with suitable power consumption.
[0081] This structure creates multiple degrees of freedom, which allows great flexibility and adaptability to many scenarios, making the coding format suitable for many use cases, including OTT transmission, live streaming, live UHD broadcast, etc.
[0082] Although the decoded output of the base codec is not intended for viewing, it is a fully decoded video at lower resolution, making the output compatible with existing decoders, and can also be used as a lower resolution output if deemed appropriate.
[0083] In this and other contemplated examples, in general, each or both enhancement streams can be encapsulated into one or more enhancement bitstreams using a set of network abstraction layer units (NALUs). The NALUs are intended to encapsulate the enhancement bitstreams so as to apply the enhancement to the correct base reconstructed frame. The NALUs may, for example, contain a reference index to a NALU that contains the base decoder reconstructed frame bitstream to which the enhancement must be applied. In this way, the enhancement can be synchronized to the base stream, and the frames of each bitstream are combined to produce the decoded output video (i.e., the residual of each frame of the enhancement level is combined with the frame of the base decoded stream). A set of pictures can represent multiple NALUs.
[0084] Returning to the initial process described above, in which the base stream is provided along with enhancements of two levels (or sub-levels) within the enhancement stream, an example of a generalized encoding process is depicted in the block diagram of Figure 6 Input full resolution video 600 is processed to produce various encoded streams 601, 602, 603. A first encoded stream (encoded base stream) is produced by feeding a base codec (e.g., AVC, HEVC, or any other codec) with a down-sampled version of the input video. The encoded base stream can be referred to as a base layer or base level. A second encoded stream (encoded level 1 stream) is produced by processing a residual obtained by taking the difference between the reconstructed base codec video and the down-sampled version of the input video. A third encoded stream (encoded level 2 stream) is produced by processing a residual obtained by taking the difference between an up-sampled version of the corrected version of the reconstructed base encoded video and the input video. In certain cases, Figure 6 The components of FIG. 1 can provide a generalized low complexity encoder. In certain cases, the enhancement streams can be produced by an encoding process that forms part of the low complexity encoder, and the low complexity encoder can be configured to control an independent base encoder and decoder (e.g., packaged as a base codec). In other cases, the base encoder and decoder can be provided as part of the low complexity encoder. In one case, Figure 6 The low complexity encoder of FIG. 1 can be considered as a form of wrapper for the base codec, in which the functionality of the base codec can be hidden from the entity implementing the low complexity encoder.
[0085] The downsampling operation shown by the downsampling component 105 can be applied to the input video to produce a downsampled video to be encoded by the base encoder 613 of the base codec. The downsampling can be done in both vertical and horizontal directions, or alternatively only in the horizontal direction. The base encoder 613 and the base decoder 614 can be implemented by a base codec (e.g., as different functions of a common codec). The base codec and / or one or more of the base encoder 613 and the base decoder 614 can comprise suitably configured electronic circuitry (e.g., a hardware encoder / decoder) and / or computer program code executed by a processor.
[0086] Each enhancement stream encoding process can not necessarily include an upsampling step. For example, in Figure 6 the first enhancement stream is conceptually a corrected stream, while the second enhancement stream is upsampled to provide an enhancement level.
[0087] Referring in more detail to the process of producing the enhancement streams, to generate the encoded level 1 stream, the encoded base stream is decoded by the base decoder 614 (i.e., a decoding operation is applied to the encoded base stream to produce a decoded base stream). The decoding can be performed by a decoding function or mode of the base codec. The difference between the decoded base stream and the downsampled input video is then produced at the level 1 comparator 610 (i.e., a subtraction operation is applied to the downsampled input video and the decoded base stream to produce a first set of residuals). The output of the comparator 610 can be referred to as a first set of residuals, e.g., a surface or frame of residual data, where a residual value is determined for each picture element at the resolution of the output of the base encoder 613, the base decoder 614, and the downsampling block 605.
[0088] The difference is then encoded by the first encoder 615 (i.e., a level 1 encoder) to produce the encoded level 1 stream 602 (i.e., an encoding operation is applied to the first set of residuals to produce a first enhancement stream).
[0089] As noted above, the enhancement streams can include a first enhancement level 602 and a second enhancement level 603. The first enhancement level 602 can be considered a corrected stream, e.g., a stream that provides a correction level to the base encoded / decoded video signal at a lower resolution than the input video 600. The second enhancement level 603 can be considered another enhancement level that converts the corrected stream to the original input video 600, e.g., it applies an enhancement or correction level to a signal reconstructed from the corrected stream.
[0090] In Figure 6In this example, a second enhancement layer 603 is generated by encoding another set of residuals. This other set of residuals is generated by a level 2 comparator 619. The level 2 comparator 619 determines the difference between the upsampled form of the decoded level 1 stream, such as the output of upsampling component 617, and the input video 600. The input to upsampling component 617 is generated by applying a first decoder (i.e., a level 1 decoder) to the output of first encoder 615. This produces a decoded set of level 1 residuals. These are then combined with the output of the base decoder 614 at summing component 620. This effectively applies level 1 residuals to the output of base decoder 614. This allows losses during level 1 encoding and decoding to be corrected by level 2 residuals. The output of summing component 620 can be considered as an analog signal representing the output of applying level 1 processing to the encoded base stream 601 and the encoded level 1 stream 602 at the decoder.
[0091] As mentioned, the upsampled stream is compared with the input video, which forms another set of residuals (i.e., the difference operation is applied to the regenerated upsampled stream to produce another set of residuals). The other set of residuals is then encoded by the second encoder 621 (i.e., the level 2 encoder) into an encoded level 2 enhanced stream (i.e., the encoding operation is then applied to the other set of residuals to produce another encoded enhanced stream).
[0092] Therefore, as Figure 6 As shown and described above, the output of the encoding process is a base stream 601 and one or more enhancement streams 602, 603, which preferably include a first enhancement level and another enhancement level. The three streams 601, 602, and 603 can be combined, with or without additional information such as control headers, to produce a combined stream representing the video coding architecture of the input video 600. It should be noted that... Figure 6 The components shown operate on blocks or coding units of data, such as 2×2 or 4×4 portions of a frame at a specific resolution level. These components operate without any inter-block dependencies, thus allowing them to be applied in parallel to multiple blocks or coding units within a frame. This differs from contrasting video coding schemes, where dependencies (e.g., spatial or temporal) exist between blocks. These dependencies limit the level of parallelism and require significantly higher complexity.
[0093] exist Figure 7 The block diagram depicts the corresponding generalized decoding process. It is said that... Figure 7 Can display corresponding Figure 6The low-complexity decoder is a low-complexity encoder. The low-complexity decoder receives three streams 601, 602, and 603 generated by the low-complexity encoder along with a header 704 containing other decoding information. The encoded base stream 601 is decoded by a base decoder 710 corresponding to the base codec used in the low-complexity encoder. The encoded Level 1 stream 602 is received by a first decoder 711 (i.e., a Level 1 decoder), which decodes streams generated by the low-complexity encoder along with a header 704 containing other decoding information. Figure 1 The first encoder 615 decodes the first residual set encoded by the first encoder 615. At the first summing component 712, the output of the base decoder 710 is combined with the decoded residual obtained from the first decoder 711. The combined video, which can be referred to as the Level 1 reconstructed video signal, is upsampled by the upsampling component 713. The encoded Level 2 stream 103 is received by the second decoder 714 (i.e., the Level 2 decoder). The second decoder 714 decodes the data as shown in the image. Figure 1 The second encoder 621 decodes the second residual set encoded by the second residual set. Although the header 704... Figure 7 The video is shown as being used by the second decoder 714, but it can also be used by the first decoder 711 and the base decoder 710. The output of the second decoder 714 is a second set of decoded residuals. These provide higher resolution for the first residual set and the input of the upsampling component 713. At the second summing component 715, the second residual set from the second decoder 714 is combined with the output of the upsampling component 713, i.e., the upsampled reconstructed Level 1 signal, to reconstruct the decoded video 750.
[0094] According to a low-complexity encoder, Figure 7 The low-complexity decoder can operate in parallel on different blocks or coding units of a given frame of the video signal. Furthermore, decoding performed by two or more of the base decoder 710, the first decoder 711, and the second decoder 714 can be executed in parallel. This is possible because there is no inter-block dependency.
[0095] During decoding, the decoder can parse header 704 (which may contain global configuration information, picture or frame configuration information, and data block configuration information) and configure the low-complexity decoder based on those headers. To regenerate the input video, the low-complexity decoder can decode each of the base stream, the first enhancement stream, and another or a second enhancement stream. Frames of the streams can be synchronized and then combined to derive decoded video 750. Depending on the configuration of the low-complexity encoder and decoder, decoded video 750 can be a lossy or lossless reconstruction of the original input video 100. In many cases, decoded video 750 can be a lossy reconstruction of the original input video 600, wherein the loss has a reduced or minimal impact on the perception of decoded video 750.
[0096] exist Figure 6 and7 In each of the 2-level and 1-level encoding operations, the steps of transformation, quantization and entropy encoding can include (e.g., in that order). It can also include residual staging, weighting and filtering. Similarly, in the decoding stage, the residuals can be passed through an entropy decoder, dequantizer and inverse transform module (e.g., in that order). Any suitable encoding and corresponding decoding operations can be used. Preferably, however, the 2-level and 1-level encoding steps can be performed in software (e.g., as performed by one or more central or graphics processing units in the encoding device).
[0097] The transforms as described herein can use directional decomposition transforms, e.g., Hadamard-based transforms. Both can comprise small kernels or matrices applied to flat encoding units of the residuals (i.e., 2x2 or 4x4 residual blocks). Further details on the transforms can be found, e.g., in patent applications PCT / EP2013 / 059847 or PCT / GB2017 / 052632, which are incorporated herein by reference. The encoder can select between different transforms to be used, e.g., between the size of the kernels to be applied.
[0098] The transforms can transform the residual information into four surfaces. For example, the transforms can yield the following components: average, vertical, horizontal and diagonal. As mentioned earlier in this disclosure, these components output by the transforms can be employed in such embodiments as coefficients to be quantized according to the described methods.
[0099] The quantization scheme can be used to produce the residual signal into quanta, such that certain variables can take on only certain discrete values.
[0100] The entropy encoding in this example can comprise run-length encoding (RLE), followed by processing the encoded output using a Huffman encoder. In some cases, only one of these schemes can be used when entropy encoding is required.
[0101] In summary, the methods and devices herein are based on an overall approach that is built via existing encoding and / or decoding algorithms (e.g., MPEG standards such as AVC / H.264, HEVC / H.265, and non-standard algorithms such as VP9, AV1, etc.), which are used as a baseline for a respective enhancement layer for different encoding and / or decoding methods. In contrast to using the block-based approach used in the MPEG family of algorithms, the example idea behind the overall approach is to encode / decode video frames in a hierarchical manner. Encoding frames in a hierarchical manner includes producing residuals for the full frame, and then producing residuals for extracted frames, etc.
[0102] Video compression residual data for full-size video frames can be referred to as LoQ-2 (e.g., 1920x1080 for HD video frames or higher for UHD frames), while video compression residual data for decimated frames can be referred to as LoQ-x, where x represents the number corresponding to the hierarchical decimation. In Figure 1 and 2 In the described examples of FIGS. 1-3, the variable x can have values 1 and 2 representing the first and second enhancement streams. Thus, there are 2 hierarchical levels that will produce compression residuals. Other naming schemes for the levels can also be applied without any functional changes (e.g., the 1st and 2nd enhancement streams described herein can alternatively be referred to as 1st and 2nd streams, denoting a count down from the highest resolution).
[0103] As noted above, the process can be applied in parallel to the coding units or blocks of the color components of the frame, as there are no inter-block dependencies. The encoding of each color component within the set of color components can also be performed in parallel (e.g., so that there are (number of frames)*(number of color components)*(number of coding units per frame) copy operations). It should also be noted that different color components can have different numbers of coding units per frame, e.g., the luminance (e.g., Y) component can be processed at higher resolution than the set of chroma (e.g., U or V) components when human vision can detect a change in luminance greater than a change in color.
[0104] Thus, as shown and described above, the output of the decoding process is the (optional) base reconstruction, as well as the original signal reconstruction at the higher level. This example is particularly well suited for producing encoded and decoded video at different frame resolutions. For example, the input signal 30 can be an HD video signal comprising frames at 1920x1080 resolution. In some cases, both the base reconstruction and the 2nd level reconstruction can be used by a display device. For example, in the case of network traffic, the 2nd level stream can be more severely interrupted than the 1st and base streams (as it can contain up to 4x the amount of data, with the downsampling reducing the dimension in each direction by 2). In this case, when traffic occurs, the display device can resume displaying the base reconstruction while the 2nd level stream is interrupted (e.g., while the 2nd level reconstruction is not available), and then resume displaying the 2nd level reconstruction when network conditions improve. A similar approach can be applied when the decoding device is subject to resource constraints, e.g., a set-top box performing system updates can have the base decoder 220 operating to output the base reconstruction, but can not have the processing capacity to compute the 2nd level reconstruction.
[0105] The encoding arrangement also enables the video distributor to distribute video to a heterogeneous set of devices; those with only the base decoder 720 view the base reconstruction, while those with the enhancement levels can view the higher quality 2-level reconstruction. In a comparative case, two complete video streams at separate resolutions would be needed to service the two sets of devices. Since the 2-level and 1-level enhancement streams encode residual data, the 2-level and 1-level enhancement streams can be more efficiently encoded, e.g., the distribution of residual data typically has most of its quality around 0 (i.e., no difference) and typically takes on small range values around 0. This can be especially the case after quantization. In contrast, complete video streams at different resolutions will have different distributions of non-zero mean or median values that need to be transmitted to the decoder at higher bit rates.
[0106] In examples described herein, residuals are encoded by the encoding pipeline. This can include transform, quantization, and entropy encoding operations. It can also include residual staging, weighting, and filtering. The residuals are then transmitted to a decoder, e.g., as L-l and L-2 enhancement streams, which can be combined with the base stream as a hybrid stream (or transmitted separately). In one case, a bit rate is set for a hybrid data stream that includes the base stream and two enhancement streams, and then different adaptive bit rates are applied to the individual streams based on the data being processed to meet the set bit rate (e.g., a high quality video perceived by a low level artifact can be constructed by adaptively assigning bit rates to different individual streams, even at a frame by frame level, such that constrained data can be used by the individual streams that are most perceptually impacted, which can change as the image data changes).
[0107] A set of residuals as described herein can be viewed as sparse data, e.g., in many cases there is no difference for a given pixel or region, and the resulting residual value is zero. When looking at the distribution of residuals, much of the probability mass is assigned to small residual values that are located close to zero, e.g., certain video values for -2, -1, 0, 1, 2, etc. occur most frequently. In certain cases, the distribution of residual values is symmetric or approximately symmetric about 0. In certain test video cases, the distribution of residual values (e.g., symmetrically or approximately symmetrically) about 0 was found to be similar in shape to a logarithmic or exponential distribution. The exact distribution of residual values can depend on the content of the input video stream.
[0108] Residuals can themselves be treated as two-dimensional images, e.g., difference magnitude images. In this way, one can see that the sparsity of the data involves features such as "dots", small "lines", "edges", "corners", etc. that are visible in the residual image. It has been found that these features are often not fully correlated (e.g., spatially and / or temporally). The features have characteristics that are different from the characteristics of the image data from which they are derived (e.g., pixel characteristics of the original video signal).
[0109] Because of the different nature of the residuals from the nature of the image data from which they are derived, it is generally not possible to apply standard encoding methods, such as those found in traditional Moving Pictures Expert Group (MPEG) encoding and decoding standards. For example, many comparison schemes use larger transforms (e.g., transforms of larger regions of pixels in a normal video frame). Due to the nature of the residuals, for example as described above, using these larger transforms on the residual image would be very inefficient. For example, it would be very difficult to encode a small dot in a residual image using a large block designed for a region of a normal image.
[0110] Certain examples described herein address these problems by instead using smaller and simple transform kernels (e.g., 2x2 or 4x4 kernels as presented herein - directional decomposition and directional decomposition square). The transforms described herein can be applied using Hadamard matrices (e.g., 4x4 matrices for flattening 2x2 encoding blocks or 16x16 matrices for flattening 4x4 encoding blocks). This moves in a different direction than comparative video encoding methods. Applying these new methods to residual blocks yields compression efficiency. For example, certain transforms yield uncorrelated coefficients (e.g., in space) that can be efficiently compressed. While correlation between coefficients can be exploited, for example for lines in a residual image, these can yield encoding complexity that makes it difficult to implement on legacy and low resource devices, and these often yield other complex artifacts that need to be corrected. Preprocessing residuals by setting certain residual values to 0 (i.e., not forwarding these for processing) can provide a controllable and flexible way to manage bit rate and stream bandwidth, as well as resource usage.
[0111] Exchanging information
[0112] As described above, the present application contemplates the principle of exchange of useful information between two (or more) encoders encoding the same content (or portions thereof), in the form of a joint module or a module implemented in one of the encoders but producing information that can be useful to the other. For example, an encoder can be Figure 6 a base encoder and an enhancement level encoder of the architecture of Figure 3 an encoder of each tier or tier index (or residual quality level) of the architecture of
[0113] It will of course be understood that there are various possible mechanisms to exchange information between multiple encoders or decoders.
[0114] In examples, information can be passed in the form of data structures to memory locations in the shared memory space where the information is stored, using a common API or as pointers. Furthermore, information can be exchanged between two encoder or decoder modules via Supplemental Enhancement Information, SEI (https: / / mpeg.chiariglione.org / tags / sei-messages, accessed April 16, 2019). Another mechanism can be passed as metadata in the bitstream, or via an API.
[0115] Depending on how the information is implemented (e.g., above as modules), the level of integration of the encoders, the allowed latency, or the storage and memory bandwidth; this information can be passed per group of pictures, per picture, per group of pixels, or per pixel. Similarly, in examples of Figure 3
[0116] Based on the information exchanged between encoders or decoders, each encoder or decoder can adapt the encoding parameters to improve the encoding operation. With such information exchange, the adaptation can be coordinated to improve the overall efficiency or quality or provide a balance across the hierarchy or layers of the hierarchy. It can be seen that each level of the hierarchy provides a specific benefit, and thus by balancing the parameters at each level, an overall improvement can be made.
[0117] In example implementations using the principles of Figure 6 Figure 3 In example implementations using the principles of
[0118] In another example, where a base encoder operates on an input video and an enhancement layer takes the output of the base encoder and operates on it (in conjunction with the input video, e.g., by upsampling or reconstructing the video and producing a set of residuals), the base encoder can include a message in the stream picked up by the enhancement layer, or can send a separate message to the enhancement layer, allowing the enhancement layer to act on (i.e., the message contains information or a pointer to the information). The message can also contain information allowing the enhancement to synchronize the exchange of information with the stream / frame, etc., depending on the granularity of the information.
[0119] Examples of information that can be exchanged between multiple encoding or decoding operations are provided below.
[0120] In a first example, the user can enter information that can be exchanged by the encoder, such as exchanged information about the type of content. For example, if the encoded data represents live sports content, it can be expected that there will be a lot of temporal activity, and in this sense, as lower frequency artifacts will be more apparent, it is determined that more bits are provided to the base encoder. In another example, if the content contains graphics, high frequencies become more important and more bits can be given to the enhancement layer. Such coordination between layers provides improved quality at reconstruction.
[0121] In certain examples, the user can insert information to be exchanged via the user interface. This information is received by a general controller that sends it simultaneously to the base and enhancement encoders / decoders, or the entire encoder can be controlled by the base encoder, in which case the information is passed from the base encoder to the enhancement layer.
[0122] In another example, in the case where the encoder contains multiple enhancement layers, such as a correction and enhancement layer, if the content is time critical and it can be determined that the base layer can be more strained, it can be decided to place more bits in the correction layer, especially in the case where high frequency details are less apparent in fast moving content. Otherwise, if there is a lot of graphics, it can be decided to allocate more bits to the enhancement layer to save more high frequency details.
[0123] Information collected from the host device can also be exchanged or passed between the encoders. Examples include device cameras, motion sensors, temperature sensors, etc. This information can be very useful to set both the base and enhancement encoders. The processing of this information can be implemented as a separate module that then passes the information to the encoders or can have been implemented in one of the encoders, and in this case, the host encoder can pass the information to the enhancement encoder.
[0124] A content analysis module can also be used to provide information to be exchanged. Various pre-processing modules can be present to analyze the picture to be encoded as a separate module or as a component of one of the encoders. Examples include perceptual information describing the regions of the image, where the human visual system (HVS) is taken into account, encoding artifacts are less apparent; motion information, including temporal activity of the frame, per group of pixels (block) information, such as motion estimation in the form of motion vectors or per pixel information, such as optical flow; and any other information in any other form about the estimated complexity of the frame to be encoded, the entropy of the frame, the estimated number of bits needed, etc.
[0125] In another example, the information exchanged between encoders can include the encoding decisions taken by any of the encoders. The encoders can pass to each other any information about the decisions they made during the encoding process. This can include the prediction type used for a frame (I, P or B frame in h264 or an autonomous frame predicted from a previously displayed frame or a frame predicted from both a previous and a subsequent frame). The decisions can also include the quantization level used for a frame or decisions made per group of pixels (motion vectors, quantization steps used, mode decisions, etc.). This information can be highly beneficial for an encoder receiving the information as it can be used for example to decide how many bits to use to correct the error of the base encoder and how many to use for enhancement, what type of upsampler to use and / or how to spatially spread the bit budget.
[0126] In a final example, the information exchanged between encoders can include statistics about the encoded content. The encoders can exchange any information collected about the group of pictures, pictures, group of pixels or pixels they have encoded. For example, how many bits each picture in a particular group of pictures took can assist the decision of the rate control module of an encoder receiving the information. In another statistics example, a target video quality metric computed in one of the encoders can help the other encoders in bit budget decisions and in estimating the final picture quality. Any statistics on the mode used for each group of pixels (intra or inter coded) can be used to assist the decision of the prediction mode used when an encoder receives the information.
[0127] As noted above, the information exchanged between encoders can be exchanged at different levels of granularity. For example: information about the whole video; information per group of pictures, i.e. a segment of the video to be encoded; information about a picture to be encoded; information about a group of pixels (block, tile of pixels) to be encoded; and / or information per pixel.
[0128] The granularity chosen can provide specific benefits, for example it can be useful information exchanged to help give more bits to specific parts of the image. In the case where the information exchanged is from a device camera, it can represent focus and exposure information, telling which parts of the image contain details. Modern cameras also have depth of field information. This information can be used to provide details in areas of the image. In addition, algorithms that can be implemented in the camera or base encoder can provide useful information to be exchanged. For example, face detection, region of interest extraction and background extraction can all be used to understand which details are important and which are not and in this way can help decide where to put more bits. Thus, the encoders (and their parameters) can be adapted based on the algorithms used between the levels of the hierarchy and the information passed. The coordination between the levels can be provided by the shared information but the information from the algorithms themselves is not shared but is based on the information provided or to be provided for the adaptation.
[0129] Depending on how the encoders have been implemented (specialized hardware, software on a CPU, GPU) and the resources available (memory management, processing speed, etc.), different types of information exchange will be able to be exchanged. For example, if shared memory is limited and one encoder is implemented in hardware and the other in software, information can not be exchanged for each pixel or even each group of pixel information, but statistics for each picture can be collected, which will be easier to pass from one encoder to the other. Sending statistics instead of raw per-pixel information is one way to address these challenges. While the precision can be lower, it can be sufficient in many cases.
[0130] The methods and processes described herein can be embodied as code (e.g., software code) and / or data at both the encoder and the decoder, for example, implemented in a streaming server or client device or a client device decoding from data storage. The encoder and decoder can be implemented in hardware or software as is well known in the art of data compression. For example, hardware acceleration using a specially programmed graphics processing unit (GPU) or a specially designed field-programmable gate array (FPGA) can provide certain efficiencies. For completeness, such code and data can be stored on one or more computer-readable media, which can include any device or medium that can store code and / or data for use by a computer system. When a computer system reads and executes the code and / or data stored on the computer-readable media, the computer system performs the methods and processes embodied as data structures and code stored within the computer-readable storage media. In certain embodiments, one or more of the steps of the methods and processes described herein can be performed by a processor (e.g., a processor of a computer system or data storage system).
[0131] In general, any of the functionality described in this text or shown in the drawings can be implemented using software, firmware (e.g., fixed logic circuitry), programmable or non-programmable hardware, or a combination of these implementations. In general, the term “component” or “functionality” as used herein refers to software, firmware, hardware, or a combination of these. For example, in the case of software implementations, the term “component” or “functionality” can refer to program code that, when executed on one or more processing devices, carries out the specified tasks. The shown separation of components and functionality into distinct units can reflect any actual or conceptual physical grouping and allocation of such software and / or hardware and tasks.
Claims
1. A method for encoding a signal, the method comprising: Receive input signals; A first encoding operation is applied to the input signal to generate a first encoded stream representing the input signal of a first quality level; as well as, A second encoding operation is applied to the input signal to generate a second encoded stream representing an input signal of a second quality level, which is higher than the first quality level. Wherein, the first encoded stream and the second encoded stream are combined at the decoder, the first encoding operation and the second encoding operation are encoding operations of a hierarchical encoding scheme, and the second encoded stream is encoded separately without considering the first encoded stream; Information is exchanged between the first encoding operation and the second encoding operation, wherein the exchange step includes sending at least a portion of the information from the second encoding operation to the first encoding operation; and The first encoding operation is adjusted based on at least some of the information.
2. The method according to claim 1, wherein the method further comprises adjusting the first or second encoding operation or both based on the information.
3. The method according to claim 1 or 2, wherein the hierarchical coding scheme comprises: The basic encoded signal is generated by feeding the input signal in a downsampled form to the encoder. The first residual signal is generated through the following operations: Obtain the decoded form of the basic encoded signal; as well as The first residual signal is generated using the difference between the decoded form of the base encoded signal and the downsampled form of the input signal; as well as, The first residual signal is encoded to generate a first encoded residual signal.
4. The method of claim 3, further comprising: The second residual signal is generated through the following operation: The first encoded residual signal is decoded to generate a first decoded residual signal; The decoded pattern of the base coded signal is corrected using the first decoded residual signal to produce a corrected decoded pattern; Upsample the corrected decoded format; as well as The second residual signal is generated by using the difference between the upsampled pattern and the input signal; The method further includes: The second residual signal is encoded to generate a second encoded residual signal. The base coded signal, the first coded residual signal, and the second coded residual signal include encoding of the input signal.
5. The method according to claim 1 or 2, wherein the hierarchical coding scheme comprises: A basic encoded signal is generated by feeding an input signal in a downsampled form to the encoder, the downsampled form undergoing one or more downsampling operations. One or more residual signals are generated through the following operations: The output of each downsampling operation is upsampled to generate one or more upsampled signals; as well as The one or more residual signals are generated using the difference between each upsampled signal and the input to the corresponding downsampled operation; as well as, The one or more residual signals are encoded to generate one or more encoded residual signals.
6. The method according to claim 1 or 2, wherein the hierarchical coding scheme comprises: The basic encoded signal is generated by feeding the input signal to the encoder in a downsampled form, which undergoes multiple sequential downsampling operations. as well as The encoded first residual signal is generated through the following operations: Upsample the input signal in the downsampled form; The first residual signal is generated using the difference between the upsampled form of the downsampled pattern and the input to the last downsampled operation of the plurality of sequential downsampled operations; and, The first residual signal is encoded; The second residual signal is generated through the following operation: Upsample the sum of the first residual signal and the output of the last downsampling operation of the plurality of sequential downsampling operations; The second residual signal is generated using the difference between the upsampled sum and the input to the aforementioned downsampled operation; and, The second residual signal is encoded.
7. The method of claim 1 or 2, wherein the step of exchanging information includes sending information having metadata in a stream.
8. The method of claim 1 or 2, wherein the step of exchanging information includes embedding the information in a stream.
9. The method of claim 1 or 2, wherein the step of exchanging information includes sending information using an application programming interface (API).
10. The method of claim 1 or 2, wherein the step of exchanging information includes sending a pointer to a shared memory space.
11. The method of claim 1 or 2, wherein the step of exchanging information includes exchanging information using supplementary enhancement information.
12. The method of claim 1 or 2, wherein the exchanged information includes encoding parameters for modifying the encoding operation.
13. The method of claim 1 or 2, wherein the exchanged information comprises one or more selected from the group consisting of: User input information; Metadata describing the content of the input video; Host device information; Content analysis information; The perceptual information describing the regions of an image, where the human visual system (HVS) is taken into account, makes the encoded artifacts less noticeable; Sports information; The complexity of the frame to be encoded; Frame entropy; The estimated number of bits required; Frame information; Decisions made during encoding or decoding; The type of prediction used for the frame to be decoded / encoded; The quantization level used for the frame to be decoded / encoded. The decisions made by each group of pixels; Statistical data; Target video quality metrics; and, Rate control information.
14. The method of claim 1 or 2, wherein the signal is video and wherein the step of exchanging information is performed according to one or more of each video, each set of pictures, each segment, each picture, each set of pixels, and each pixel.
15. A method for decoding a signal, the method comprising: Receive the first encoded signal and the second encoded signal; The first decoding operation is applied to the first encoded signal to generate a first output signal representing the input signal of the first quality level; The second decoding operation is applied to the second encoded signal to generate a second output signal representing the input signal of a second quality level, wherein the second quality level is a quality level higher than the first quality level; as well as, The first output signal and the second output signal are combined to reconstruct the input signal. Wherein, the first decoding operation and the second decoding operation are decoding operations of a hierarchical coding scheme, the second encoded signal is decoded separately without considering the first encoded signal, and the method further includes: Information is exchanged between the first decoding operation and the second decoding operation, wherein the exchange step includes sending at least a portion of the information from the second encoding operation to the first encoding operation; and The first encoding operation is adjusted based on at least some of the information.
16. The method of claim 15, the method further comprising adjusting the first or second decoding operation or both based on the information.
17. The method according to claim 15 or 16, wherein the hierarchical coding scheme comprises: Receive a basic encoded signal and instruct the decoding of the basic encoded signal to generate a basic decoded signal; Receive the first encoded residual signal and decode the first encoded residual signal to generate the first decoded residual signal; The first decoded residual signal is used to correct the basic decoded signal to generate a corrected form of the basic decoded signal; The corrected form of the basic decoded signal is upsampled to generate an upsampled signal; Receive the second encoded residual signal and decode the second encoded residual signal to generate the second decoded residual signal; as well as The upsampled signal is combined with the second decoded residual signal to generate a reconstructed form of the input signal.
18. The method of claim 15 or 16, wherein the first coded signal and the second coded signal respectively comprise first and second component sets, the first component set corresponding to an image resolution lower than the second component set, the method comprising: For each of the first and second component sets: The component set is decoded to obtain the decoded set. The method further includes: Upgrade the decoded first component set to increase the corresponding image resolution of the decoded first component set to be equal to the corresponding image resolution of the decoded second component set, and The decoded first and second component sets are combined to produce a reconstructed set.
19. The method of claim 18, further comprising receiving one or more additional component sets, each of the one or more additional component sets corresponding to an image resolution higher than the second component set, and each of the one or more additional component sets corresponding to progressively higher image resolutions, the method comprising decoding the component set for each of the one or more additional component sets to obtain a decoded set, the method further comprising, for each of the one or more additional component sets, in ascending order of corresponding image resolutions: Upgrade the reconstructed set with the highest corresponding image resolution so that the corresponding image resolution of the reconstructed set is increased to be equal to the corresponding image resolution of the other component set, and The reconstructed set is combined with the other component set to produce another reconstructed set.
20. The method of claim 15 or 16, wherein the step of exchanging information includes sending information having metadata in a stream.
21. The method of claim 15 or 16, wherein the step of exchanging information includes embedding the information in a stream.
22. The method of claim 15 or 16, wherein the step of exchanging information includes sending information using an API.
23. The method of claim 15 or 16, wherein the step of exchanging information includes sending a pointer to a shared memory space.
24. The method of claim 15 or 16, wherein the step of exchanging information includes exchanging information using supplementary enhancement information.
25. The method of claim 15 or 16, wherein the exchanged information includes decoding parameters for modifying the decoding operation.
26. The method of claim 15 or 16, wherein the exchanged information comprises one or more selected from the group consisting of: User input information; Metadata describing the content of the input video; Host device information; Content analysis information; The perceptual information describing the regions of an image, where the human visual system (HVS) is taken into account, makes the encoded artifacts less noticeable; Sports information; The complexity of the frame to be encoded; Frame entropy; The estimated number of bits required; Frame information; Decisions made during encoding or decoding; The type of prediction used for the frame to be decoded / encoded; The quantization level used for the frame to be decoded / encoded. The decisions made by each group of pixels; Statistical data; Target video quality metrics; and, Rate control information.
27. The method of claim 15 or 16, wherein the signal is video and wherein the step of exchanging information is performed according to each video, each set of pictures, each segment, each picture, each set of pixels, and each pixel.
28. An encoding device configured to perform the method according to any one of claims 1 to 14.
29. A decoding device configured to perform the method according to any one of claims 15 to 27.
30. A computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 27.
Citation Information
Patent Citations
Decomposition of residual data during signal encoding, decoding and reconstruction in a tiered hierarchy
WO2013171173A1
Hybrid backward-compatible signal encoding and decoding
WO2014170819A1
Video compression using differences between a higher and a lower layer
WO2018046940A1
Methods and apparatuses for encoding and decoding a bytestream
WO2019111004A1
Apparatus and methods for video compression using multi-resolution scalable coding
US20170223368A1