Applications of tiered improvements in embedded signaling for backward compatibility and super-resolution signaling

By embedding signal processing information into the encoded data stream and selectively performing signal processing operations, the problem of insufficient signal reconstruction quality in the prior art is solved, signal enhancement and backward compatibility of decoders are achieved, and the detail resolution of image and video signals is improved.

CN114788283BActive Publication Date: 2026-03-10V NOVA INT LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-02
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing hierarchical coding formats struggle to effectively utilize signal processing information for enhancement during signal decoding, resulting in insufficient reconstructed signal quality. This is especially problematic when resources are limited, as they cannot achieve backward-compatible signal enhancement.

Method used

By embedding signal processing information, such as adaptive filters and neural network methods, into the encoded data stream, signal processing operations are selectively performed to enhance higher resolution levels. Flexible signal enhancement is achieved by utilizing embedded transform coefficients and supplementary enhancement information messages to transmit signal processing parameters.

Benefits of technology

It improves signal reconstruction quality, enables signal enhancement when resources allow, and maintains backward compatibility, ensuring that older decoders can decode normally and new decoders can apply additional signal processing to improve the detail resolution of image or video signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114788283B_ABST
    Figure CN114788283B_ABST
Patent Text Reader

Abstract

Some examples described herein relate to methods for encoding and decoding signals. Some examples relate to the control of signal processing operations performed at the decoder. These operations may include optional signal processing operations to provide an enhanced output signal. For video signals, the enhanced output signal may include a so-called "super-resolution" signal, such as a signal with improved detail resolution compared to a reference signal. Some examples described herein provide communication for enhancement operations (e.g., so-called super-resolution modes) within user data of one or more hierarchical encoding and decoding schemes. The user data may be embedded within values ​​of the enhancement stream, for example, replacing one or more values ​​with a predefined set of transform coefficients, and / or embedded within supplementary enhancement information messages. The user data may have a defined syntax including header and payload portions. The syntax may differ for different data frames; for example, for video encoding, an instant-decoded refresh frame may carry different information than a non-instant-decoded refresh frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to methods for processing signals, such as by means of non-limiting examples of video, images, hyperspectral images, audio, point clouds, 3DoF / 6DoF, and volumetric signals. The processed data may include, but is not limited to, acquiring, deriving, encoding, outputting, receiving, and reconstructing signals within the context of a hierarchical (layer-based) coding format, wherein the signals are decoded into subsequent layers of higher quality, thereby utilizing and combining subsequent layers of reconstructed data (“ladders”). Different layers of the signal may be encoded using different coding formats by means of different elementary streams that may or may not be multiplexed in a single bitstream (e.g., by means of non-limiting examples, conventional single-layer DCT-based codecs, SO / IEC MPEG-5 Part 2 Low Complexity Enhanced Video Coding SMPTE VC-62117, etc.). Background Technology

[0002] In hierarchical coding formats, such as ISO / IEC MPEG-5 Part 2 LCEVC (hereinafter referred to as "LCEVC") or SMPTE VC-62117 (hereinafter referred to as "VC-6"), the signal is decomposed into multiple "steps" (also known as "hierarchical levels") of data, each step corresponding to a "quality level" ("LoQ") of the signal, ranging from the highest step at the original signal's sampling rate to the lowest step at a sampling rate typically lower than that of the original signal. In a non-limiting instance, when the signal is frames of a video stream, the lowest step can be a thumbnail of the original frame, or even just a single pixel. Other steps contain information about corrections to be applied to the reconstructed version to produce the final output. Steps can be based on residual information, such as the difference between the pattern of the original signal at a particular quality level and the reconstructed pattern of the signal at the same quality level. The lowest step may not include residual information, but may include the lowest sampled value of the original signal. The decoded signal at a given quality level is reconstructed by first decoding the lowest quality level (thus reconstructing the signal at the first lowest quality level), then predicting the reproduction of the signal at the next higher quality level, then decoding the corresponding second quality level of the reconstructed data (also referred to as the "residual data" at the second quality level), then combining the predicted and reconstructed data to reconstruct the reproduction of the signal at the next higher quality level, and so on, until the given quality level is reconstructed. The reconstructed signal may include the decoded residual data, which is used to correct the pattern at a specific quality level derived from the pattern of the signal at lower quality levels. Different encoding formats can be used to encode different quality levels of the data, and different quality levels can have different sampling rates (e.g., resolution for image or video signals). Subsequent quality levels can refer to the same signal resolution (i.e., sampling rate) or to progressively higher signal resolutions.

[0003] U.S. Patent 8,948,248 B2 discloses a decoder for decoding a first set of data. The first set of decoded data is used to reconstruct a signal according to a first quality level. The decoder further decodes a second set of data and identifies an upsampling operation specified by the second set of decoded data. The decoder applies the upsampling operation identified in the second set of decoded data to the reconstructed signal at the first quality level to reconstruct the signal at a second, higher quality level. To enhance the reconstructed signal, the decoder retrieves residual data from the second set of decoded data. The residual data indicates how the reconstructed signal at the second quality level should be modified after applying the upsampling operation as described above. As specified by the residual data, the decoder then modifies the reconstructed signal at the second quality level.

[0004] A switchable scalable video coding (SVC) upsampling filter is described in a proposal submitted to the ISO / IEC MPEG and ITU-T VCEG Joint Video Group (ISO / IEC JTC1 / SC29 / WG11 and ITU-T SG16 0.6 - document: JVT-R075 associated with the 18th meeting held in Bangkok, Thailand, January 14-20, 2006). In this proposal, different upsampling filters are selected based on either a quantization parameter (QP) threshold or a rate distortion value sent to the decoder.

[0005] The paper "Sample Adaptive Offset in the HEVC Standard," published in December 2012 in IEEE Transactions on Circuits and Systems for Video Technology, Volume 22, Issue 12, by Chih-Ming Fu et al., describes the in-loop sample adaptive offset (SAO) adaptive filtering technique used in HEVC. The sample adaptive offset parameters for each coding tree unit (CTU) can be interleaved into the slice data. These parameters can be applied to each coding tree unit. Summary of the Invention

[0006] Aspects and variations of the invention are set forth in the appended claims.

[0007] Certain unclaimed aspects are further described in the detailed embodiments described below.

[0008] In one described aspect, a method for decoding a signal includes: obtaining an encoded data stream, the encoded data stream being encoded by an encoder according to a hierarchical format based on a hierarchy; analyzing the encoded data stream to determine signal processing information conveyed by the encoder; and reconstructing a higher resolution level of the signal from a lower resolution level of the signal, including selectively performing one or more signal processing operations to enhance the higher resolution level based on the determined signal processing information. According to this method, the functionality of the standardized hierarchical format based on a hierarchy can be flexibly extended to include more advanced signal processing operations, such as adaptive filters and neural network methods.

[0009] In some of the described instances, at least a portion of the data corresponding to the signal processing information is embedded in one or more values ​​received in one or more coded data layers transmitted within the coded data stream, wherein the values ​​are associated with transform coefficients of elements processed to derive the signal during the decoding process. These values ​​may be values ​​of predefined transform coefficients within a different set of transform coefficients generated by the coded transform (e.g., A, H, V, and D values ​​for a 2×2 Hadamard transform). Embedding the signal processing information in this manner allows for advantageous compression of the information using normalization methods for residual data values ​​(e.g., run-length and prefix coding (alternatively also called Huffman coding)). The signal processing information embedded in this manner also allows for the transmission of local parameters associated with a particular coding unit or data block along with the transform data values ​​of the coding unit or data block. Transform coefficients (e.g., H or HH coefficients) that have a negligible effect on the reconstructed signal may be selected. Furthermore, it is unlikely that the absence of transform coefficient values ​​at the beginning of a picture frame in a video or audio track will be noticeable.

[0010] In other or complementary instances, at least a portion of the data corresponding to the signaling processing information is encoded within the Supplemental Enhancement Information (SEI) message. The SEI message provides a simple way to access global signaling information without interfering with regular processing according to defined encoding standards.

[0011] In yet another example, the signal processing information may be determined at least in part based on a predefined set of values ​​for configuration data for the signal, which configures one or more signal processing operations not intended to enhance the signal processing operations at the higher resolution level. In this way, parameters for non-standard enhancements can be conveyed using data fields as defined in standardized signal coding methods.

[0012] In some instances, the one or more signal processing operations are selectively performed before adding residual data for a higher-resolution hierarchy of the signal. This can be considered "in-loop" enhancement. This allows the residual data to correct rare artifacts generated by the one or more signal processing operations. Therefore, less-than-perfect signal processing operations that were previously unusable due to occasional image quality degradation (e.g., producing good results less than 100% of the time) become available. In this case, one or more signal processing operations can be performed within the frame decoding loop for the hierarchical format.

[0013] In some instances, the one or more signal processing operations provide a super-resolution signal, i.e., an upsampling operation that enhances or compares to upsampling operations (e.g., those operations that at least copy lower-level data to multiple pixels in higher-level data) to provide improved detail. In some instances, the one or more signal processing operations are implemented as part of an upsampling operation that generates the higher-resolution level of the signal from the lower-resolution level of the signal.

[0014] In some instances, selectively performing one or more signal processing operations to enhance the higher resolution level includes: determining operating parameters of the decoder performing the decoding; performing the one or more signal processing operations in response to a first set of operating parameters to enhance the higher resolution level using signal processing parameters within the determined signal processing information; and omitting the one or more signal processing operations in response to a second set of operating parameters. This allows for the performance of (optional) enhancement operations based on local decoder conditions. In one instance, the method includes: determining a resource usage metric of the decoder performing the decoding; comparing the resource usage metric to a resource usage threshold; performing the one or more signal processing operations to enhance the higher resolution level based on the determined signal processing information in response to the comparison indicating no restriction on resource usage of the decoder; and omitting the one or more signal processing operations during the reconstruction in response to the comparison indicating a restriction on resource usage of the decoder. For example, many enhancement operations are more resource-intensive than comparing default or standard decoding methods; by applying these methods, they are applied only if the decoder has the resources available to apply them. This provides a simplified implementation of a relatively complex adaptive enhancement system.

[0015] In one example, the method includes: using the determined signal processing information to identify signal processing operations used to enhance the higher resolution level; determining whether a decoder performing the decoding is capable of implementing the identified signal processing operations; ignoring the determined signal processing information in response to the decoder's inability to implement the identified signal processing operations; and performing the determined signal processing operations parameterized by the determined signal processing information in response to the decoder's ability to implement the identified signal processing operations. In this way, backward compatibility can be maintained. For example, older legacy decoders constructed according to existing decoding standards can ignore the signal processing information and still decode data according to the standard; while newer decoders can modularly implement the latest available advancements in signal enhancement while remaining compatible with the standard. For example, decoders of different brands and models can implement different enhancement operations, which can be flexibly communicated and applied while maintaining compatibility with encoding and broadcasting systems.

[0016] In some instances, the one or more signal processing operations include a sharpening filter applied in addition to the upsampling operation used for the reconstruction, the upsampling operation generating the higher resolution level of the signal from the lower resolution level of the signal. In these instances, the determined signal processing information may indicate at least one coefficient value of the desharpening mask, and this coefficient value may be applicable to local content (or global application). In one instance, the determined signal processing information indicates the center integer coefficient value of the desharpening mask. The desharpening mask can reduce the bit rate required for residual data in higher levels by providing an upsampled signal that is closer to the original signal. Improvements in the visibility of the video signal can also be provided even when residual data is unavailable (e.g., during network congestion).

[0017] In one instance, one or more signal processing operations form part of a cascade of linear operations applied to data from a lower resolution level of the signal. This cascade of linear operations may include an increase in a predicted average modifier.

[0018] In some instances, the one or more signal processing operations include a neural network upsampler. The neural network upsampler can be a small and efficient implementation capable of operating at a real-time signal frame rate. The methods described herein allow for flexible communication of different neural network configurations, thereby allowing for upgrades to decoder functionality during use. In one case, the determined signal processing information indicates the coefficient values ​​of one or more linear layers of a convolutional neural network. In this case, the neural network can adaptively upsample the coding units of the signal based on the local context.

[0019] In some instances, the one or more signal processing operations include an additional upsampling operation applied to the output of the final layer of residual data within the hierarchical format. For example, the methods described herein allow for both modifications to the normalized upsampling process and the provision of upgraded or “additional” upsampling for communication purposes.

[0020] In some instances, the method includes applying dithering to the output of the reconstructed higher-resolution level after reconstructing the higher-resolution level. Applying dithering to the higher-resolution signal is advantageous for obtaining optimal visual results. In the current case, this can be achieved by switching the enhancement operation under normalized dithering operation.

[0021] In some instances, the signal processing information includes header data and payload data, and the method includes: analyzing a first set of values ​​received in one or more coded data layers to extract the header data; and analyzing a second set of subsequent values ​​received in one or more coded data layers to extract the payload data. Therefore, the signal processing information can be defined according to a shared syntax applicable to both embedded transform coefficients and SEI messages in user data. This syntax can be extended by linking additional non-enhanced user data before or after the described signal processing information, i.e., additional user data can be embedded in a third set of values ​​following the second set of values. In some cases, the header data and payload data can be separated by different signaling methods (e.g., the header data can be sent via SEI messages, and the payload data can be sent via embedded user data). Header data can be provided to configure global aspects of signal processing, while payload data can be used to parameterize local (e.g., intra-frame) processing. In some instances, depending on the enhancement mode identified in the header data, analysis of the second set of values ​​is performed selectively; for example, payload data can be ignored and / or omitted if no enhancement is specified, thereby allowing limited interruptions to transform coefficient processing. In some instances, the embedded data is set as a series of n-bit values, such as 2-bit, 6-bit, or 8-bit (byte) values. In some instances, the signal includes a video signal, and a first header structure is used for Instant Decoding Refresh (IDR) picture frames, while a second header structure is used for non-IDR picture frames, wherein the second header structure indicates whether the configuration indicated in the first header structure has changed. Therefore, the signal processing information accompanying the IDR picture frame can be used to configure enhancement operations for multiple future frames, where frame-by-frame adaptation can be signaled within non-IDR frames. In some cases, the payload data of the non-IDR picture frame includes values ​​that instantiate the changed configuration indicated in the first header structure. This allows for efficient signaling of variable values.

[0022] In one set of examples, the signal processing information is embedded in one or more values ​​received in an encoded data layer that provides transformed residual data of the signal at the lower resolution level. For example, the signal processing information may be embedded within a Level 1 encoded stream for LCEVC. This stream can be received more reliably (due to its smaller size), and allows for a longer timeframe to be configured for enhancement operations prior to decoding of the Level 2 encoded stream (which is typically received before the Level 2 encoded stream).

[0023] In some instances, one or more signal processing operations may be performed on the data output from the frame decoding loop for the hierarchical format. This so-called "out-of-loop" enhancement may be advantageous when enhancement can be performed on the fully reconstructed signal (e.g., advanced statistical jitter or methods that utilize information from the fully reconstructed signal, such as post-processing filters and "additional" upgrades).

[0024] In some instances, the hierarchical format is either MPEG-5 Part 2 Low Complexity Enhanced Video Coding (“LCEVC”) or SMPTE VC-6 ST-2117.

[0025] A decoder can be configured to perform the methods described in this paper.

[0026] According to another described aspect, a method for encoding a signal is provided. The method includes: encoding a lower-resolution layer in a hierarchical, layer-based format; encoding a higher-resolution layer in a hierarchical, layer-based format, the higher-resolution layer being encoded using data generated during the encoding of the lower-resolution layer; and generating an encoded data stream using the output of the encoding of the lower-resolution layer and the output of the encoding of the higher-resolution layer; the method further includes: determining signal processing information for performing one or more signal processing operations to enhance data within the higher-resolution layer, the one or more signal processing operations being performed to reconstruct a portion of the higher-resolution layer using the data generated during the encoding of the lower-resolution layer; and encoding the signal processing information as a portion of the encoded data stream.

[0027] The above method thus provides a complementary encoding method that can be performed at the encoder to generate signal processing information analyzed and determined in the decoding method. The one or more signal processing operations may form part of the encoder upsampling operation and / or encoder post-processing following upsampling.

[0028] In some decoding methods, for example, the signal processing information replaces one or more quantization symbols of predefined transform coefficients within one or more of the lower resolution level and the higher resolution level. The predefined transform coefficients include one of a plurality of transform coefficients generated by transforming residual data within one or more of the lower resolution level and the higher resolution level. The signal processing information may replace one or more quantization symbols of the predefined transform coefficients within the lower resolution level. The one or more signal processing operations may include an optional set of signal processing operations comprising applying one or more of a sharpening filter and a convolutional neural network; and / or the one or more signal processing operations may include a set of cascaded linear filters sampling data from the lower resolution level to the higher resolution level, wherein the signal processing information includes parameters of at least one of the cascaded linear filters.

[0029] An encoder may also be provided to perform this encoding method.

[0030] Further features and advantages will become apparent from the following description, given by way of example only and with reference to the accompanying drawings. Attached Figure Description

[0031] Figure 1 A high-level schematic diagram of the hierarchical encoding and decoding process is shown;

[0032] Figure 2 A high-level schematic diagram of the hierarchical deconstruction process is shown;

[0033] Figure 3 An alternative high-level diagram of the hierarchical deconstruction process is shown;

[0034] Figure 4 A high-level schematic diagram is shown of the encoding process applicable to encoding the residuals of hierarchical outputs;

[0035] Figure 5 It shows the applicability to the source Figure 4 A high-level diagram illustrating the hierarchical decoding process that decodes each output level;

[0036] Figure 6 A high-level schematic diagram of the encoding process of hierarchical coding technology is shown; and

[0037] Figure 7 It shows the applicability to... Figure 6 A high-level diagram illustrating the decoding process of decoding the output;

[0038] Figure 8A A block diagram of an example decoding system with optional enhancement operations is shown;

[0039] Figure 8B A block diagram of another example decoding system with enhancement operations at two levels in a hierarchical decoding process is shown;

[0040] Figure 9 A high-level schematic diagram of several example decoders performing different reconstruction operations is shown;

[0041] Figure 10A A block diagram of the sampler on the example is shown;

[0042] Figure 10B A block diagram of an example conversion process for a floating-point on-sampler is shown;

[0043] Figure 11A and 11B A block diagram of an example enhanced upsampler is shown;

[0044] Figure 12A and 11B A block diagram of a sampler on an example neural network is shown;

[0045] Figure 13 A high-level schematic diagram showing the switching between normal and enhanced on-sampling operations based on an example is provided.

[0046] Figure 14 An example sharpening filter is shown; and

[0047] Figure 15 A block diagram of an example of a device according to an embodiment is shown. Detailed Implementation

[0048] Some examples described herein relate to methods for encoding and decoding signals. Processing data may include, but is not limited to, acquiring, exporting, outputting, receiving, and reconstructing data. Examples of the invention relate to the control of signal processing operations performed at a decoder. These operations may include optional signal processing operations to provide an enhanced output signal. For video signals, the enhanced output signal may include a so-called "super-resolution" signal, such as a signal with improved detail resolution compared to a reference signal. The reference signal may include video sequence encoding at a first resolution, and the enhanced output signal may include a video sequence decoding pattern at a second resolution higher than the first resolution. The first resolution may include the original resolution of the video sequence, such as the resolution at which the video sequence is acquired for encoding.

[0049] Some of the examples described herein provide signaling for enhancement operations (e.g., so-called super-resolution modes) within user data of one or more hierarchical encoding and decoding schemes. The user data may be embedded within values ​​of the enhancement stream, for example, replacing one or more predefined sets of transform coefficients with values, and / or embedded within supplementary enhancement information messages. The user data may have a defined syntax including header and payload portions. This syntax may differ for different data frames; for example, for video encoding, an instant-decoded refreshed image frame may carry different information than a non-instant-decoded refreshed image frame.

[0050] introduction

[0051] The examples described herein relate to signal processing. A signal can be viewed as a sequence of samples (i.e., a two-dimensional image, a video frame, a video field, an audio frame, etc.). In the description, the terms “image,” “picture,” or “plane” (meaning the broadest sense of a “hyperplane,” i.e., an array of elements with arbitrary dimensions and a given sampling grid) will generally be used to identify a digital reproduction of a signal sampled along a sequence of samples, where each plane has a given resolution for each of its dimensions (e.g., X and Y) and comprises a set of planar elements (or “elements,” or “pixels,” or display elements commonly referred to as “pixels” for two-dimensional images, commonly referred to as “volumetric cells” for volumetric images, etc.), characterized by one or more “values” or “settings” (e.g., by means of non-limiting examples, color settings in a suitable color space, settings indicating density levels, settings indicating temperature levels, settings indicating audio pitch, settings indicating amplitude, settings indicating depth, settings indicating alpha channel transparency levels, etc.). Each planar element is identified by a set of suitable coordinates representing the integer position of the element in the image sampling grid. The signal dimension may include only the spatial dimension (e.g., in the case of an image) or also include the temporal dimension (e.g., in the case of a signal that evolves over time, such as a video signal).

[0052] For example, the signal can be an image, an audio signal, a multi-channel audio signal, a telemetry signal, a video signal, a 3DoF / 6DoF video signal, a volumetric signal (e.g., medical imaging, scientific imaging, holographic imaging, etc.), a volumetric video signal, or even a signal with more than four dimensions.

[0053] For simplicity, the examples described herein generally refer to signals displayed as a defined 2D plane (e.g., a 2D image in a suitable color space), such as video signals. The terms "frame" or "field" will be used interchangeably with the term "image" to indicate a sample of a video signal over time: any concepts and methods described for video signals composed of frames (progressive video signals) can also be readily applied to video signals composed of fields (interlaced video signals), and vice versa. Although the embodiments described herein focus on image and video signals, those skilled in the art will readily understand that the same concepts and methods are also applicable to any other type of multidimensional signal (e.g., audio signals, volumetric signals, stereo video signals, 3DoF / 6DoF video signals, all-optical signals, point clouds, etc.).

[0054] Some of the hierarchical, layered formats described in this paper use corrections to the amount of variation (e.g., also in the form of “residual data” or simply “residuals”) to reconstruct (or even reconstruct without loss) a signal most similar to the original signal at a given quality level. The correction amount can be based on the fidelity of the predicted reproduction at a given quality level.

[0055] To achieve high-fidelity reconstruction, coding methods can upsample a signal from a lower-resolution reconstruction to a higher-resolution reconstruction. In some cases, different methods may be best suited for different signals; that is, the same method may not be ideal for all signals.

[0056] Furthermore, it has been established that nonlinear methods can be more efficient than more conventional linear kernels (especially separable linear kernels), but at the cost of increased processing power requirements. In most cases, due to processing power limitations, various sizes of linear upsampling kernels (e.g., bilinear, bicubic, multi-lobe Lanczos, etc.) have been used to date, but in recent years even more sophisticated nonlinear techniques, such as the use of convolutional neural networks in VC-6, have shown to produce higher quality initial reconstructions, thereby reducing the entropy of residual data added for high-fidelity final reconstructions.

[0057] In formats such as LCEVC, it is possible to send the coefficients of the upsampling kernel to be used to the decoder before adding the "prediction residual" to the LCEVC nonlinearity. Simultaneously, it is proposed to extend the coding standard's ability to embed reconstructed metadata into the encoded stream, which is ignored by an unobservable decoder but processed by a decoder capable of decoding the user data.

[0058] In some instances, signal processing information is transmitted using one or more of embedded transform coefficient values, supplementary enhancement information (SEI) messages, and custom configuration settings. In this way, signaling is optional and backward compatibility is maintained (e.g., a decoder compliant with LCEVC or VC-6 standards but unable to implement additional signal processing may simply ignore the additional signaling and decode as usual).

[0059] The example method described in this paper uses user data to transmit information to a decoder about more complex hierarchical operations to be performed by the decoder, which is capable of decoding the user data and has sufficient computational and / or battery power resources to perform more complex signal reconstruction tasks.

[0060] Certain instances described in this paper allow for the efficient generation, signaling, and decoding of optional enhanced upsampling method information (signal processing information), which the decoder can use along with residual data to appropriately refine the signal reconstruction to improve the quality of the reconstructed signal. In one set of instances described, this information is effectively embedded in the coefficients of the residual data for one or more steps of the encoded signal, thereby allowing for the avoidance of additional signaling overhead and the effective differentiation of signals that can benefit from a range of quality enhancement operations. Furthermore, the signal processing operations can be optional, and decoders that cannot decode user data or have stricter processing constraints can still decode the signal, albeit with lower quality reproduction due to less-than-ideal upsampling. This then maintains backward compatibility (e.g., the methods proposed in this paper complement rather than “break” existing defined coding standards).

[0061] In some of the instances described herein, optional signal processing operations include sharpening filters, such as desharpening masks or modified desharpening masks. Emitterable signals represent the purpose and strength of these filters. These sharpening filters can be cascaded after standard separable upsampling, or before (i.e., within the loop) or after (i.e., outside the loop) the application of residuals. In some instances, the use of a sharpening kernel is associated with modifications to the coefficients of the linear upsampling kernel to reduce ringing impairment while maintaining sharper edge reconstruction.

[0062] In some instances described herein, optional signal processing operations include upsampling via a neural network. For example, the method may include signaling an indication to upsample using a super-resolution simplified convolutional neural network (“minConv”) instead of conventional separable upsampling filtering, the topology of which is known to both the encoder and decoder. In some instances, user data signaling includes values ​​of the neural network coefficients that allow the decoder to configure the upsampling of a particular signal for better customization. In some implementations utilizing LCEVC, signaling to the “sensory” decoder indicates upsampling using a simplified convolutional neural network. When such a signal is detected, the decoder—in some cases, when sufficient processing resources are available—performs upsampling using the simplified convolutional neural network instead of a typical separable upsampling filter. A prediction residual is then added after the enhanced upsampling.

[0063] Examples of hierarchical coding schemes or formats

[0064] In preferred embodiments, the encoder or decoder is part of a hierarchical coding scheme or format. Examples of hierarchical coding schemes include LCEVC: MPEG-5 Part 2 LCEVC (“Low Complexity Enhanced Video Coding”) and VC-6: SMPTE VC-6ST-2117, the former described in PCT / GB2020 / 050695 (and associated standard documents), and the latter described in PCT / GB2018 / 053552 (and associated standard documents), all of which are incorporated herein by reference. However, the concepts described herein are not intended to be limited to these specific hierarchical coding schemes.

[0065] Figures 1 to 7 An overview of different exemplary hierarchical coding formats is provided. These formats are provided as a context for adding additional signal processing operations, which are then... Figure 7 The following diagrams will provide further details. Figures 1 to 5 Examples of implementation schemes similar to SMPTE VC-6 ST-2117 are provided, while Figure 6 and 7 Examples of implementation schemes similar to MPEG-5 Part 2 LCEVC are provided. It can be seen that both sets of examples utilize common basic operations (e.g., downsampling, upsampling, and residual generation) and may share modular implementation techniques.

[0066] Figure 1The hierarchical coding scheme is illustrated in large form. The hierarchical encoder 102, which outputs encoded data 103, retrieves the data to be encoded 101. Subsequently, the hierarchical decoder 104, which decodes the data and outputs decoded data 105, receives the encoded data 103.

[0067] Typically, the hierarchical coding scheme used in the examples herein creates a base or core layer, representing the raw data at a lower quality level, and one or more residual layers that can be used to recreate the raw data at a higher quality level using the decoded form of the base layer data. Generally, as used herein, the term "residual" refers to the difference between the values ​​of a reference array or reference frame and the actual array or frame of data. The array can be a one-dimensional or two-dimensional array representing coding units. For example, a coding unit can be a set of 2×2 or 4×4 residual values ​​corresponding to a region of similar size to an input video frame.

[0068] It should be noted that the generalized examples are independent of the nature of the input signal. The reference to "residual data" as used herein refers to data derived from the residual set, such as the residual set itself or the output of a set of data processing operations performed on the residual set. Throughout this specification, in general, the residual set contains multiple residuals or residual elements, each of which corresponds to a signal element, i.e., an element of the signal or original data.

[0069] In certain instances, the data may be an image or a video. In these instances, the residual set corresponds to an image or frame of the video, where each residual is associated with a pixel of the signal, the pixel being a signal element.

[0070] The method described herein can be applied to so-called data planes that reflect different color components of a video signal. For example, the method can be applied to different planes reflecting YUV or RGB data of different color channels. Different color channels can be processed in parallel. The components of each stream can be compared in any logical order.

[0071] A hierarchical coding scheme in which the concepts of the present invention can be deployed will now be described. The scheme is conceptually illustrated in… Figures 2 to 5In this context, the term generally corresponds to VC-6 as described above. In this type of coding technique, residual data is used for progressively higher quality levels. In the proposed technique, the core layer represents the image at a first resolution, and subsequent layers in the hierarchical structure are residual data or adjustment layers required by the decoding side to reconstruct the image at higher resolutions. Each layer or level may be referred to as a ladder index, such that the residual data is the data needed to correct low-quality information present in lower ladder indices. Each layer or ladder index (specifically each residual layer) in this hierarchical technique is typically a relatively sparse set of data with many zero-valued elements. When referring to a ladder index, it collectively refers to the set of all ladders or components under the stated level, for example, all subsets produced by the transformation steps performed at the stated quality level.

[0072] By employing this particular hierarchical approach, the described data structure eliminates any requirements or dependencies on the aforementioned or preceding quality levels. Quality levels can be encoded and decoded independently without reference to any other layers. Therefore, compared to many other known hierarchical coding schemes that require decoding of the lowest quality level in order to decode any higher quality level, the described method does not require decoding of any other layers. However, the principles of information exchange described below can also be applied to other hierarchical coding schemes.

[0073] like Figure 2 As shown, the encoded data represents a set of layers or levels, which is broadly referred to here as a ladder index. The base or core layer represents the original data frame 210, but at the lowest quality level or resolution, and subsequent residual data ladders can be combined with data under the core ladder index to recreate the original image at progressively higher resolutions.

[0074] To create the core hierarchical index, the input data frame 210 can be downsampled using several downsampling operations 201, each corresponding to the number of levels or hierarchical indices to be used in the hierarchical coding operation. One fewer downsampling operation 201 than the number of levels in the hierarchical structure is required. In all the examples shown herein, although there are four levels or hierarchical indices in the output encoded data and therefore three downsampling operations, it will be understood that these are for illustrative purposes only. Here, n indicates the number of levels, and the number of downsamplers is n-1. Core Hierarchy R 1-n This is the output of the third downsampling operation. As mentioned above, the core level R... 1-n This corresponds to the representation of the input data frame at the lowest quality level.

[0075] To distinguish downsampling operation 201, each operation will be referred to in the order in which it is performed on the input data 210 or the data represented by its output. For example, in this instance, the third downsampling operation 201...1-n It can also be called a core downsampler because its output generates a core tier index or tier 1-n, meaning that all tiers at this level are indices from 1 to n. Therefore, in this example, the first downsampling operation 201... -1 Corresponding to R -1 Lower sampler, second lower sampling operation 201 -2 Corresponding to R -2 The sampler was used for the third sampling operation, 201. 1-n Corresponding to core or R -3 Lower sampler.

[0076] like Figure 2 As shown, this represents the core quality level R. 1-n Data undergoes upsampling operation 202 1-n This is referred to here as the core upsampler. In the second downsampling operation 201 -2 The output (R) -2 The output of the downsampler (i.e., the input to the core downsampler) and the core upsampler 202 1-n The difference between the outputs is 203 -2 The output is the first residual data R. -2 This first residual data R -2 Correspondingly, this represents the core level R. -3 The error between the signal used to create the hierarchy and the signal itself. Since the signal itself undergoes two downsampling operations in this example, the first residual data R... -2 As an adjustment layer, the adjustment layer can be used to recreate the original signal at a quality level higher than the core quality level but lower than the input data frame 210.

[0077] exist Figure 2 and 3 The concept illustrates the changes in how residual data representing higher quality levels is created.

[0078] exist Figure 2 In the second downsampling operation 201 -2 (or R) -2 The downsampler is used to create the first residual data R. -2 The output of the signal is upsampled 202 times. -2 And according to the creation of the first residual data R -2 The same method is used to calculate up to the second downsampling operation 201. -2 (or R) -2 The lower sampler, i.e., R -1 The difference between the input and output of the downsampler (203) -1 This difference is correspondingly the second residual data R. -1It also indicates that it can be used to recreate the original signal at a higher quality level using data from lower levels.

[0079] However, in Figure 3 In the changes, the second downsampling operation 201 -2 (or R) -2 The output of the lower sampler and the first residual data R -2 Combining or summing 304 -2 To recreate the core sampler 202 1-n The output. In this change, the recreated data is upsampled instead of the downsampled data. 202 -2 The sampled data from the previous sampling operation will be compared with the data from the second sampling operation (or R). -2 The lower sampler, i.e., R -1 A similar comparison is made between the input of the downsampler (output) and the input of 203. -1 To create the second residual data R -1 .

[0080] Figure 2 and 3 Changes between implementation schemes cause minor changes in the residual data between the two implementation schemes. Figure 2 Benefiting from greater parallelization possibilities.

[0081] The process or loop is repeated to create a third residual R0. Figure 2 and 3 In this example, the output residual data R0 (i.e., the third residual data) corresponds to the highest level and is used at the decoder to recreate the input data frame. At this level, the difference operation is based on the same input data frame as the input to the first subsampling operation.

[0082] Figure 4 An example encoding process 401 is shown for encoding each of the hierarchical or tiered indices of data to produce an encoded tiered set of data with tiered indices. This encoding process is only an example of a suitable encoding process for encoding each of the hierarchies, but it should be understood that any suitable encoding process can be used. The input to the process is from... Figure 2 Or the corresponding level of the residual data output by 3, and the output is a set of tiers of encoded residual data, the tiers of encoded residual data together represent encoded data in a hierarchical manner.

[0083] In the first step, transformation 402 is performed. This transformation can be a directional decomposition transform, wavelet transform, or discrete cosine transform as described in WO2013 / 171173. If a directional decomposition transform is used, a set of four components (also referred to as transform coefficients) is output. When referring to the ladder index, these collectively refer to all directions (A, H, V, D), i.e., the four ladders. The set of components is then quantized 403 before entropy encoding. In this example, the entropy encoding operation 404 is coupled to a sparsification step 405, which utilizes the sparsity of the residual data to reduce the total data size and involves mapping data elements to a sorted quadtree. This coupling of entropy encoding and sparsification is further described in WO2019 / 111004, but the precise details of this process are not relevant to the understanding of this invention. Each array of residuals can be considered a ladder.

[0084] The process described above corresponds to the encoding process suitable for encoding data used for reconstruction according to the SMPTE ST 2117, VC-6 multiplane image format. VC-6 is a flexible, multi-resolution, intrinsic bitstream-only format capable of compressing any ordered set of integer-element grids, each grid having an independent size and designed for image compression. It employs data-independent techniques for compression and can compress low- or high-bit-depth images. The bitstream header can contain various metadata about the image.

[0085] It should be understood that each step or step index can be implemented using a separate encoder or encoding operation. Similarly, the encoding module can be divided into downsampling and comparison steps to generate residual data, and the residuals can then be encoded. Alternatively, each step of the step can be implemented in a combined encoding module. Thus, the process can be implemented, for example, using four encoders: one encoder for each step index, one encoder and multiple encoding modules operating in parallel or serially, or one encoder repeatedly operating on different data sets.

[0086] The following describes an example of reconstructing an original data frame encoded using the exemplary process described above. This reconstruction process may be referred to as pyramid reconstruction. Advantageously, the method provides an efficient technique for reconstructing an image encoded in a received dataset, which can be received by means of a data stream, for example by individually decoding different component sets corresponding to different image sizes or resolution levels and combining image details from one decoded component set with upgraded decoded image data from a lower-resolution component set. Thus, by performing this process for two or more component sets, a digital image at the structure or detail in said component sets can be reconstructed for progressively higher resolutions or a larger number of pixels, without needing to receive the complete or all image details of the highest resolution component set. Specifically, the method facilitates the progressive addition of increasingly higher resolution details while reconstructing the image from lower-resolution component sets in a hierarchical manner.

[0087] Furthermore, decoding each component set individually facilitates parallel processing of the received component sets, thereby improving reconstruction speed and efficiency in implementations where multiple processes are available.

[0088] Each resolution level corresponds to a quality level or step index. This is a collective term associated with the plane describing all new inputs or the set of received components (in this instance, a representation of a grid of integer-valued elements) and the output reconstructed image of the loop used for index -m. For example, the reconstructed image at step index zero is the output of the final loop of the pyramidal reconstruction.

[0089] Pyramid reconstruction can be a process of reconstructing an inverted pyramid from an initial tier index and using loops with new residuals to derive higher tier indices up to maximum quality, quality zero, at tier index zero. A loop can be considered a step in this type of pyramid reconstruction, identified by index -m. Steps typically involve upsampling data from possible previous steps, such as upgrading the decoded first component set, and using new residual data as additional input to obtain output data to be upsampled in possible subsequent steps. In the case of only receiving the first and second component sets, the number of tier indices will be two, and no further steps are possible. However, in instances where the number of component sets or the tier index is three or greater, the output data can be progressively upsampled in the following steps.

[0090] The first component set typically corresponds to the initial ladder index, which can be represented by ladder indices 1-N, where N is the number of ladder indices in the plane.

[0091] Typically, upgrading the decoded first set of components involves applying an upsampler to the output of the decoding procedure for the initial ladder index. In this example, this involves making the resolution of the reconstructed image from the decoded output of the initial ladder index set of components consistent with the resolution of the second set of components corresponding to 2-N. Typically, the upgraded output from the lower ladder index set of components corresponds to the predicted image at the higher ladder index resolution. Due to the lower resolution initial ladder index image and the upsampling process, the predicted image typically corresponds to a smoothed or blurred image.

[0092] Adding higher-resolution details from the above-mentioned tiered indexes to this predicted image provides a combined set of reconstructed images. Advantageously, when the received component set for one or more higher-tiered index component sets includes residual image data or data indicating the pixel value difference between the upgraded predicted image and the original, uncompressed, or pre-coded image, the amount of received data required to reconstruct an image or data set of a given resolution or quality can be significantly less than the amount of data or rate required to receive images of the same quality using other techniques. Therefore, by combining low-detail image data received at lower resolutions with progressively larger-detail image data received at increasingly higher resolutions according to the method, the data rate requirement is reduced.

[0093] Typically, the encoded data set includes one or more additional component sets, each of which corresponds to a higher image resolution than a second component set, and each of which corresponds to a progressively higher image resolution. The method includes decoding the component sets for each of the one or more additional component sets to obtain a decoded set. The method further includes, for each of the one or more additional component sets, in ascending order of corresponding image resolution: upgrading the reconstructed set with the highest corresponding image resolution to increase the corresponding image resolution of the reconstructed set to be equal to the corresponding image resolution of the additional component sets; and combining the reconstructed set with the additional component sets to produce another reconstructed set.

[0094] In this manner, the method may involve upgrading the reconstructed image output with a given component set hierarchy or tier index, and combining it with the decoded output of the aforementioned component set or tier index to produce a new, higher-resolution reconstructed image. It should be understood that, depending on the total number of component sets in the received set, the method may be repeatedly performed for progressively higher tier indices.

[0095] In a typical example, each of the component sets corresponds to a progressively higher image resolution, where each progressively higher image resolution corresponds to a fourfold increase in the number of pixels in the corresponding image. Typically, therefore, the image size corresponding to a given component set is four times the size or number of pixels of the image corresponding to the following component set, or twice the height and twice the width of the image, where the following component set is a component set having a ladder index one smaller than the ladder index discussed. The received set of component sets can, for example, facilitate a simpler upgrade operation, where the linear size of each corresponding image in the component set is twice the size of the following image.

[0096] In the example shown, the number of additional component sets is two. Therefore, the total number of component sets in the received set is four. This corresponds to the initial ladder index for ladder-3.

[0097] The first component set may correspond to image data, and the second and any other component sets correspond to residual image data. As noted above, the method provides a particularly advantageous reduction in data rate requirements for a given image size when the lowest gradient index (i.e., the first component set) contains a low-resolution or downsampled form of the image being transmitted. In this way, in each cycle where reconstruction begins with a low-resolution image, the image is upgraded to produce a high-resolution, but smoothed, form, and then improved by adding the difference between the upgraded, predicted image and the actual image to be transmitted at said resolution, and this improvement can be repeated for each cycle. Therefore, each component set above the initial gradient index requirement contains only residual data to reintroduce information that may have been lost when downsampling the original image to the lowest gradient index.

[0098] The method, for example, provides a way to obtain image data, which may be residual data, after receiving a set of data that has been compressed, for example by means of decomposition, quantization, entropy encoding, and sparsification.

[0099] The sparsification step is particularly advantageous when used in conjunction with a sparse set of raw or pre-transmitted data, which may typically correspond to residual image data. The residual may be the difference between elements of a first image and elements of a second image that are typically in the same position. Such residual image data may typically be highly sparsity. This can be considered as corresponding to an image in which detailed regions are sparsely distributed in regions where detail is minimal, negligible, or nonexistent. Such sparse data can be described as a data array where the data is organized in at least a two-dimensional structure (e.g., a grid), and where a majority of such organized data is zero (logically or numerically) or considered below a certain threshold. Residual data is just one example. Additionally, metadata may be sparse and thus significantly reduced in size through this process. Sending the already sparsified data allows for a significant reduction in the desired data rate by not sending such sparse regions and instead reintroducing them at appropriate positions within the received byte set at the decoder.

[0100] Typically, entropy decoding, dequantization, and directional synthesis transformation steps are performed according to parameters defined by the encoder or node, and the received encoded data set is transmitted from the encoder or node. For each tier index or component set, these steps are used to decode the image data to obtain a set that can be combined with different tier indices according to the techniques disclosed above, while allowing for efficient transmission of the set at each level.

[0101] A method for reconstructing an encoded data set according to the method disclosed above may also be provided, wherein decoding of each of the first and second component sets is performed according to the method disclosed above. Therefore, the advantageous decoding method of this disclosure can be used for each component set or ladder index in the received image data set and reconstructed accordingly.

[0102] refer to Figure 5 The decoding example is described below. A set of encoded data 501 is received, wherein the set comprises four tier indices, each tier comprising four tiers: from tier 0 (i.e., the highest resolution or quality level) to tier 1. -3 (i.e., the initial tier). In the tiers -3 The image data carried in the component set corresponds to the image data, and the other component sets contain the residual data of the transmitted image. Although each of the layers can output data that can be regarded as residuals, the initial ladder layer (i.e., the ladder) -3 The residuals in the ) actually correspond to the actual reconstructed image. At stage 503, each of the component sets is processed in parallel to decode the encoded set.

[0103] Referring to the initial tier index or the core tier index, for each component set tier... -3 The following decoding steps are performed at step 0.

[0104] At step 507, the component set is desparsed. Desparsing can be an optional step not performed in other hierarchical, layered formats. In this example, desparsing causes a sparse two-dimensional array to be recreated from the set of encoded bytes received at each tier. This process refills the zero values ​​grouped at positions within the two-dimensional array that were not received (since these are omitted from the transmitted byte set to reduce the amount of data transmitted). Non-zero values ​​in the array retain their correct values ​​and positions within the recreated two-dimensional array, with the desparsing step refilling the transmitted zero values ​​at appropriate positions or groups of positions.

[0105] At step 509, a range decoder is applied to the desparse set at each step to replace the encoded symbols in the array with pixel values. The configuration parameters of the range decoder correspond to those used to encode the transmitted data prior to transmission. Based on an approximation of the pixel value distribution of the image, the encoded symbols in the received set are replaced with pixel values. Using an approximate distribution—the relative frequency of each value across all pixel values ​​in the image—rather than the true distribution, allows for a reduction in the amount of data required to decode the set, since the range decoder needs distribution information to perform this step. As described in this disclosure, the desparse and range decoding steps are interdependent rather than sequential. This is indicated by the loop formed by the arrows in the flowchart.

[0106] At step 511, the array of values ​​is dequantized. This process is performed again based on the parameters quantized before the decomposed image was transmitted.

[0107] After dequantization, at step 513 the set is transformed by a synthetic transformation that includes applying an inverse directional decomposition operation to the dequantized array. This reverses the directional filtering based on the set of operators including averaging, horizontal, vertical, and diagonal operators, resulting in an array suitable for laddering. -3 Image data and used for ladder -2 The residual data up to grade 0.

[0108] Stage 505 illustrates the several loops involved in reconstruction using the output of the synthetic transform for each of the stepwise component sets 501. Stage 515 indicates the reconstructed image data for the initial stepwise stage output from the decoder 503. In this example, the reconstructed image 515 has a resolution of 64×64. At 516, this reconstructed image is upsampled to quadruple the number of its pixel components, resulting in a predicted image 517 with a resolution of 128×128. At stage 520, in the stepwise stage... -2Next, the predicted image 517 is added from the decoder output to the decoded residual 518. The addition of these two 128×128 images produces a 128×128 reconstructed image, which contains data from the ladder sequence. -2 The higher resolution details of the residuals enhance the smoothness of the initial gradient image details. If the desired output resolution corresponds to the gradient... -2 The resulting reconstructed image 519 can then be output or displayed. In this example, the reconstructed image 519 is used in another loop. At step 512, the reconstructed image 519 is upsampled in the same manner as at step 516 to produce a predicted image 524 of size 256×256. This is then followed by the decoded ladder at step 528. -1 Output 526 is combined, resulting in a reconstructed image 527 of size 256×256, which is an upgraded form of prediction 519 enhanced by the higher resolution details of residual 526. At 530, this process is repeated at the final time, and the reconstructed image 527 is upgraded to a resolution of 512×512 for combination with the tier 0 residual at stage 532. This yields a reconstructed image 531 of 512×512.

[0109] Another hierarchical coding technique is shown in Figure 6 and 7 The principles of this invention can be utilized through the aforementioned technology. This technology is a flexible, adaptable, efficient, and computationally inexpensive coding format that combines different video coding formats, underlying codecs (e.g., AVC, HEVC, or any other current or future codecs) with encoded data at least two enhancement layers.

[0110] The general structure of the encoding scheme uses a downsampled source signal encoded by a base codec, adds first-level correction data to the decoded output of the base codec to generate a corrected picture, and then adds data from another enhancement level to the upsampled form of the corrected picture. Thus, the stream is considered as a base stream and an enhancement stream, which can be further multiplexed or otherwise combined to generate an encoded data stream. In some cases, the base stream and enhancement stream can be transmitted separately. References to encoded data as described herein may refer to the enhancement stream or a combination of the base and enhancement streams. The base stream can be decoded by a hardware decoder, while the enhancement stream can be adapted to software processing implementations with suitable power consumption. This general encoding structure creates multiple degrees of freedom, allowing for great flexibility and adaptability to many situations, making the encoding format suitable for many use cases, including OTT transmission, live streaming, live UHD broadcasting, etc. Although the decoded output of the base codec is not intended for viewing, it is fully decoded video at a lower resolution, making the output compatible with existing decoders and usable as a lower-resolution output where appropriate.

[0111] In some instances, each enhancement stream or two enhancement streams can be encapsulated into one or more enhancement bitstreams using a set of Network Abstraction Units (NALUs). A NALU is intended to encapsulate the enhancement bitstreams so that the enhancements are applied to the correct underlying reconstructed frames. A NALU may, for example, contain a reference index containing the underlying decoder-reconstructed frame bitstreams to which the enhancements must be applied. In this way, the enhancements can be synchronized to the underlying stream, and frames from each bitstream are combined to produce the decoded output video (i.e., the residual of each frame from the enhancement layer combined with frames from the underlying decoded stream). A picture group can represent multiple NALUs.

[0112] Returning to the initial process described above, where the base stream is provided along with two enhancement layers (or sub-layers) within the enhancement stream, an instance of the generalized coding process is depicted... Figure 6 In the block diagram, the input video 600 at the initial resolution is processed to generate various encoded streams 601, 602, and 603. A first encoded stream (encoded base stream) is generated by feeding the input video in a downsampled form to a base codec (e.g., AVC, HEVC, or any other codec). The encoded base stream may be referred to as the base layer or base level. A second encoded stream (encoded level 1 stream) is generated by processing the residual obtained by taking the difference between the reconstructed base codec video and the input video in a downsampled form. A third encoded stream (encoded level 2 stream) is generated by processing the residual obtained by taking the difference between the upsampled form of the reconstructed base decoded video in a corrected form and the input video. In some cases, Figure 6The components can provide a general low-complexity encoder. In some cases, enhanced streams can be generated through an encoding process that forms part of the low-complexity encoder, and the low-complexity encoder can be configured to control independent base encoders and decoders (e.g., encapsulated as a base codec). In other cases, the base encoder and decoder can be provided as part of the low-complexity encoder. In one case, Figure 6 A low-complexity encoder can be viewed as a form of wrapper for a base codec, where the functionality of the base codec is hidden from the entity implementing the low-complexity encoder.

[0113] The downsampling operation illustrated by downsampling component 105 can be applied to input video to produce downsampled video to be encoded by the base encoder 613 of the base codec. Downsampling can be performed in both the vertical and horizontal directions, or alternatively only in the horizontal direction. The base encoder 613 and the base decoder 614 can be implemented by the base codec (e.g., as different functions of a common codec). One or more of the base codec and / or the base encoder 613 and the base decoder 614 may include appropriately configured electronic circuitry (e.g., hardware encoder / decoder) and / or computer program code executed by a processor.

[0114] Each enhanced stream coding process may not necessarily include an upsampling step. For example, in Figure 6 In this context, the first enhancement flow is conceptually a correction flow, while the second enhancement flow is upsampled to provide an enhancement level.

[0115] For more details, see the process of generating the enhanced stream. To generate the encoded Level 1 stream, the encoded base stream is decoded by the base decoder 614 (i.e., a decoding operation is applied to the encoded base stream to generate the decoded base stream). Decoding can be performed by the decoding function or mode of the base codec. Then, at the Level 1 comparator 610, the difference between the decoded base stream and the undersampled input video is created (i.e., a subtraction operation is applied to the undersampled input video and the decoded base stream to generate a first set of residuals). The output of comparator 610 may be referred to as the first set of residuals, such as a surface or frame of residual data, where a residual value is determined for each pixel at the resolution of the outputs of the base encoder 613, the base decoder 614, and the undersampling block 605.

[0116] The difference is then encoded by the first encoder 615 (i.e., the level 1 encoder) to generate an encoded level 1 stream 602 (i.e., the encoding operation is applied to the first residual set to generate the first enhanced stream).

[0117] As described above, the enhancement stream may include a first enhancement level 602 and a second enhancement level 603. The first enhancement level 602 may be considered as a corrected stream, for example, a stream that provides a correction level to the underlying encoded / decoded video signal at a lower resolution than the input video 600. The second enhancement level 603 may be considered as another enhancement level that transforms the corrected stream back to the original input video 600, for example, by applying an enhancement level or correction to the signal reconstructed from the corrected stream.

[0118] exist Figure 6 In this example, a second enhancement layer 603 is created by encoding another set of residuals. This other set of residuals is generated by a layer 2 comparator 619. The layer 2 comparator 619 determines, for example, the difference between the upsampled form of the decoded layer 1 stream, such as the output of the upsampling component 617, and the input video 600. The input to the upsampling component 617 is generated by applying a first decoder (i.e., a layer 1 decoder) to the output of the first encoder 615. This generates a decoded set of layer 1 residuals. These residuals are then combined with the output of the base decoder 614 at a summing component 620. This effectively applies the layer 1 residuals to the output of the base decoder 614. This allows losses during layer 1 encoding and decoding to be corrected by layer 2 residuals. The output of the summing component 620 can be considered as an analog signal representing the output of the encoded base stream 601 and the encoded layer 1 stream 602 at the decoder after applying layer 1 processing.

[0119] As mentioned, the upsampled stream is compared with the input video, which creates another set of residuals (i.e., the difference operation is applied to the recreated upsampled stream to generate another set of residuals). The other set of residuals is then encoded by the second encoder 621 (i.e., the level 2 encoder) into a coded level 2 enhanced stream (i.e., the encoding operation is then applied to the other set of residuals to generate another coded enhanced stream).

[0120] Therefore, as Figure 6 As explained and described above, the output of the encoding process is a base stream 601 and one or more enhancement streams 602, 603, which preferably include a first enhancement level and another enhancement level. The three streams 601, 602, and 603 can be combined, with or without additional information such as a control header, to generate a combined stream representing video encoded frames of the input video 600. It should be noted that... Figure 6The components shown operate on blocks or coding units of data, such as 2×2 or 4×4 portions of a frame at a specific resolution level. These components operate without any inter-block dependencies, thus allowing them to be applied in parallel to multiple blocks or coding units within a frame. This differs from contrasting video coding schemes, where dependencies (e.g., spatial or temporal) exist between blocks. These dependencies limit the level of parallelism and require significantly higher complexity.

[0121] exist Figure 7 The block diagram depicts the corresponding generalized decoding process. It is said that... Figure 7 It can show the corresponding Figure 6 The low-complexity decoder is a low-complexity encoder. The low-complexity decoder receives three streams 601, 602, and 603 generated by the low-complexity encoder along with a header 704 containing other decoded information. The encoded base stream 601 is decoded by a base decoder 710 corresponding to the base codec used in the low-complexity encoder. The encoded level 1 stream 602 is received by a first decoder 711 (i.e., a level 1 decoder), which decodes streams generated by the low-complexity encoder along with a header 704 containing other decoded information. Figure 1 The first encoder 615 decodes the first set of residuals encoded by the first residual. At the first summing component 712, the output of the base decoder 710 is combined with the decoded residuals obtained from the first decoder 711. The combined video, which can be referred to as the layer 1 reconstructed video signal, is upsampled by the upsampling component 713. The encoded layer 2 stream 103 is received by the second decoder 714 (i.e., the layer 2 decoder). The second decoder 714 decodes the data as shown by the first encoder 615. Figure 1 The second encoder 621 decodes the second residual set encoded by the second residual set. Although the header 704... Figure 7 The image shows the second decoder 714, but it can also be used by the first decoder 711 and the base decoder 710. The output of the second decoder 714 is a second decoded residual set. The second decoded residual set can have a higher resolution relative to the first residual set and the input to the upsampling component 713. At the second summing component 715, the second residual set from the second decoder 714 is combined with the output of the upsampling component 713, i.e., the upsampled reconstructed level 1 signal, to reconstruct the decoded video 750.

[0122] According to a low-complexity encoder, Figure 7 The low-complexity decoder can operate in parallel on different blocks or coding units of a given frame of a video signal. Furthermore, decoding performed by two or more of the base decoder 710, the first decoder 711, and the second decoder 714 can be executed in parallel. This is possible because there is no inter-block dependency.

[0123] During decoding, the decoder can analyze header 704 (which may contain global configuration information, picture or frame configuration information, and data block configuration information) and configure the low-complexity decoder based on those headers. To recreate the input video, the low-complexity decoder can decode each of the base stream, the first enhancement stream, and another enhancement stream or a second enhancement stream. The frames of the streams can be synchronized and then combined to derive the decoded video 750. Depending on the configuration of the low-complexity encoder and decoder, the decoded video 750 can be a lossy or lossless reconstruction of the original input video 100. In many cases, the decoded video 750 can be a lossy reconstruction of the original input video 600, wherein the loss has a reduced or minimal impact on the perception of the decoded video 750.

[0124] exist Figure 6 and 7 In each of these operations, the level 2 and level 1 encoding operations may include transformation, quantization, and entropy encoding steps (e.g., in the order described). These steps can be combined with... Figure 4 and 5 The operations shown are implemented in a similar manner. The encoding operation may also include residual grading, weighting, and filtering. Similarly, in the decoding stage, the residual can be passed through the entropy decoder, dequantizer, and inverse transform module (e.g., in the order described). Any suitable encoding and corresponding decoding operations can be used. However, preferably, the level 2 and level 1 encoding steps can be performed in software (e.g., by one or more central or graphics processing units in the encoding device).

[0125] The transforms described herein can use directional decomposition transforms, such as Hadamard-based transforms. Both can include small kernels or matrices applied to the flattened coding units (i.e., 2×2 or 4×4 residual blocks) of the residuals. Further details regarding the transforms can be found, for example, in patent applications PCT / EP2013 / 059847 or PCT / GB2017 / 052632, which are incorporated herein by reference. The encoder can select between different transforms to be used, such as between kernel sizes to be applied.

[0126] The transformation can transform the residual information onto four surfaces. For example, the transformation can produce components or transformation coefficients such as average, vertical, horizontal, and diagonal. A particular surface can include all values ​​for a particular component; for example, a first surface can include all average values, a second surface can include all vertical values, and so on. As mentioned earlier in this disclosure, these components output by the transformation can be used as coefficients, for example, to be quantized according to the described method, in such embodiments. A quantization scheme can be used to create quantities from the residual signal such that a particular variable can take only specific discrete values. Entropy encoding in this example can include run-length encoding (RLE), followed by processing the encoded output using a Huffman encoder. In some cases, when entropy encoding is required, only one of these schemes may be used.

[0127] In summary, the methods and apparatus described in this paper are based on a general approach that is constructed using existing encoding and / or decoding algorithms (e.g., MPEG standards such as AVC / H.264, HEVC / H.265, etc.; and non-standard algorithms such as VP9, ​​AV1, etc.), which serve as baselines for enhancement layers correspondingly used for different encoding and / or decoding methods. The idea behind this general approach is to encode / decode video frames in a hierarchical manner, in contrast to the block-based approach used in the MPEG family of algorithms. Encoding frames hierarchically involves generating residuals for the entire frame, and then generating residuals for extracted frames, and so on.

[0128] As noted above, the process can be applied in parallel to the coding units or blocks of the color components of a frame because there is no inter-block dependency. Encoding of each color component within the set of color components can also be performed in parallel (e.g., such that a copy operation is performed according to (number of frames) * (number of color components) * (number of coding units per frame)). It should also be noted that different color components can have different numbers of coding units per frame; for example, the luminance (e.g., Y) component can be processed at a high resolution of the chromaticity (e.g., U or V) component set when the change in illuminance is greater than the change in color that is detectable by human vision.

[0129] Therefore, as shown and described above, the output of the decoding process is (optionally) a basic reconstruction, as well as a reconstruction of the original signal at a higher level. This example is particularly well-suited for creating encoded and decoded video at different frame resolutions. For example, the input signal 30 could be an HD video signal comprising frames at a resolution of 1920×1080. In some cases, both the basic reconstruction and the level 2 reconstruction can be used by the display device. For example, in the case of network traffic, the level 2 stream can be interrupted more severely than the level 1 and basic streams (because it can contain up to 4× data, where downsampling reduces the dimension in each direction by 2). In this case, when traffic occurs, the display device can resume displaying the basic reconstruction while the level 2 stream is interrupted (e.g., and the level 2 reconstruction is unavailable), and then resume displaying the level 2 reconstruction when network conditions improve. A similar approach can be applied when the decoding device is resource-constrained; for example, a set-top box performing a system update may have an operating basic decoder 220 to output the basic reconstruction, but may not have the processing capacity to compute the level 2 reconstruction.

[0130] The encoding arrangement also allows the video distributor to distribute video to a non-homogeneous set of devices; only those devices with the base decoder 720 view the base reconstruction, while those with enhancement layers view the higher-quality layer 2 reconstruction. In contrast, two complete video streams at separate resolutions are needed to serve the two sets of devices. Since the layer 2 and layer 1 enhancement streams encode residual data, they can be encoded more efficiently, for example, the residual data distribution is typically mostly around quality 0 (i.e., no difference) and usually takes a small range of values ​​around 0. This may be especially true after quantization. In contrast, the complete video streams at different resolutions will have different distributions with non-zero means or medians, requiring higher bit rates for transmission to the decoder.

[0131] In the examples described herein, the residuals are encoded by an encoding pipeline. This may include transform, quantization, and entropy coding operations. It may also include residual grading, weighting, and filtering. The residuals are then passed to a decoder, for example as L-1 and L-2 enhancement streams, which may be combined with the base stream as a hybrid stream (or transmitted separately). In one case, a bit rate is set for the hybrid data stream comprising the base stream and the two enhancement streams, and then different adaptive bit rates are applied to individual streams based on the data being processed to meet the set bit rate (e.g., high-quality video perceived with low-level artifacts can be constructed by adaptively allocating bit rates to different individual streams (even at the frame-by-frame level) so that constrained data can be used by the individual stream most affected, which may change as the image data changes).

[0132] The set of residuals described in this paper can be considered sparse data, for example, in many cases there is no difference for a given pixel or region, and the resulting residual value is zero. When looking at the distribution of the residuals, many probability masses are assigned to small residual values ​​located close to zero, such as those occurring most frequently for certain video values ​​like -2, -1, 0, 1, 2, etc. In some cases, the distribution of residual values ​​is symmetric or approximately symmetric about 0. In some test video cases, the distribution of residual values ​​is found to have a shape similar to a logarithmic or exponential distribution about 0 (e.g., symmetric or approximately symmetric). The exact distribution of the residual values ​​can depend on the content of the input video stream.

[0133] The residual can be processed into a two-dimensional image, such as a difference image. In this way, the sparsity of the data can be seen to involve features visible in the residual image, such as “points,” small “lines,” “edges,” “corners,” etc. These features have been found to be generally not perfectly correlated (e.g., spatially and / or temporally). These features have properties that differ from those of the image data from which they originate (e.g., the pixel characteristics of the original video signal).

[0134] Because the characteristics of residuals differ from those of the image data from which they originate, it is generally impossible to apply standard coding methods, such as those found in the traditional Moving Picture Experts Group (MPEG) coding and decoding standards. For example, many contrast schemes use large transforms (e.g., transforms of large pixel regions in a normal video frame). Due to the characteristics of residuals, such as those described above, using these large transforms on residual images would be extremely inefficient. For example, encoding small points in a residual image using large blocks of regions designed for normal images would be very difficult.

[0135] Some of the examples described in this paper address these problems by alternatively using smaller and simpler transform kernels (e.g., 2×2 or 4×4 kernels—directed decomposition and directional decomposition squared, as presented in this paper). The transforms described in this paper can be applied using Hadamard matrices (e.g., a 4×4 matrix for flattened 2×2 coded blocks or a 16×16 matrix for flattened 4×4 coded blocks). This moves in a different direction than the contrasting video coding methods. Applying these new methods to residual blocks generates compression efficiency. For example, some transforms generate uncorrelated transform coefficients (e.g., in space) that can be compressed efficiently. While the correlation between transform coefficients can be utilized, for example, for lines in the residual image, these correlations can introduce coding complexity, making implementation difficult on legacy and low-resource devices, and these correlations often generate other complex artifacts that require correction. Preprocessing residuals by setting some residual values ​​to 0 (i.e., not forwarding these residual values ​​for processing) provides a controlled and flexible way to manage bit rates and stream bandwidth, as well as resource usage.

[0136] Examples related to enhancement of higher resolution levels

[0137] In some instances described in this article, upsampling operations, such as Figure 2 and 3 Operation 202 and Figure 5 Operations 526, 522, or 530 in the middle Figure 6 Operation 617 or Figure 7 One or more of operations 713 (and other upsampling operations not shown) include optional enhancement operations. These optional enhancement operations can be sent from the encoder to the decoder using signals. The operations may include one or more signal processing operations for enhancing the output of a specific level in a hierarchical, layered format. In a video example, the output may include a reconstructed video signal at a specific resolution (e.g., Figure 5 The outputs are 520, 528, or 531 or Figure 7 (Decoded video 750 in the example). Optional enhancement operations can provide a so-called super-resolution mode. Optional enhancement operations can be performed in place of existing default upsampling operations and / or in addition to existing default upsampling operations. Existing default upsampling operations may include upsampling operations as defined in standard hierarchical coding schemes (e.g., as defined in one or more of the LCEVC or VC-6 standards). Therefore, the decoder can decode the signal quite adequately without optional enhancement operations that provide optional additional functionality, such as a sharper-looking image desired using this functionality. For example, optional enhancement operations may only be available at the decoder if available computational resources exist and / or the decoder is configured to apply these operations.

[0138] In some instances, signaling for these optional enhancement operations can be provided using user data within a hierarchical, layered format of the bitstream. This user data may include a configurable data stream carrying data not directly used to reconstruct the output signal (e.g., not the base encoded stream or the residual / enhanced encoded stream). In some instances, the user data may be embedded within values ​​directly used to reconstruct the output signal, such as within the residual / enhanced encoded stream. In other instances, or beyond those described above, the user data may also be embedded within supplementary enhancement information messages of the bitstream.

[0139] Figure 8A An example is shown where signal processing information for one or more enhancement operations is embedded in one or more values ​​received in one or more coded data layers transmitted within a coded data stream. In this example, the values ​​are associated with transform coefficients processed to derive signal elements during decoding. These transform coefficients may include those derived from... Figure 4Transformation 402 or formation Figure 6 The L1 and L2 in the encoded stream are transformed into A, V, H, and D by one or more portions of the encoded stream. In some instances, the transformation is applied as a linear transformation (e.g., a linear transformation of the form y = Ax, where x is the flattened input derived from an n×n block of residuals, and y is the set of transform coefficients—typically having the same length as x). As described above, the transformation can be implemented using a 4×4 or 16×16 Hadamard matrix (depending on whether n is 2 or 4). In this example, the input, signal processing information is embedded in one or more values ​​of predefined transform coefficients in different sets of transform coefficients generated by the encoded transformation, such as the value of a specific element or index in the output vector y. In some instances, H (n = 2) or HH (n = 4) elements are preferred for this embedding because the substitution of these values ​​has minimal impact on the reconstructed output.

[0140] refer to Figure 8A This illustrates an example of a method implemented within a decoding system. A set of quantized symbols 800-1 to 800-N is received and processed. These quantized symbols include quantized transform coefficients, where quantization may be optional and / or vary in degree of quantization based on the encoding configuration. Quantized symbols may include those generated by L1 and L2 via one or more of the encoded stream and / or correspond to those generated via... Figure 4 The symbol of the data generated by quantization block 403 in the code. Figure 8A In this instance, one of the symbols is configured to carry user data. Therefore, the selected symbol (e.g., a symbol derived from the H or HH transform coefficients) is called a "reserved symbol." Depending on whether symbol 800-1 is intended to be a reserved symbol, the decoder follows two different approaches.

[0141] If symbol 800-1 is not intended as a reserved symbol, for example, if it is intended to carry residual data for signal reconstruction, its decoding follows the normal process performed for other symbols in the set: dequantization and inverse transform are performed according to method 810, resulting in a set of decoded data 830. This is illustrated by comparison block 805. For example, method 810 may include at least... Figure 5 Blocks 511 and 513 and / or blocks forming part of one or more of the L-1 and L-2 decoding processes 711 and 714. The decoded data can then be further processed by decoding operation 850 to produce a decoded signal 860. In one set of examples, decoding operation 850 may include, according to... Figure 5 Phase 505 reconstruction and / or via Figure 7 The reconstruction is implemented in 715. In this case, the decoded signal 860 may include the reconstructed image 531 or 750. In other cases, the decoding operation 850 may include inputs performed to generate an upsampling operation to a hierarchical, hierarchical format (e.g., to...). Figure 5 530 or Figure 7 The operation of input 713 in the above sample (in this case, the decoded signal 860 includes this input for the upsampling operation).

[0142] If symbol 800-1 is intended to be used as a reserved symbol, its decoding follows a different process, as indicated by comparison block 805. At block 820, decoding method 820 is applied to embedded signal processing information, such as user data within symbol 800-1, to extract signal processing information 840. This signal processing information may include information about enhancement operations to be performed at block 870. For example, the signal processing information may include one or more flags to indicate one or more signal processing operations to be performed. In some cases, the signal processing information may also include parameters for the signal processing operations, such as coefficients of an adaptive filter. In one case, the parameters for the signal processing operations may vary with coding units or data blocks (e.g., the n×n data set described above). For example, the parameters for the signal processing operations may vary with each coding unit or data block or with consecutive groups of coding units or data blocks. In these cases, the reserved symbol for a given coding unit or data block may include parameters for the signal processing operations to be performed with respect to said unit or block.

[0143] exist Figure 8A At block 870, one or more signal processing operations are performed according to signal processing information 840 as part of enhancement operation 870 to generate signal enhancement reconstruction 880. This signal enhancement reconstruction 880 can be used to replace... Figure 5 and 7 The output in the text is 531 or 750, or it can be replaced with... Figure 5 and 7 The output of the sampled 530 or 713.

[0144] In some instances, a bit (not shown in the diagram) in the decoded byte stream signals to the decoder that symbol 800-1 will be processed as a reserved symbol. For example, this bit could include a "User Data" flag that is toggled to "Enabled" or "Disabled" in global configuration information.

[0145] While examples have been provided in the context of hierarchical, layered formats, the methods described herein can be used in non-hierarchical and / or non-layered formats in other instances. For example, they can be applied to data streams that do not include different stream outputs with varying quality levels but still embed enhancement operation information in the transform coefficients. Figure 8A The operation.

[0146] Figure 8B It shows Figure 8AVariations of the instance, where enhancement operations are optionally performed on multiple quality levels within a hierarchical, layered format. Similar reference numerals are used to refer to similar components, where variations in the last digit of the reference numeral indicate possible variations within the instance.

[0147] refer to Figure 8B This illustrates an example of a method implemented within a decoding system that employs a hierarchical encoding approach based on levels. Figure 8A As shown, quantization symbol 800-1 is received and processed along with other quantization symbols 800-2...800-N. In a preferred embodiment, the quantization symbol represents the residual data stream of the first quality level. At block 805, the decoder checks whether symbol 800-1 should be intended as a reserved symbol. Depending on whether symbol 800-1 is intended as a reserved symbol, the decoder follows two different methods.

[0148] If symbol 800-1 is not intended to be used as a reserved symbol, its decoding follows the normal process implemented for other symbols in the set: according to the dequantization and inverse transformation of method 810, a set of decoded residual data 832 is produced.

[0149] exist Figure 8B There are two sets of enhancement operations: a first set 872 for the signal at the first quality level (LOQ1 or L-1), and a second set 874 for the signal at the second quality level (LOQ2 or L-2). These sets of enhancement operations can be flexibly applied based on decoder configuration and / or signal processing information; for example, only the lower quality level or only the higher quality level may be applied in different situations. The first set of enhancement operations 872 can be applied to... Figure 5 The data in the dataset includes one or more of 515, 518, 526, etc., or the output of the basic decoder 710. The second set of enhancement operations 874 can be applied after the reconstruction by the reconstructor 852, for example, in... Figure 5 or Figure 7 After adding as shown in 712.

[0150] If symbol 800-1 is intended to be used as a reserved symbol, its decoding follows a different process via block 805. At block 822, a method is implemented to decode the embedded information within the reserved symbol, for example, to analyze the data of the reserved symbol to extract signal processing information 842, 844, and 846. Reserved symbols may include data configured according to a prescribed syntax. This syntax may include a header portion and a payload portion. Figure 8BIn this process, signal processing information is extracted for the first set of enhancement operations 872, the reconstructor 852, and the second set of enhancement operations 874. However, in other instances, any one or more of this data may be extracted; for example, in one case, the initial reconstruction 808 of the signal may not be performed or enhancement operations may not be applied at the reconstructor 852, such that the retained symbol 400-1 at the first quality level includes signal processing information for the higher quality level. This may be advantageous because the first quality level is typically small (due to its reduced resolution) and is usually received before the second quality level.

[0151] At block 822, the reserved symbol 800-1 is processed to generate signal processing information 842, 844, and 846. Residual data 832 (e.g., at the first quality level—e.g.) Figure 7 The output of L-1 decoding at block 711 or Figure 5 One of the steps from -1 downwards) is further processed by reconstructor 852 (e.g., together with the remaining samples of the signal or other residual data of the frame) to produce a reconstructed reproduction 834 of the signal at the first quality level (e.g., Figure 7 (LOQ#1 or L-1 in the original text). Figure 8B In this process, for example, a first set of enhancement operations 872 may be applied to the initial reproduction 808 of the signal at a first quality level based on signal processing information 842. As discussed above, this may include enhancing the reconstructed base signal. The signal processing information may also include information for performing one or more signal processing operations at the reconstructor 852, as described in 844.

[0152] Once the reconstructor 852 outputs a possible enhanced reconstruction 834 of the signal at the first quality level, for example after adding residual data 832 to the data derived from the initial reconstruction 808 of the signal at the first quality level, the reconstruction 834 is further processed by decoding operation 852 to produce a reconstruction 862 of the signal at the second quality level. In these instances, it is assumed that the second quality level is at a higher resolution than the first quality level, i.e., a higher-level signal compared to the lower-level signal at the first quality level. The difference in resolution can be a self-defined factor in one or more dimensions of a multidimensional signal (e.g., the horizontal and vertical dimensions of a video frame). Decoding operation 852 may include the operation at stage 505 and / or Figure 7 One or more of the operations at blocks 713 and 715. Reproduction of the signal 862 at the second quality level may include... Figure 7 The output is 750, or Figure 5 The higher-level output is either 528 or 531. Figure 8BIn the second quality level, the signal reproduction 862, together with signal processing information 846, is processed by a second enhancement operation set 874 to produce an enhanced final reproduction 890 of the signal at the second quality level. Output 890 may include a signal suitable for reproduction, for example, for display or output to a user via an output device. In some instances, the second enhancement operation set 874 may be applied during decoding operation 852. For example, the second enhancement operation set 874 may be added to or replace an upsampling operation application conforming to one of the LCEVC or VC-6 standards.

[0153] In the examples described herein, one or more signal processing operations to enhance higher resolution levels can be performed "inside the loop" or "outside the loop," such as forming Figure 8A and 8B One or more signal processing operations are part of enhancement operations 870 or 874 in the code. "In-loop" signal processing operations are operations applied as part of a decoding method for a higher resolution level, such as coding units or data blocks that may be iteratively processed within the decoding loop (as described above in serial and parallel processing), and the signal processing operations may be applied to the data of a specific coding unit or data block within the decoding loop. "Out-of-loop" signal processing operations are operations applied to the reconstructed signal output by the decoding method (i.e., the decoding loop). This reconstructed signal may include a viewable frame sequence for a video signal. In one case, "in-loop" processing involves applying the signal processing operations before adding residual data for a higher resolution level, for example, in... Figure 5 Adding or at block 532 Figure 7 Before the addition at block 715. "In-loop" processing may include applying one or more signal processing operations as an alternative enhanced upsampling operation performed in place of a standard upsampling operation. This is described in more detail below. Furthermore, both "in-loop" and "out-of-loop" signal processing operations can be represented and applied using signals; for example, a convolutional neural network upsampler may be applied "in-loop," while a sharpening filter may be applied "out-of-loop."

[0154] The "in-loop" signal processing operation prior to adding residual data offers the advantage that the residual data itself can correct for artifacts introduced by the signal processing operation. For example, if the signal processing operation used to enhance higher-level signals is... Figure 2 Or the upsampling process 202 in 3 or Figure 6If a portion of the upsampling in step 617 is applied, then a subsequent comparison at block 203 or 619 generates residual data indicating the difference between the output of the enhancement operation and the original input signal (e.g., data frame 210 or input video 600). Therefore, the signal processing operation does not always need to produce a high-quality, artifact-free output; if the signal processing operation does produce visible artifacts, these can be mitigated by... Figure 7 The residual data applied at blocks 532 or 715 in the super-resolution upscalorie is used to correct these artifacts. This becomes particularly advantageous when implementing unpredictable neural network enhancers and / or statistical processing where the output form cannot be guaranteed (e.g., due to the complexity of the process and / or variations in the statistical process). For example, a super-resolution upscalorie only needs to produce high-quality predictions 80% of the time; the remaining 20% ​​of pixels that may appear to be artifacts can be corrected using residuals. Another benefit may be that better predictive upsampling helps reduce the number of non-zero bytes required in higher-level encoded data streams (e.g., good predictions may have many residuals with values ​​equal to or close to zero), thereby reducing the number of bits required in higher-level encoded streams.

[0155] In some instances, the process of encoding signals at the first quality level (e.g., Figure 6 615) may include detecting one or more impairments that cannot be properly corrected with residual data at the target bit rate (e.g., one or more of coding level 1 stream 602 and coding level 2 stream 603). In this case, the coding operation for the first quality level (e.g., Figure 6 The encoder (615) generates an encoded data stream (e.g., encoded level 1 stream 602) that utilizes the reserved symbol set in the encoded residual data as described above to signal to the decoder the type and / or location of the desired attenuation. Thus, the decoder can apply appropriate corrections to reduce the attenuation (e.g., reduce visual effects). In some instances, the encoding operation for the first quality level switches specific locations in the encoded byte stream (e.g., encoded level 1 stream 602 or a multiplexed stream including two or more of encoded base stream 601, encoded level 1 stream 602, and encoded level 2 stream) to signal to the decoder whether a given symbol set in the encoded data should be interpreted as actual residual data or as additional contextual information (i.e., signal processing information) to inform signal enhancement operations. In some instances, the encoder performs "in-loop" decoding (e.g., ...) on the output encoded at the first quality level. Figure 6 The L-1 decoder (618) can also use signal processing information to simulate the reconstruction to be generated by the decoder.

[0156] As described above, in this current instance, when decoding a specific set of data within an encoded data stream and finding a specific set of quantized symbols, the decoder does not interpret the symbols as residual data. Instead, it performs signal enhancement operations based on the received symbols. This use of reserved symbols can be instructed to send signals to... Figure 7 Bits from one or more of the decoded byte streams in L-1 decoding 711 and L-2 decoding 714 are transmitted, for example, within control header 714. In this case, the bits indicate that a specific set of quantization symbols in a particular set of residual data should not be interpreted as actual residual data, but rather as contextual information used to inform signal enhancement operations. In some instances, some reserved symbols may correspond to a specific type of attenuation, informing the decoder of post-processing operations applicable to the corresponding region of the signal (whether in the loop or at the end of the decoding process) to improve the quality of the final signal reconstruction.

[0157] Enhanced conditions

[0158] In the examples described herein, one or more signal processing operations may be selectively applied based on determined signal processing information to enhance data associated with higher levels of the hierarchically encoded signal. The phrase "selectively" applying or performing one or more signal processing operations indicates that these operations may be optional. In some cases, these operations may replace and / or supplement defined encoding procedures, such as the decoding procedures specified in the LCEVC and VC-6 standards. In these cases, the signal processing information may include one or more flags indicating whether one or more corresponding signal processing operations are applied. If the signal processing information is absent and / or has a specific value (e.g., a flag value of "false" or 0), the encoded data stream may be decoded according to the defined encoding procedure. If the signal processing information is present and / or has a specific value (e.g., a flag value of "true" or 1), the encoded data stream may be decoded according to the signal processing operation. It should be noted that, in the examples, "enhancement" at higher resolution levels is enhancement other than adding residual data to correct the upsampled reproduction of the signal. For example, signal processing operations may include optional sharpening filters and / or samplers on neural networks.

[0159] In some instances, selectively performing signal processing operations is further based on the operating conditions or parameters of the decoder used to perform decoding. For example, where signal processing information exists and indicates one or more optional signal processing operations, these operations may be performed only if further criteria are met. For example, selectively performing one or more signal processing operations to enhance a higher resolution level may include determining the operating parameters of the decoder used to perform decoding. These operating parameters may include one or more of the following: resource utilization (e.g., CPU or GPU utilization or memory utilization); environmental conditions (e.g., processing unit temperature); power and / or battery conditions (e.g., whether the decoder is plugged into mains power and / or remaining battery power); network conditions (e.g., congestion and / or download speed), etc. In this case, in response to a first set of operating parameters, one or more signal processing operations may be performed to enhance the higher resolution level using the signal processing parameters within the determined signal processing information. In response to a second set of operating parameters, one or more signal processing operations may be omitted, for example, although signaling and / or one or more signal processing operations in the signal processing information may be replaced by default signal processing operations. In the latter case, a default or predefined set of decoding procedures can be applied (e.g., procedures defined in one of the LCEVC or VC-6 standards). Therefore, two decoders with a shared architecture (e.g., two mobile phones of the same brand) can implement different signal processing operations using the same signaling depending on their current operating conditions. For example, a decoder plugged into mains power or with a remaining battery level above a predefined threshold can apply signal processing operations that are more resource-intensive than the contrasting default decoding procedure (i.e., using more resources compared to the case where no signal processing operation is applied).

[0160] In one scenario, a method for decoding a signal may include determining a resource usage metric for the decoder. This resource metric may be a metric related to the operating parameters described above, such as CPU / GPU utilization, available memory, and / or battery percentage. The method may include comparing the resource usage metric to a resource usage threshold. The resource usage threshold may be predefined and based on a utilization test. In response to the comparison indicating no restriction on the decoder's resource usage, one or more signal processing operations may be performed to enhance higher resolution levels based on the determined signal processing information. In response to the comparison indicating a restriction on the decoder's resource usage, one or more signal processing operations may be omitted during reconstruction.

[0161] Signal processing operations for higher-level enhancements can also be performed based on the decoder's performance; these may include post-processing operations. For example, a conventional decoder may lack suitable software, hardware, and / or available resources to implement certain signal processing operations. In these cases, the determined signal processing information can be used to identify signal processing operations for enhancing higher-resolution levels. For example, header data within coefficient embedded and / or SEI user data may include m-bit or byte values ​​indicating the signal processing operation to be performed from a plurality of signal processing operations, or flags for each of the plurality of signal processing operations. Once the user data has been analyzed and the signal processing operations have been identified, the decoder can determine whether it is capable of implementing the identified signal processing operation. For example, the decoder may include a lookup table listing the signal processing operations that the decoder can perform. In response to the decoder's inability to implement the identified signal processing operation, the determined signal processing information can be ignored, and a similar procedure can be followed. Figures 1 to 7 The decoding process illustrated decodes the encoded data stream. In response to the decoder's ability to implement the identified signal processing operation, the decoder can execute the determined signal processing operation parameterized by the determined signal processing information. In some cases, in response to a positive determination, checks on operating parameters and / or resource utilization, as described above, can be further implemented. Therefore, multiple criteria can be cascaded to determine whether one or more signal processing operations should be applied.

[0162] Therefore, in the above examples, the decoder may perform signal enhancement operations in different ways at any time based on the nature of the decoder device and / or the conditions at the decoder device (including sometimes not performing the operation at all).

[0163] Figure 9An example of an encoding and decoding system utilizing the innovative methods described herein is illustrated. Encoder 910 processes raw signal 900 to produce data stream 920. Data stream 920 is processed by two decoders. Decoder 930-0 implements a signal enhancement method based on information transmitted by encoder 910, thereby decoding reconstructed signal 940-0. Decoder 930-1 ignores information transmitted by encoder 910 and reconstructs reconstructed signal 940-1. Reconstructed signal 940-1 may include a fully feasible reconstruction of the signal for a given purpose. For example, it may include normal or standard decoding using an option scheme defined as part of the LCEVC or VC-6 standard, making the enhancement operation performed by decoder 930-0 entirely optional. As described above, decoder 930-0 may sometimes decide to ignore partial information transmitted by encoder 910 without considering the signal processing information transmitted by encoder 910. In some cases, decoder 930-0 defines whether to ignore partial information transmitted by encoder 910 based on one or more of the following: signal resolution and frame rate, processing power load during decoding, and battery power state. The encoder 910 can use user data embedded in the transform coefficient data set, user data in the SEI message, and / or use specific combinations of predefined parameter values ​​defined in the signal coding standard to transmit signal processing information for enhanced operation.

[0164] Example Enhancement Operation

[0165] In the examples described herein, the method for decoding a signal includes: obtaining an encoded data stream; analyzing the encoded data stream to determine signal processing information transmitted by the encoder; and reconstructing a higher resolution level of the signal from a lower resolution level, including selectively performing one or more signal processing operations to enhance the higher resolution level based on the determined signal processing information. Two sets of example signal processing operations are described in this section. These include a sharpening filter for video signals and an efficient neural network upsampler. Generally, these two sets of signal processing operations can be viewed as a cascade of linear filtering operations within a configurable (and optional) intermediate nonlinearity.

[0166] In an example of this section, signal processing operations (which may include...) Figure 8A and 8B The enhancement operations 870, 872, and / or 874 in the above process form part of the upsampler. As discussed above, this upsampler may include... Figure 2 Or the upper sampler 202 in 3, Figure 5 The upper samplers 522, 526 and 530 in the middle, Figure 6 The top sampler 617 and Figure 7One of the upsamplers 713. An example assuming symmetric upsampling is performed at both the encoder and decoder will be described, but such examples may not always be as described; in some cases, it may be possible to apply upsampling at the decoder differently than that applied at the encoder. In some instances, the upsampling described herein may also be applied as an additional "extra" upsampling stage applied to the output of a standard decoding process (e.g., decoding video 531 or 750). In this case, there may be no corresponding encoder upsampling process, and the additional upsampling stage may be considered a post-processing upgrade stage.

[0167] Figure 10A A first instance 1000 based on the sampler configuration is shown. The upsampler can be used to convert signal data at a first level (n-1) and signal data at a second level n. In the context of the current instance, the upsampler can convert data processed at enhancement level 1 (i.e., quality level (LoQ)-1) and data processed at enhancement level 2 (i.e., quality level (LoQ)-2), for example... Figure 7 The upsampler 713 is described above. In another case, the upsampler may include an additional upsampling stage applied to data processed at enhancement level 2 (i.e., quality level (LoQ)-2) – for example, decoded video 750 – to generate a third quality level (e.g., LoQ 3). In one case, the first level (n-1) may have a first resolution (e.g., size_1 × size_2 elements), and the second level n may have a second resolution (e.g., size_3 × size_4 elements). The number of elements in each dimension at the second resolution may be a multiple of the number of elements in each dimension at the first resolution (e.g., size_3 = F1 * size_1 and size_4 = F2 * size_2). In the described examples, the multiple may be the same in both dimensions (e.g., F1 = F2 = F, and in some instances, F = 2).

[0168] In some instances, using enhancement operations during upsampling may involve converting element data (e.g., pixel values, such as values ​​of a color plane) from one data format to another. For example, element data (e.g., as input to the upsampler in a non-neural context) may be in the form of 8- or 16-bit integers, while neural networks or other adaptive filtering operations may operate on floating data values ​​(e.g., 32- or 64-bit floating-point values). The element data can thus be converted from integer to floating format before upsampling, and / or from floating format to integer format after neural enhancement upsampling. This is in... Figure 10B As shown in the image.

[0169] exist Figure 10BIn this process, an enhanced upsampler 1005 is used. The input to the enhanced upsampler 1005 is first processed by a first conversion component 1010. The first conversion component 1010 converts the input data from integer format to floating-point format. The floating-point data is then input to the enhanced upsampler 1005, which freely performs floating-point operations. The output from the neural enhanced upsampler 1005 includes data in floating-point format. Figure 10B The data is then processed by a second conversion component 1020, which converts the data from floating-point format to integer format. The integer format can be the same as the original input data or a different integer format (e.g., the input data can be provided as an 8-bit integer, but the output can be provided as a 10, 12, or 16-bit integer). The output of the second conversion component 1020 can format the output data to be suitable for operations at higher enhancement levels (e.g., level 2 enhancement as described herein).

[0170] In some instances, as an alternative to or supplement to data format conversion, the first conversion component 1010 and / or the second conversion component 1020 may also provide data scaling. Data scaling can place the input data in a form more suitable for applications involving artificial neural network architectures. For example, data scaling may include a normalization operation. Example normalization operations are described below:

[0171] norm_value=(input_value-min_int_value) / (max_int_value-min_int_value)

[0172] Where `input_value` is the input value, `min_int_value` is the minimum integer value, and `max_int_value` is the maximum integer value. Additional scaling can be applied by multiplying by the scaling divisor (i.e., dividing by the scaling factor) and / or subtracting the scaling offset. The first transformation component 1010 provides positive data scaling, and the second transformation component 1020 applies the corresponding inverse operation (e.g., inverse normalization). The second transformation component 1020 can also round the values ​​to generate an integer representation.

[0173] Figure 11AA first example of an enhanced upsampler 1105, which can be used to apply enhancement operations as described herein (e.g., applying one or more signal processing operations to enhance the hierarchy of a signal), is shown. The enhanced upsampler 1105 includes an upsampling kernel 1110, a predicted average modification component 1120, and a post-processing filter 1130. The upsampling kernel 1110 may include known upsampling kernels, such as one of the following: the closest sample upsampler kernel, bilinear upsampler kernel, and cubic upsampler kernel described in the LCEVC standard specification entitled “Decoding processing for the upscaling” and international patent application PCT / GB2019 / 052152, both of which are incorporated herein by reference. The upsampling kernel 1110 transforms a lower-level representation of a signal into a higher-level representation of the signal (e.g., by enhancing the hierarchy of a signal as described herein). Figure 10A (The resolution described). The predicted average modification component 1120 can add modifiers to the output of the upsampler kernel, as described in the section entitled "Predicted residual process description" in the LCEVC standard specification and in international patent application PCT / GB2020 / 050574, both of which are incorporated herein by reference.

[0174] In a brief summary of the predicted average modification, values ​​derived from elements of the first residual set of blocks in the derived upsampled video are added to blocks in the upsampled second output video. Modification terms are added by the predicted average modification component 1120 and represent the difference between the values ​​from the lower-resolution representation and the average of the values ​​in the blocks of the upsampled video. The predicted average modification component 1120 can be enabled and disabled based on flags in control signaling.

[0175] exist Figure 11A In this context, post-processing filter 1130 includes higher-level signal processing operations for enhancing the signal (e.g., as output by the predictive averaging modification component 1120). Post-processing filter 1130 may differ from another jitter filter that can be applied after adding any residual data (e.g., different from a jitter filter applied as a final stage before outputting the final reconstructed video signal). In one example, post-processing filter 1130 includes a sharpening filter. This is in... Figure 11BAs shown in the diagram, a sharpening filter is configured to sharpen the signal after upsampling. For example, as resolution increases from limited lower-level information, the output of upsampling may include a relatively blurred signal. A sharpening filter can help sharpen the output of upsampling in a way that modifies the data distribution of the residual data set to be added to the upsampled sample (e.g., reducing the number of non-zero values ​​and / or modifying values ​​so that the resulting distribution can be compressed more efficiently through a combination of run-length and Huffman coding). A sharpening filter may include a modified desharpening mask. This is discussed below regarding... Figure 14 It was described in more detail.

[0176] Figure 11B This illustrates how the enhanced upsampler 1155 with a sharpening filter can be viewed as a linear operation or a cascade of filters. Figure 11B The diagram illustrates a separable upsampling kernel (though a non-separable kernel may also be used in other instances). The separable upsampling kernel has two stages, 1112 and 1114, where one-dimensional convolutions are used to process each dimension of the frame, resulting in a two-dimensional convolution. The sharpening filter 1132 can also be applied as a two-dimensional convolution (or a series of one-dimensional convolutions). The coefficient values ​​for the upsampling kernel (e.g., 1110, 1112, or 1114) can be signaled according to a hierarchical coding standard. Each stage of the separable upsampling kernel may include a 4-tap upsampling filter. The coefficient values ​​for the sharpening filter 1132 can be signaled by the encoder using user data as described herein (e.g., using embedded coefficient values ​​and / or SEI user data). The coefficients used for the sharpening filter 1132 can be adjusted as different coding units or data blocks are upsampled. This can be implemented by extracting coefficient values ​​of predefined transform coefficients used as preserved symbols. Therefore, the sharpening filter 1132 can be adapted based on the image content. The coefficients used for the sharpening filter 1132 can be determined by the encoder and then transmitted to the decoder.

[0177] In some instances, upsampling can be enhanced using artificial neural networks. For example, a convolutional neural network can be used as part of an upsampling operation to predict the values ​​of upsampled pixels or signal elements. The use of artificial neural networks to enhance upsampling operations is described in WO 2019 / 111011 A1, which is incorporated herein by reference. In this case, a neural network upsampler can be used to perform signal processing operations to enhance higher levels of the signal. The neural network upsampler described herein is a specific, efficient “minConv” implementation that has been tested to operate at speeds sufficient to allow processing at common video frame rates (e.g., 30 Hz).

[0178] Figure 12AAn enhanced upsampler 1205 is shown, comprising a simple neural network upsampler 1210. An optional post-processing operation 1230 is also present, which can be similar to... Figure 11A The post-processing operation in 1130. In the enhanced upsampler 1205, the neural network upsampler 1210 is used as an alternative upsampler to the upsampler kernel 1110 (e.g., the upsampler kernel defined in the standard decoding process).

[0179] Figure 12B The enhanced upsampler 1205 is shown in more detail. In this example, the neural network upsampler 1210 includes two layers 1212 and 1216 separated by a nonlinearity 1214. By simplifying the neural network architecture to have this structure, upsampling is enhanced while still allowing real-time video decoding. For example, processing a frame may take about 1 ms, which allows decoding at frame rates of 30 Hz and 60 Hz (e.g., several frames every 33 ms and 16 ms, respectively).

[0180] Convolutional layers 1212 and 1216 may include two-dimensional convolutions. Each convolutional layer may apply one or more filter kernels of predefined size. In one case, the filter kernel may be 3×3 or 4×4. The convolutional layer may apply filter kernels that can be defined by a set of weight values, and may also apply a bias. The bias has the same dimension as the output of the convolutional layer. Figure 12BIn the example, the two convolutional layers 1212 and 1216 may share a common structure or function but have different parameters (e.g., different filter kernel weight values ​​and different bias values). Each convolutional layer can operate in different dimensions. The parameters of each convolutional layer can be defined as a four-dimensional tensor with size (kernel_size1, kernel_size2, input_size, output_size). The input of each convolutional layer may include a three-dimensional tensor with size (input_size_1, input_size_2, input_size). The output of each convolutional layer may include a three-dimensional tensor with size (input_size_1, input_size_2, output_size). The first convolutional layer 1212 may have an input_size of 1, so that it receives a two-dimensional input similar to a non-neural sampler as described herein. Example values ​​for these sizes are as follows: kernel_size1 and kernel_size2 = 3; for the first convolutional layer 1212, input_size = 1 and output_size = 16; and for the second convolutional layer 1216, input_size = 16 and output_size = 4. Other values ​​may be used depending on the implementation and empirical performance. With an output size of 4 (i.e., four channels output for each input element), this can be reconstructed as a 2×2 block representing the upsampled output of a given pixel. The parameters of each convolutional layer containing one or more of the layer size, filter kernel weight values, and bias values ​​can be signaled using the signaling methods described herein (e.g., via embedded coefficient signaling and / or SEI messages).

[0181] The input to the first convolutional layer 1212 can be a two-dimensional array, similar to other upsampler implementations described herein. For example, the neural network upsampler 1210 can receive a portion and / or the entire reconstructed frame (e.g., the base layer plus the decoded output enhanced by layer 1). The output of the neural network upsampler 1210 can include a portion and / or the entire reconstructed frame at a higher resolution, as in other upsampler implementations described herein. The neural network upsampler 1210 can therefore be used as a modular component, as in other available upsampling methods described herein. In one case, the selection of the neural network upsampler, for example at the decoder, can be signaled within the user data as described herein, for example in a flag within the header portion of the user data.

[0182] The nonlinear layer 1214 may include any known nonlinearity, such as the sigmoid function, tanh function, modified linear unit (ReLU), or exponential linear unit (ELU). Variations of common functions, such as so-called leaky ReLU or scaled ELU, may also be used. In one instance, the nonlinear layer 1214 includes a leaky ReLU—in which case the layer's output is equal to the input for input values ​​greater than 0 (or equal to 0) and equal to a predefined scale of the input for input values ​​less than 0, such as a * input. In one case, a may be set to 0.2.

[0183] exist Figure 12B In the example, convolutional layers 1212 and 1216 and post-processing operation 1230 can be viewed as a cascade of linear operations (with intermediate nonlinear operations). Therefore, the general configuration can be similar to Figure 11B The cascade of linear operations is shown. In both cases, filter parameters (e.g., filter coefficients) can be transmitted via the signal processing information described herein.

[0184] In one scenario, the sampler 1210 on the neural network may be incompatible with the predicted average modification performed by component 1120. Therefore, the use of the sampler 1210 on the neural network can be signaled by the encoder by setting the predicted_residual_mode_flag in the global configuration header of the encoded data stream to 0 (e.g., when the predicted residual mode is off). In another scenario, the use of the sampler 1210 on the neural network can be signaled by adding a predicted_residual_mode_flag value of 0 to a set of layer coefficient values ​​transmitted via user data (e.g., embedded transform coefficients and / or SEI user data).

[0185] In a variant of the onsampler in a neural network, post-processing operation 1230 may include an inverse transform operation. In this case, the second convolutional layer 1216 may output a tensor of size (size1, size2, number of coefficients), i.e., the same size as the input but with channels representing each direction within the directed decomposition. The inverse transform operation may be similar to the inverse transform operation performed in the layer 1 enhancement layer. In this case, the second convolutional layer 1216 may be viewed as outputting coefficient estimates of the onsampled coding unit (e.g., for a 2×2 coding block, the 4-channel output represents the A, H, V, and D coefficients). The inverse transform step then converts the multi-channel output into a two-dimensional set of pixels, e.g., the [A, H, V, D] vector of each input pixel is converted into a 2×2 pixel block in layer n. The inverse transform may include setting the values ​​of the coefficients carrying user data (e.g., H or HH) to zero before performing the transformation.

[0186] In the above example, the parameters of the convolutional layer can be trained based on a pair of layer (n-1) and layer n data. For example, the input during training may include reconstructed video data at a first resolution generated by applying one or more of the encoder and decoder paths, while the ground truth output used for training may include the actual corresponding content from the original signal (e.g., higher or second resolution video data, rather than upsampled video data). Thus, the neural network upsampler is trained to predict the input layer n video data (e.g., input video enhancement layer 2) as closely as possible given a lower resolution representation. If the neural network upsampler can generate an output that is closer to the input video than the contrasting upsampler, this will have the benefit of reducing the layer 2 residual, which will further reduce the number of bits required to transmit for the encoded layer 2 enhancement stream. Training can be performed offline based on various test media contents. The parameters generated by training can then be used in online prediction modes. These parameters can be transmitted to the decoder as part of an encoded byte stream (e.g., in header information) for picture groups and / or during over-the-air or wired updates. In one scenario, different video types may have different sets of parameters (e.g., film versus live sports). In another scenario, different parameters may be used for different parts of the video (e.g., the period of action versus a relatively static scene).

[0187] Figure 13 A schematic diagram is shown illustrating how an enhanced upsampler 1105 or 1205 can be implemented based on signal processing information (SPI) extracted from the encoded data stream. Figure 13 A switching arrangement is shown in which different forms of upsampling operations are performed based on signal processing information (i.e., enhanced upsampling operations are selectively performed based on said information).

[0188] exist Figure 13 In the diagram, the upsampling operation is shown as block 1305. Upsampling operation 1305 receives data to... Figures 1 to 7Upsampling is performed during the upsampling operation. Within upsampling operation 1305, there are at least two possible upsampling configurations—standard upsampler 1312 and enhanced upsampler 1314. Enhanced upsampler 1314 may be enhanced upsampler 1105 or 1205. Switch 1320 then receives signal processing information, which may include a flag value indicating whether to use enhanced upsampler 1314 (e.g., as signaled from the encoder or additionally determined based on the current operating conditions as described above). The default mode may be to use standard upsampler 1312 (e.g., as shown). The arrow indicates that when appropriate signal processing information is received, switch 1320 may be activated to switch to upsampling via enhanced upsampler 1314. As shown, enhanced upsampler 1314 may further receive signal processing information to configure enhanced upsampler 1314 (e.g., on one or more local or global bases relative to the coding units of a frame). Both the standard upsampler 1312 and the enhanced upsampler 1314 provide outputs for upsampling operation 1305.

[0189] In this example, at block 1320, residual data (R) is added after upsampling operation 1305 (i.e., after any enhancement operation). Dithering 1330 can be applied to the final output as a final operation before display. In some cases or configurations, such as if network congestion renders the residual data unacceptable and / or if upsampling operation 1305 is implemented as an “additional” upsampling operation applied to the output of the standard decoding process, residual data cannot be added at block 1320 (or block 1320 can be omitted). If upsampling operation 1305 is implemented as “additional” upsampling, enhanced upsampler 1314 can provide super-resolution output. In these cases, image quality is improved by adding dithering at the highest possible output resolution (e.g., an upgraded resolution beyond the standard output resolution produced by enhanced upsampler 1314).

[0190] Figure 14 An example desharpening mask 1400 is shown that can be used to implement a sharpening filter, for example... Figure 11A and 11B Post-processing operations 1130 or 1132 are performed within the video frame. In a preferred embodiment, the sharpening filter may be applied only to one color component of the video frame, i.e., the luminance or Y plane, to generate a filtered luminance plane. The sharpening filter may not be applied to the U or V chrominance planes.

[0191] Figure 14 The sharpening filter can be implemented as a convolution of the input image f and a weighted Laplacian kernel L:

[0192] z = f * L

[0193] Where f is the input image, z is the output (filtered) image, and L is... Figure 14 The filter kernel shown. In Figure 14 In this context, S and C are parameters that control the effect of the sharpening filter. In one case, S can be 1 and only C is controllable. In other cases, C = 4S + 1. In both cases, it may be necessary to signal only one parameter value (S or C). In these cases, the signaled parameter can include integer or floating-point values. In some cases, 0 ≤ S ≤ 1, where S = 0 corresponds to no filtering effect (i.e., z = f) and s = 1 produces the strongest filtering effect. The value of S (and / or C) can be selected by the user configuration or set depending on what is being processed. In the latter case, the value of S or C can vary in each coding block and therefore can be signaled in the embedded transform coefficient signaling used for the coding block (e.g., within the user data used for the coding block). The filter can be called an unsharpening mask because it uses a negative blur (or unsharpening) form of the image as a mask to perform sharpening, which is then combined with the original image. In other instances, the sharpening filter can include any linear and / or nonlinear sharpening filter.

[0194] Examples of user data messaging

[0195] As described in the examples herein, a signal processor (e.g., computer processor hardware) is configured to receive data and decode it (“decoder”). The decoder obtains a reproduction of the signal at a first (lower) quality level and detects user data that specifies optional upsampling and signal enhancement operations. The decoder reconstructs a reproduction of the signal at a second (next higher) quality level based at least in part on the user data. Some examples of user data will now be described in more detail.

[0196] In the first set of instances, signal processing information is embedded in one or more values ​​received in one or more encoded data layers transmitted within the encoded data stream. These values ​​are associated with transform coefficients of elements processed to derive the signal during decoding; for example, they may include values ​​of predefined transform coefficients from different sets of transform coefficients generated by the encoded transform.

[0197] For example, bits in the bitstream of the encoded data stream can be used to signal the presence of user data in lieu of one of the coefficients associated with the transform block (e.g., specifically, the HH coefficients in the case of a 4×4 transform). These bits may include the user_data_enabled bit, which may be present in the global configuration header of the encoded data stream.

[0198] In some instances, the encoding of user data, replacing one of the coefficients, can be configured as follows: If the bit is set to "0", the decoder will interpret the data as correlation transform coefficients. If the bit is set to "1", the data contained in the correlation coefficients is considered user data, and the decoder is configured to ignore the data—that is, decode the correlation coefficients as zero.

[0199] User data transmitted in this manner can be used to enable the decoder to obtain supplementary information, including, for example, various feature extractions and guides. Although the instances claimed herein involve optional upsampling and signal enhancement operations, user data may also be used to transmit other optional parameters related to implementations outside the standardized implementation.

[0200] In one case, the user_data_enabled variable can be a k-bit variable. For example, user_data_enabled can include a 2-bit variable with the following values:

[0201] user_data_enabled Type value 0 Discontinued 1 Enable 2 bits 2 Enable 6-bit 3 reserve

[0202] In this case, user data specifying optional upsampling and signal enhancement operations can be embedded into the last u significant bits of one or more of the decoded coefficient data set (e.g., within the encoded residual coefficient data).

[0203] When user data is enabled—for example, to transmit signal processing information as described in the examples in this article—the “in-loop” processing of the transform coefficients can be modified. Two examples are shown below. Figure 8A and 8B Additionally, the decoding of the transform coefficients can be adjusted so that when user data is enabled, the value of a particular transform coefficient (e.g., H or HH) is set to 0 before the inverse transform is performed. In the cases described in the table above, if 2 bits are used (e.g., user_data_enabled = 1), the value of the transform coefficient used to carry user data can be right-shifted (e.g., shifted by bits) by 2 bits (>>2), or if 6 bits are used (e.g., user_data_enabled = 1), the value can be right-shifted (e.g., shifted by bits) by 6 bits (>>6). In one case, if the length of the transform coefficient value is b bits, where b > u, u is the length of the user data in bits (e.g., 2 or 6 in the table above), the remaining bu bits of the transform coefficient can be used to carry the transform coefficient value (e.g., a more quantized integer value compared to a full b-bit representation). In this case, the user data and the transform coefficient value can be divided across b bits. In other simpler cases, user data can be extracted, and the transform coefficient value can be set to 0 (i.e., so that the transform coefficient value has no effect on the output of the inverse transform).

[0204] In some instances, user data can be formatted according to a defined syntax. This defined syntax divides the user data into header data and payload data. In this case, decoding the user data may include analyzing a first set of values ​​received in one or more encoded data layers to extract the header data, and analyzing a second set of subsequent values ​​received in one or more encoded data layers to extract the payload data. The header data may be set to a first set of defined bits. For example, in the instance where the user data is defined as a 2- or 6-bit value, the first x values ​​may include the header data. In one case, x may be equal to 1, such that the first value of the user data (e.g., the transform coefficient value of the first coding unit or data block of a given frame or plane of video) defines the header data (e.g., 2 or 6 bits of the first value define the header data).

[0205] In some instances, the header data may indicate at least whether optional upsampling and signal enhancement operations are enabled, and whether any other user data is signaled. In the latter case, after user data associated with optional upsampling and signal enhancement operations has been signaled, the remaining values ​​within the defined transform coefficients can be used to transmit other data (e.g., data unrelated to optional upsampling and signal enhancement operations). With 2-bit user data values, two 1-bit flags can be used to signal these two variables. With 6-bit user data values, one or more types of optional upsampling and signal enhancement operations can be signaled (e.g., using 3-bit integers to index lookup table values), and a 1-bit flag can indicate whether the user data also includes additional post-processing operations. In this case, the type can indicate what type of neural network upsampler will be used, and the 1-bit flag can indicate whether a sharpening filter will be applied. It should be understood that different format combinations can be used, for example, a 6-bit value can be constructed from three consecutive 2-bit values.

[0206] Generally, header data indicates global parameters of the signal processing information, while payload data indicates local parameters. The separation between global and local parameters can also be implemented in other ways; for example, global parameters can be set within the SEI message user data, while local parameters can be set within the embedded transform coefficient values. In this case, header data may not be present within the embedded transform coefficient values, as it may be carried alternatively within the SEI message user data.

[0207] The following describes some user data implementation examples related to the LCEVC standard. It should be noted that similar syntax can be used with other standards and implementations. In these examples, the optional signal augmentation operation is referred to as a “super-resolution” mode. For example, if the described neural network upsampler is used, this can be described as producing a “super-resolution” upgrade, where the level of detail in the higher-resolution image frame is greater than that of simple comparative upsampling (e.g., the neural network is configured to predict additional detail in the higher-resolution image frame).

[0208] In some instances, the signal includes a video signal, and a first header structure is used for Instant Decoding Refresh (IDR) picture frames and a second header structure is used for non-IDR picture frames. In this case, the IDR picture frame may carry a global user data configuration, while subsequent non-IDR picture frames may carry locally applicable user data (e.g., data associated with a particular non-IDR picture frame). An IDR picture frame includes a picture frame in which the encoded data stream contains a block of global configuration data, wherein the picture frame does not refer to any other picture used in the decoding process of the picture frame, and wherein subsequent picture frames in decoding order do not refer to any picture frame preceding the IDR picture frame in decoding order. The IDR picture should appear at least when the IDR picture used for the base decoder appears. In one embodiment, locally applicable user data may be signaled as one or more changes or differentials from information signaled within the global user data configuration.

[0209] In a 6-bit user data implementation scheme compatible with LCEVC, the first few bits of the user data can be constructed as follows to make signaling suitable for embedding user data in 6-bit groups (in the table, u(n) indicates the number of unsigned bits used for the variables indicated in bold):

[0210]

[0211] Table 1 – 6-bit User Data for LCEVC – Global User_Data_Configuration for IDR Frames

[0212] In a 2-bit user data implementation scheme compatible with LCEVC, the first few bits of the user data can be constructed as follows to make signaling suitable for embedding user data in 2-bit groups:

[0213]

[0214]

[0215] Table 2 – 2-bit User Data for LCEVC – Global User_Data_Configuration for IDR Frames

[0216] In the above examples where user data is embedded in the LCEVC stream according to the LCEVC embedded user data syntax, the user data configuration information, as shown in the examples in Table 1 or Table 2, is extracted by the decoder from the user data bits of the first few coefficients of the IDR frame. In some cases, the user data configuration defined for the picture frame (e.g., the User_Data_Configuration mentioned above) is maintained until subsequent IDR frames. In other cases, it is possible to signal changes to the user data configuration of non-IDR frames using flag bits in the first few user data bits of non-IDR frames (e.g., for LCEVC, the user data bits of the first few coefficients within the embedded user data). Table 3 below shows examples in the context of the two-bit cases in Table 2:

[0217]

[0218] Table 3 – 2-bit User Data for LCEVC – User_Data_Picture_Configuration for Non-IDR Frames

[0219] Although the encoding format for residual data and embedded context information is LCEVC in the above example, in other examples, the encoding format for residual data and embedded context information can be VC-6 or another signal encoding standard.

[0220] In the above example, the value in the "optional_super-resolution_type" variable of the first user data byte can be set to allow for the optional use of a sharpening filter cascaded with a signaling and separable upsampling filter (e.g., a modified desharpening mask filter as described above) and the application of prediction residuals (e.g., as described above). Figure 11B (As shown). Sharpening filters can be applied before applying residual data from higher residual sublayers and before applying statistical jitter (e.g., as shown). Figure 13 (As shown) is applied before. Another value in "optional_super-resolution_type" can be set for optional use of the same sharpening filter, but after the residual data of the higher residual sublayer is applied. For example, in one mode, it can be... Figure 11B After block 1120, but Figure 11BThe residual addition at block 1320 is performed before block 1132. In this mode, the sharpening filter can still be applied (optionally) before statistical jitter. This mode allows for backward compatibility with decoders that cannot understand signaling or process filters, for example, allowing the sharpening filter to be applied modularly outside the loop. In other instances, the sharpening filter can be applied with default configuration and intensity in the absence of any "super-resolution_configuration_data" specified as described above (i.e., s_configuration_data_signalled == 0), while if "super-resolution_configuration_data" is signaled, the data may include information about the configuration and intensity of the filter to be applied.

[0221] Similarly, in some instances, another value in the "optional_super-resolution_type" of the first user data byte above may correspond to a convolutional neural network as an alternative to a separable upsampling filter (e.g., as referenced). Figure 12A and 12B The optional use and application of the predicted residuals (described herein) occurs before the application of residual data from higher residual sublayers but before the application of optional and signaled statistical jitter. The mode can be used in conjunction with the modes described above (e.g., part of a plurality of available modes to be signaled) and independent of those modes. In some instances, convolutional neural network filters (e.g., as described herein) can be applied with the default configuration and coefficient set in the absence of any specified "super-resolution_configuration_data" (i.e., s_configuration_data_signalled == 0), whereas if "super-resolution_configuration_data" is signaled, additional data within the user data includes information about the configuration and coefficient set to be applied.

[0222] In some instances, the upsampling of the convolutional neural networks described herein can be used for multiple upsampling passes. For example, LCEVC can define a scaling mode that indicates whether upsampling is to be used for multiple levels in a hierarchical, layered format (e.g., more similar to...). Figure 2 and 3(VC-6 example). To communicate this, the “scaling_mode” parameter can be defined as part of the LCEVC standard. The value of this parameter can indicate whether scaling will be applied (e.g., 0 = no scaling) and whether scaling will be applied in one or two dimensions (e.g., 1 = one dimension and 2 = two dimensions, or 1 = horizontal scaling, 2 = vertical scaling and 4 = scaling in both horizontal and vertical dimensions). In this case, if for the LCEVC implementation, “scaling_mode_level1” = 2 and “scaling_mode_level2” = 2 (e.g., indicating two-dimensional scaling), then sampling on the convolutional neural network can be used to first reconstruct the initial image at level 1 resolution (in this non-limiting example, scaling 2:1 in both directions), and then—after adding sub-level 1 residual data correction (if present)—reconstruct the initial image at level 2 resolution. In these cases, the default configuration of the network used for each upgrade process (i.e., tier 1 relative to tier 2) may be different, where the transmitted “super-resolution_configuration_data” specifies the different configuration data used for the two upgrade processes.

[0223] As an alternative to the embedded transform coefficient example above, or in combination with those examples, it is specified that user data for optional upsampling and signal enhancement operations can be encapsulated in an SEI (Supplemental Enhancement Information) message.

[0224] In video coding implementations, SEI messages are typically used to convey information related to, for example, the color and luminance levels of the reconstructed video to be displayed. While SEI messages can be used to aid processes related to decoding, display, or other purposes, these processes may not be necessary for constructing luminance or chrominance samples through standard decoding procedures. The use of SEI messages can therefore be considered an optional variation to allow for increased functionality.

[0225] In this instance, the SEI message can be configured to carry signaling information for optional enhancements to the signaling operation. For example, one or more of the "Reserved" or "User Data" sections of the defined SEI message syntax can be used to carry this signaling information. The SEI message can reside in a bitstream of encoded data and / or be transmitted in a manner other than that found in the example bitstream described herein.

[0226] When used with LCEVC, the example syntax for decoding the SEI payload is as follows (where u(n) indicates an n-bit unsigned integer as described above, and f(n) indicates a bit string of a fixed pattern):

[0227]

[0228] Table 4 – General SEI Message Syntax

[0229] In this scenario, signaling for the current instance can be carried within one or more of the registered user data, unregistered user data, and reserved data in the SEI message. An example of the syntax for unregistered user data and reserved data is shown below:

[0230]

[0231] Table 5 – Syntax for User Data Not Registered with SEI Messages

[0232]

[0233] Table 6 – Preserving SEI Message Syntax

[0234] User data not registered with the SEI message may be preferred. In some cases, the header can be used to identify signal processing information related to the enhancement operation. For example, a Universally Unique Identifier (UUID) can be used to identify a specific type of signal processing information. In one case, the sharpening filter or sampler on the neural network to be applied may have its own UUID, which may be a 16-byte value. Following the UUID, the payload data described below may exist.

[0235] If used within LCEVC, subsequent syntax within LCEVC can be used to process SEI messages:

[0236]

[0237] Table 7 – Payload for processing additional information

[0238] The advantage of SEI messages is that they can be processed before the decoding loop used for the received data. Therefore, SEI messages may be preferred when transmitting a global configuration of optional enhancement operations (e.g., because there may be more time to configure these enhancement operations before receiving frame data). For example, SEI messages can be used to indicate the use of sharpening filters as described herein. In some cases, if local signal processing information is also required, this can be advantageously carried within embedded transform coefficients, where the signal processing information can be decoded and accessed within the loop (e.g., for one or more coding units or data blocks). In some cases, the combination of SEI messages and embedded coefficient data can have a synergistic effect, providing advantages over using these messages and data alone, combining the benefits of global and local processing and availability. For example, the use of sharpening filters can be indicated by means of SEI messages and used for... Figure 14 The encoding unit dependency value S of the sharpening filter (where C = 4S + 1) can be transmitted within the transform coefficient value.

[0239] In addition to or replacing the embedded transform coefficients and SEI method described above, another signaling method can be to signal an optional upsampling method to the decoder using a specific combination of signals from the standard upsampling method. For example, the decoder can be configured to apply the optional upsampling method based on a specific combination of parameters defined within a standard bitstream (e.g., LCEVC or VC-6). In one case, the optional upsampling method can be signaled to the decoder by signaling a specific custom configuration of disabling the prediction residual mode and the kernel coefficients of the standard upsampling method. For example, a simplified neural network upsampling method can be implemented by setting the prediction residual mode flag to 0 and signaling the coefficients (or other parameters) of a simplified neural network upsampling method within the syntax specified for a non-neural network upsampling method that forms part of the LCEVC.

[0240] In some implementations, the payload of data configured for one or more signal processing operations may be independent of the method used to transmit such data within the bit stream. For example, the payload may be transmitted in a similar manner within embedded transform coefficient values ​​and / or SEI messages.

[0241] In an LCEVC instance, the payload transmission frequency can be equal to the frequency of the "Global Configuration" block in the LCEVC bitstream. This allows certain aspects of signal processing operations to be updated per group of pictures (GOP). For example, the sharpening filter strength and / or the type of sharpening filter to be applied can be updated at the per-GOP update frequency, including the ability to deactivate sharpening filters used for the entire GOP. A GOP can comprise a group of frames associated with a given IDR picture frame.

[0242] In some instances, if no payload carrying signal processing information for one or more signal processing operations is transmitted, it can be assumed that the one or more signal processing operations are disabled and / or that default operations will be applied instead. For example, if no payload is present, it can be assumed that a sharpening filter will not be used and / or that a standard upsampler will be used one-to-one instead of a neural network upsampler. This then allows the encoded data stream to behave according to standard specifications (e.g., LCEVC or VC-6) without unintended signal modifications.

[0243] The syntax for a sample payload used for the sharpening filter is described below. This payload is one byte (8 bits), with the first 3 bits used for type definition and the last 5 bits used for configuration data.

[0244]

[0245] Table 8 -- Sharpening (S) - Filter Configuration

[0246] In this example, the `super_resolution_type` variable defines the benchmark for the sharpening filter relative to its default value, and where the filter is applied during decoding and encoding. Instances of the `super_resolution_type` set are listed in the table below.

[0247]

[0248] Table 10 -- Sharpening Filter Types

[0249] For type 2 and earlier types in the examples above, the last 5 bits of the payload data specify the strength of the sharpening filter to be applied. The application of the sharpening filter can use real numbers to determine the weighting used for the filter. For the cases of 0 and 1 above, no strength is transmitted, and a default real value of 0.15 can be used. In this example, the last 5 bits of the payload data may include the variable `super_resolution_configuration_data`, which defines the strength of the sharpening filter. In one case, the last 5 bits may define an unsigned integer value ranging from 0 to 31 (inclusive). This value can then be converted into a real number used to configure the strength of the sharpening filter using the following:

[0250] S-filter strength = (super_resolution_configuration_data + 1) * 0.1.

[0251] When the sharpening filter intensity changes, it can be transmitted as an embedded transform coefficient value as described herein. The first-level configuration can be set using variables transmitted for the IDR image frames maintained by the GOP. This configuration can be assumed to apply unless overridden by values ​​transmitted within one or more embedded transform coefficients. For example, a new super_resolution_configuration_data value can be transmitted, or a signed change to the GOP super_resolution_configuration_data value can be transmitted (e.g., original GOP super_resolution_configuration_data + / - m, where m is transmitted in user data).

[0252] In LCEVC, SEI messages can be encapsulated within an "Extra Information" block within the LCEVC bitstream (e.g., as shown in Table 7 relative to SEI messages). Within the LCEVC standard, the Extra Information block can carry both SEI data and Video Availability Information (VUI) data. In one case, for example as an alternative to using SEI messages, signal processing information can be carried within this "Extra Information" block. In the LCEVC standard, it can be defined that the decoder can skip the "Extra Information" block if it does not know the type of data within the block. This is possible by defining the block as a predefined size (e.g., 1 byte). The size of the "Extra Information" block can be reduced compared to SEI messages (e.g., if the SEI message uses a 16-byte UUID, the overhead information is 3 bytes instead of 21 bytes). One approach can be configured based on one or more of the following: the overall data rate of the encoded data stream and the GOP length.

[0253] Other variations

[0254] Some other variations of the instances described in this article will now be described.

[0255] In the case of optional super-resolution mode, this can be selectively performed based on a metric of the processing power available at the decoder, as described above. In this case, the decoder decodes the optional super-resolution configuration (e.g., from user data as described above), but based on a less complex separable upsampling method (e.g., ...). Figure 13 Switch 1320 in the middle is set to perform upgrade and preliminary signal reconstruction operations as a conventional upsampler 1312 rather than an enhanced upsampler 1314, wherein the main syntax of the stream (e.g., according to) is followed. Figure 11A and 11B The application of predictive residuals is specified in block 1120 of the code. In this way, when decoding is performed on more powerful hardware or when the hardware has more resources available for processing (e.g., no other extremely power-intensive applications running in parallel, relatively sufficient battery power), the same stream sent to the same decoder can be decoded seamlessly with high quality. Conversely, when decoding is performed on hardware with lower processing resource availability, processing power and battery consumption can be automatically saved by defaulting to simpler upsampling methods and decoding with lower processing power requirements. This allows for the use of more complex methods while ensuring appropriate backward compatibility with lower-power devices by providing a suitable backup or default kernel upgrade for replacing more complex super-resolution methods.

[0256] In another instance, a signal processor (e.g., computer processor hardware) is configured to receive data and encode it (i.e., configured as an "encoder"). The encoder generates a downsampled reproduction of the source signal at a first (lower) quality level according to a first downsampling method. The encoder then generates a predicted reproduction of the signal at a second (higher) quality level according to a first upsampling method based on the downsampled reproduction of the signal at the first quality level, and correspondingly analyzes the residual data (e.g., at a predefined difference level, which may be a difference of 0 representing a "perfect" reconstruction) required to properly reconstruct the source signal. Based on a metric generated at least in part by processing the residual data, the encoder selects a second combination of the downsampling and upsampling methods for processing the signal. In some non-limiting embodiments, when an ideal upsampling method is not supported in the list of standard upsampling methods provided by the encoding format, it is optional for the encoder to signal to the decoder a default upsampling method for backward compatibility and an upsampling method in the user data.

[0257] In some instances, the process involves iteratively selecting undersampling and upsampling methods, based on a process aimed at optimizing a metric generated at least in part by processing the residual data produced at each iteration. In some instances, the metric to be optimized may also depend at least in part on the bit rate available for encoding the residual data.

[0258] In some instances, the encoder may produce a reconstructed signal at a first (lower) quality level according to a first downsampling method, and further encode the reconstructed signal at a second (higher) quality level according to a first upsampling method using a first encoding method to produce a more accurate metric generated at least partially from the residual data required to properly reconstruct the source signal. In one case, the process is iterated multiple times to optimize the metric generated at least partially from the residual data.

[0259] In some specific instances, the downsampling method may comprise a nonlinear downsampling method obtained by cascading a linear downsampling method (e.g., for example, a separable 12-tap filter with custom kernel coefficients) using at least one image processing filter. For example, the method may be corresponding to a reference... Figures 10A to 14 The described cascaded linear upsampling method includes a downsampling method. In other instances, such as at the encoder, the downsampling method may incorporate a method utilizing a convolutional neural network based on the described upsampling method. As previously discussed, the upsampling and downsampling methods at the encoder can be asymmetric because the residual data can be used to compensate for the difference between the output generated by downsampling followed by upsampling and the original signal fed to the downsampling. In this case, upsampling may be preferable to a simpler method that can be implemented on a less resource-intensive decoder.

[0260] In some instances, methods for encoding signals include encoding lower-resolution layers of a hierarchical, hierarchical format (e.g., Figure 6 Level 1 encoding in the hierarchy); encoding a higher resolution level in a hierarchical format based on hierarchy, where the higher resolution level is encoded using data generated during the encoding of the lower resolution level (e.g., ...). Figure 6 The method may further include: determining signal processing information for performing one or more signal processing operations to enhance data within the higher resolution layer, the one or more signal processing operations being performed to reconstruct a portion of the higher resolution layer using data generated during encoding at the lower resolution layer; and encoding the signal processing information as a portion of the encoded data stream.

[0261] In some instances, signal processing information that determines one or more signal processing operations includes: a reduced-resolution frame for processing the signal; and determining an ideal signal processing operation for the frame based on the reduced-resolution frame. For example, video frames may be reduced (e.g., according to...). Figure 2 and Figure 3 The image is decimated (or otherwise passed through a downsampling pyramid), and then frame metrics are calculated based on the reduced-resolution format of the frame. For example, the metrics may indicate the complexity of the frame and thus indicate a reduced-bit-rate image processing operation that can generate a higher-level set of residuals. In one case, a fast bisection search can be performed using the image's decimation format and one or more metrics as references. By applying tests to the calculated metrics, the ideal image processing method for a particular frame can be determined.

[0262] In some instances, bits in the decoded byte stream may be embedded in some residual data coefficients to signal additional information to the decoder. Furthermore, a specific set of symbols in a particular residual data set should not be interpreted as actual residual data, but rather as contextual information used to inform signal enhancement operations. In some cases, instead of parameters used for enhancement operations, some reserved symbols can be used to signal specific types of attenuation, informing the decoder of post-processing operations applicable to the corresponding regions of the signal to improve the quality of the final signal reconstruction. In these instances, when one or more impairments are detected in the encoding of the signal at the first quality level that cannot be properly corrected with residual data at the target bit rate, the encoder may use a set of reserved symbols from the residual data set of the second quality level's residual data sequence to signal to the decoder the type and / or location of the desired attenuation.

[0263] While the examples described herein depict signaling embedded within a single transform coefficient, in other instances, signaling may be embedded within the values ​​of more than one transform coefficient. For instance, user data, as described herein, can be multiplexed across a set of transform coefficient values ​​used for one or more initial pixel rows that may not be visible in the reproduced output under certain circumstances. Therefore, in some instances, contextual information may be embedded within more than one step of the residual data.

[0264] In addition to signaling parameters related to the sharpening filter and the sampler on the convolutional neural network, contextual signaling information (e.g., information embedded in the residual data) may also include data corresponding to blocking attenuation. For example, the decoder may perform deblocking post-processing operations in the signal region corresponding to the residual coefficients containing the preserved symbols. In some cases, the contextual signaling information may indicate different strengths of the decoder's deblocking filter. The decoder may deblock the signal using, for example, the deblocking method described in patent US9445131B1, "De-blocking and de-banding filter with adjustable filter strength for video and image processing," where QP information for a given adjacent region is embedded in the symbols (the patent is incorporated herein by reference). In these variations, the decoder may apply the deblocking method within the loop before applying residual data decoded sequentially from data containing embedded information about blocking attenuation. In other cases, the decoder may apply the deblocking method after the initial reproduction of the signal at a second quality level has been combined with the decoded residual data.

[0265] Similar to the deblocking variants described above, in some variants, contextual signaling information (e.g., information embedded in the residual data) includes parametric filtering to correct for banding, ringing, and softening attenuation. In these cases, the decoder can perform signal enhancement operations in the signal region corresponding to the residual coefficients containing the preserved symbols, including debanding, deranging, edge enhancement, range equalization, and sharpening post-processing operations.

[0266] In some variations, contextual signaling information (e.g., information embedded in the residual data) contains data corresponding to the risk of chromaticity inversion loss in the case of color conversion from a wide color gamut to a standard color gamut. For example, this loss might be due to limitations of the transformation LUT (for a "lookup table"). In one case, before applying the color conversion method, the decoder preserves the color values ​​in a signal region corresponding to the contextual signaling information contained within the preservation symbols.

[0267] In some variations, contextual signaling information (e.g., information embedded in the residual data) includes data corresponding to quantization noise reduction. In some cases, the decoder applies a denoising method to the signal region corresponding to the residual coefficients containing the preserved sign. The denoiser can be applied either inside or outside the loop. Similarly, in some variations, contextual signaling information embedded in the residual data includes data corresponding to film grain loss and / or camera noise. In some cases, the decoder applies a statistical jitter method to the signal region corresponding to the residual coefficients containing the preserved sign. In some implementations, statistical jitter is applied within the loop at multiple levels of a hierarchical structure, for example, at both the resolution of a given quality level and the resolution of subsequent (higher) quality levels.

[0268] Depending on certain variations, the embedded information may include watermarking information. In one case, the watermarking information can be used to identify and verify the encoder that generated the data stream. In another case, the watermarking information may contain information about the time and location of the encoding. In some cases, the watermarking information may be used, for example, to identify the nature of the signal. The watermarking information may instruct the decoder to begin applying watermarking to the decoded signal.

[0269] In some variations, user data as described herein (including possible additional user data following signal processing information related to enhanced operation) may indicate compatibility information, which may include any of the following: the method of signal generation, the specific encoder type used to generate the signal, and licensing information associated with the signal and / or the encoder type that generated the signal. Compatibility information can be used by the decoder to initiate compatibility actions when it detects a mismatch between the compatibility information and a record (e.g., a valid license used to generate the signal). In this case, the decoder may, for example, initiate a compatibility process for the signal, such as interrupting the display or playback of the signal, sending a request to the source of the transmitted signal to obtain a valid license, etc.

[0270] In other variations, user data may identify objects in the signal, such as unique identifiers known to the decoder. User data may also include tags associated with one or more elements of the signal. For example, tags may include identifying whether an end user of the signal can select an element of the signal. In other cases, tags may include identifying whether an element of the signal can be linked to an action the end user of the signal is to take, such as clicking the element and / or linking to a different signal / webpage. In yet another case, tags may include identifying elements of the signal as belonging to a category, such as a video category or an object category. For example, an element may represent a person, and a tag would identify who that person is. Alternatively, an element may represent an object, and a tag may identify what that object is. Alternatively, the tag may identify what category the object belongs to. Generally, a category may include the association between the element and an identifier category, such as the category to which the element belongs.

[0271] In some variations, reserved symbols can be used to embed different secondary signals as part of an encoded stream, which are encoded with a given public key and can only be decoded by a decoder that knows the existence of the secondary signals and the private key corresponding to the public key used to encrypt the secondary signals.

[0272] Although examples have been described in the context of hierarchical coding formats, contextual signal information can also be embedded in encoded data generated using non-hierarchical coding formats. In these cases, signal processing information can be embedded at the macroblock level using the reserved symbol set in the quantization coefficients.

[0273] Example device for implementing a decoder or encoder

[0274] refer to Figure 15 A schematic block diagram of an example of device 1500 is shown.

[0275] Examples of device 1500 include, but are not limited to, mobile computers, personal computer systems, wireless devices, base stations, telephone devices, desktop computers, laptop computers, notebook computers, netbook computers, mainframe computer systems, handheld computers, workstations, network computers, application servers, storage devices, consumer electronic devices such as cameras, portable camcorders, mobile devices, video game consoles, handheld video game devices, peripheral devices such as switches, modems, routers, vehicles, etc., or generally any type of computing or electronic device.

[0276] In this example, device 1500 includes one or more processors 1501 configured to process information and / or instructions. The one or more processors 1501 may include a central processing unit (CPU). The one or more processors 1501 are coupled to a bus 1511. Operations performed by the one or more processors 1501 may be implemented by hardware and / or software. The one or more processors 1501 may include multiple processors located in the same location or multiple processors located in different locations.

[0277] In this example, device 1501 includes computer-usable memory 1512 configured to store information and / or instructions for the one or more processors 1501. Computer-usable memory 1512 is coupled to bus 1511. Computer-usable memory may include one or more of volatile memory and non-volatile memory. Volatile memory may include random access memory (RAM). Non-volatile memory may include read-only memory (ROM).

[0278] In this example, device 1500 includes one or more external data storage units 1580 configured to store information and / or instructions. The one or more external data storage units 1580 are coupled to device 1500 via I / O interface 1514. The one or more external data storage units 1580 may include, for example, a magnetic disk or optical disk and a disk drive or solid-state drive (SSD).

[0279] In this example, device 1500 further includes one or more input / output (I / O) devices 1516 coupled via I / O interface 1514. Device 1500 also includes at least one network interface 1590. Both I / O interface 1514 and network interface 1517 are coupled to system bus 1511. The at least one network interface 1517 enables device 1500 to communicate via one or more data communication networks 1590. Examples of data communication networks include, but are not limited to, the Internet and local area networks (LANs). The one or more I / O devices 1516 enable a user to provide input to device 1500 via one or more input devices (not shown). The one or more I / O devices 1516 enable information to be provided to the user via one or more output devices (not shown).

[0280] exist Figure 15In the diagram, (signal) processor application 1540-1 is shown loaded into memory 1512. This can be executed as (signal) processor processing 1540-2 to implement the methods described herein (e.g., to implement a suitable encoder or decoder). Device 1500 may also include additional features, not shown for clarity, including an operating system and additional data processing modules. (Signal) processor processing 1540-2 can be implemented by means of computer program code stored in a memory location within a computer-available non-volatile memory, a computer-readable storage medium within one or more data storage units, and / or other tangible computer-readable storage media. Examples of tangible computer-readable storage media include, but are not limited to, optical media (e.g., CD-ROM, DVD-ROM, or Blu-ray), flash memory cards, floppy disks, or hard disks, or any other medium capable of storing computer-readable instructions, such as firmware or microcode or application-specific integrated circuit (ASIC) in at least one ROM or RAM or programmable ROM (PROM) chip.

[0281] The device 1500 may therefore include a data processing module executable by one or more processors 1501. The data processing module may be configured to contain instructions for implementing at least some of the operations described herein. During operation, one or more processors 1501 initiate, run, execute, interpret, or otherwise perform the instructions.

[0282] While at least some aspects of the examples described herein with reference to the figures include computer processes executed in a processing system or processor, the examples described herein also extend to computer programs suitable for putting the examples into practice, such as computer programs on or within a carrier. The carrier can be any entity or device capable of carrying a program. It should be understood that device 1500 may include, compared to... Figure 15 The device 1500 may contain more, fewer, and / or different components as described. It may be located in a single location or distributed across multiple locations. These locations may be local or remote.

[0283] The techniques described herein may be implemented in software or hardware, or in a combination of software and hardware. They may include configuration devices to implement and / or support any and all of the techniques described herein.

[0284] The above embodiments should be understood as illustrative examples. Other embodiments are also contemplated.

[0285] It should be understood that any feature described with respect to any embodiment may be used alone or in combination with other described features, and may also be used in combination with one or more features of any other embodiment, or in any combination of any other embodiment. Furthermore, equivalents and modifications not described above may be employed without departing from the scope of the invention as defined by the appended claims.

Claims

1. A method of decoding a signal, comprising: obtaining an encoded data stream, the encoded data stream encoded by an encoder according to a hierarchical format based on levels; analyzing the encoded data stream to determine signal processing information signaled by the encoder; and reconstructing a higher resolution level of the signal from a lower resolution level of the signal, characterized in that: one or more signal processing operations are selectively performed to enhance the higher resolution level based on the determined signal processing information, wherein selectively performing one or more signal processing operations to enhance the higher resolution level comprises: determining operational parameters of a decoder performing the decoding, wherein the operational parameters are determined based on properties and / or conditions of the decoder; in response to a first set of operational parameters, performing the one or more signal processing operations to enhance the higher resolution level using signal processing parameters within the determined signal processing information; and in response to a second set of operational parameters, omitting the one or more signal processing operations or replacing the one or more signal processing operations with default signal processing operations.

2. The method of claim 1, wherein at least part of data corresponding to the signal processing information is embedded in one or more values received in one or more encoded data layers transmitted within the encoded data stream, wherein the values are associated with transform coefficients processed to derive elements of the signal during the decoding.

3. The method of claim 2, wherein the signal processing information is embedded in one or more values of predefined transform coefficients within different sets of transform coefficients generated by an encoding transform.

4. The method of claim 1, wherein at least part of data corresponding to the signal processing information is encoded within an additional information payload.

5. The method of claim 4, wherein the additional information payload comprises one or more supplemental enhancement information messages.

6. The method of claim 1, wherein at least part of data corresponding to the signal processing information is determined based at least in part on a predefined set of values of configuration data for the signal, the configuration data configuring one or more signal processing operations other than the signal processing operations used to enhance the higher resolution level of the signal.

7. The method of any of claims 1 to 6, wherein the one or more signal processing operations are selectively performed prior to adding residual data for the higher resolution level of the signal; and / or the one or more signal processing operations provide a super-resolution signal; and / or the one or more signal processing operations are implemented as part of an upsampling operation that generates the higher resolution level of the signal from the lower resolution level of the signal.

8. The method of any of claims 1 to 6, comprising: using the determined signal processing information to identify signal processing operations used to enhance the higher resolution level; determining whether a decoder performing the decoding is able to implement the identified signal processing operations; ignoring the determined signal processing information in response to the decoder being unable to implement the identified signal processing operation; and performing the determined signal processing operation parameterized by the determined signal processing information in response to the decoder being able to implement the identified signal processing operation.

9. The method of any of claims 1-6, comprising: determining a resource usage metric of a decoder performing the decoding; comparing the resource usage metric to a resource usage threshold; performing the one or more signal processing operations to enhance the higher resolution tier based on the determined signal processing information in response to the comparison indicating no limit on resource usage of the decoder; and omitting the one or more signal processing operations during the reconstruction in response to the comparison indicating a limit on resource usage of the decoder.

10. The method of any of claims 1-6, wherein the one or more signal processing operations comprise a sharpening filter applied in addition to an upsampling operation used for the reconstruction that generates the higher resolution tier of the signal from the lower resolution tier of the signal.

11. The method of claim 10, wherein the determined signal processing information indicates at least one coefficient value of an unsharp mask.

12. The method of claim 11, wherein the determined signal processing information indicates a center integer coefficient value of an unsharp mask.

13. The method of any of claims 1-6, wherein the one or more signal processing operations comprise a neural network upsampler.

14. The method of claim 13, wherein the determined signal processing information indicates coefficient values of one or more linear layers of a convolutional neural network.

15. The method of any of claims 1-6, wherein the one or more signal processing operations comprise an additional upsampling operation applied to an output of a last tier having residual data within the tier-based hierarchical format.

16. The method of any of claims 1-6, comprising, after reconstructing a higher resolution tier: applying dithering to an output of the reconstructed higher resolution tier.

17. The method of any of claims 1-6, wherein the tier-based hierarchical format is one of MPEG-5 Part 2 LCEVC and SMPTE VC-6 ST-2117.

18. A decoder configured to perform the method of any of claims 1-17.

19. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any of claims 1-17.

Citation Information

Patent Citations

  • Tiered signal decoding and signal reconstruction

    US8948248B2

  • Decomposition of residual data during signal encoding, decoding and reconstruction in a tiered hierarchy

    WO2013171173A1

  • Methods and apparatuses for encoding and decoding a bytestream

    WO2019111004A1

  • Processing signal data using an upsampling adjuster

    WO2019111011A1