Aspect Ratio Management in Layered Coding Schemes
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-11
- Publication Date
- 2026-08-14
Smart Images

Figure CN116547967B_ABST
Abstract
Description
Technical Field
[0001] This application describes methods and apparatus for managing the aspect ratio of an output signal when encoding using a layered coding scheme. In particular, but not exclusively, this application relates to encoding a base layer and an enhancement layer using two or more separate and / or independent coding schemes in a layered coding scheme, such as MPEG-5 Part 2 Low Complexity Enhanced Video Coding (as further described in patent application PCT / GB2020 / 050695, the contents of which are incorporated herein by reference). Background Technology
[0002] In layered coding schemes, such as MPEG-5 Part 2 LCEVC (as further described in patent application PCT / GB2020 / 050695, published as WO2020188273 and entitled "Low Complexity Enhancement Video Coding"), the input video signal is first downscaled using a downsampling / reduction process. The resulting downscaled video is then encoded using a first coding scheme (via a base encoder) to produce a coded base layer. The resulting coded base layer is then decoded using a corresponding decoder (implemented according to a first decoding scheme corresponding to the first coding scheme), and upscaled using a upsampling / reduction process to produce a preliminary reconstructed video signal. The preliminary reconstructed video signal is then subtracted from the source video signal to produce a set of residual data, which is optionally encoded using a second coding scheme (via an enhancement encoder).
[0003] On the decoder side, a corresponding decoder (implementing a first decoding scheme corresponding to a first encoding scheme) is also used to decode the encoded base layer to produce an encoded base layer suitable for display on a monitor, and sometimes output for display on a monitor. The encoded base layer is also amplified using a scaling / upsampling process to produce a preliminary reconstructed video signal. The preliminary reconstructed video signal is then combined with the encoded residual data and decoded using a decoder (implementing a second decoding scheme corresponding to a second encoding scheme) to produce the final reconstructed video.
[0004] Downsampling can change the ratio of the input video in a "square" manner (i.e., 1:1) or a non-square manner (e.g., 2:1, or other ratios, for a one-dimensional "1D" downsampling only in the horizontal dimension). Similarly, upsampling can change the ratio of the input video in a "square" manner (i.e., 1:1) or a non-square manner (e.g., 2:1, or other ratios, for a one-dimensional "1D" upsampling only in the horizontal dimension).
[0005] In non-square aspect ratios, one or both of the encoded base layer and the final reconstructed video are frequently displayed incorrectly. This disclosure considers solutions to this problem. Furthermore, in certain layered coding schemes, such as LCEVC, the base layer bitstream differs from the enhancement layer bitstream, and these two bitstreams are generated according to two separate coding schemes. Therefore, controlling how the base layer is configured can be challenging, making it even more important to ensure that the base encoder and enhancement encoder are configured appropriately. Summary of the Invention
[0006] Methods, apparatus, and computer programs as outlined in the appended claims are provided.
[0007] One object of this disclosure is to provide a solution to the problem of changing the aspect ratio in an encoding pipeline, particularly in the context of hierarchical coding schemes that change the aspect ratio of a signal during the encoding process. Such instances occur when a hierarchical coding scheme operates on signals at different resolutions and uses downsampling operations to vary the resolution from a first resolution to a second, lower resolution. If the downsampling operation changes the aspect ratio of the signal during the downsampling process, for example by using non-square downsampling techniques, the aspect ratio of the signal ultimately displayed at the end of the encoding pipeline may be undesirable and may not match the corresponding aspect ratio of the signal initially input to the encoding pipeline.
[0008] In a first aspect, this disclosure provides a technique for aspect ratio management signaling that enables the pixel aspect ratio of a signal to be modified during the encoding process to address any resolution aspect ratio changes that occur in the signal during the encoding process, such as through a non-square downsampling process.
[0009] In a second aspect, this disclosure provides a technique for aspect ratio management correction or modification that allows enhancement signals in a layered coding scheme to be corrected or modified to address any pixel aspect ratio modifications signaled by the aforementioned signaling techniques. Such correction includes signaling the pixel aspect ratio or similar information (such as display aspect ratio) of the original input signal to the coding pipeline via the coding pipeline, enabling the decoding system to use this information to overlay corresponding information contained in the decoded signal at the enhancement level. Alternatively, such correction includes determining information in the decoding system itself to alter the decoded output signal, the scaling factor being determined from, for example, upsampling or other signal modification operations and any resulting changes in the resolution aspect ratio of the signal obtained from such processes.
[0010] In a first aspect, a method is provided for signaling signal adjustment when encoding an input signal using a hierarchical encoding scheme to manage display aspect ratio, wherein the hierarchical encoding scheme includes encoding a downsampled version of the input signal using a first encoding method to generate a first encoded signal. The method includes signaling adjustment when the downsampling operation of the hierarchical encoding scheme is a non-square downsampling operation, such that the pixel aspect ratio of the first encoded signal is adjusted according to a scaling factor, wherein the scaling factor is determined from the non-square downsampling operation. Pixel aspect ratio is the ratio of the width to the height of each pixel in the signal.
[0011] Non-square downsampling results in a change in resolution aspect ratio from the input signal to a downsampled version of the input signal, where resolution aspect ratio is the ratio between the width and height (usually measured in pixels) of each frame in the signal.
[0012] The scaling factor is a typical implementation; it is the ratio of the aspect ratio of the input signal's resolution to the aspect ratio of the downsampled version of the input signal.
[0013] The pixel aspect ratio of the first coded signal can be determined by the following equation:
[0014] PAR e =PAR s ×width s / width e × Height e / high e
[0015] Among them PAR e It is the pixel aspect ratio of the first encoded signal, PAR s It is the pixel aspect ratio of the input signal, and the scaling factor is the width of the input signal. s Width of the first encoded signal e The ratio multiplied by the height of the first coded signal e Highly similar to the input signal s The ratio.
[0016] When the encoding system operates in 1D mode, downsampling occurs only in the horizontal dimension of the signal at an X:1 ratio. The scaling factor increases the pixel aspect ratio of the first encoded signal by scaling the horizontal dimension of each pixel by a factor of X, without scaling the height dimension. Typically, 1D mode operates at a 2:1 ratio, but other ratios, such as 3:1, 4:1, and non-integer ratios, can also be used.
[0017] This method describes a way of signaling an adjustment so that the display aspect ratio of the first coded signal is substantially the same as that of the input signal.
[0018] In a particularly common alternative embodiment, the step of signaling the adjustment includes signaling to set the pixel aspect ratio of the first encoded signal to the encoder or encoding module performing the first encoding method. However, the adjustment can be performed earlier in the signal pipeline, and can be performed before downsampling, after downsampling, at the encoder or encoding module, or after encoding by the encoder or encoding module.
[0019] In a more detailed embodiment, the layered encoding scheme may further include: upsampling a decoded version of the first encoded signal to generate an upsampled decoded signal, wherein the first encoded signal is decoded using a first decoding method corresponding to the first encoding method; generating a residual signal based on a comparison between the input signal and the upsampled decoded signal; and outputting the residual signal. The method further includes outputting metadata for the decoding system, the metadata including information related to the pixel aspect ratio of the input signal.
[0020] This metadata may include information related to the display aspect ratio of the input signal. This metadata may include the pixel aspect ratio of the input signal.
[0021] Typically, when the upsampling operation corresponding to the downsampling operation in a hierarchical coding scheme is a non-square upsampling operation, this method outputs metadata only if the upsampling operation in the hierarchical coding scheme is non-square.
[0022] This method typically, but not always, encodes the residual signal using a second encoding method before output.
[0023] In some cases, it can be useful to transmit metadata along with residual signals.
[0024] In a second aspect, a method for adjusting a decoded signal is provided, which decodes the signal using a layered coding scheme, wherein the layered coding scheme includes upsampling a decoded version of the encoded signal to generate an upsampled version of the signal, decoding the decoded version of the signal using a first decoding method, and combining the upsampled version of the signal with a residual signal to generate an output decoded signal, the encoded signal being derived from the input signal. The method includes adjusting the pixel aspect ratio of the output decoded signal such that the pixel aspect ratio of the output decoded signal matches the pixel aspect ratio of the input signal, the adjustment using one of: a pixel aspect ratio received as metadata from an encoding system or a desired display aspect ratio, wherein the display aspect ratio is the aspect ratio when the signal is presented on a display and is derived from the pixel aspect ratio and resolution aspect ratio; and a scaling factor derived from the upsampling operation.
[0025] The metadata can include the pixel aspect ratio of the input signal, and this adjustment can match the pixel aspect ratio of the output decoded signal to the pixel aspect ratio of the input signal.
[0026] The scaling factor is typically the ratio of the aspect ratio of the decoded version of the signal to the aspect ratio of the upsampled version of the signal.
[0027] In one embodiment, metadata or scaling factors are used to adjust the output decoded signal only if the upsampling operation is non-square and the aspect ratio of the output decoded signal changes as the output decoded signal passes through the upsampling operation.
[0028] The residual signal is typically a separate decoded component of the signal, and the second decoding method is used to decode the separate decoded component of the signal.
[0029] An encoding module is provided, which is configured to perform the above encoding steps.
[0030] A decoding module is provided, which is configured to perform the above decoding steps.
[0031] A computer program is provided that includes instructions that cause the computer to perform the methods described above when the computer executes the program.
[0032] A method for adjusting a signal in a hierarchical coding scheme is provided, wherein the hierarchical coding scheme includes: at an encoding system: downsampling an input signal to generate a downsampled version; passing the downsampled version to an encoder, causing the encoder to generate a first encoded signal; receiving a first decoded version of the first encoded signal from a decoder in the encoding system; upsampling the decoded version to generate an encoder-side reconstruction of the input signal; comparing the encoder-side reconstruction with the input signal to generate a residual signal; and outputting the residual signal for use by the decoding system with the first encoded signal. Furthermore, the hierarchical coding scheme includes: at a decoding system: receiving a second decoded version of the first encoded signal from a decoder in the decoding system; upsampling the second decoded version to generate a decoder-side reconstruction of the input signal; receiving the residual signal and adding the residual signal to the decoder-side reconstruction to generate a decoded output signal. The method further includes: at the encoding system: signaling an adjustment in the encoding system to adjust the pixel aspect ratio of the first encoded signal according to a scaling factor, wherein the scaling factor is determined from the resolution aspect ratio change caused by the downsampling operation. Pixel aspect ratio is the ratio between the width and height of each pixel in the signal, and resolution aspect ratio is the ratio between the width and height of each image in the signal. The method also includes output metadata for use by the decoding system when using the residual signal. The metadata includes the pixel aspect ratio of the input signal or information that allows the decoding system to derive the pixel aspect ratio. The method includes: at the decoding system: adjusting the pixel aspect ratio of the output decoded signal using the metadata such that the corresponding display aspect ratio for the output decoded signal when presented on a display matches the display aspect ratio of the input signal, wherein the display aspect ratio is the aspect ratio of the signal when presented on the display and is derived from the pixel aspect ratio multiplied by the resolution aspect ratio of the signal.
[0033] In a particularly common optional embodiment, the step of signaling the adjustment includes signaling to set the pixel aspect ratio of the first encoded signal to the encoder. However, the adjustment can be performed earlier in the signal pipeline, and can be performed before downsampling, after downsampling, at the encoder or encoding module, or after encoding by the encoder or encoding module.
[0034] An encoding system is provided, which is configured to perform the above-described methods. Attached Figure Description
[0035] The invention will now be described by way of example only with reference to the accompanying drawings, in which:
[0036] Figure 1 This is a block diagram illustrating an exemplary layered signal coding system with the aid of background information.
[0037] Figure 2 It shows the corresponding Figure 1A block diagram of an exemplary layered signal decoding system for an encoding system.
[0038] Figure 3 It shows that it will be by Figure 1 A schematic diagram illustrating the generation of an exemplary video bitstream by the encoding system.
[0039] Figure 4 This is a schematic diagram illustrating different portions of data that can be generated as part of exemplary video encoding.
[0040] Figure 5 This is a block diagram showing side-by-side examples of an encoding system and a decoding system to illustrate background information suitable for understanding this disclosure and the problems that this disclosure aims to solve.
[0041] Figure 6 It is shown Figure 5 The block diagram of the encoding and decoding system is modified according to the first aspect of the invention to control the aspect ratio of the basic encoded signal.
[0042] Figure 7 It is shown Figure 6 The encoding and decoding systems and their components are described. Figure 6 The diagram illustrates another problem introduced by the modification.
[0043] Figure 8 It is shown Figure 6 The block diagram of the encoding and decoding system is modified according to the second aspect of the invention to control the aspect ratio of the enhanced decoded signal.
[0044] Figure 9 This is a flowchart of a method for adjusting the aspect ratio of a basic coded signal at an encoder according to a first aspect of the present invention.
[0045] Figure 10 This is a flowchart of a method for controlling the aspect ratio of the enhancement layer output at the decoder according to a second aspect of the present invention at the encoder.
[0046] Figure 11 This is a flowchart of a method at a decoder for controlling the aspect ratio of the enhancement layer output according to a second aspect of the present invention.
[0047] Figure 12 yes Figure 1 The block diagram of the encoding system is modified according to the first aspect and the second aspect of the present invention.
[0048] Figure 13 yes Figure 2 The block diagram of the decoding system is modified according to the second aspect of the present invention. Detailed Implementation
[0049] With the help of background information, see reference Figures 1 to 5 An exemplary layered coding system is described.
[0050] This document presents examples with reference to signals as sample sequences (i.e., two-dimensional images, video frames, video fields, audio frames, etc.). For simplicity, the non-limiting embodiments described herein generally refer to signals displayed as a set 2D plane (e.g., 2D images in a suitable color space), such as, for example, video signals. In a preferred embodiment, the signal includes a video signal. Reference Figure 4 The exemplary video signal is described in more detail.
[0051] The terms “picture,” “frame,” or “field” will be used interchangeably with the term “image” to indicate a temporal sample of a video signal: any concepts and methods described for a video signal consisting of frames (progressive video signals) can also be readily applied to video signals consisting of fields (interlaced video signals) and vice versa. Although the embodiments described herein focus on image and video signals, those skilled in the art will readily understand that the same concepts and methods are also applicable to any other type of multidimensional signal (e.g., audio signals, volumetric signals, stereo video signals, 3DoF / 6DoF video signals, all-optical signals, point clouds, etc.). Although examples of image or video encoding are provided, the same methods can be applied to signals with dimensions less than two (e.g., audio or sensor streams) or greater than two (e.g., volumetric signals).
[0052] In this specification, the terms “image,” “picture,” or “plane” (meaning in the broadest sense of a “hyperplane,” i.e., an array of elements with arbitrary dimensions and a given sampling grid) are generally used to identify digital reproductions of signal samples along a sample sequence, wherein each plane has a given resolution for each of its dimensions (e.g., X and Y) and comprises a set of planar elements (or “elements” or “pixels,” or commonly referred to as “pixels” for display elements of two-dimensional images, commonly referred to as “voxels” for display elements of volumetric images, etc.) characterized by one or more “values” or “settings” (e.g., by way of non-limiting example, color settings in a suitable color space, settings indicating density levels, settings indicating temperature levels, settings indicating audio pitch, settings indicating amplitude, settings indicating depth, settings indicating alpha channel transparency levels, etc.). Each planar element is identified by a set of suitable coordinates indicating the integer position of the element in the sampling grid of the image. The signal dimension may contain only a spatial dimension (e.g., in the case of an image) or a temporal dimension (e.g., in the case of a signal that evolves over time, such as a video signal). In one scenario, a video signal frame may be a two-dimensional array with three color component channels, or a three-dimensional array with two spatial dimensions (e.g., a spatial dimension indicating resolution—its length equal to the corresponding width and height of the frame) and one color component dimension (e.g., a length of 3). In some cases, the processing described herein is performed separately for each plane that constitutes the color component values of the frame. For example, a plane representing the pixel values of each of the Y, U, and V color components can be processed in parallel using the methods described herein.
[0053] Exemplary coding system
[0054] Figures 1 to 4 An encoding scheme is shown that uses a downsampled source signal encoded by a base codec, adds first-level correction or enhancement data to the decoding output of the base codec to generate a corrected picture, and then adds another level of correction or enhancement data to an upsampled version of the corrected picture. Therefore, this encoding scheme can generate an enhanced stream with two spatial resolutions (higher and lower), which can be combined with the base stream at the lower spatial resolution.
[0055] In this encoding scheme, the method and apparatus can be based on a holistic algorithm constructed from existing encoding and / or decoding algorithms (e.g., MPEG standards such as AVC / H.264 and HEVC / H.265, and non-standard algorithms such as VP9 and AV1), which serve as the baseline for enhancement layers. Enhancement layers operate according to different encoding and / or decoding algorithms. In contrast to the block-based approach used in MPEG family algorithms, the idea behind the holistic algorithm is to encode / decode video frames in a hierarchical manner. Hierarchical frame encoding involves generating residuals for the entire frame, and then generating residuals for reduced or extracted frames, etc.
[0056] Figure 1 A system configuration of an exemplary encoder 100 is shown. The encoding process is split into two halves, as indicated by the dashed lines. Below the dashed lines is the base level, and above the dashed lines is the enhancement level, which can be usefully implemented in software. Encoder 100 may include enhancement-level-only processes or a combination of base-level and enhancement-level processes as needed. The encoder 100 topology at a typical level is as follows. Encoder 100 includes an input I for receiving an input signal 10. Input I is connected to a downsampler 105D. Downsampler 105D outputs to a base encoder 120E at the base level of encoder 100. Downsampler 105D also outputs to a residual generator 110-S. The encoded base stream is created directly by the base encoder 120E and can be quantized and entropy-encoded as needed according to the base encoding scheme. The encoded base stream may be referred to as the base layer or base level.
[0057] To generate the encoded sublayer 1 enhanced stream, the encoded base stream is decoded via a decoding operation applied at the base decoder 120D. In a preferred embodiment, the base decoder 120D may be a decoding component that complements the encoding component within the base codec, which takes the form of the base encoder 120E. In other embodiments, the base decoding block 120D may alternatively be part of the enhancement layer. A difference is created between the decoded base stream output from the base decoder 120D and the downsampled input video via residual generator 110-S (i.e., subtraction 110-S is applied to frames of the downsampled input video and frames of the decoded base stream to generate a first set of residuals). Here, the residual represents the error or difference between a reference signal or frame and a desired signal or frame. The residuals used in the first enhancement layer can be considered as correction signals because they are capable of 'correcting' future frames of the decoded base stream. This is useful because it can correct for odd patterns or other characteristics of the base codec. These features include, in particular, motion compensation algorithms applied by the underlying codec, quantization and entropy coding applied by the underlying codec, and block adjustment applied by the underlying codec.
[0058] exist Figure 1In the first residual set, transformation, quantization, and entropy encoding are performed to produce the encoded sub-layer 1 stream. Figure 1 In this process, transform operation 110-1 is applied to the first residual set; quantization operation 120-1 is applied to the transformed residual set to generate a quantized residual set; and entropy coding operation 130-1 is applied to the quantized residual set to generate a coded sublayer 1 stream at the first enhancement level. However, it should be noted that in other instances, only quantization step 120-1 or only transform step 110-1 may be performed. Entropy coding may not be used, or it may optionally be used as a supplement to one or both of transform step 110-1 and quantization step 120-1. The entropy coding operation can be any suitable type of entropy coding, such as a Huffmann coding operation or a run-length encoding (RLE) operation, or a combination of both Huffmann coding and RLE operations (e.g., RLE followed by Huffmann or prefix coding).
[0059] To generate the encoded sub-layer 2 stream, another set of residuals is generated and encoded via residual generator 100-S to create additional enhancement layer information. This additional set of residuals is the difference between an upsampled version (via upsampler 105U) of the corrected version of the decoded base stream (reference signal or frame) and the input signal 10 (desired signal or frame).
[0060] To achieve reconstruction of the corrected version of the decoded base stream generated at the decoder (e.g., ... Figure 2 As shown), at least some of the encoding operations in sublayer 1 are undone to simulate the decoder process, taking into account at least some losses and oddities in the transform and quantization processes. For this purpose, the first residual set is processed by a decoding pipeline including inverse quantization block 120-1i and inverse transform block 110-1i. The quantized first residual set is dequantized at inverse quantization block 120-1i in encoder 100 and inverse transformed at inverse transform block 110-1i to regenerate the decoder-side form of the first residual set. The decoded base stream from decoder 120D is then combined with the decoder-side form of the first residual set (i.e., a summation operation 110-C is performed on the decoded base stream and the decoder-side form of the first residual set). The summation operation 110-C generates a reconstruction of the input video, likely to be generated at the decoder—i.e., the reconstructed base codec video. The reconstructed base codec video is then upsampled by upsampler 105U. The processing in this example is typically performed frame-by-frame.
[0061] The upsampled signal (i.e., the reference signal or frame) is then compared with the input signal 10 (i.e., the desired signal or frame) to create another set of residuals (i.e., the difference operation is applied by the residual generator 100-S to the upsampled recreated frame to generate another set of residuals). This other set of residuals is then processed via a coding pipeline that serves as a mirror image of the coding pipeline used for the first set of residuals to become a coded sublayer 2 stream (i.e., coding operations are then applied to the other set of residuals to generate another encoded enhanced stream). Specifically, the other set of residuals is transformed (i.e., transform operation 110-0 is performed on the other set of residuals to generate another transformed set of residuals). The transformed residuals are then quantized and entropy encoded in the manner described above with respect to the first set of residuals (i.e., quantization operation 120-0 is applied to the transformed set of residuals to generate another quantized set of residuals; and entropy encoding operation 120-0 is applied to the quantized set of residuals to generate a coded sublayer 2 stream containing another level of enhanced information). In some cases, the operation can be controlled, for example, to make only quantization step 120-1, or only the transformation and quantization steps, possible. Entropy coding may optionally be used as a supplement. Preferably, the entropy coding operation may be a Huffman coding operation or a run-length encoding (RLE) operation, or both (e.g., RLE followed by Huffman coding). The transformation applied to blocks 110-1 and 110-0 may be a Hadamard transform applied to 2×2 or 4×4 residual blocks.
[0062] Figure 1 The coding operations in this approach do not introduce dependencies between local blocks of the input signal (e.g., compared to many known coding schemes that apply inter-frame or intra-frame prediction to macroblocks and thus introduce macroblock dependencies). Therefore, Figure 1 The operations shown can be executed in parallel on 4×4 or 2×2 blocks, which greatly improves coding efficiency on multi-core central processing units (CPUs) or graphics processing units (GPUs).
[0063] like Figure 1 As shown, the output of the encoding process is one or more enhancement streams at an enhancement level, which preferably includes a first enhancement level and another enhancement level. This can then be combined with the base stream at the base level (e.g., via multiplexing or other means). The first enhancement level (sub-layer 1) can be considered as implementing the corrected video at the base level, that is,, for example, correcting encoder glitches. The second enhancement level (sub-layer 2) can be considered as another enhancement level used to convert the corrected video into the original input video or an approximation thereof. For example, the second enhancement level can add fine details lost during downsampling and / or help correct errors introduced by one or more of the transform operation 110-1 and quantization operation 120-1.
[0064] Figure 2An exemplary decoder 200 corresponding to the encoding scheme is shown. The encoded base stream is decoded at the base decoder 220 to produce a base reconstruction of the input signal 10. This base reconstruction can be used in practice to provide a visual reproduction of the signal 10 at a lower quality level. However, the primary purpose of this base reconstruction signal is to provide a basis for a higher quality reproduction of the input signal 10. For this purpose, a decoded base stream is provided for sublayer 1 processing (i.e., sublayer 1 decoding). Figure 2 The sublayer 1 processing includes entropy decoding 230-1, inverse quantization 220-1, and inverse transform 210-1. Optionally, one or more of these steps may be performed depending on the operation performed at the corresponding block 100-1 at the encoder. By performing these corresponding steps, the encoded sublayer 1 stream, including the first residual set, becomes available at the decoder 200. The first residual set is combined with the decoded base stream from the base decoder 220 (i.e., a summation operation 210-C is performed on frames of the decoded base stream and frames of the decoded first residual set to generate a reconstruction of the downsampled form of the input video—i.e., the reconstructed base codec video). The frames of the reconstructed base codec video are then upsampled by the upsampler 205U.
[0065] Additionally, optionally in parallel, the encoded sublayer 2 stream is processed to produce another set of decoded residuals. Similar to sublayer 1 processing, sublayer 2 processing includes entropy decoding process 230-0, inverse quantization process 220-0, and inverse transform process 210-0. These operations will, of course, correspond to the operations performed at block 100-0 in encoder 100, and one or more of these steps may be omitted as needed. Block 200-0 produces an encoded sublayer 2 stream including another set of residuals, and these are summed at operation 200-C with the output from upsampler 205U to create a sublayer 2 reconstruction of input signal 10, which can be provided as the output of the decoder. Therefore, as... Figure 1 and Figure 2 As shown, the output of the decoding process can include up to three outputs: the basic reconstruction, the corrected lower-resolution signal, and the higher-resolution reconstruction of the original signal.
[0066] Figure 3 An alternative representation of the scheme is shown in the form of an exemplary signal encoding system 300. Signal encoding system 300 is a multi-layer or layer-based encoding system because the signal is encoded via multiple bitstreams, each bitstream representing a different encoding of the signal at a different quality level (e.g., different spatial resolution). Figure 3 In this example, there is a base layer 301 and an enhancement layer 302. Enhancement layer 302 (and...) Figure 1 and Figure 2Enhancement layers can implement enhanced coding schemes such as LCEVC. LCEVC is described in PCT / GB2020 / 050695 and related standard specifications, the latter being the ISO / IEC DIS23094-2 draft of Low Complexity Enhanced Video Coding, presented at the MPEG 129 conference held in Brussels from Monday, January 13, 2020 to Friday, January 17, 2020. Both documents are incorporated herein by reference. Figure 1 and Figure 2 In the example, Figure 3 In this context, enhancement layer 302 comprises two sublayers: a first sublayer 303 and a second sublayer 304. Each layer and sublayer can be associated with a specific quality level. As used herein, a quality level can refer to one or more of the following: sampling rate, spatial resolution, and bit depth, etc. In LCEVC, base layer 301 is located at a base quality level, first sublayer 303 is located at a first quality level, and second sublayer 304 is located at a second quality level. The base quality level and the first quality level can include common (i.e., shared or identical) quality levels or different quality levels. In cases where quality levels correspond to different spatial resolutions, such as in LCEVC, the input for each quality level can be obtained by downsampling and / or upsampling from another quality level. For example, the first quality level can be located at a first spatial resolution, and the second quality level can be located at a second higher spatial resolution, where the signal can be converted between quality levels by downsampling from the second quality level to the first quality level and by upsampling from the first quality level to the second quality level.
[0067] exist Figure 3 The diagram illustrates the corresponding encoder 305 and decoder 306 portions of the signal encoding system 300. It should be noted that encoder 305 and decoder 306 can be implemented as separate products, and the encoder and decoder do not need to originate from the same manufacturer or be provided as a single combined unit. Encoder 305 and decoder 306 are typically implemented in different geographical locations to generate encoded data streams for transmitting input signals between said two locations. Each of encoder 305 and decoder 306 can be implemented as part of one or more codecs—hardware and / or software entities capable of encoding and decoding signals. Reference to signal transmission as described herein also covers the encoding and decoding of files, where transmission can occur at a time on a common machine (e.g., by generating an encoded file and accessing it at a later point in time) or via physical transmission over a medium between two devices.
[0068] In some preferred embodiments, components of the base layer 301 may be provided separately to components of the enhancement layer 302; for example, the base layer 301 may be implemented by a hardware-accelerated codec, while the enhancement layer 302 may include a software-implemented enhancement codec. The base layer 301 includes a base encoder 310. The base encoder 310 receives a version of the input signal to be encoded 306 (e.g., a signal downsampled one or two times) and generates a base bitstream 312. The base bitstream 312 is transmitted between the encoder 305 and the decoder 306. At the decoder 306, the base decoder 314 decodes the base bitstream 312 to generate a reconstruction of the input signal at a base quality level 316.
[0069] Both enhancement sublayers 303 and 304 include a common set of encoding and decoding components. First sublayer 303 includes a first sublayer transform and quantization component 320 that outputs a set of first sublayer transform coefficients 322. The first sublayer transform and quantization component 320 receives data 318 derived from an input signal of a first quality level and applies a transform operation. This data may include the first residual set as described above. The first sublayer transform and quantization component 320 may also apply variable-level quantization to the output of the transform operation (including configurations that do not apply quantization). Quality scalability can be applied by varying the quantization applied in one or more enhancement sublayers. The set of first sublayer transform coefficients 322 is encoded by a first sublayer bitstream encoding component 324 to generate a first sublayer bitstream 326. This first sublayer bitstream 326 is transmitted from encoder 305 to decoder 306. At decoder 306, the first sublayer bitstream 326 is received and decoded by first sublayer bitstream decoder 328 to obtain a set of decoded first sublayer transform coefficients 330. The first sub-layer transform coefficient set 330 of the decoded signal is passed to the first sub-layer inverse transform and inverse quantization component 332. The first sub-layer inverse transform and inverse quantization component 332 applies further decoding operations, which include applying an inverse transform operation to at least the first sub-layer transform coefficient set 330 of the decoded signal. If the encoder 305 has already applied quantization, the first sub-layer inverse transform and inverse quantization component 332 may apply an inverse quantization operation before the inverse transform. Further decoding is used to generate a reconstruction of the input signal. In one case, the output of the first sub-layer inverse transform and inverse quantization component 332 is a first set of reconstructed residuals 334, which can be combined with the reconstructed base stream 316 as described above.
[0070] Similarly, the second sublayer 304 also includes a second sublayer transform and quantization component 340 that outputs a set of second sublayer transform coefficients 342. The second sublayer transform and quantization component 340 receives data derived from the input signal of the second quality level and applies a transform operation. In some embodiments, this data may also include residual data 338, although this residual data may be different from the residual data received by the first sublayer 303; for example, it may include another set of residuals as described above. The transform operation may be the same transform operation applied at the first sublayer 303. The second sublayer transform and quantization component 340 may also apply variable-level quantization prior to the transform operation (including configuration to not apply quantization). The set of second sublayer transform coefficients 342 is encoded by a second sublayer bitstream encoding component 344 to generate a second sublayer bitstream 346. This second sublayer bitstream 346 is transmitted from the encoder 305 to the decoder 306. In one case, at least the first sublayer bitstream 326 and the second sublayer bitstream 346 may be multiplexed into a single encoded data stream. In one scenario, the three bitstreams 312, 326, and 346 can all be multiplexed into a single encoded data stream. The single encoded data stream can be received at decoder 306 and demultiplexed to obtain each individual bitstream.
[0071] At decoder 306, the second sub-layer bitstream 346 is received and decoded by second sub-layer bitstream decoder 348 to obtain a set of decoded second sub-layer transform coefficients 350. As described above, this decoding involves bitstream decoding and can form part of a decoding pipeline (i.e., the decoded transform coefficient set 330 and the decoded transform coefficient set 350 can represent a set of partially decoded values that are further decoded through further operations). The decoded second sub-layer transform coefficient set 350 is passed to the second sub-layer inverse transform and inverse quantization component 352. The second sub-layer inverse transform and inverse quantization component 352 applies further decoding operations, which include applying at least an inverse transform operation to the decoded second sub-layer transform coefficient set 350. If the encoder 305 has already applied quantization at the second sub-layer, the second sub-layer inverse transform and inverse quantization component 352 can apply an inverse quantization operation before the inverse transform. Further decoding is used to generate a reconstruction of the input signal. This may include reconstructing another set of residuals 354 for use in combination with the reconstruction of the first set of residuals 334 and the upsampling combination of the underlying stream 316 (e.g., as described above).
[0072] Bitstream encoding components 324 and 344 can implement a configurable combination of one or more of entropy encoding and run-length encoding. Similarly, bitstream encoding components 328 and 348 can implement a configurable combination of one or more of entropy encoding and run-length encoding.
[0073] Further details and examples of the two sub-layer enhanced encoding and decoding systems can be found in the published LCEVC documentation.
[0074] Typically, the examples described herein operate within an encoding and decoding pipeline that includes at least a transform operation. The transform operation may include a DCT or a variant of a DCT, a Fast Fourier Transform (FFT), or a Hadamard transform such as that implemented by LCEVC. The transform operation may be applied block-by-block. For example, the input signal may be divided into multiple distinct continuous signal portions or blocks, and the transform operation may include matrix multiplication (i.e., a linear transform) applied to data from each of these blocks (e.g., as shown in a 1-dimensional vector). In this specification and in the art, it may be said that the transform operation produces a set of values for a predefined number of data elements, which, for example, represent positions in the transformed composite vector. These data elements are called transform coefficients (or sometimes simply “coefficients”).
[0075] As described herein, when the signal data includes residual data, a set of reconstructed coefficient bits may include the residual data of the transform, and the decoding method may further include instructing the combination of the residual data obtained from further decoding of the reconstructed coefficient bit set with a reconstruction of the input signal generated from a representation of the input signal of a lower quality level, to generate a reconstruction of the input signal of a first quality level. The representation of the lower quality level input signal may be the base signal to be decoded (e.g., from the base decoder 314), and the base signal to be decoded may optionally be upgraded before being combined with the residual data obtained from further decoding of the reconstructed coefficient bit set, the residual data being at a first quality level (e.g., a first resolution). Decoding may further include receiving and decoding the residual data associated with the second sublayer 304, for example, obtaining the output of the inverse transform and inverse quantization component 352, and combining it with data derived from the reconstruction of the input signal of the first quality level described above. This data may include data derived from an upgraded version of the reconstruction of the input signal of the first quality level, i.e., upgraded to a second quality level.
[0076] Although examples have been described with reference to layer-based hierarchical coding schemes in the form of LCEVC, the methods described herein can also be applied to other layer-based hierarchical coding schemes, such as VC-6:SMPTE VC-6ST-2117, as described in PCT / GB2018 / 053552 and / or related published standard documents, both of which are incorporated herein by reference.
[0077] Figure 4 This demonstrates how to decompose a video signal into different components and then encode them. Figure 4In this example, video signal 402 is encoded. Video signal 402 includes multiple frames or images 404, for example, where the multiple frames represent motion over time. In this example, each frame 404 consists of three color components. The color components can be in any known color space. Figure 4 In this context, the three color components are Y (luminance), U (first chromaticity opposite color), and V (second chromaticity opposite color). Each color component can be considered as a plane 408 of values. Plane 408 can be decomposed into a set of n×n signal data blocks 410. For example, in LCEVC, n can be 2 or 4; in other video coding techniques, n can be 8 to 32.
[0078] In LCEVC and some other coding techniques, the video signal fed back to a base layer such as 301 is a downgraded version of the input video signal 302. In this case, the signals fed back to both sub-layers include residual signals containing residual data. The residual data plane can also be organized into an n×n group of signal data blocks 410. The residual data can be generated by comparing data derived from the encoded input signal (e.g., video signal 402) with data derived from the reconstruction of the input signal, which is generated from a lower-quality representation of the input signal. Figure 3 In an example, the reconstruction of the input signal may include decoding of the coded base bitstream 312 available at encoder 305. Such decoding of the coded base bitstream 312 may include a lower-resolution video signal, which is then compared to a video signal downsampled from the input video signal 402. The comparison may include subtracting the reconstruction from the downsampled version. The comparison may be performed on a frame-by-frame (and / or block-by-block) basis. The comparison may be performed at a first quality level; if the base quality level is lower than the first quality level, the reconstruction from the base quality level may be upgraded before the comparison. Similarly, the input signal of the second sublayer, such as the input of the second sublayer transform and quantization component 340, may include residual data generated by comparing the input video signal 402 of the second quality level (which may include a full-quality original version of the video signal) with the reconstruction of the video signal of the second quality level. As previously described, the comparison may be performed on a frame-by-frame (and / or block-by-block) basis and may include subtraction. The reconstruction of the video signal may include a reconstruction generated from the decoding of the coded base bitstream 312 and the decoded version of the first sublayer residual data stream. Reconstruction can be generated at the first quality level and can be upsampled to the second quality level.
[0079] Therefore, the data plane 408 of the first sub-layer 303 can include residual data arranged in n×n signal blocks 410. One such 2×2 signal block... Figure 4The diagram illustrates this in more detail (n is chosen as 2 for ease of explanation), where for a color plane, the block can have a value of 412 and a set bit length (e.g., 8 or 16 bits). Each n×n signal block can be represented as a block of length n. 2 The flattened vector 414 represents the signal data block. To perform the transformation operation, the flattened vector 414 can be multiplied by the transformation matrix 416 (i.e., the taken dot product). This then generates a vector of length n. 2 Another vector 418 represents the different transform coefficients of a given signal block 410. Figure 4 An example similar to LCEVC is shown, where the transformation matrix 416 is a 4×4 Hadamard matrix, resulting in a transformation coefficient vector 418 with four elements, each having its own value. These elements are sometimes referred to by the letters A, H, V, and D, as they can represent the mean, horizontal difference, vertical difference, and diagonal difference. This transformation operation can also be called directional decomposition. When n=4, the transformation operation can use a 16×16 matrix and is called the square of the directional decomposition.
[0080] like Figure 4 As shown, a set of values for each data element in the entire group of signal blocks 410 of plane 408 can itself be represented as a plane or surface 420 of coefficient values. For example, the values of the “H” data elements in this group of signal blocks can be combined into a single plane, where the original plane 408 is then represented as four separate coefficient planes 422. For example, the coefficient plane 422 shown contains all the “H” values. These values are stored in a predefined bit length (e.g., bit length B), which can be 8, 16, 32, or 64 depending on the bit depth. A 16-bit example is given below, but this is not limiting. Therefore, coefficient plane 422 can be represented as a sequence 424 of 16-bit or 2-byte values (e.g., in memory) representing the value of a data element from the transform coefficients. These values can be called coefficient bits.
[0081] Aspect Ratio Management in Layered Coding Schemes
[0082] The following terms are used in the following description:
[0083] A pixel refers to a base tile (also called a "sample") of a solid-color rectangle.
[0084] Pixel aspect ratio (PAR, or sometimes called sample aspect ratio - SAR) refers to the ratio of the width (w) to the height (h) of a pixel or sample, usually expressed as a fraction w / h or w:h.
[0085] An image refers to a rectangular grid of pixels.
[0086] Resolution refers to a pair of positive integers, representing the number of pixels in the width and height of an image, respectively.
[0087] The aspect ratio (RAR) of an image refers to the ratio between the number of pixels in the width (w) and height (h) directions, usually expressed as a fraction w / h or w:h.
[0088] Display aspect ratio (DAR) refers to the ratio of the width (w) to the height (h) of an image displayed on a monitor along its linear dimension.
[0089] The triples PAR, RAR, and DAR have two degrees of freedom, which means that given two of the parameters, the third parameter is given.
[0090] DAR = PAR × RAR (Equation 1)
[0091] For example, consider a video source with the following parameters (the subscript 's' indicates that it refers to the "source"):
[0092] PAR s =1:1
[0093] resolution s =[width] s ,high s ] = [4,3]
[0094] This results in the displayed aspect ratio:
[0095]
[0096] [Issuance of base layer non-square downsampling modification]
[0097] Figure 5 This is a block diagram showing side-by-side examples of an encoding system and a decoding system to illustrate background information suitable for understanding this disclosure and the problems that this disclosure aims to solve.
[0098] Figure 5 Similar to Figure 1 Layered coding system and Figure 2 The corresponding decoding system, but in a more general form, and intended to illustrate the common problems with aspect ratio control in layered coding schemes. Figure 5 An exemplary base level for encoding and decoding, as well as an exemplary enhancement level for encoding and decoding, are shown, which together form a hierarchical encoding structure.
[0099] although Figure 5 The signaling and information transmission between the encoding system 510 and the decoding system 550 are illustrated in a simplified manner. Those skilled in the art will recognize that a network and / or storage system will be used between the encoding system 510 and the decoding system 550, as is known in the prior art.
[0100] The encoding system 510 includes a downsampler 512, a basic encoder 514-E, a basic decoder 514-D, an upsampler 516, and a comparator 518.
[0101] Encoding system 510 operates within a hierarchical coding scheme and is configured to receive input signal 510-In and pass it to downsampler 512 to generate a downsampled version of input signal 510-DS. The downsampled version of input signal 510-DS is then passed to base encoder 514-E to generate a first encoded signal 510-En. The first encoded signal 510-En is sometimes referred to as the base layer signal in the hierarchical coding scheme, and this base layer signal forms the basis of the encoded signal output from encoding system 510. Additional coding techniques can be applied to this signal, as described in reference [reference needed]. Figure 1 As described above or as known in the prior art, as part of the layered coding scheme, the first coded signal 510-En is also decoded using a base decoder 514-D to produce a decoded version 510-De. The decoded version 510-De is passed to an upsampler 516 to produce an upsampled signal 510-US. The upsampled signal 510-US is compared with the input signal 510-In at a comparator 518 to produce a residual signal R. In this example, the residual signal R is the difference between the input signal 510-In and the reconstructed version of the input signal produced by the base decoder 514-D and the upsampler 516, and is generated frame by frame. As those skilled in the art will know, the residual signal R can be further processed or encoded, or output in its original format. This type of layered coding scheme advantageously allows for parallel processing of frames within the input signal and allows for flexibility in data processing. The base layer signal and the residual signal in its original or processed / encoded form together form the layered coded signal output from the coding system 510. Decoding systems such as Decoding System 550 can partially decode layered coded signals by decoding the base layer signals or completely decode them by supplementing the base layer signals with residual signals.
[0102] Figure 5 An exemplary decoding system 550 includes a basic decoder 552-D, an upsampler 554, and a combination module 556.
[0103] Decoding system 550 is configured to receive layered encoded signals output from encoding system 510, either directly or indirectly via a network or other storage or transmission device. Decoding system 550 is configured to operate according to a layered encoding scheme to decode the first encoded signal 510-En using a base decoder 552-D to produce a decoded version of the first encoded signal. Base decoder 552-D is configured to output the decoded version of the base signal 550-De to be passed to upsampler 554 to produce an upsampled version of the decoded base signal 550-US. At combining module 556, the upsampled version 550-US is combined with a residual signal R obtained from the received layered encoded signal to produce an output decoded signal 550-OE, which is at an enhanced quality level higher than the base quality level. The residual signal R can be received in its original or encoded form, and if received in encoded form, decoding system 550 is configured to decode the residual signal R, typically using a different decoding scheme than that employed by base decoder 552-D.
[0104] Additionally, if using the enhanced quality level signal 550-OE is unsuitable, for example because a particular decoding system does not have enhancement level capabilities, or because the bandwidth limitation of the signals received by the decoding system means that enhancement level information cannot be transmitted to or received by the decoding system, then the decoding system 550 is arranged to output a basic output signal 550-OB, which is a decoded version of the first encoded signal 510-En using the basic decoder 552-D, at the basic quality level, to be optionally presented on the display.
[0105] like Figure 5 As can be observed from the representations of the various signals 510-xx and 550-xx, the aspect ratio of the base output signal 550-OB differs from that of the input signal 510-In. In this example, the display aspect ratio differs due to the non-square downsampling operation performed at downsampler 512. Therefore, when the base output signal 550-OB is presented on the display, the presented signal is likely to be displayed incorrectly (i.e., with an incorrect aspect ratio), resulting in a poor viewing experience.
[0106] Typically, downsampling operations can alter the resolution aspect ratio of an input signal either in a "square" manner by preserving the original aspect ratio of the signal (i.e., 1:1) or in a non-square manner (where one dimension is downsampled disproportionately to the other dimensions) (e.g., for a one-dimensional "1D" downsampling, a 2:1 ratio, where samples or pixels are reduced by half in the horizontal dimension but are fully preserved in the height dimension). Other non-square downsampling ratios exist, and the 2:1 ratio is only one example. Similarly, upsampling operations can alter the aspect ratio of an input video either in a "square" manner (i.e., 1:1) or in a non-square manner (e.g., for a one-dimensional "1D" upsampling only in the horizontal dimension, a 2:1 ratio, or other ratios).
[0107] exist Figure 5 In a specific instance, the individual signals 510-xx and 550-xx are represented by a single box within a box matrix, with each pixel representing a single box. In this instance, the input signal 510-In has 12 pixels arranged at a 4×3 resolution, thus having a 4:3 aspect ratio (RAR). The pixel aspect ratio (PAR) of the input signal is 1:1. RAR and PAR can be used together to determine the display aspect ratio (DAR) when the signal is presented on the display, which is 4:3 according to Equation 1. Ideally, the output signals 550-OB and 550-OE from 550 should each have a 4:3 DAR, such that they will be presented by the display to match the expected DAR of the input signals to the encoding system 510. Figure 5 In this example, due to the non-square downsampling operation at downsampler 512, the basic output signal 550-OB has a 2:3 DAR, therefore the DAR of the basic output signal 550-OB is not the same as that of the input signal 510-In. This is clearly problematic when presenting the basic output signal 550-OB.
[0108] As described above, the aspect ratio mismatch (RAR) of the output decoded signal 550-OB is caused by downsampler 512, which in this instance performs downsampling in a non-square manner. A practical example of non-square downsampling is when the encoding and decoding systems operate in a so-called 1D mode, where downsampling occurs only in the horizontal dimension of the signal at a ratio of 2:1. Therefore, the downsampled signal 510-DS has a 2:3 RAR (unlike the input signal 510-In, which has a 4:3 RAR). This altered RAR, cascaded through the encoding and decoding pipelines, causes an aspect ratio mismatch between the underlying output signal 550-OB and the input signal 510-In, which can only be corrected by upsampling operations at upsamplers 516 and 554, assuming these upsamplers operate in the corresponding mode (e.g., in 1-D mode). Sometimes, a particular upsampler 554 in the decoding system 550 may not operate in the correct mode, and in this case, the output signal 550-OE may also have a presentation DAR that does not match the input signal 510-In.
[0109] The first aspect of this invention focuses on ensuring that the aspect ratio of the base output signal 550-OB matches the corresponding aspect ratio of the input signal 510-In. The second aspect of this invention focuses on ensuring that the aspect ratio of the enhanced output signal 550-OE matches the corresponding aspect ratio of the input signal 510-In.
[0110] Figure 6 It is shown Figure 5 A block diagram of the encoding and decoding system, modified according to the first aspect of the invention, to control the aspect ratio of the underlying encoded signal. (Description only) Figure 5 and Figure 6 The differences between them, and Figure 6 Use the same reference numerals for similar components and signals. Figure 6 This illustrates how to match the aspect ratio of the base encoded signal 550-OB at the decoding system 550 and the input signal 510-In at the encoding system 510.
[0111] exist Figure 6 In this embodiment, the aspect ratio of the base level signal 610-En leaving the encoding system 510 has been adjusted according to a scaling factor. In this way, the aspect ratio of the base level signal can be controlled via the encoding pipeline to ultimately reproduce it reliably and predictably at the display via the decoding system 550. In this exemplary embodiment, the adjusted aspect ratio is the pixel aspect ratio.
[0112] from Figure 6 As can be seen from the signal representation 610-En, the encoded signal 610-En generated at the basic encoder 514-E is similar to... Figure 5The first coded signal 510-En generated in the process has different PAR. Figure 5 The encoded signal 510-En has a PAR ratio of 1:1, while Figure 6 The adjusted base encoded signal 610-En has a PAR of 2:1 (i.e., 2 units width to 1 unit height). This PAR modification produces an encoded base layer signal 610-En with a corresponding modified DAR. This modification is then fed through the encoding pipeline to the decoding system 550 and to the encoded base layer signal 650-OB, which shares the modified PAR with signal 610-En. Signal 650-OB prompts the display to present signal 650-OB to display such signals with a DAR that substantially matches that of the input signal 510-In, even though the RARs of the two signals are different (4:3 vs. 2:3). This is because the modified PAR requires the display to present each pixel at a size of 2 units width to 1 unit height, thus compensating for the difference in RAR.
[0113] More specifically, in the 1D LCEVC encoding of the input signal 510-In, the base code will have a RAR with half the horizontal width and the same vertical height as the source. To maintain the same DAR, this will cause the PAR to change by a factor of 2.
[0114] This translates to the following parameters of the basic coded signal (the subscript "e" stands for "coded"):
[0115] DAR e =DAR s =4∶3
[0116] resolution e =[width] e ,high e ] = [2,3]
[0117]
[0118] As a more general relationship between source PARs and encoded PARs:
[0119]
[0120] in
[0121] width e =width s / N (Equation 3)
[0122] high e =Height s / M (Equation 4)
[0123] therefore
[0124] PARe =PAR s ×width s / width e × Height e / high s =PAR s ×N / M (Equation 5)
[0125] Therefore, when using a non-square downsampling / upsampling ratio, the base encoder should adjust the base stream's PAR by a scaling factor (e.g., N / M) relative to the PAR value of the source signal. Specifically, this should be set to N / M times the PAR value of the source signal. For example, in the above example (i.e., in the case of a 2:1 downsampling / upsampling ratio), it should be twice the PAR value of the source video. If there is a double downsampling process (i.e., in the case of a 2:1 downsampling / upsampling ratio twice), the factor should be four.
[0126] Figure 7 It is shown Figure 6 The encoding and decoding systems and their components are described. Figure 6 A diagram illustrating another problem introduced by the modification. Only the differences are described, using similar figure labels. Figure 7 Continue Figure 6 The results show the results of modifying the aspect ratio of the base coded signal 610-En at the enhancement level.
[0127] Specifically, at encoding system 510, since the PAR of the basic encoded signal 610-En has been modified or scaled, the decoded version of the basic encoded signal 710-De is generated by the basic decoder 514-D, and thus also has a modified PAR value. After upsampling at upsampler 516, due to the modified PAR, an upsampled modified decoded signal 710-US with a DAR different from that of the input signal 510-In is generated. However, since comparator 518 only compares the pixel values of the upsampled modified decoded signal 710-US with the input signal 510-In pixel by pixel and ignores any PAR value, the fact that the two signals have different PARs and therefore different DARs is not a problem, and the generation of the residual R is unaffected.
[0128] In decoding system 550, base decoder 552-D decodes the modified base encoded signal 610-En to produce a modified decoded base signal 750-De, which has a modified PAR value, for example, set at the encoder base level. The decoded version of the modified base signal 750-De is then upsampled at upsampler 554 to produce an upsampled version of the modified decoded base signal 750-US. The upsampled signal may also retain the modified PAR value. At combining module 556, the upsampled version of the modified decoded base signal 750-US is combined with the residual R to produce a modified decoded enhanced output signal 750-OE. The enhanced output signal 750-OE may retain the modified PAR values from signals 750-De and 750-US, and in this case, due to the different PAR values of the two signals, the enhanced output signal will have a different DAR than the input signal 510-In at encoding system 510.
[0129] Therefore, when the DAR of the base layer output 650-OB of the decoding system 550 is modified by altering the PAR of the base encoded signal to match the aspect ratio of the input signal 510-In, the DAR of the enhancement layer output signal 750-OE of the decoding system 550 does not match the input signal 510-In.
[0130] The objective of the second aspect of the invention is to make the output DAR, PAR, and RAR at the enhancement level output of the decoding system 550 identical to the source signal or input signal 510-In. In one instance, this means that the following relationship applies to PAR (the subscript "o" represents the output):
[0131] PAR o =PAR s =PAR e / (N / M)=PAR e ×M / N (Equation 6)
[0132] Therefore, in this example, when using a non-square downsampling / upsampling ratio, the decoding system should adjust the PAR of the final reconstructed video according to a scaling factor (e.g., N / M) of the PAR of the encoding base. Specifically, it should be set to M / N times the PAR value of the encoding base. For example, in the above example (i.e., in the case of a 2:1 downsampling / upsampling ratio), the PAR of the final reconstructed video should be divided by 2, which is signaled on the base stream, and the resulting value should be used in the final rendering stage to produce the final reconstructed video.
[0133] Figure 8 It is shown Figure 6 or Figure 7The block diagrams of the encoding and decoding systems are modified according to the second aspect of the invention to control the aspect ratio of the enhanced decoded signal. Only the differences are described, and similar reference numerals are used.
[0134] exist Figure 8 In this embodiment, encoding system 510 corrects the aspect ratio of the enhancement layer output signal to achieve enhanced output signal 850-OE by sending the residual signal R and some metadata. The metadata includes data that enables the decoding system to generate an enhanced reconstruction of the input signal with an aspect ratio matching the input signal. The metadata can be sent as a data packet along with the residual signal R or separately. In this example, the metadata includes the pixel aspect ratio of the input signal 510-In; however, the metadata may include scaling factors, such as those mentioned in the preceding paragraph, or may include the DAR of the input signal 510-In. The residual signal R and the metadata are received at decoding system 550. The metadata is used to generate a corrected output decoded signal 850-OE with the same DAR, PAR, and RAR as the input signal 510-In.
[0135] In this way, the input signal can be reconstructed at the enhancement level output in the decoding system with the same or substantially the same aspect ratio.
[0136] Figure 9 This is a flowchart of a method for adjusting the aspect ratio of a basic coded signal using a coding system according to a first aspect of the present invention. In a particular example, the coding system is referenced... Figures 6 to 8 The coding system 510 described by any of the above.
[0137] At step 910, the method includes receiving an input signal, such as input signal 510-In. The input signal has a first resolution aspect ratio and a first pixel aspect ratio, both as previously defined in this specification, which together define the display aspect ratio of the input signal. At step 920, the method includes downsampling the input signal to produce a downsampled version of the input signal. At step 930, the method includes sending the downsampled version to an encoder in an encoding system to encode the downsampled version of the input signal to produce a first encoded signal. At step 940, the method includes signaling an adjustment to the pixel aspect ratio of the first encoded signal. The signaling includes a scaling factor for adjusting the pixel aspect ratio of the first encoded signal.
[0138] The scaling factor is derived from the aspect ratio of the input signal and the aspect ratio of the downsampled version. In this exemplary method, the scaling factor is the ratio of the aspect ratio of the input signal to the aspect ratio of the downsampled version of the input signal. More specifically, the pixel aspect ratio of the first coded signal is determined by the following calculation:
[0139] PARe =PAR s ×width s / width e × Height e / high s (Equation 7)
[0140] Among them PAR e It is the pixel aspect ratio of the first encoded signal, PAR s It is the pixel aspect ratio of the input signal, and the scaling factor is the ratio of the input signal width to the width of the first encoded signal multiplied by the ratio of the height of the first encoded signal to the height of the input signal, wherein the height and width related to the resolution aspect ratio are measured and given in pixels.
[0141] In this exemplary method, any adjustment to the pixel aspect ratio of the first encoded signal occurs, or in some cases, is signaled, only if the downsampled version of the input signal has a different resolution aspect ratio than the input signal. This occurs when using non-square downsampling operations (such as the horizontal 1D mode sampling operation in the typical example described earlier in this document). When the encoding system is running in 1D mode, where downsampling occurs only in the horizontal dimension of the signal at a 2:1 ratio, a scaling factor increases the pixel aspect ratio of the first encoded signal by scaling the horizontal dimension by a factor of 2 without scaling the height dimension.
[0142] In this exemplary method, the pixel aspect ratio is adjusted. In this way, a convenient and easily signalable aspect ratio is used. Alternatively, the display aspect ratio can be signaled and adjusted. In practice, signaling the adjustment makes the display aspect ratio of the first coded signal substantially the same as the display aspect ratio of the input signal.
[0143] The step of signaling the adjustment includes signaling to set the pixel aspect ratio of the first encoded signal to the encoding module performing the first encoding method. Alternatively, the method may adjust the downsampled version before sending the downsampled version to the encoder to create the first encoded signal.
[0144] In this example, the scaling factor is the ratio of the resolution aspect ratio of the input signal to the resolution aspect ratio of the downsampled version of the input signal, see, for example, Equation 7. In this way, the display aspect ratio of the input signal can be maintained through the encoding pipeline even after downsampling and encoding.
[0145] In this example, the method may further include upsampling a decoded version of the first encoded signal to generate an upsampled decoded signal. The first encoded signal may be received from a first decoding method corresponding to the first encoding method and decoded using the first decoding method to generate a decoded version. The method also includes generating a residual signal based on a comparison between the input signal and the upsampled decoded signal, and the method may include the pixel aspect ratio of the output residual signal and the input signal, or information that allows this information to be derived for use by the decoding system.
[0146] In this way, the residual signal can be decoded in a decoding system (such as a reference). Figures 6 to 8 The decoding system 550 described above adds enhancements to the first encoded signal. Furthermore, information including the pixel aspect ratio of the input signal can be used at the decoding system to manage or correct any problems the decoding system may encounter due to non-square downsampling or upsampling operations or other aspect ratio management signaling and adjustments that may occur at the encoding or decoding system.
[0147] Information or metadata that allows us to know the pixel aspect ratio of the input signal can be the display aspect ratio of the input signal, because the resolution aspect ratio of the upsampled decoded signal in a normally operating decoding system will likely match that of the input signal.
[0148] In the hierarchical encoding scheme, the upsampling operation is a non-square upsampling operation corresponding to the downsampling operation. The method outputs metadata only if one of the downsampling or upsampling operations in the hierarchical encoding scheme is a non-square upsampling operation. However, a 1:1 scaling factor or ratio can be applied to square downsampling / upsampling settings.
[0149] In a typical variation of the above method, the method may include encoding the residual signal using a second encoding method before outputting the residual signal.
[0150] Metadata is typically transmitted along with the residual signal, but can be transmitted independently of the residual signal or the first encoded signal, or in any other manner. The signaling notification section at the end of this specification discusses appropriate signaling aspect ratios.
[0151] Figure 10 This is a flowchart of a method for managing the aspect ratio of an enhancement layer output at a decoding system according to a second aspect of the present invention at an encoding system. In a particular instance, the encoding system is referenced... Figure 8 The encoding system described is 510. Figure 10 The method largely follows Figure 9 The method is as shown by the connection symbol A, but this is not required.
[0152] At step 1010, the method includes a decoded version of the received signal. At step 1020, the method includes upsampling the decoded version of the signal to generate an upsampled decoded signal. At step 1030, the method includes generating a residual signal based on a comparison between the input signal and the upsampled decoded signal. At step 1140, the method includes outputting the aspect ratio of the residual signal and the input signal.
[0153] In the exemplary method described, the aspect ratio of the input signal is the pixel aspect ratio. However, other data, such as the display aspect ratio of the input signal, can be signaled to the decoding system to allow the decoding system to more faithfully reproduce the aspect ratio of the input signal.
[0154] Figure 11 This is a flowchart of a method for controlling the aspect ratio of an enhancement layer output in a decoding system according to a second aspect of the present invention. In a particular example, the decoding system is referenced... Figure 8 The decoding system described is 550.
[0155] At step 1110, the method includes upsampling the decoded version of the signal to produce an upsampled version of the signal. At step 1120, the method includes combining the upsampled version of the signal with the residual signal to produce an output decoded signal. At step 1130, the method includes adjusting the aspect ratio of the output decoded signal according to a scaling factor. Specifically, the adjustment is made such that the output decoded signal matches as closely as possible to the overall shape and aspect ratio of the original encoded signal on the encoding system side of the encoding pipeline.
[0156] The aspect ratio of the output decoded signal is typically the pixel aspect ratio. In this way, the pixel aspect ratio of the output decoded signal can be modified so that the output decoded signal has a display aspect ratio similar to the input signal. However, other data can be signaled to the decoding system to allow it to more faithfully reproduce the aspect ratio of the input signal, such as its display aspect ratio.
[0157] Typically, this adjustment uses the pixel aspect ratio or desired display aspect ratio received as metadata from the encoding system. Display aspect ratio is the aspect ratio of the signal when it is presented on a display, and it is derived from pixel aspect ratio and resolution aspect ratio, as previously described with reference to Equation 1 in this disclosure. This adjustment matches the pixel aspect ratio or display aspect ratio of the output decoded signal to the received information. However, this adjustment can alternatively use a scaling factor derived from the upsampling operation.
[0158] The scaling factor is the ratio of the aspect ratio of the decoded version of the signal to the aspect ratio of the upsampled version of the signal, and can be derived using the scaling factors previously described in this disclosure.
[0159] In some variations, metadata or scaling factors are used to adjust the output decoded signal only when the upsampling operation is non-square and changes the aspect ratio of the upsampled signal.
[0160] In this exemplary method, the residual signal is the decoded component of the signal, and the decoded component of the signal is decoded using a second decoding method.
[0161] Figure 12 yes Figure 1 The block diagram of the encoding system is modified according to the first and second aspects of the present invention. Only the differences are described.
[0162] exist Figure 12 According to a first aspect of the invention, the basic encoded signal is modified by signaling a scaling factor N / M to scale the aspect ratio of the basic encoded signal. As indicated by the three arrows derived from the scaling factor N / M, the scaling factor can be signaled, and then applied at one of three locations in the encoding pipeline as follows: 1) by adjusting the pixel aspect ratio of the output of the downsampler 105D according to the scaling factor before the output is sent for encoding at the basic encoder 120E; 2) by signaling the adjustment and scaling factor to the basic encoder 120E, such that the basic encoded signal, which generates the pixel aspect ratio, is scaled according to the scaling factor at the basic encoder 120E; or 3) by scaling the aspect ratio of the basic encoded signal after it has been generated at the basic encoder 120E. As described above, the modified basic encoded signal, along with the required enhancement level information, is then sent to the decoding system.
[0163] Furthermore, metadata is transmitted from the encoding system 100 according to the second aspect of the invention to compensate for the modified underlying encoded signal. The metadata includes data that enables the decoding system to generate an enhanced reconstruction of the input signal that matches the aspect ratio of the input signal, even though a scaling factor has been applied at the underlying layer. The metadata can be transmitted as a data packet along with the encoded signal data (i.e., sublayer 2, sublayer 1, and the underlying layer). In one instance, the metadata can be signaled using sublayer 2. In another instance, the metadata is signaled separately from the encoded signal data and can be transmitted, for example, on a separate communication channel or stored separately from the encoded signal data on a storage medium. The metadata may include the pixel aspect ratio of the input signal or, in some cases, the display aspect ratio of the input signal.
[0164] Figure 13 yes Figure 2 The block diagram of the decoding system is modified according to the second aspect of the present invention. Only the differences are described.
[0165] The decoding system 200 receives a modified base encoded signal and the base decoder decodes the modified base encoded signal to produce a base reconstruction that, when displayed, has a display aspect ratio that matches or substantially matches the display aspect ratio of the input signal of the encoding system.
[0166] In addition, the decoding system 200 receives metadata sent from the encoding system, such as reference... Figure 12 The process involves generating a sublayer 2 reconstruction with a display aspect ratio that matches or substantially matches the display aspect ratio of the input signal. Alternatively, the decoding system 200 derives a scaling factor, as per [the description of the scaling factor]. Figure 11 The subject of discussion.
[0167] Send signal notification
[0168] In the example above, when using a non-square downsampling / upsampling ratio, the underlying encoder sets the underlying Video Usability Information (VUI) in the Sequence Parameter Set (SPS) to an aspect ratio that is N / M times that of the source video.
[0169] In the example above, when using a non-square downsampling / upsampling ratio, the decoding system sets the final reconstructed video (i.e., the resulting output image) to a PAR with an M / N value that signals the underlying stream.
[0170] Examples of MPEG-5 LCEVC related implementations described in the following documents: F. Maurer, S. Battista, L. Ciccarelli, G. Meardi, S. Ferrara, “Overview of MPEG-5 Part 2 – Low Complexity Enhancement Video Coding (LCEVC)”, ITU Journal: ICT Discoveries, Vol. 3(1), June 8, 2020; and “MPEG-5 Part 2: Low Complexity Enhancement Video Coding (LCEVC): Overview and performance evaluation”, Proc. SPIE 11510, Applications of Digital Image Processing XLIII, 115101C (August 21, 2020); https: / / doi.org / 10.1117 / 12.2569246; and the international standard ISO / IEC 23094-2 (its specification “Draft Text of ISO / IEC DIS 23094-2 Low Complexity Enhancement Video Coding”). The “Coding” is incorporated herein by reference (ISO / IEC WG11, w18986, Brussels, January 2020). When encoding with scaling_mode_level1 or scaling_mode_level2 equal to 1, for a one-dimensional 2:1 scaling only in the horizontal dimension, in order to maintain the source display aspect ratio, it is recommended that the bitstream signal the sample aspect ratio in the Video Availability Information (VUI), and as signaled in the Video Availability Information (VUI), for each scaling_mode_level equal to 1, the base encoder doubles the horizontal value of the sample aspect ratio.
[0171] Examples of MPEG-5 LCEVC related implementations described in the following documents: F. Maurer, S. Battista, L. Ciccarelli, G. Meardi, S. Ferrara, “Overview of MPEG-5 Part 2 – Low Complexity Enhancement Video Coding (LCEVC)”, ITU Journal: ICT Discoveries, Vol. 3(1), June 8, 2020; and “MPEG-5 Part 2: Low Complexity Enhancement Video Coding (LCEVC): Overview and performance evaluation”, Proc. SPIE 11510, Applications of Digital Image Processing XLIII, 115101C (August 21, 2020); https: / / doi.org / 10.1117 / 12.2569246; and the international standard ISO / IEC 23094-2 (its specification “Draft Text of ISO / IEC DIS 23094-2 Low Complexity Enhancement Video Coding”). The aspect ratio of the enhanced image output for the decoding system is the aspect ratio indicated by reference in this document, ISO / IEC WG11, w18986, Brussels, January 2020. This aspect ratio is specified in the bitstream VUI (as indicated in Annex E of the “Draft Text of ISO / IEC DIS 23094-2 Low Complexity Enhancement Video Coding”, ISO / IEC WG11, w18986, Brussels, January 2020) and carried in the VUI parameter payload_type equal to 5 (Section 7.3.3, Table 7) or additional type equal to 1 (Sections 7.3.10 and 7.4.3.8). If no additional information or VUI parameter or aspect ratio information is available, the decoding system should assume an aspect_ratio_idc value of 1 for a 1:1 sample aspect ratio (“square” sample).
[0172] Furthermore, in the ISO basic media file format (also known as MP4), the aspect ratio can be signaled in the form of an unsigned integer numerator and denominator within an atom named "pasp". Therefore, the encoding system can signal the aspect ratio in the form of an unsigned integer numerator and denominator within the atom named "pasp".
[0173] Furthermore, in MPEG-TS, the aspect ratio can be signaled in the "Target Background Grid Descriptor," which is defined as an enumeration from the MPEG-2 video specification. Therefore, the encoding system can signal the aspect ratio that can be signaled in the "Target Background Grid Descriptor," which is defined as an enumeration from the MPEG-2 video specification.
[0174] In an embodiment, if the decoding system receives an aspect ratio at the container level (e.g., MPEG-TS or ISO BMFF) that is different from the aspect ratio indicated in the underlying bitstream, the decoder may choose to use one of the underlying bitstreams.
[0175] Computer programs and computer-readable storage media are also disclosed, which, when implemented on a general-purpose computer system performing the functions of an encoding system or encoder or a decoding system or decoder, can perform any of the methods described above and can provide the functions described herein as enhancement-level functions or both enhancement-level and basic-level functions.
[0176] Generally, a particular instance is described with reference to an exemplary video signal, in which pixels and frames exist, as those skilled in the art will understand. Of course, the signal may involve non-video signals, where the aspect ratio of the displayed or other aspects is important for signal reproduction. In such cases, those skilled in the art are taught to manage and scale the sample aspect ratio or other equivalent aspect ratio, rather than the pixel aspect ratio, in the same manner disclosed herein throughout the encoding pipeline.
[0177] The above embodiments should be understood as illustrative examples. Other embodiments are envisioned. It should be understood that any feature described with respect to any embodiment may be used alone or in combination with other described features, and may also be used in combination with one or more features of any other embodiment, or in any combination of any other embodiment. Furthermore, equivalents and modifications not described above may be employed without departing from the scope of the invention as defined by the appended claims.
[0178] Additional Statement
[0179] A method is provided for encoding an input signal using a hierarchical coding scheme, wherein the scheme includes encoding a downsampled version of the input signal using a first coding method to generate a first coded signal having a first aspect ratio, the method comprising: adjusting the aspect ratio of the first coded signal according to a scaling factor with respect to the aspect ratio value of the input signal.
[0180] Optionally, the adjustment is performed when the downsampled version of the input signal has a different aspect ratio than the input signal.
[0181] Optionally, the adjustment step includes setting the aspect ratio of the first coded signal by scaling the aspect ratio of the input signal according to a scaling factor.
[0182] An encoding module is provided, which is configured to perform any of the steps described above in the encoding steps.
[0183] A method for decoding a signal using a hierarchical coding scheme is provided, wherein the scheme includes upsampling a decoded version of the signal to generate an upsampled version of the signal, the decoded version of the signal being decoded using a first decoding method, and combining the upsampled version of the signal with the decoded components of the signal to generate an output decoded signal, and decoding the decoded version of the signal using a second decoding method, the signal having a first aspect ratio and the decoded version of the signal having a second aspect ratio, the method including: adjusting the aspect ratio of the output decoded signal according to a scaling factor with respect to the second aspect ratio value.
[0184] Optionally, adjustments are made when the first aspect ratio differs from the second aspect ratio.
[0185] Optionally, the adjustment step includes setting the aspect ratio of the output decoded signal by scaling the aspect ratio of the second aspect ratio according to the scaling factor.
[0186] A decoding module is provided, which is configured to perform any of the steps described above in the decoding steps.
Claims
1. A method for signaling signal adjustment when using a hierarchical encoding scheme to encode an input signal to manage the aspect ratio of a display, wherein the hierarchical encoding scheme includes encoding a downsampled version of the input signal using a first encoding method to generate a first coded signal, the method comprising: When the downsampling operation of the hierarchical coding scheme is a non-square downsampling operation, a signal is sent to notify an adjustment so that the pixel aspect ratio of the first coded signal is adjusted according to a scaling factor, wherein the scaling factor is determined from the non-square downsampling operation, and the pixel aspect ratio is the ratio of the width to the height of each pixel in the signal. The aspect ratio of the pixels in the first encoded signal is determined as follows: Among them PAR e The aspect ratio of the pixels in the first encoded signal, PAR s The aspect ratio of the input signal is the pixel width, and the scaling factor is the input signal width. With the width of the first encoded signal The ratio, multiplied by the height of the first coded signal With the height of the input signal The ratio; The layered coding scheme further includes: The decoded version of the first encoded signal is upsampled to generate an upsampled decoded signal, wherein the first encoded signal is decoded using a first decoding method corresponding to the first encoding method; A residual signal is generated based on a comparison between the input signal and the upsampled decoded signal; and Output the residual signal; The method further includes outputting metadata for the decoding system, the metadata including information related to the pixel aspect ratio of the input signal.
2. The method of claim 1, wherein the non-square downsampling operation results in a change in resolution aspect ratio from the input signal to the downsampled version of the input signal, wherein the resolution aspect ratio is the ratio between the width in pixels and the height in pixels for each frame of the signal.
3. The method of claim 1, wherein the scaling factor is the ratio of the aspect ratio of the input signal to the aspect ratio of the downsampled version of the input signal.
4. The method according to claim 1, wherein, When the encoding system is running in 1D mode, downsampling occurs only in the horizontal dimension of the signal at a ratio of X:1, and the scaling factor increases the pixel aspect ratio of the first encoded signal by scaling the horizontal dimension of each pixel by a factor of X, without scaling the height dimension.
5. The method of claim 1, wherein the adjusted signal notification makes the display aspect ratio of the first coded signal substantially the same as the display aspect ratio of the input signal.
6. The method according to claim 1, wherein the step of signaling the adjustment includes signaling the encoding module performing the first encoding method to set the pixel aspect ratio of the first encoded signal.
7. The method of claim 1, wherein the metadata includes information relating to the display aspect ratio of the input signal.
8. The method of claim 1, wherein the metadata includes the pixel aspect ratio of the input signal.
9. The method according to claim 1, wherein in the hierarchical coding scheme, the upsampling operation of the hierarchical coding scheme is a non-square upsampling operation corresponding to the downsampling operation.
10. The method of claim 1, wherein the method outputs the metadata only when one of the downsampling operation and the upsampling operation of the hierarchical encoding scheme is not square.
11. The method of claim 1, wherein the method includes encoding the residual signal using a second encoding method before output.
12. The method of claim 1, wherein the metadata is transmitted together with the residual signal.
13. The method according to claim 1, wherein the first encoding method is one of AVC / H.264, HEVC / H.265 or AV1.
14. A method for adjusting a decoded signal, the decoded signal being decoded using a layered coding scheme, wherein the layered coding scheme includes upsampling a decoded version of an encoded signal to generate an upsampled version of the signal, the decoded version of the signal being decoded using a first decoding method, and combining the upsampled version of the signal with a residual signal received from an encoding system to generate the decoded signal, the encoded signal being derived from an input signal having a first adjusted pixel aspect ratio, and the encoded signal being received from the encoding system, the method comprising: The aspect ratio of the decoded signal is adjusted to match the aspect ratio of the input signal. This adjustment uses one of the following methods: The pixel aspect ratio or desired display aspect ratio received from the encoding system as metadata, wherein the display aspect ratio is the aspect ratio when the signal is presented on the display, and can be derived from the pixel aspect ratio and resolution aspect ratio; and The scaling factor derived from the upsampling operation; The aspect ratio of the pixels in the decoded signal is determined as follows: Among them PAR e PAR is the pixel aspect ratio of the decoded signal. s The aspect ratio of the input signal is the pixel width, and the scaling factor is the input signal width. With the width of the decoded signal The ratio, multiplied by the height of the decoded signal. With the height of the input signal The ratio.
15. The method of claim 14, wherein the metadata includes the pixel aspect ratio of the input signal, and the adjustment matches the pixel aspect ratio of the decoded signal to the pixel aspect ratio of the input signal.
16. The method of claim 14, wherein the scaling factor is the ratio of the aspect ratio of the decoded version of the signal to the aspect ratio of the upsampled version of the signal.
17. The method of claim 14, wherein the metadata or scaling factor is used to adjust the decoded signal only when the upsampling operation is non-square and the aspect ratio of the decoded signal changes as the decoded signal passes through the upsampling operation.
18. The method of claim 14, wherein the residual signal is a separate decoded component of the signal, and the separate decoded component of the signal is decoded using a second decoding method.
19. The method according to claim 14, wherein the first decoding method is one of AVC / H.264, HEVC / H.265 or AV1.
20. An encoding module comprising means configured to perform the method of any one of claims 1 to 13.
21. A decoding module comprising means configured to perform the method of any one of claims 14 to 19.
22. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 13 or any one of claims 14 to 19.
23. A method for adjusting a signal in a hierarchical coding scheme, wherein the hierarchical coding scheme comprises: At the encoding system: The input signal is downsampled to produce a downsampled version; The downsampled version is passed to the encoder, causing the encoder to generate a first encoded signal; Receive a first decoded version of the first encoded signal from the decoder in the encoding system; The decoded version is upsampled to generate encoder-side reconstruction of the input signal; The encoder end reconstruction is compared with the input signal to generate a residual signal; The residual signal is output for use by the decoding system in conjunction with the first encoded signal; At the decoding system: Receive a second decoded version of the first coded signal from the decoder in the decoding system; Upsample the second decoded version to generate a decoder-side reconstruction of the input signal; The residual signal is received and added to the decoder for reconstruction to generate a decoded output signal; The method includes, at the encoding system: In the encoding system, a signal is sent to notify an adjustment so that the pixel aspect ratio of the first encoded signal is adjusted according to a scaling factor, wherein the scaling factor is determined from the resolution aspect ratio change caused by the downsampling operation, wherein the pixel aspect ratio is the ratio between the width and height of each pixel in the signal, and the resolution aspect ratio is the ratio between the width and height of each image in the signal; Output metadata for use by the decoding system when using the residual signal, the metadata including the pixel aspect ratio of the input signal or information that allows the decoding system to derive the pixel aspect ratio; The method includes, at the decoding system: The pixel aspect ratio of the output decoded signal is adjusted using the metadata so that the corresponding display aspect ratio when the output decoded signal is presented on the display matches the display aspect ratio of the input signal, wherein the display aspect ratio is the aspect ratio when the signal is presented on the display and can be derived from the pixel aspect ratio multiplied by the resolution aspect ratio of the signal. The aspect ratio of the pixels in the first encoded signal is determined as follows: Among them PAR e The aspect ratio of the pixels in the first encoded signal, PAR s The aspect ratio of the input signal is the pixel width, and the scaling factor is the input signal width. With the width of the first encoded signal The ratio multiplied by the height of the first coded signal With the height of the input signal The ratio.
24. The method of claim 23, wherein the encoder that generates the first encoded signal uses one of AVC / H.264, HEVC / H.265, or AV1.
25. A codec system comprising means configured to perform the method of any one of claims 23 to 24.
Citation Information
Patent Citations
Low complexity enhancement video coding
WO2020188273A1