Enhanced interlaced scanning
By introducing LCEVC technology into the encoding system, it integrates into the existing ecosystem and integrating low-complexity enhanced video encoding into the TV hardware legacy ecosystem, solving the problem of converting from SDR interlaced video to HDR progressive video, realizing efficient conversion and unified format of video, and improving the reproduction fidelity of video.
Patent Information
- Application Number
- CN202380056095.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-27
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-23
AI Technical Summary
It is difficult for the prior art to efficiently integrate low-complexity enhanced video encoding (LCEVC) into existing ecosystems, especially within TV hardware legacy ecosystems, to convert from SDR interlaced video to HDR progressive video.
By introducing an enhanced encoder into the encoding system, the available interlaced 1080i SDR stream is enhanced to a progressive 1080p HDR stream using LCEVC technology, which enables the conversion from SDR to HDR and simplifies the technical workflow of video.
The process of converting input video from interlaced to line by line is realized, simplifying the technical workflow of video, and unifying it into a single format, such as HDR line by line 4K or higher, improving the fidelity of video reproduction.
Smart Images

Figure CN120035996A_ABST
Abstract
Description
Background Art
[0001] The coding technology in the following specification is particularly suitable for use with the existing Low Complexity Enhancement Video Coding (LCEVC) technology.
[0002] The standard specification of LCEVC is provided in the text of ISO / IEC 23094-2, version 1, "Low Complexity Enhancement Video Coding", published in November 2021, and many possible implementation details of LCEVC are described in patent publications WO 2020 / 188273 and WO 2020 / 188229. Each of these earlier documents is incorporated herein by reference.
[0003] Broadly speaking, LCEVC improves the fidelity of the reproduction of decoded video after encoding and decoding using an existing codec. This is achieved by combining a base layer with an enhancement layer, where the base layer contains video encoded using an existing codec and the enhancement layer indicates the residual between the original video and the predicted decoded video produced by decoding the base layer using the existing codec. The enhancement layer can be combined with the decoded base layer to more accurately reproduce the original video.
[0004] Color conversion within a scalable video coding scheme has been previously described in WO2020 / 074896, the contents of which are incorporated herein by reference.
[0005] Examples of implementations of LCEVCs may be described in WO2022 / 023747 and WO2022 / 023739, which are incorporated herein by reference.
[0006] Effectively and efficiently integrating enhancement coding into the existing ecosystem remains a goal. DETAILED DESCRIPTION
[0007] Various aspects of the invention are set out in the accompanying claims.
[0008] Examples according to the present disclosure provide integration of enhanced coding (i.e., LCEVC) in current TV specifications to enhance available interlaced 1080i SDR streams to progressive 1080p HDR streams to improve current TV use cases, particularly within the TV hardware legacy ecosystem. One of the key advantages of enabling conversion from SDR interlaced video to HDR progressive video is the ability to convert the generation of input video from interlaced to progressive, thereby simplifying the generation of video for those services that currently require the generation of both interlaced and progressive sequences (since many legacy systems still support interlaced, while more modern systems are progressive-based). In this sense, this solution simplifies the technical workflow of the input video and unifies it into a single format (typically HDR progressive 4K or higher).
[0009] It should be noted that the scheme proposed herein is also suitable for enhancing lower resolution interlaced video (eg, 576i or others).
[0010] In a specific current use case, as of today, in most of Brazil, the TV 2.0 standard is used for broadcasting. While TV 3.0 will replace it in the near future, the goal is to distribute an enhancement layer on top of TV 2.0 to provide 10-bit HDR streams to supported receivers. TV 2.0 currently uses AVC 1080i30 SDR (BT.709 color space) and may need to be backwards compatible.
[0011] Examples disclosed herein enable TV 2.0 infrastructure to be enhanced with LCEVC for playback as 1080p60 HDR with BT.2020 color space.
[0012] Initial research activity around deinterlacing and HDR / color space conversion has investigated the feasibility of using enhanced encoding and taking advantage of the fact that deinterlacing and tone mapping blocks are not necessarily standardized and the available algorithms can vary depending on the chipset.
[0013] The number of supported decoding devices (TV, Set Top Box (STB)) and their similarity (ie, de-interlacing process, location of color space conversion blocks within the framework) may influence the development of the decoder.
[0014] Figure 1 An example of an encoding system according to the present disclosure is shown.
[0015] In some instances, an input sequence or input frame is obtained. The input sequence can be obtained at various resolutions, frame rates, SDR / HDR characteristics, bit depths, and color gamuts. In some instances, the input sequence is obtained at a resolution of 2160p60 (i.e., at a 4K resolution). The input sequence can be a 10-bit HDR input signal using an HDR color space such as BT.2020.
[0016] Throughout this specification, we will refer to color conversions and conversions from one definition or type of color to another. Sometimes we will use the terms gamut, space, and range. These are not necessarily interchangeable, but where the text describes a conversion (e.g., a color space conversion, a color range conversion, or a gamut conversion), it should be understood that all conversions can be provided separately, independently, or simultaneously in the same step. For example, where the text describes a color conversion, the conversion can be from HDR to SDR, from 10 bits to 8 bits, a narrowing or widening of a color space (e.g., BT202 to 709), or a narrowing / widening of the gamut. For brevity, only one term is used in some places, but it should be understood that all conversions can be performed simultaneously or separately, or at different locations in the process. For example, the range can be changed before scaling, and the space can be changed at another point in the process.
[0017] The input sequence is optionally provided to a downsampler that converts the resolution of the input frames to a lower resolution. In some instances, the resolution can be downsampled from 2160p to 1080p. However, these are merely examples, and it should be understood that the downsampler converts the input frames from a relatively high resolution to a relatively low resolution.
[0018] The downsampled signal may then optionally be HDR mapped. This may include providing HDR metadata to an enhancement encoder. For example, the enhancement encoder may be an LCEVC encoder and the HDR metadata may be static HDR10 metadata.
[0019] The input frame is then converted. This may include a color conversion to a different color gamut. This may additionally or alternatively include converting the dynamic range, for example the downsampled input signal may be converted from an HDR color space to an SDR color space. In some instances, this may involve converting the downsampled signal from BT.2020 to BT.709. In general, this may be considered a "down conversion" because it converts a first color range to a second color range, where the first color range has a larger volume (e.g., a "wider" color space and / or a larger dynamic range) than the second color range. Thus, "down conversion" may be considered to reduce the volume of the color range of a signal.
[0020] Optionally, the bit depth of the signal may be reduced. For example, the bit depth may be reduced from 10 bits to 8 bits.
[0021] In an example, an input frame is provided to a color conversion block for color conversion (i.e., "down conversion"), and in response, a color converted frame is received. In this example, the input frame has a wider color gamut (i.e., a larger area of color space) than the color converted frame. The conversion may be provided by a separate entity or module to the entity or module that provides the encoding process and method.
[0022] After color conversion, the color converted signal is interlaced. In an example, the color converted signal is provided to an interlacer for interlaced scanning, and in response, an interlaced signal is received.
[0023] This processing (ie conversion and then interlacing) may provide a high quality final reconstructed output image without utilizing the enhancement layer, and also provide a final reconstructed output image with utilizing the enhancement layer.
[0024] The interlaced signal is provided to a base codec for encoding and decoding. In some instances, the base codec may be an AVC 8-bit encoder. Of course, different base codecs may be used to achieve the same effect. The base codec may generate a base signal that is transmitted as part of a bitstream. Example base codecs include MPEG standards (such as AVC / H.264, HEVC / H.265, etc.) and non-standard algorithms (such as VP9, AV1, etc.). In an example, the method may include performing base encoding and decoding. However, base encoding and decoding are typically performed by dedicated hardware blocks, so we describe this step as providing the signal to a base codec for encoding and decoding.
[0025] The decoded presentation of the downsampled, color converted and interlaced signal is then deinterlaced. Functionally, this can be the inverse of the interlaced operation performed before encoding. Thus, the decoded signal is converted from 1080i30 to 1080p60. Thus, in an example, the described example may include deinterlacing the decoded presentation of the downsampled, color converted and interlaced signal. However, deinterlacing can be performed by a dedicated (e.g., hardware) block, and therefore in the described example, we instead describe providing the decoded presentation of the downsampled, color converted and interlaced signal to a deinterlacer, and in response, receiving the decoded deinterlaced presentation of the downsampled, color converted and interlaced signal.
[0026] It should be noted that each block shown in the figures may be performed by a coding process or an independent or separable module or entity. For example, modules present on a pre-existing SoC may be utilized in the pipeline and their functions managed by indicating their functions or by passing appropriate data to the module and receiving output in response.
[0027] After deinterlacing, the decoded signal is color converted. Functionally, the color range conversion can be the inverse of the previous color range conversion. That is, the color range of the decoded signal is converted back to the color range of the original input frame. In some instances, the signal is converted from an SDR signal (e.g., BT.709) to an HDR signal (e.g., BT.2020). Generally speaking, this can be considered "up-conversion" because it converts the third color range to a fourth color range, where the fourth color range has a larger volume than the third color range. Therefore, "up-conversion" can be considered to increase the color range of the signal.
[0028] Many different techniques for converting from one color range to another color space (i.e., "up-conversion" and "down-conversion") are known, for example, Recommendation ITU-R BT.2087 https: / / www.itu.int / rec / R-REC-BT.2087-0-201510-I / en, Recommendation ITU-R BT.709 https: / / www.itu.int / rec / R-REC-BT.709, Recommendation ITU-R BT.2020 https: / / www.itu.int / rec / R-REC-BT.2020, Recommendation ITU-R BT.2100 https: / / www.itu.int / rec / R-REC-BT.2100 are all methods of converting one color range to another color range.
[0029] The deinterlaced presentation and color converted presentation of the decoded signal may be referred to as a reconstructed signal. Thus, in an example, the described example may include color converting a downsampled, color converted, and decoded deinterlaced presentation of an interlaced signal. However, the color conversion may be performed by a dedicated (e.g., hardware) block, and thus in the described example, we instead describe providing the downsampled, color converted, and decoded deinterlaced presentation of an interlaced signal to a color space converter, and in response, receiving a color converted decoded deinterlaced presentation of the downsampled, color converted, and interlaced signal.
[0030] In an example, de-interlacing and color conversion may be asymmetric, i.e., they may use different methods or parameters. This may introduce different errors or artifacts to be corrected.
[0031] Optionally, the bit depth of the encoded signal may be converted (eg, increased). In an example, the bit depth is converted from 8 bits to 10 bits, which may introduce different errors or artifacts to be corrected.
[0032] A residual signal is then generated by taking the difference between the representation of the input signal and the reconstructed signal. The residual signal is provided to an enhancement encoder which may also be configured to receive HDR metadata in an optional configuration.
[0033] The enhancement encoder generates an enhancement signal which can be transmitted together with the base signal in a bitstream. The enhancement signal may optionally further include supplementary HDR data, such as HDR10 data as SEI / VUI data.
[0034] It is important to note that the steps of downsampling, HDR mapping and base encoding can be found in conventional encoding systems, so these modules can be reused. However, the color conversion and interlacing of the signal before passing it to the base encoder, the subsequent deinterlacing and inverse color space conversion, as well as the generation of the residual and the enhancement encoding are specific to the enhancement encoding pipeline. In particular, these steps may be specific to LCEVC processing.
[0035] In an example, the color space dialogue occurs before interlacing in the "downscaling" path and after deinterlacing in the "upscaling" or "reconstruction" path, i.e., after the underlying encoding and decoding. In this way, the color range conversion can maximally correct the deinterlaced signal to improve the overall reconstructed video quality.
[0036] It should be understood that throughout, where a resolution, bit depth or frame rate is mentioned, this is exemplary only and further resolutions are possible.
[0037] It should be noted that the examples herein do not include the typical LCEVC spatial upscaling step in the encoding pipeline. Of course, it should be understood that the spatial upscaling included in LCEVC can also be used with the examples herein, so that the color space converted de-interlaced signal is further upscaled to be combined with the higher resolution input signal to produce a set of residuals. This can be for example Figure 1 A complement or alternative to the set of residuals encoded as shown.
[0038] The bitstream is transmitted and received by a decoding system. An example of a decoding system is Figure 2 . The bitstream typically includes a base signal, and the enhanced signal is received by a decoding system. The base signal is passed to a base decoder, and the enhanced signal is passed to an enhanced decoder. In some instances, the base decoder may be an 8-bit AVC decoder, and the enhanced decoder may be a 10-bit LCEVC decoder.
[0039] The base signal is decoded and deinterlaced by a base decoder. For example, the decoded representation of the signal may be temporally scaled from a resolution of 1080i30 to 1080p60. In an example, this is accomplished by passing the decoded representation of the signal to a deinterlacer (e.g., a deinterlacing hardware block) and receiving a deinterlaced decoded signal in response.
[0040] After deinterlacing, the signal undergoes a color range conversion. For example, the signal can be converted from an SDR color space to an HDR color space. Additionally or alternatively, the signal can be converted from a narrow color gamut to a wide color gamut. In a specific example, the color gamut can be converted from BT.709 to BT.2020 or BT.2087. In an example, this is achieved by passing a decoded deinterlaced presentation of the signal to a color range converter (e.g., a color gamut conversion hardware block and / or a dynamic range converter block) and receiving a color range converted deinterlaced decoded signal in response.
[0041] Optionally, the bit depth of the signal may be upscaled, for example from 8 bits to 10 bits.
[0042] In some instances, spatial upsampling can be performed after the signal is converted to the HDR color space. It should be noted that the scheme proposed herein is therefore also suitable for enhancing lower resolution interlaced video (e.g., 576i or other). Spatial upsampling can be based on existing LCEVC technology and can be located at various different points of the pipeline to achieve the desired effect. In these instances, the upsampled frame is compared with a frame of a similar resolution input video to generate a set of residuals for encoding using an enhanced coding technique (e.g., LCEVC). The color conversion presentation presented by the decoding of the base signal is delivered to the enhanced decoder. The enhanced decoder combines the color conversion presentation presented by the decoding of the base signal with the enhanced signal to produce an enhanced video signal, which, in an instance, provides an enhancement layer that can correct artifacts / errors introduced during color range conversion and / or deinterlacing. In other instances, color range conversion is performed after the enhancement layer is applied to the base layer (i.e., by a decoder), which can provide a process compatible with a wide range of decoding devices.
[0043] The enhancement decoder may also generate an HDR metadata stream, such as static HDR10 metadata (if such metadata is included with the enhancement stream).
[0044] The video signal and metadata are delivered to an output device. For example, the video signal and metadata are delivered to a TV for display. In some instances, the TV is a TV that supports the TV 2.5 specification. That is, the TV supports certain aspects of HDR.
[0045] The enhanced video signal may have a 1080p resolution, HDR10 image quality, and a BT.2020 color gamut.
[0046] To ensure backward compatibility, in some instances, the decoded base signal may be passed to the output device without upsampling or color gamut conversion. This may be necessary, for example, for a TV that complies with the TV 2.0 specification. In this case, the output signal is provided at a resolution of 1080i30 for SDR and a color gamut of BT.709 without using the enhanced signal.
[0047] It should be noted that the base codec, upsampling and color conversion may be part of an existing video decoder pipeline, eg implemented by a SoC. The enhancement decoder is specific to an enhancement decoding pipeline, such as the LCEVC pipeline.
[0048] Certain notable features of examples of the present disclosure include:
[0049] - Interlacing and deinterlacing.
[0050] - In the example, there is upsampling / downsampling within the enhancement encoding pipeline. In the above example, downsampling is optionally done before the color conversion / interlacing / deinterlacing / conversion pipeline, but not within that pipeline. This is a bit unusual because normally if downsampling occurs in the pipeline, then upsampling would be expected elsewhere in the pipeline (which would then be compared to get a set of residuals).
[0051] - The combination and order of deinterlacing and SDR to HDR conversion is noteworthy, in particular (at the decoder) deinterlacing is configured before color range conversion.
[0052] In summary, the following are the exemplary steps of the proposed example:
[0053] a. Perform color conversion on the input frame, and optionally perform color conversion on the rendering of the input frame (e.g. spatial downsampling rendering). The conversion can be, for example, from HDR to SDR, and further optionally from 10 bits to 8 bits.
[0054] b. Interlace the (color converted) frames.
[0055] c. Sending the interlaced color converted (ie SDR) frame to the underlying codec (for encoding and decoding). Further, and in response to said sending, a decoded encoded representation of the interlaced color converted frame may be received.
[0056] d. Deinterlace the decoded encoded presentation of the interlaced color converted frame.
[0057] e. Perform color conversion on the deinterlaced frame to produce a "reconstructed" frame. In fact, this color conversion can be functionally considered as the reverse process of the previous color conversion.
[0058] f. Generate a residual between the "reconstructed" frame and the input frame (rendering).
[0059] g. Encode the residual, for example, using an enhanced encoder such as LCEVC.
[0060] The residual may be encoded in a step of generating an encoded enhancement signal for the encoded video signal, the encoded enhancement signal comprising one or more layers of residual data generated based on a comparison of data derived from the decoded video signal with data derived from the input video signal.
[0061] It should be noted that although the above steps describe conversion from SDR to HDR at the encoder after deinterlacing and before residual calculation, in an embodiment, the conversion may be performed after residual calculation (e.g., a hardware block in the decoder that performs SDR to HDR conversion may occur on the output). This may provide a process that is compatible with a wide range of decoding devices.
[0062] In an example, there may be further optional steps, such as:
[0063] h. Encode HDR metadata together with the residual.
[0064] i. Downsample the input frame to obtain a "rendering" of the input frame.
[0065] In an example, the input frame may be 10 bits, and the color range conversion may include converting from 10 bits to 8 bits, and optionally vice versa. The first color range conversion may convert from a first space (e.g., BT2020) to a second space (e.g., BT709), and the second conversion may convert from a second space (e.g., BT709) to a first space (e.g., BT2020). In this way, the conversion may be symmetrical, but in other examples, it may be asymmetrical.
[0066] The implementation described above is non-trivial for a number of reasons, not least because in practical embodiments the de-interlacer at the decoder side is often different from the interlacer at the encoder.
[0067] Figures 3 to 11 Shows Figure 1 and Figure 2 Instance variant of the step.
[0068] exist Figure 3 In , HDR metadata is sent from the input sequence to the enhancement encoder for assembly into the bitstream. Figure 3In , the color range conversion is a SL-HDR1 decomposition. The input sequence undergoes a SL-HDR1 decomposition, where SL-HDR1 metadata is provided to the AVC encoding and decoding steps, and a SL-HDR1 reconstruction step after deinterlacing. Figure 1 Likewise, after the color range conversion of the SL-HDR1 decomposition, the signal is deinterlaced and provided to the AVC encoder. The encoded version of the signal is output and decoded. The decoded version is provided to the deinterlacer and SL-HDR1 reconstruction step to reconstruct the video before comparison with the input sequence to create a residual for LCEVC encoding. Figure 4 The corresponding decoding process is shown, where SL-HDR1 metadata is provided from an AVC decoder to a color range conversion step performed after de-interlacing, i.e. SL-HDR1 reconstruction.
[0069] exist Figure 5 and Figure 6 In the variant of , the post-generation step occurs after downsampling by providing static HDR10 metadata to the LCEVC encoder. Color range conversion is an HDR decomposition where metadata is provided to the base encoder. On the decoder side, as Figure 4 As shown, the AVC decoder provides HDR metadata to the HDR reconstruction step that occurs after de-interlacing.
[0070] Figure 7 A decoding variant is shown in which the base decoded signal is HDR reconstructed and provided to the TV alone.
[0071] Figure 8 An 8-bit method is shown. A rendering of the input signal after color range conversion is combined with a deinterlaced base decoded rendering to produce a residual. In this example, during the reconstruction process, the color range is not recreated, but the output of the color range conversion is used before the interlacing step to form the residual.
[0072] Fig. 9 Shows Figure 8 8-bit method decoding, where the color range conversion occurs after the enhanced decoding.
[0073] Fig.10 and Fig.11 Demonstrates a decoding mode that selectively disables deinterlacing and color space conversion.
[0074] Color depth can also be called bit depth. Color depth can represent the total number of bits used to indicate the color of a single pixel (bpp), such as in a bitmap image or video frame buffer. Color depth can represent the number of bits used to make up each color component (such as red, green, and blue) of a single pixel. Color depth determines the total / maximum amount of different colors that can be displayed. This does not mean that the image must use all of these available colors, but it can specify colors with that level of precision.
[0075] A color space may also be called a color model or a color system.
[0076] Color range may be referred to as color volume. Thus, color range may include and / or be associated with one or more of: dynamic range (i.e., color range), color gamut (i.e., color range), color depth (i.e., color range). Thus, color range may determine which specific colors may be displayed and at what brightness, which will affect their "brightness" and "color" (intensity and saturation). Color range may further determine how many different color variations (hues / chromas) may be seen. A larger color gamut may be referred to as a wider color gamut.
[0077] In an example, we describe a system including a modified downscaler / downsampler and a modified upscaler / upsampler. The modified downsampler / downscaler is configured to: downscale the color range (e.g., color volume) of a signal (e.g., a video frame) to produce a downscaled signal (e.g., an input video); and temporally downsample the downscaled signal to produce a downscaled downsampled signal. The modified downscaler can be configured to input the downscaled downsampled signal to a base codec, which in turn returns a decoded encoded representation of the content input into the base codec. The modified upsampler / upscaler is configured to: upscale the color range (e.g., color volume) of the signal to produce an upscaled signal; and temporally upsample the upscaled signal to produce an upscaled upsampled signal. The system can be configured to compare the upscaled upsampled signal with a representation of the (e.g., input) signal to produce a residual. The system can be configured to optionally encode the residual using an LCEVC encoding scheme.
[0078] We describe a corresponding system that is in communication with the system and is configured to receive an encoded stream. The corresponding system includes a modified upscaler / upsampler. The corresponding system includes a decoder configured to: decode a base signal using a base decoder; upsample / upscaling the decoded base signal using the modified upscaler / upsampler; decode an encoded residual enhancement signal; and combine the decoded residual with the upsampled / upscaling base signal.
[0079] In all examples described herein, the base video may be output with the LCEVC encoded signal or may be output separately from the LCEVC encoded signal. That is, the base may be encoded with the LCEVC signal, or LCEVC may encode the residual as a separate signal for independent transmission, and the LCEVC encoder does not include the base signal in the output enhancement sequence.
[0080] In some cases, a base codec may be used. The base codec may include an independent codec controlled in a modular or "black box" manner. The methods described herein may be implemented with the aid of a computer program code that is executed by a processor and makes function calls on a hardware and / or software implemented base codec.
[0081] In general, the term "residual" as used herein refers to the difference between the value of a reference array or reference frame and the actual array or frame of data. The array may be a one-dimensional or two-dimensional array representing a coding unit. For example, a coding unit may be a 2x2 or 4x4 group of residual values corresponding to a similarly sized region of an input video frame. It should be noted that this generalized example is agnostic to the nature of the encoding operation performed and the input signal. Reference to "residual data" as used herein refers to data derived from a set of residuals, such as a set of residuals themselves or the output of a set of data processing operations performed on a set of residuals. Throughout this specification, in general, a set of residuals includes a plurality of residuals or residual elements, each of which corresponds to a signal element, i.e., an element of a signal or raw data. The signal may be an image or video. In these examples, the set of residuals corresponds to an image or frame of a video, wherein each residual is associated with a pixel of a signal, which is a signal element. The examples disclosed herein describe how these residuals can be modified (i.e., processed) to affect the encoding pipeline or the final decoded image while reducing the overall data size. The residuals or sets may be processed per residual element (or residual), or on a group basis, such as per tile or per coding unit, where a tile or coding unit is a contiguous subset of a set of residuals. In one case, a tile may comprise a group of smaller coding units. A tile may comprise a 16x16 set of pixels or residuals (e.g., an 8 by 8 set of 2x2 coding units, or a 4 by 4 set of 4x4 coding units). It should be noted that the processing may be performed on every video frame, or on only a set number of frames in a sequence.
[0082] In general, each enhancement stream or both enhancement streams may be encapsulated into one or more enhancement bitstreams using a set of Network Abstraction Layer Units (NALUs). A NALU is intended to encapsulate an enhancement bitstream in order to apply the enhancement to the correct base reconstructed frame. A NALU may, for example, contain a reference index to a NALU containing the base decoder reconstructed frame bitstream to which the enhancement must be applied. In this way, the enhancement may be synchronized to the base stream and the frames of each bitstream combined to produce the decoded output video (i.e., the residual of each frame of the enhancement level is combined with the frames of the base decoded stream). A group of pictures may represent multiple NALUs.
[0083] In some instances, a bit sequence representing an encoding of a video signal is provided, the bit sequence comprising a plurality of encoded streams, the bit sequence comprising: a base layer encoded stream, encoded by a base encoder, the stream comprising an interlaced representation of a color converted representation of an initial representation of an input signal, the initial representation of the input signal having a first color range, and the color converted representation having a second color range; an enhancement layer encoded stream, encoded by an enhancement encoder, the stream comprising a set of residuals representing differences between a reconstructed decoded representation of the initial representation and a version of the initial representation, the reconstructed decoded representation comprising a deinterlaced representation of a decoded representation of an encoded representation of an interlaced color converted representation from the base encoder, the residuals being used to be combined with a decoded version of the base encoded signal to reconstruct a representation of the initial representation of the input signal.
[0084] The following particularly preferred examples of the present disclosure are described as a group of numbered clauses. It should be understood that these are examples that are helpful for understanding the present invention.
[0085] Numbered Clause 1. A method of encoding a signal for reproducing interlaced or non-interlaced video, the method comprising:
[0086] obtaining an initial representation of an input signal having a first color range;
[0087] converting the initial representation of the input signal from the first color range to a second color range to produce a color converted representation;
[0088] processing the color converted presentation to produce an interlaced color converted presentation, the interlaced color converted presentation being an interlaced presentation of the color converted presentation;
[0089] sending the interlaced color converted presentation to an underlying encoder;
[0090] obtaining a decoded representation of an encoded representation of the interlaced color converted representation from the base encoder;
[0091] processing the decoded presentation to produce a reconstructed decoded presentation, the reconstructed decoded presentation being a deinterlaced presentation of the decoded presentation;
[0092] comparing the reconstructed decoded representation to a version of the initial representation to produce a set of residuals; and
[0093] Instructing to encode the set of residuals using an enhancement encoder to produce an encoded enhancement signal for later decoding of the residuals for combination with a decoded version of a base encoded signal to reconstruct a representation of the initial representation of the input signal.
[0094] Numbered clause 2. A method according to numbered clause 1, wherein the reconstructed decoded presentation and / or the initial presentation
[0095] The version uses a different color range than the decoded rendering.
[0096] Numbered clause 3. A method according to numbered clause 1 or 2, wherein the step of comparing comprises comparing the reconstructed decoded
[0097] Now compare it with the color conversion presentation.
[0098] Clause 4. A method according to clause 1 or 2, wherein the step of processing the decoded presentation includes: converting the deinterlaced decoded presentation from a third color range to a fourth color range to produce the reconstructed decoded presentation.
[0099] Item 5. The method according to Item 4, wherein the second color range is
[0100] same.
[0101] Item 6. The method according to item 4 or 5, wherein the first color range is
[0102] Same scope.
[0103] Numbered clause 7. A method according to any one of numbered clauses 4 to 6, wherein from the third color range to the
[0104] The conversion of the fourth color range is opposite to the conversion from the first color range to the second color range.
[0105] Numbered clause 8. A method according to any one of numbered clauses 4 to 7, wherein the fourth color range includes
[0106] The third color range has a larger color volume.
[0107] Clause 9. A method according to any one of clauses 4 to 8, wherein the deinterlaced decoding is presented from
[0108] Converting the third color range to the fourth color range includes increasing the color volume.
[0109] Item 10. The method of any preceding item, wherein the input signal is a frame of an input video. Item 11. The method of any preceding item, wherein the step of initially presenting the input signal having a first color range comprises:
[0110] receiving the input signal; and
[0111] The input signal is spatially downsampled.
[0112] Numbered clause 12. A method according to any preceding numbered clause, wherein the method further comprises:
[0113] HDR mapping the initial rendering to generate a set of HDR metadata; and
[0114] The HDR metadata is provided to an enhancement decoder for encoding with the encoded enhancement signal.
[0115] Numbered clause 13. A method according to any preceding numbered clause, wherein the enhancement encoder is LCEVC.
[0116] Numbered clause 14. A method according to any preceding numbered clause, wherein the initial presentation is changed from the first color
[0117] The step of converting the range to a second color range is a conversion that reduces the dynamic range of the first color space, optionally wherein the step of converting the initial presentation from the first color range to the second color range is a conversion from HDR to SDR.
[0118] Numbered clause 15. A method according to any preceding numbered clause, wherein the initial presentation is changed from the first color
[0119] The step of converting the range into a second color range is a conversion from a first color space to a second color space, wherein the first color space comprises a wider color gamut than the second color space.
[0120] Numbered clause 16. A method according to any of the preceding numbered clauses, wherein the first color range has a greater
[0121] The two color ranges have a higher color depth, in particular, wherein the first color range is a 10-bit color range and the second color range is an 8-bit color range.
[0122] Numbered clause 17. A method according to any of the preceding numbered clauses, wherein the first color range includes more than the second color range.
[0123] 2. Larger color range and volume.
[0124] Clause 18. The method of clause 17, wherein the initial presentation of the input signal is changed from the
[0125] The first color range is converted to the second color range to produce a color conversion presentation including one or more of the following:
[0126] The color depth of the first color space is reduced; the color gamut of the first color space is reduced; and the dynamic range of the first color space is reduced.
[0127] Numbered clause 19. A method according to any preceding numbered clause, wherein the color range includes color space data. Numbered clause 20. A method according to any preceding numbered clause, wherein the color range includes dynamic range data. Numbered clause 21. A method according to any preceding numbered clause, wherein the color range includes color gamut data.
[0128] Clause 22. A method of decoding a coded enhancement signal for use in reproducing a presentation of a video signal, the method comprising:
[0129] The law includes:
[0130] Obtaining a coded signal comprising a coded enhancement signal and a base coded signal;
[0131] passing the base encoded signal to a base decoder to produce an interlaced base decoded video signal;
[0132] processing the base decoded video signal to produce a reconstructed representation of the video signal, the reconstructed representation being a deinterlaced representation of the base decoded video signal;
[0133] combining the reconstructed representation with a set of residuals decoded from an encoded enhancement stream to produce a reconstructed output video,
[0134] Wherein the reconstructed output video is in a sixth color range different from the base decoded video signal which is in a fifth color range.
[0135] Numbered clause 23. The method according to numbered clause 22, wherein the step of combining comprises combining the reconstructed representation from the
[0136] The fifth color range is converted into the sixth color range.
[0137] Numbered clause 24. The method according to numbered clause 22, wherein the method further comprises:
[0138] A deinterlaced decoded representation of the video signal is converted from the fifth color range to the sixth color range to produce the reconstructed representation of the video signal.
[0139] Numbered clause 25. A method according to any one of numbered clauses 22 to 24, wherein the sixth color range includes
[0140] One or more of: a greater color depth than the fifth color range; a greater color gamut than the fifth color range; and a greater dynamic range than the fifth color range.
[0141] Numbered clause 26. The method according to any one of numbered clauses 22 to 25, further comprising:
[0142] Outputs interlaced based decoded video.
[0143] Numbered clause 27. A method according to any one of numbered clauses 22 to 26, wherein the fifth color range is
[0144] The three color ranges are identical, and the third color range is used to reconstruct a presentation of the video signal during encoding.
[0145] Numbered clause 28. A method according to any one of numbered clauses 22 to 27, wherein the sixth color range is
[0146] The four color ranges are the same, with the fourth color range being used to convert the deinterlaced base decoded presentation of the video signal during encoding.
[0147] Numbered clause 29. The method according to any one of numbered clauses 22 to 28, further comprising:
[0148] HDR metadata is obtained from the encoded enhancement signal and provided together with the reconstructed output video.
[0149] Clause 30. A method according to any one of clauses 22 to 29, wherein the encoding is performed using LCEVC.
[0150] Enhanced signal encoding.
[0151] Numbered clause 31. A method according to any one of numbered clauses 22 to 30, wherein the step of converting comprises converting the
[0152] The de-interlaced decoding presentation of the video signal is converted from SDR to HDR.
[0153] Numbered clause 32. A method according to any one of numbered clauses 22 to 31, wherein the step of converting comprises converting the
[0154] The deinterlaced decoded presentation of the video signal is converted from one color space to another color space.
[0155] Numbered clause 33. A method according to any one of numbered clauses 22 to 32, wherein the color range includes colors
[0156] Spatial data.
[0157] Clause 34. A method according to any one of clauses 22 to 33, wherein the color range includes dynamic
[0158] Range data.
[0159] Numbered clause 35. A method according to any one of numbered clauses 22 to 34, wherein the color range includes a color gamut
[0160] data.
[0161] Clause 36. An encoder configured to perform a method according to any one of clauses 1 to 21. Clause 37. A decoder configured to perform a method according to any one of clauses 22 to 35. Clause 38. A non-transitory computer-readable storage medium storing instructions, the instructions being operable when executed by one or more
[0162] When executed by a processor, the processor is caused to perform a method according to any one of numbered clauses 1 to 21 or 22 to 35.
[0163] No. 39. A method substantially as herein described or as Figures 1 to 11 Any method or product shown in .
Claims
1. A method for encoding a signal for reproducing interlaced or non-interlaced video, the method include: obtaining an initial representation of an input signal having a first color range; converting the initial representation of the input signal from the first color range to a second color range to produce a color converted representation; processing the color converted presentation to produce an interlaced color converted presentation, the interlaced color converted presentation being an interlaced presentation of the color converted presentation; sending the interlaced color converted presentation to an underlying encoder; obtaining a decoded representation of an encoded representation of the interlaced color converted representation from the base encoder; processing the decoded presentation to produce a reconstructed decoded presentation, the reconstructed decoded presentation being a deinterlaced presentation of the decoded presentation; comparing the reconstructed decoded representation to a version of the initial representation to produce a set of residuals; as well as Instructing to encode the set of residuals using an enhancement encoder to produce an encoded enhancement signal for later decoding of the residuals for combination with a decoded version of a base encoded signal to reconstruct a representation of the initial representation of the input signal.
2. The method of claim 1, wherein the reconstructed decoded presentation and / or the version of the initial presentation uses a different color range than the decoded presentation.
3. A method according to claim 1 or 2, wherein the step of comparing comprises comparing the reconstructed decoded representation with the color converted representation.
4. The method according to claim 1 or 2, wherein the step of processing the decoded presentation include: The de-interlaced decoded representation is converted from the third color range to a fourth color range to produce the reconstructed decoded representation. The method of claim 4 , wherein the second color range is the same as the third color range.
6. The method according to claim 4 or 5, wherein the first color range is the same as the fourth color range.
7. The method according to any one of claims 4 to 6, wherein the conversion from the third color range to the fourth color range is the opposite of the conversion from the first color range to the second color range.
8. A method according to any one of claims 4 to 7, wherein the fourth color range comprises a larger color volume than the third color range, preferably wherein converting the deinterlaced decoded presentation from the third color range to the fourth color range comprises increasing the color volume.
9. A method according to any preceding claim, wherein the step of initially presenting an input signal having a first colour range include: receiving the input signal; as well as The input signal is spatially downsampled.
10. A method according to any preceding claim, wherein the method further comprises include: performing HDR mapping on the initial rendering to generate a set of HDR metadata; as well as The HDR metadata is provided to an enhancement decoder for encoding with the encoded enhancement signal.
11. A method according to any preceding claim, wherein the step of converting the initial presentation from the first color range to the second color range is a conversion that reduces the dynamic range of the first color space, optionally wherein the step of converting the initial presentation from the first color range to the second color range is a conversion from HDR to SDR.
12. A method according to any preceding claim, wherein the step of converting the initial presentation from the first color range to a second color range is a conversion from a first color space to a second color space, wherein the first color space comprises a wider color gamut than the second color space.
13. The method according to any preceding claim, wherein the first color range has a higher color depth than the second color range, in particular wherein the first color range is a 10-bit color range and the second color range is an 8-bit color range.
14. A method according to any preceding claim, wherein the first colour range comprises a larger colour volume than the second colour range.
15. The method of claim 14, wherein converting the initial presentation of the input signal from the first color range to a second color range to produce a color converted presentation comprises one or more of: reducing the color depth of the first color space; reducing the color gamut of the first color space; and reducing the dynamic range of the first color space.
16. A method of decoding a coded enhancement signal for rendering a video signal, the method include: Obtaining a coded signal comprising a coded enhancement signal and a base coded signal; passing the base encoded signal to a base decoder to produce an interlaced base decoded video signal; processing the base decoded video signal to produce a reconstructed representation of the video signal, the reconstructed representation being a deinterlaced representation of the base decoded video signal; combining the reconstructed representation with a set of residuals decoded from an encoded enhancement stream to produce a reconstructed output video, Wherein the reconstructed output video is in a sixth color range different from the base decoded video signal which is in a fifth color range.
17. The method of claim 16, wherein the step of combining comprises converting the reconstructed representation from the fifth color range to the sixth color range.
18. The method according to claim 16, wherein the method further comprises include: A deinterlaced decoded representation of the video signal is converted from the fifth color range to the sixth color range to produce the reconstructed representation of the video signal.
19. The method of any one of claims 16 to 18, wherein the sixth color range comprises one or more of: a greater color depth than the fifth color range; a larger color gamut than the fifth color range; and a greater dynamic range than the fifth color range.
20. The method according to any one of claims 16 to 19, further comprising: include: Outputs interlaced based decoded video.
21. The method according to any one of claims 16 to 20, wherein The fifth color range is the same as a third color range, the third color range being used to reconstruct a presentation of the video signal during encoding.
22. The method according to any one of claims 16 to 21, wherein The sixth color range is the same as the fourth color range used to convert a deinterlaced based decoded presentation of a video signal during encoding.
23. An encoder configured to perform the method according to any one of claims 1 to 15.
24. A decoder configured to perform the method according to any one of claims 16 to 22.
25. A non-transitory computer-readable storage medium storing instructions which, when executed by one or more processors, cause the processors to perform the method of any one of claims 1 to 15 or 16 to 22.
26. A bit sequence representing an encoding of a video signal, the bit sequence comprising a plurality of encoded streams, the bit sequence include: a base layer encoded stream encoded by a base encoder, the stream comprising an interlaced representation of a color converted representation of an initial representation of an input signal, the initial representation of the input signal having a first color range and the color converted representation having a second color range; An enhancement layer coded stream, encoded by an enhancement encoder, the stream comprising a set of residuals representing differences between a reconstructed decoded presentation of the initial presentation and a version of the initial presentation, the reconstructed decoded presentation comprising a deinterlaced presentation of a decoded presentation of an interlaced color converted presentation from the base encoder, the residuals being used to be combined with a decoded version of a base coded signal to reconstruct a presentation of the initial presentation of the input signal.
Citation Information
Patent Citations
Colour conversion within a hierarchical coding scheme
WO2020074896A1
Processing of residuals in video coding
WO2020188229A1
Low complexity enhancement video coding
WO2020188273A1
Integrating a decoder for hierachical video coding
WO2022023739A1
Integrating an encoder for hierachical video coding
WO2022023747A1