Signaling low dynamic range for image and video coding

By explicitly or implicitly representing LDR in video decoding, the problem of low encoding efficiency in video decoding is solved, and more efficient encoding and decoding is achieved.

CN120345248APending Publication Date: 2025-07-18TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005229.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2024-06-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively represent and apply low dynamic range (LDR) in video decoding, resulting in low encoding efficiency.

Method used

By explicitly or implicitly signaling in a video or image bitstream, and selecting a specific LDR, encoding and decoding is performed using predefined LDR indexes and syntax elements.

Benefits of technology

It improves the encoding efficiency of video decoding, reduces signaling overhead, and adapts to the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345248A_ABST
    Figure CN120345248A_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to video coding, and in particular to methods and systems for signaling low dynamic range (LDR) in a video or image bitstream. For example, this disclosure describes implementations for explicitly or implicitly signaling, at various signaling levels, whether to apply LDRs and which of the LDRs is applied to a block.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporated by reference

[0001] This application claims the benefit of priority to U.S. Non - Provisional Patent Application No. 18 / 756,282, filed on June 27, 2024, and U.S. Provisional Patent Application No. 63 / 542,073, filed on October 2, 2023, both entitled "Signaling of Low Dynamic Range for Image and Video Coding", which are hereby incorporated by reference in their entirety. Technical Field

[0002] The present disclosure generally relates to video coding and, more particularly, to methods and systems for signaling low dynamic range (LDR) in a video or image bitstream. Background Art

[0003] Uncompressed digital video can include a series of pictures and can be associated with specific bitrate requirements for storage, data processing, and transmission bandwidth in streaming applications. One objective of video encoding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques while reducing signaling overhead. In some applications, the bit - depth of some video / image samples can be reduced to further improve coding efficiency. Summary of the Invention

[0004] The present disclosure generally relates to video coding and, more particularly, to methods and systems for signaling low dynamic range (LDR) in a video or image bitstream. For example, the present disclosure describes implementations of ways to explicitly or implicitly signal at various signaling levels whether LDR is applied and which LDR in the LDRs is applied to a block.

[0005] In some example implementations, a method for decoding a block in a bitstream of a video or image is disclosed. The method can include: receiving the bitstream; determining, based on the bitstream, that low dynamic range is applied to the block. The method can further include, in the case where low dynamic range is applied to the block: determining an LDR selected from a set of LDRs for encoding the block; decoding the block to generate reconstructed LDR samples of the block; and generating reconstructed samples from the reconstructed LDR samples according to the selected LDR.

[0006] In the above example implementation, the set of LDRs is predefined, and the LDR index is explicitly signaled by an LDR syntax element in the bitstream or implicitly derived from the bitstream.

[0007] In any of the above example implementations, each LDR in the set of LDRs is predefined as a sampling value range 2 N , where N is a non-negative integer and represents the bit depth of each LDR in the set of LDRs.

[0008] In any of the above example implementations, the LDR index is explicitly signaled by the numerical value 7 - N.

[0009] In any of the above example implementations, the LDR index is explicitly signaled by LDR syntax elements as high-level syntax in a sequence header, picture header, sub-picture header, frame header, slice header, or tile header.

[0010] In any of the above example implementations, the LDR index is explicitly signaled by LDR syntax elements as one or more block-level syntaxes.

[0011] In any of the above example implementations, the LDR index is signaled in one or more Supplemental Enhancement Information (SEI) messages, each SEI message containing additional data inserted into the bitstream to convey extra information.

[0012] In any of the above example implementations, the application of LDR or the LDR signaling indicating the LDR applied to a block depends at least on the value of additional syntax elements in the bitstream.

[0013] In any of the above example implementations, LDR is applied or LDR signaling exists in the bitstream only if the bitstream is intended for at least machine vision as indicated by additional syntax elements.

[0014] In any of the above example implementations, LDR is applied or LDR signaling exists in the bitstream only if the quantization parameter of the block is greater than a predefined quantization parameter threshold.

[0015] In any of the above example implementations, LDR is applied or LDR signaling exists in the bitstream only if the picture resolution associated with the block is greater than or less than a predefined resolution value.

[0016] In any of the above example implementations, LDR is applied or LDR signaling exists in the bitstream only if the frame rate associated with the block is greater than or less than a predefined frame rate value.

[0017] In any of the example implementations above, the LDRs for different color components of a block are signaled separately in the bitstream.

[0018] In any of the example implementations above, the LDR is applied only to the luminance component of a block and the LDR is signaled.

[0019] In any of the example implementations above, the LDR selected for encoding a block is implicitly derived based on the encoded information of the block.

[0020] Aspects of the present disclosure also provide an encoding method corresponding to the decoding method above.

[0021] Aspects of the present disclosure also provide an electronic device or apparatus including circuitry or a processor configured to perform any of the method implementations above.

[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by an electronic device, cause the electronic device to perform any of the method implementations above.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable recording medium for storing the bitstream above.

[0024] Aspects of the present disclosure also provide a method for generating the bitstream above. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the drawings, in which:

[0026] Figure 1 A schematic illustration showing a simplified block diagram of a communication system (100) according to an example embodiment;

[0027] Figure 2 A schematic illustration showing a simplified block diagram of a communication system (200) according to an example embodiment;

[0028] Figure 3 A schematic illustration showing a simplified block diagram of a video decoder according to an example embodiment;

[0029] Figure 4 A schematic illustration showing a simplified block diagram of a video encoder according to an example embodiment;

[0030] Figure 5 A block diagram showing a video encoder according to another example embodiment;

[0031] Figure 6 Shows a block diagram of a video decoder according to another exemplary embodiment;

[0032] Figure 7 Shows a scheme for coding block partitioning according to an exemplary embodiment of the present disclosure;

[0033] Figure 8 Shows another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure;

[0034] Figure 9 Shows another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure;

[0035] Figure 10 Shows an example logic flow of a decoding method.

[0036] Figure 11 Shows an example logic flow of an encoding method.

[0037] Figure 12 Shows a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. Detailed Description

[0038] Throughout the specification and claims, terms may have nuanced meanings that are presented or implied in a context beyond the explicitly stated meaning. As used herein, the phrase "in one embodiment / implementation" or "in some embodiments / implementations" does not necessarily refer to the same embodiment / implementation, and the phrase "in another embodiment / implementation" or "in other embodiments" as used herein does not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include combinations of all or parts of the exemplary embodiments / implementations.

[0039] Generally, terms can be understood at least in part according to their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can include various context-dependent meanings. Generally, "or" when used to relate a list such as A, B, or C is intended to mean: A, B, and C, used herein in an inclusive sense; and A, B, or C, used herein in an exclusive sense. Additionally, depending at least in part on the context, terms such as "one or more", "at least one", "a", "an", or "the" as used herein can be used in a singular or plural sense. Further, the terms "based on" or "determined by" can be understood to not necessarily convey an exclusive set of factors, but can alternatively allow for the existence of additional factors that are not necessarily explicitly described, which also depends at least in part on the context.

[0040] Although the following description may focus on video encoding and decoding, various disclosed embodiments may be applicable to processing still images.

[0041] Figure 1 A simplified block diagram of a communication system (100) in accordance with an embodiment of the present disclosure is shown. The communication system (100) includes a plurality of terminal devices, such as 110, 120, 130, and 140, that may communicate with each other via, for example, a network (150). In Figure 1 an example, a first pair of terminal devices (110) and (120) may perform unidirectional transmission of data. For example, the terminal device (110) may encode video data in the form of one or more encoded bitstreams (e.g., a video picture stream captured by the terminal device (110)) for transmission via the network (150). The terminal device (120) may receive the encoded video data from the network (150), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission may be implemented in media service applications and the like.

[0042] In another example, a second pair of terminal devices (130) and (140) may perform bidirectional transmission of encoded video data, such as during a video conferencing application. For bidirectional transmission of data, in an example, each of the terminal devices (130) and (140) may encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (130) and (140), and may also receive encoded video data from the other of the terminal devices (130) and (140) to recover and display the video pictures.

[0043] In Figure 1 an example, the terminal devices may be implemented as servers, personal computers, and smart phones, but the applicability of the basic principles of the present disclosure is not limited thereto. Embodiments of the present disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and the like. The network (150) represents any number or type of network that conveys encoded video data between the terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) may exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0044] As an example of an application for the disclosed subject matter, Figure 2Shows the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter can be equally applicable to other video applications, including, for example, video conferencing, digital TV, broadcasting, gaming, virtual reality, storing compressed video on digital media including CD (Compact Disc), DVD (Digital Versatile Disc), memory sticks, etc.

[0045] As Figure 2 shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201) such as a digital imaging device for creating an uncompressed video picture or image stream (202). In an example, the video picture stream (202) includes samples recorded by the digital imaging device of the video source 201. The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (204) (or encoded video bitstream), and the video picture stream (202) may be processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof to implement or enforce aspects of the disclosed subject matter described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume when compared to the uncompressed video picture stream (202), and the encoded video data (204) (or encoded video bitstream (204)) may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems such as Figure 2 the client subsystems (206) and (208) in may access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include a video decoder (210) in an electronic device (230), for example. The video decoder (210) decodes the incoming copy (207) of the encoded video data and creates an outgoing video picture stream (211) that is uncompressed and can be presented on a display (212) (e.g., a display screen) or other rendering device (not depicted).

[0046] Figure 3 Shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may be used to replace Figure 2 the video decoder (210) in the example of

[0047] As Figure 3 shown, a receiver (331) can receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle playback timing, a buffer memory (315) can be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) can reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include information for managing the operation of a video decoder (310), and potentially include information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) can parse / entropy decode the encoded video sequences. The parser (320) can extract a set of subgroup parameters for at least one subgroup in a subgroup of pixels in the video decoder from the encoded video sequences. The subgroups can include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (320) can also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequences. The reconstruction of the symbols (321) can involve multiple different processing or functional units. The units involved and how they are involved can be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.

[0048] The first unit can include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive quantized transform coefficients as symbols (321) and control information from the parser (320), including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) can output blocks including sample values that can be input into an aggregator (355).

[0049] In some cases, the output samples of the scaler / inverse transform (351) may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use the surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (358) caches the partially reconstructed current picture and / or the fully reconstructed current picture. In some implementations, the aggregator (355) may add the prediction information that the intra-prediction unit (352) has generated to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0050] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (353) may access the reference picture memory (357) based on the motion vector to obtain samples for inter-picture prediction. After motion-compensating the obtained reference samples according to the sign (321) of the block to which they belong, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or a residual signal) to generate output sample information.

[0051] The output samples of the aggregator (355) may undergo various loop filtering techniques in the loop filter unit (356) that includes several types of loop filters. The output of the loop filter unit (356) may be a sample stream that may be output to the rendering device (312) and stored in the reference picture memory (357) for future inter-picture prediction.

[0052] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., transmission circuitry). The video encoder (403) may be used in place of Figure 4 the video encoder (403) in the example of

[0053] A video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Performing an appropriate encoding speed constitutes a function of a rate controller (450). In some embodiments, the controller (450) may be functionally coupled to other functional units as described below and control these functional units. The parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization technique,...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0054] In some example embodiments, the video encoder (403) may be configured to operate in a decoding loop. The decoding loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols in a manner similar to the way a (remote) decoder would create sample data to create sample data, although the embedded decoder 433 processes the encoded video stream of the source encoder 430 without performing entropy encoding (because in the video compression techniques contemplated in the disclosed subject matter, any compression between symbols and the encoded video bitstream in entropy encoding can be lossless). At this point, it can be observed that any decoder technique other than parsing / entropy decoding that may only exist in the decoder may also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations related to the decoding part of the encoder. Thus, the description of the encoder technology can be simplified because the encoder technology is reciprocal to the fully described decoder technology. A more detailed description of the encoder is provided only in certain areas or aspects below.

[0055] In some example implementations, during operation, the source encoder (430) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures of the video sequence designated as "reference pictures".

[0056] The local video decoder (433) may decode the encoded video data of a picture that may be designated as a reference picture. The local video decoder (433) duplicates the decoding process that may be performed by the video decoder on the reference picture, and may cause the reconstructed reference picture to be stored in the reference picture cache (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference picture, which has the same content (no transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0057] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) may search in the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata such as a reference picture motion vector, block shape, etc. that may be used as an appropriate prediction reference for the new picture.

[0058] The controller (450) may manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0059] The outputs of all the above-mentioned functional units may undergo entropy encoding in the entropy encoder (445). The transmitter (440) may buffer the encoded video sequence created by the entropy encoder (445) in preparation for transmission via the communication channel (460), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) may merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0060] The controller (450) may manage the operations of the video encoder (403). During encoding, the controller (450) may assign a specific encoded picture type to each encoded picture, which may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may typically be assigned to one of the following picture types: an intra picture (I picture), a predictive picture (P picture), a bi-predictive picture (B picture), a multi-predictive picture. The source picture may typically be spatially subdivided into a plurality of sample coding blocks, as described in further detail below.

[0061] Figure 5A diagram showing a video encoder (503) according to another example embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence, and encode the processing block into an encoded picture that is part of an encoded video sequence. The example video encoder (503) can be used in place of Figure 4 the video encoder (403) in the example.

[0062] For example, the video encoder (503) receives a matrix of sample values of the processing block. The video encoder (503) then uses, for example, Rate-Distortion Optimization (RDO) to determine whether to best encode the processing block using an intra mode, an inter mode, or a bi-prediction mode.

[0063] In Figure 5 the example, the video encoder (503) includes an inter encoder (530), an intra encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the example arrangement in Figure 5 .

[0064] The inter encoder (530) is configured to: receive samples of a current block (e.g., the processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order), generate inter prediction information (e.g., a description of redundant information according to inter coding techniques, a motion vector, merge mode information); and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique.

[0065] The intra encoder (522) is configured to: receive samples of a current block (e.g., the processing block), compare the block with already encoded blocks in the same picture, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques).

[0066] The general controller (521) can be configured to determine general control data, and control other components of the video encoder (503) based on the general control data to, for example, determine the prediction mode of a block, and provide a control signal to the switch (526) based on the prediction mode.

[0067] The residual calculator (523) can be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) can be configured to encode the residual data to generate transform coefficients. Then, the transform coefficients undergo quantization processing to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) can be configured to format the bitstream to include the encoded blocks and perform entropy encoding.

[0068] Figure 6 FIG. shows an example video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (610) can be used in place of Figure 4 the video decoder (410) in the example of.

[0069] In Figure 6 the example of, the video decoder (610) includes an entropy decoder (671), an inter decoder (680), a residual decoder (673), a reconstruction module (674), and an intra decoder (672) coupled together as shown in the example arrangement of Figure 6 .

[0070] The entropy decoder (671) can be configured to reconstruct certain symbols according to the encoded picture, and the symbols represent the syntax elements that make up the encoded picture. The inter decoder (680) can be configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information. The intra decoder (672) can be configured to receive intra prediction information and generate a prediction result based on the intra prediction information. The residual decoder (673) can be configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) can be configured to combine the residual output by the residual decoder (673) with the prediction result (output by the inter prediction module or the intra prediction module as appropriate) in the spatial domain to form a reconstructed block, and the reconstructed block forms part of the reconstructed picture that is part of the reconstructed video.

[0071] Note that any suitable technology may be used to implement video encoders (203), (403), and (503) and video decoders (210), (310), and (610). In some example embodiments, one or more integrated circuits may be used to implement video encoders (203), (403), and (503) and video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions may be used to implement video encoders (203), (403), and (503) and video decoders (210), (310), and (610).

[0072] Turning to block partitioning for encoding and decoding, the general partitioning may start from a basic block and may follow a predefined set of rules, a particular pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning may be hierarchical and recursive. After splitting or partitioning the basic block following any of the example partitioning processes described below or other processes or a combination thereof, a final set of partitions or coding blocks may be obtained. Each of these partitions may be at one of various partitioning levels in the partitioning hierarchy and may have various shapes. Each of the partitions may be referred to as a Coding Block (CB). For the various example partitioning implementations described further below, each resulting CB may have any allowed size and partitioning level. Such partitions are called coding blocks because they may form units for which some basic encoding / decoding decisions may be made, encoding / decoding parameters may be optimized and determined, and encoding / decoding parameters may be signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. Coding blocks may be luminance coding blocks or chrominance coding blocks. The CB tree structure for each color may be referred to as a Coding Block Tree (CBT). The coding blocks for all color channels may be collectively referred to as Coding Units (CUs). The hierarchical structures for all color channels may be collectively referred to as Coding Tree Units (CTUs). The partitioning patterns or structures for the various color channels in a CTU may be the same or may be different.

[0073] In some implementations, the partitioning tree scheme or structure for the luminance channel and the chrominance channel may not need to be the same. In other words, the luminance channel and the chrominance channel may have separate coding tree structures or patterns. Additionally, whether the luminance channel and the chrominance channel use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used may depend on whether the slice being coded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel may have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luminance channel and the chrominance channel may share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partitioning tree structure, and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0074] Figure 7 An example predefined 10-way partitioning structure / pattern that allows recursive partitioning to form a partitioning tree is shown. The root block can start at a predefined level (e.g., starting from a base block at the 128×128 level or the 64×64 level). Figure 7 Example partitioning structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example implementations, Figure 7 none of the rectangular partitions are allowed to be further subdivided. The coding tree depth can be further defined to indicate the split depth from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after splitting the root block further once according to Figure 7 the coding tree depth increases by 1. In some implementations, only the full square partitions in 710 are allowed to follow the Figure 7 pattern to recursively partition to the next level of the partitioning tree.

[0075] In some other example implementations of coding block partitioning, a quadtree structure can be used. Such quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether a base block or an intermediate block or partition is further quadtree split can be adapted to various local characteristics of the base block or intermediate block / partition.

[0076] In yet some other examples, a ternary partitioning scheme can be used to partition a base block or any intermediate block, as Figure 8 shown. The ternary pattern can be implemented vertically as shown in 802, or horizontally as shown in 804. Although the example split ratio in Figure 8 is shown as 1:2:1, other ratios can also be predefined. In some implementations, two or more different ratios can be predefined. In some implementations, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transforms.

[0077] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the quadtree partitioning scheme and the binary partitioning scheme described above can be combined to partition basic blocks into a Quadtree-Binary-Tree (QTBT) structure. In such a scheme, according to a set of predefined conditions (if specified), a basic block or an intermediate block / partition can be split either quadtreesplit or binarysplit. In Figure 9 a specific example is shown, where a basic block is first quadtreesplit into four partitions, as shown by 902, 904, 906, and 908. Thereafter, each of the resulting partitions is either quadtreesplit into four additional partitions (such as 908) at the next level, or binarysplit into two additional partitions (horizontally or vertically, such as 902 or 906, both being symmetric), or not split (such as 904). For square-shaped partitions, binarysplit or quadtreesplit can be recursively allowed, as shown by the overall example partitioning pattern of 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtreesplits and dashed lines represent binarysplits. A flag can be used for each binarysplit node (non-leaf binary partition) to indicate whether the binarysplit is horizontal or vertical. For example, as shown in 920 and consistent with the partitioning structure of 910, the flag "0" can represent a horizontal binarysplit, and the flag "1" can represent a vertical binarysplit. For quadtreesplit partitions, there is no need to indicate the split type, since a quadtreesplit always splits a block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some implementations, the flag "1" can represent a horizontal binarysplit, and the flag "0" can represent a vertical binarysplit.

[0078] In some example implementations of QTBT, the quadtree and binary split rule sets can be represented by the following predefined parameters and their associated corresponding functions: – CTU size: the size of the root node of the quadtree (the size of the basic block) – MinQTSize: the minimum allowed quadtree leaf node size – MaxBTSize: the maximum allowed binary tree root node size – MaxBTDepth: the maximum allowed binary tree depth – MinBTSize: the minimum allowed binary tree leaf node size In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128×128 luminance samples together with two corresponding 64×64 chrominance sample blocks (when considering and using example chrominance subsampling), the MinQTSize can be set to 16×16, the MaxBTSize can be set to 64×64, the MinBTSize (for both width and height) can be set to 4×4, and the MaxBTDepth can be set to 4. The quadtree partitioning can first be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from the minimum size allowed for them, which is 16×16 (i.e., MinQTSize), to 128×128 (i.e., CTU size). If a node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), the node will not be split first by the binary tree. Otherwise, nodes that do not exceed MaxBTSize can be partitioned by the binary tree. In Figure 9 the example of Figure 9 , the basic block is 128×128. According to a predefined rule set, only the basic block can be split by the quadtree. The partitioning depth of the basic block is 0. Each of the four resulting partitions is 64×64 - not exceeding MaxBTSize - and can be further split by the quadtree or the binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of a binary tree node equals MinBTSize, further vertical splitting is not considered.

[0079] In some example implementations, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure for luminance and chrominance or separate QTBT structures. For example, for P slices and B slices, the luminance CTB and chrominance CTB in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be partitioned into CBs by one QTBT structure, and the chrominance CTB can be partitioned into chrominance CBs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can consist of coded blocks of the luminance component or coded blocks of two chrominance components, and a CU in a P slice or B slice can consist of coded blocks of all three color components.

[0080] The above various CB partitioning schemes and the further partitioning of CBs into PBs can be combined in any way. The following specific implementations are provided as non - restrictive examples.

[0081] Partitioning video frames into partitions of various levels is applicable to still images or images in a non-video background to some extent. Other partitioning methods can be applied to still images (referred to as "images"). The term "block" is used below for the purpose of representing encoding / decoding units in low dynamic range (LDR) encoding / decoding implementations. A block can be one or more of the above partitions or one or more partitions determined using any other partitioning scheme. For example, a block can include various color components, such as a luminance component and a chrominance component, or RGB components, etc. The various implementations described below for LDR are applicable to blocks in video frames or still images. The implementations proposed below can be used alone or in any order of combination. In addition, each of the implementations in the encoder, decoder, and bitstream can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors can execute a program stored in a memory or a non-transitory computer-readable medium to implement the various implementations below. Alternatively, dedicated circuitry can be configured to implement such an encoder-decoder.

[0082] By way of illustration of LDR, for machine consumption in computer vision, automation, and other autonomous applications, and for some other scenarios, a lower bit depth of a video or an image may be sufficient. For these applications, an original bit depth of, for example, 8 bits per pixel for each color component of the video or image is more than sufficient. For these applications, before compressing or encoding the original data into a bitstream, the original data with a higher bit depth can be reduced to a lower bit depth. Any compression and encoding scheme can be used to compress the video / image data with the reduced bit depth. The decoder can be configured to reconstruct video / image samples with the reduced bit depth. The decoder can also implement a bit depth up-conversion or recovery process to generate decoded samples with, for example, the original bit depth. The up-conversion / recovery step can be optional. Due to the nature and use of the image / video in these applications, information loss (not recovered by the decoder) during the mapping process to LDR by the encoder may not have a negative impact on these applications.

[0083] In the following context, the reduced bit depth can be referred to as the lower dynamic range (LDR) of a video or an image. Thus, LDR can represent a range of sample values that is less than the dynamic range of the original sequence of samples. For example, LDR can have 2 N1A number of different sample values, where N1 is an integer value less than the bit depth N2 of the original sequence. For example, if the original sequence has a bit depth N2 = 8, then N1 can be any integer value less than 8, such as 1, 2, 3, 4, 5, 6, or 7. Thus, N1 represents the number of bits required to represent all sample values associated with the LDR. As another example, if the original sequence has a bit depth N2 = 10, then in addition to a 1, 2, 3, 4, 5, 6, or 7-bit dynamic range, a 9-bit or 8-bit dynamic range will also be considered an LDR.

[0084] In some example implementations, the encoder can first determine whether to apply LDR to a block of an input video frame or image (for simplicity, the term "image" is used below to represent a video frame or a still image). If LDR is to be applied to the block, the encoder can further determine the LDR to be applied. For example, the encoder can determine the original bit depth N2 and the reduced bit depth N1 for the LDR. The reduced bit depth N1 will provide the value range of the samples with the bit depth reduction before compression. Then, the encoder can perform an LDR mapping process to map the original input samples (having a value range determined by the original bit depth N2) to the mapped sample values in the reduced value range (LDR) characterized by the bit depth N1. Then, the mapped sample values of the block can be encoded by any of the encoding processes described above.

[0085] On the other hand, the decoder can first receive the bitstream and determine whether LDR has been applied to the block. The decoder can also determine the LDR applied to the block. For example, the decoder can determine the above bit depths N1 and N2 based on the bitstream or based on other information. Then, the decoder can decode the block according to the bitstream to generate reconstructed samples with an LDR bit depth of N1 for the block. Then, the decoder can perform a mapping of the reconstructed samples to the mapped samples with the original bit depth of N2 for the block.

[0086] In some example implementations, once the encoder determines to apply LDR, it can explicitly signal the syntax to indicate whether the current image or video has been encoded using LDR. Alternatively, an implicit indication of whether LDR has been used to encode the current image or video can be derived from the bitstream. From the decoder's perspective, if there is an explicit or implicit indication that LDR has been used for the block, the decoder can perform one or more operations related to LDR. Otherwise, the decoder can decode the block without performing any LDR operations.

[0087] In some example implementations, a set of low dynamic range (LDR) is predefined. These predefined LDRs can be identified by LDR indices. These LDR indices can be predefined. An LDR for encoding a block can be selected by an encoder, and the index of the selected LDR can be signaled in the bitstream. Alternatively, a selection bitmap or other type of flag can be used in the bitstream to indicate the LDR selected from a set of predefined LDRs for encoding a block.

[0088] In some example implementations, signaling indicating whether to apply LDR can be signaled first. If such signaling indicates applying LDR, additional signaling can be included in the bitstream to indicate which LDR from a set of predefined LDRs is applied to the block. In some example implementations, the predefined LDR indices can include a value for indicating not to apply LDR, and can include other values for indicating which LDR in a set of LDRs.

[0089] In some example implementations, the number of sample values supported by the low dynamic range can always be a power of 2, i.e., 2 N , as N1 described above. For an 8-bit original bit depth, example values of N can include but are not limited to 7 (128 sample values), 6 (64 sample values), 5 (32 sample values), or 4 (16 sample values). Thus, the allowed values of N form a set of possible LDRs.

[0090] In some of the above example implementations, the encoder can select a specific value of N for the LDR, and the selected dynamic range can be signaled by the numerical value 7 - N, which is a non-negative integer value. In this way, the higher the N for the selected LDR, the smaller the value signaled to indicate the selected LDR. Considering that LDR often only involves a small reduction in bit depth when adopted, this can help improve the encoding efficiency of such signaling in the bitstream.

[0091] In some example implementations, the above signaling regarding whether to perform LDR and / or which LDR to select can be provided by one or more high-level syntaxes, such as high-level syntax at a sequence header, picture header, sub-picture header, frame header, slice header, or tile header. Examples of the sequence header include but are not limited to a Sequence Parameter Set (SPS). Examples of the picture header include but are not limited to a Picture Parameter Set (PPS). Thus, LDR signaling can be applicable at different levels. For example, if such LDR signaling is provided in the bitstream at the picture level, LDR is applied to all blocks in the picture, or the selected LDR is applied to all LDR blocks in the picture.

[0092] In some example implementations, the signaling above regarding whether to perform LDR and / or which LDR to select can be provided by one or more block-level grammars in the bitstream. Examples of block-level grammars include, but are not limited to, the maximum coded block-level grammar (e.g., CTU or superblock-level grammar), coded block-level grammar, prediction block-level grammar, transform block-level grammar, or predefined fixed block size (e.g., 512×512, 256×256, 128×128, 64×64) level grammar.

[0093] In some example implementations, the signaling above regarding whether to perform LDR and / or which LDR to select can be provided in one or more supplementary enhancement information (SEI) messages, which can contain additional data inserted into the bitstream to convey additional information and can be received in exact synchronization with the associated image and video content.

[0094] In some example implementations, the application and / or selection of low dynamic range is signaled conditionally. For example, whether to signal the grammar related to low dynamic range can depend on the value of another grammar in the bitstream. For example, the grammar related to low dynamic range can be signaled only if the bitstream is intended for at least machine vision or some other related application. An indication of such an intention for machine vision can be specified by another grammar in the bitstream (e.g., included in one or more SEI messages). As another example, the signaling related to low dynamic range can be included (and / or applied) only if one or more quantization parameters of a block are greater than (or alternatively, less than) a predefined quantization parameter value. As another example, the signaling related to low dynamic range can be included (and / or applied) only if the picture resolution associated with the block is greater than (or alternatively, less than) a predefined picture resolution value. As yet another example, the signaling related to low dynamic range can be included (and / or applied) only if the frame rate is greater than (or less than) a predefined value.

[0095] In some example implementations, different low dynamic ranges are signaled for different color components. Alternatively, low dynamic range is signaled only for a specific color component. In one example, low dynamic range is signaled only for the luminance component.

[0096] In some example implementations, the selected LDR can be implicitly derived based on the encoded information, which includes but is not limited to the grammar value indicating whether the bitstream is intended for at least machine tasks Figure 10FIG. 1000 shows an example logical flow for a decoding method according to the above implementation. The logical flow 1000 starts at S1001. At S1010, a bitstream is received. At S1020, it is determined, based on the bitstream, that a low dynamic range (LDR) is applied to a block. At S1030, in the case where the low dynamic range is applied to the block: determine an LDR selected from a set of LDRs for encoding the block; decode the block to generate reconstructed LDR samples of the block; and generate reconstructed samples from the reconstructed LDR samples according to the selected LDR. The logical flow 1000 ends at S1099.

[0097] Figure 11 FIG. 1100 shows an example logical flow for an encoding method according to the above implementation. The logical flow 1100 starts at S1101. At S1110, it is determined that a block is to be encoded by applying an LDR. At S1120, in the case where it is determined that a low dynamic range is to be applied to the block: select an LDR from a set of LDRs for encoding the block; process the block using the selected LDR to generate encoded LDR samples of the block in the bitstream; and signal, either explicitly or implicitly, in the bitstream whether the LDR is applied to the block and the selected LDR. The logical flow 1100 stops at S1199.

[0098] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 12 FIG. 1200 shows a computer system (1200) suitable for implementing certain embodiments of the disclosed subject matter.

[0099] The computer software can be encoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0100] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0101] Figure 12The components shown for the computer system (1200) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependencies or requirements related to any of the components or combinations thereof shown in the exemplary embodiments of the computer system (1200).

[0102] The computer system (1200) may include certain human-machine interface input devices. The human-machine interface input devices may include one or more of the following (only one of each is depicted): keyboard (1201), mouse (1202), touchpad (1203), touch screen (1210), data glove (not shown), joystick (1205), microphone (1206), scanner (1207), camera device (1208).

[0103] The computer system (1200) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, haptic output, sound, light, and smell / taste. Such human-machine interface output devices may include: haptic output devices (e.g., haptic feedback through the touch screen (1210), data glove (not shown), or joystick (1205), but there may also be haptic feedback devices that do not serve as input devices); audio output devices (e.g., speakers (1209), headphones (not depicted)); visual output devices (e.g., screen (1210), including CRT (Cathode Ray Tube) screens, LCD (Liquid Crystal Display) screens, plasma screens, OLED (Organic Light-Emitting Diode) screens, each screen having or not having touch screen input capabilities, each screen having or not having haptic feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted), holographic displays, and fog machines (not depicted)) and printers (not depicted).

[0104] The computer system (1200) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory) / RW (1220) with media such as CD / DVD (1221), thumb drives (1222), removable hard disk drives or solid state drives (1223), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit) / PLD (Programable Logic Device) such as security dongles (not depicted), etc.

[0105] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.

[0106] The computer system (1200) may also include an interface (1254) to one or more communication networks (1255). The network can be, for example, wireless, wired, optical. The network can also be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet, wireless LAN (Local Area Network); cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network - BUS), etc.

[0107] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1240) of the computer system (1200).

[0108] The core (1240) may include one or more central processing units (CPUs) (1241), a graphics processing unit (GPU) (1242), a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) (1243), a hardware accelerator for certain tasks (1244), a graphics adapter (1250), etc. These devices, together with a read-only memory (ROM) (1245), a random access memory (1246), an internal mass storage device (1247) such as an internal non-user-accessible hard disk drive, a solid-state drive (SSD), etc., may be connected via a system bus (1248). In some computer systems, the system bus (1248) may be accessible in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus (1248) of the core or may be attached to the system bus (1248) of the core via a peripheral bus (1249). In an example, a screen (1210) may be connected to the graphics adapter (1250). The architecture of the peripheral bus includes PCI (Peripheral Component Interconnect), USB (Universal Serial Bus), etc.

[0109] Computer-readable media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code specially designed and constructed for the purposes of this disclosure, or they may be of the type well-known and available to those of skill in the computer software art.

[0110] Although the present disclosure has described several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to envision many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for decoding a block in a bitstream of a video or an image, comprising: Receiving the bitstream; Determining, based on the bitstream, that a low dynamic range (LDR) is applied to the block; And In the case where the low dynamic range is applied to the block: Determining an LDR selected from a set of LDRs for encoding the block; Decoding the block to generate reconstructed LDR samples of the block; And Generating reconstructed samples from the reconstructed LDR samples according to the selected LDR.

2. The method according to claim 1, wherein The set of LDRs is predefined, and the LDR index is explicitly signaled by an LDR syntax element in the bitstream or implicitly derived from the bitstream.

3. The method according to claim 2, wherein, Each LDR in the set of LDRs is predefined as a sample value range 2 N , where N is a non-negative integer and represents the bit depth of each LDR in the set of LDRs.

4. The method according to claim 3, wherein, The LDR index is explicitly signaled by a numerical value 7 - N.

5. The method according to claim 2, wherein The LDR index is explicitly signaled by the LDR syntax element as high - level syntax in a sequence header, a picture header, a sub - picture header, a frame header, a slice header, or a tile header.

6. The method according to claim 2, wherein, The LDR index is explicitly signaled by the LDR syntax element as one or more block - level syntaxes.

7. The method according to claim 2, wherein Signaling the LDR index in one or more supplementary enhancement information (SEI) messages, each SEI message containing additional data inserted into the bitstream to convey additional information.

8. The method according to any one of claims 1 to 7, wherein Applying the LDR or the LDR signaling indicating the LDR applied to the block depends at least on the value of another syntax element in the bitstream.

9. The method according to claim 8, wherein Applying the LDR or having the LDR signaling in the bitstream only when the bitstream is intended for at least machine vision as indicated by the other syntax element.

10. The method according to claim 8, wherein Applying the LDR or having the LDR signaling in the bitstream only when the quantization parameter of the block is greater than a predefined quantization parameter threshold.

11. The method according to claim 8, wherein, Applying the LDR or having the LDR signaling in the bitstream only when the picture resolution associated with the block is greater than or less than a predefined resolution value.

12. The method according to claim 8, wherein, Applying the LDR or having the LDR signaling in the bitstream only when the frame rate associated with the block is greater than or less than a predefined frame rate value.

13. The method according to any one of claims 1 to 7, wherein Signaling the LDRs for different color components of the block separately in the bitstream.

14. The method according to any one of claims 1 to 7, wherein, Applying the LDR only to the luminance component of the block and signaling the LDR.

15. The method according to any one of claims 1 to 7, implicitly deriving the LDR selected for encoding the block based on the encoded information of the block.

16. A method for encoding a block in a bitstream of a video or an image, comprising: Determining to apply a low dynamic range (LDR) to the block; And In the case of determining to apply the low dynamic range to the block: Selecting an LDR from a set of LDRs for encoding the block; Processing the block using the selected LDR to generate encoded LDR samples of the block in the bitstream; And Explicitly or implicitly signaling in the bitstream whether the LDR is applied to the block and the selected LDR.

17. The method according to claim 16, wherein, The set of LDRs is predefined, and the LDR index is explicitly signaled by the LDR syntax element in the bitstream.

18. The method according to claim 17, wherein, Each LDR in the set of LDRs is predefined as a sample value range 2 N , where N is a non-negative integer and represents the bit depth of each LDR in the set of LDRs.

19. The method according to claim 18, wherein, The LDR index is explicitly signaled by the value 7-N in the bitstream.

20. A method for processing blocks of video or images, including converting the blocks into a bitstream, wherein, The bitstream includes: An encoded block generated by: Determining to apply a low dynamic range (LDR) to the block; and In the case of determining to apply the low dynamic range to the block: Selecting an LDR from a set of LDRs for encoding the block; Using the selected LDR to process the block to generate the encoded LDR samples of the block in the bitstream; and Explicitly or implicitly signaling in the bitstream whether the LDR is applied to the block and the selected LDR; and At least one syntax element for explicitly or implicitly indicating the LDR.