Encoder, decoder and corresponding method
By introducing multiple APS types and overlapping value spaces for APS IDs, the inefficiencies in signaling ALF and LMCS parameters are addressed, enhancing coding efficiency and reducing resource usage in video compression systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2026-04-13
AI Technical Summary
Existing video compression systems face inefficiencies in signaling Adaptive Loop Filter (ALF) and Chroma Mapping with Chroma Scaling (LMCS) parameters, leading to redundant coding and wastage of network, memory, and processing resources due to frequent changes across multiple pictures.
Implementing multiple types of Adaptive Parameter Sets (APS) such as ALF APS, Scaling List APS, and LMCS APS, each containing specific parameters, and using overlapping value spaces for APS IDs to reduce redundant coding and resource usage.
Significantly reduces redundant coding of LMCS parameters by encoding them only when necessary, thereby optimizing network, memory, and processing resources in both encoder and decoder.
Smart Images

Figure 0007844696000018 
Figure 0007844696000019 
Figure 0007844696000020
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to video coding, and more particularly, to efficient signaling of coding tool parameters used to compress video data in video coding.
Background Art
[0002] Even the amount of video data required to depict a relatively short video can be substantial, which can cause difficulties when the data is streamed or communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, the size of the video can be a problem when the video is stored on a memory device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable because network resources are limited and the demand for higher video quality is constantly increasing.
Summary of the Invention
[0003] In one embodiment, the Disclosure includes a method implemented in a decoder, comprising: the decoder's receiver receiving a bitstream containing an LMCS Adaptive Parameter Set (APS) containing chroma mapping (LMCS) parameters with chroma scaling associated with a coded slice; a processor determining that the LMCS APS is referenced in the data relating to the coded slice; the processor decoding the coded slice using the LMCS parameters from the LMCS APS; and the processor transferring the decoding result for display as part of the decoded video sequence. The APS is used to maintain data relating to multiple slices across multiple pictures. The Disclosure describes various modifications relating to different APS. In this example, the LMCS parameters are included in the LMCS APS. The LMCS / reshaper parameters may change approximately once per second. The video sequence may display 30 to 60 pictures per second. Therefore, the LMCS parameters do not need to change for 30 to 60 frames. Including LMCS parameters in LMCS APS instead of picture-level parameter sets significantly reduces the redundant coding of LMCS parameters (e.g., by 30-60 times). Slice headers and / or picture headers associated with slices can reference the associated LMCS APS. In this way, LMCS parameters are encoded only when the LMCS parameters for a slice change. Therefore, using LMCS APS to encode LMCS parameters improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0004] Optionally, in any of the above embodiments, another implementation of the embodiment provides that the bitstream further includes a slice header, the coded slice refers to the slice header, the slice header includes data relating to the coded slice, and refers to the LMCS APS.
[0005] Optionally, in any of the above embodiments, another implementation of the embodiment provides that the bitstream further includes an ALF APS including adaptive loop filter (ALF) parameters and a scaling list APS including APS parameters.
[0006] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters included in each APS.
[0007] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of each APS.
[0008] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS is identified by a combination of parameter type and APS ID.
[0009] Optionally, in any of the above embodiments, another implementation of the embodiment provides that the bitstream further includes a sequence parameter set (SPS) which includes a flag set to indicate that the LMCS is valid for an encoded video sequence containing a slice that has been coded, and the LMCS parameters from the LMCS APS are obtained based on the flag.
[0010] In one embodiment, the Disclosure includes a method implemented in an encoder, comprising: the encoder's processor determining LMCS parameters for application to a slice; the processor encoding the slice into a bitstream as a coded slice; the processor encoding the LMCS parameters into a bitstream in an LMCS APS; the processor encoding data relating to the coded slice referencing the LMCS APS into a bitstream; and memory coupled to the processor storing the bitstream for communication toward a decoder. The APS is used to maintain data relating to multiple slices across multiple pictures. The Disclosure describes various modifications relating to different APSs. In this example, the LMCS parameters are included in the LMCS APS. The LMCS / reshaper parameters may change approximately once per second. A video sequence may display 30 to 60 pictures per second. Therefore, the LMCS parameters do not need to change for 30 to 60 frames. Including LMCS parameters in LMCS APS instead of picture-level parameter sets significantly reduces the redundant coding of LMCS parameters (e.g., by 30-60 times). Slice headers and / or picture headers associated with slices can reference the relevant LMCS APS. In this way, LMCS parameters are encoded only when the LMCS parameters for a slice change. Therefore, using LMCS APS to encode LMCS parameters improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0011] Optionally, in any of the above embodiments, another implementation of the embodiment further includes encoding a slice header into a bitstream by the processor, wherein the coded slice refers to the slice header, the slice header contains data relating to the coded slice, and refers to the LMCS APS.
[0012] Optionally, in any of the above embodiments, another implementation of the embodiment further includes the processor encoding an ALF APS containing the ALF parameter and a scaling list APS containing the APS parameter into a bitstream.
[0013] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS includes an aps_params_type code set to a predefined value indicating the type of parameter included in each APS.
[0014] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS includes an APS ID selected from a predefined range, the predefined range being determined based on the parameter type of each APS.
[0015] Optionally, in any of the above embodiments, another implementation of the embodiment provides that each APS is identified by a combination of parameter type and APS ID.
[0016] Optionally, in any of the above embodiments, another implementation of the embodiment further includes encoding the SPS into a bitstream by the processor, wherein the SPS includes a flag set to indicate that the LMCS is valid for an encoded video sequence containing the coded slice.
[0017] In one embodiment, the disclosure includes a video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, receiver, memory, and transmitter are configured to perform the method described in any of the above embodiments.
[0018] In one embodiment, the Disclosure relates to a non-temporary computer-readable medium including a computer program product used by a video coding device, wherein the computer program product includes computer-executable instructions stored on the non-temporary computer-readable medium, which, when executed by a processor, causes the video coding device to perform the method described in any of the above embodiments.
[0019] In one embodiment, the disclosure includes a decoder comprising: receiving means for receiving a bitstream including an LMCS APS containing LMCS parameters related to a coded slice; determining means for determining that the LMCS APS is referenced in data related to the coded slice; decoding means for decoding a coded slice using the LMCS parameters from the LMCS APS; and transferring means for transferring the decoding result for display as part of a decoded video sequence.
[0020] Optionally, in any of the above embodiments, another implementation of the embodiment provides that the decoder is further configured to perform any of the methods of the above embodiments.
[0021] In one embodiment, the present disclosure includes determination means for determining LMCS parameters for application to slices, encoding means for encoding a slice as a coded slice into a bitstream, encoding the LMCS parameters into the bitstream in an LMCS Adaptive Parameter Set (APS), and encoding data related to the coded slice that refers to the LMCS APS into the bitstream, and storage means for storing the bitstream for communication towards a decoder.
[0022] Optionally, in any of the above aspects, another implementation aspect of the aspect provides that the encoder is further configured to execute any of the methods of the above aspects.
[0023] For clarity, any one of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0024] These and other features will be more clearly understood from the following detailed description, which is to be construed in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0025] To more fully understand the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0026] [Figure 1] It is a flowchart of an exemplary method for coding a video signal.
[0027] [Figure 2] It is a schematic diagram of an example of an exemplary video coding and decoding (codec) system for video coding.
[0028] [Figure 3] It is a schematic diagram showing an exemplary video encoder.
[0029] [Figure 4] It is a schematic diagram showing an exemplary video decoder.
[0030] [Figure 5] It is a schematic diagram showing an exemplary bitstream including multiple types of adaptive parameter sets including different types of coding tool parameters.
[0031] [Figure 6] It is a schematic diagram showing an exemplary mechanism for assigning APS identifiers (IDs) to different APS types on different value spaces.
[0032] [Figure 7] It is a schematic diagram of an exemplary video coding device.
[0033] [Figure 8] It is a flowchart of an exemplary method for encoding a video sequence into a bitstream by using multiple APS types.
[0034] [Figure 9] It is a flowchart of an exemplary method for decoding a video sequence from a bitstream by using multiple APS types.
[0035] [Figure 10] It is a schematic diagram of an exemplary system for coding a video sequence of an image into a bitstream by using multiple APS types.
Embodiments for Carrying Out the Invention
[0036] Firstly, while exemplary implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or existing. This disclosure is not to be limited in any way to the exemplary implementations, drawings and techniques shown below, including the exemplary designs and implementations illustrated and described herein, and may be modified within the scope of the appended claims, along with the entire scope of their equivalents.
[0037] The following abbreviations are used herein: Adaptive Loop Filter (ALF), Adaptive Parameter Set (APS), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coated Video Sequence (CVS), Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH), Intra-Random Access Point (IRAP), Joint Video Expert Team (JVET), Motion Constraint Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Low Byte Sequence Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).
[0038] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or eliminate data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or part of a video picture) may be partitioned into video blocks, which may also be called tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. In an intra-coded (I) slice of a picture, video blocks are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. In an inter-coding one-way prediction (P) or two-way prediction (B) slice of a picture, video blocks may be coded by using spatial prediction with respect to reference samples in adjacent blocks within the same picture, or by using temporal prediction with respect to reference samples in other reference pictures. A picture may be called a frame and / or image, and a reference picture may be called a reference frame and / or reference image. Spatial or temporal predictions result in predicted blocks representing image blocks. Residual data represents the pixel difference between the original image blocks and the predicted blocks. Thus, intercoded blocks are encoded according to a motion vector pointing to a block of reference samples that forms the predicted block and residual data indicating the difference between the coded block and the predicted block. Intracoded blocks are encoded according to the intracoded mode and residual data. For further compression, the residual data may be transformed from the pixel domain to the transformation domain. These result in residual transformation coefficients, which can be quantized. The quantized transformation coefficients may first be placed in a two-dimensional array. The quantized transformation coefficients may be scanned to generate a one-dimensional vector of transformation coefficients.Entropy coding can be applied to achieve even greater compression. Such video compression techniques are discussed in more detail below.
[0039] To ensure that encoded video is accurately decoded, video is encoded and decoded according to the corresponding video coding standard. Video coding standards include Advanced Video Coding (AVC), also known as International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding Plus Depth (MVC+D), and 3D AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The ITU-T and ISO / IEC Joint Video Experts Team (JVET) has begun development of a video coding standard called Versatile Video Coding (VVC). VVC is included in working drafts (WDs) including JVET-M1001-v5 and JVET-M1002-v1, which provide algorithm descriptions, encoder-side descriptions of the VVC WDs, and reference software.
[0040] Video sequences are coded using various coding tools. The encoder selects parameters for the coding tools with the aim of improving compression with minimal quality loss when the video sequence is decoded. Coding tools may relate to different parts of the video at different scopes. For example, some coding tools relate at the video sequence level, some at the picture level, and some at the slice level. APS can be used to signal information that may be shared across multiple pictures and / or slices. Specifically, APS can carry Adaptive Loop Filter (ALF) parameters. ALF information may not be suitable for signaling at the sequence level in the Sequence Parameter Set (SPS), the picture level in the Picture Parameter Set (PPS) or picture header, or the slice level in the tile group / slice header for various reasons.
[0041] If ALF information is signaled in the SPS, the encoder must generate a new SPS and a new IRAP picture every time the ALF information changes. IRAP pictures significantly reduce coding efficiency. Therefore, placing ALF information in the SPS is particularly problematic in low-latency application environments that do not frequently use IRAP pictures. Also, including ALF information in the SPS may disable out-of-band transmission of the SPS. Out-of-band transmission refers to the transmission of corresponding data in a transport data flow different from the video bitstream (e.g., in sample descriptions or sample entries of media files, in session description protocol (SDP) files, etc.). Signaling ALF information in the PPS can also be problematic for similar reasons. Specifically, including ALF information in the PPS may disable out-of-band transmission of the PPS. Signaling ALF information in the picture header can also be problematic. Picture headers are sometimes not used. Furthermore, ALF information may be applied to multiple pictures. Thus, signaling ALF information in the picture header causes redundant information transmission and therefore wastes bandwidth. Signaling ALF information in the tile group / slice header is also problematic because ALF information may apply to multiple pictures and therefore multiple slices / tile groups. Therefore, signaling ALF information in the slice / tile group header causes redundant information transmission and therefore wastes bandwidth.
[0042] Based on the above, APS may be used to signal the ALF parameter. However, a video coding system may use APS only to signal the ALF parameter. Exemplary APS syntax and semantics are as follows: [Table 1]
[0043] `adaption_parameter_set_id` provides an identifier for the APS for reference by other syntactic elements. The APS can be shared between pictures and can be different in different tile groups within a picture. `aps_extension_flag` is set to 0 to specify that the `aps_extension_data_flag` syntactic element does not exist in the APS RBSP syntactic structure. `aps_extension_flag` is set to 1 to specify that the `aps_extension_data_flag` syntactic element exists in the APS RBSP syntactic structure. `aps_extension_data_flag` may have any value. The presence and value of `aps_extension_data_flag` may not affect the decoder's conformance to the profile specified in the VVC. A VVC-compliant decoder may ignore all `aps_extension_data_flag` syntactic elements.
[0044] The following is an example of tile group header syntax related to the ALF parameter: [Table 2]
[0045] tile_group_alf_enabled_flag is set to equal to 1 to specify that adaptive loop filtering is enabled and may be applied to luma (Y), blue chroma (Cb), or red chroma (Cr) color components within a tile group. tile_group_alf_enabled_flag is set to equal to 0 to specify that adaptive loop filtering is disabled for all color components within a tile group. tile_group_aps_id specifies the adaptation_parameter_set_id of the APS referenced by the tile group. The TemporalId of an APS NAL unit having an adaptation_parameter_set_id equal to tile_group_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. If multiple APSs with the same adaptation_parameter_set_id value are referenced by multiple tile groups of the same picture, the multiple APSs with the same adaptation_parameter_set_id value may contain the same content.
[0046] Reshaper parameters are parameters used in the Adaptive Loop Reshaper video coding tool, also known as Chroma Mapping with Chroma Scaling (LMCS). Exemplary SPS reshaper syntax and semantics are as follows: [Table 3]
[0047] The sps_reshaper_enabled_flag is set to equal to 1 to specify that the reshaper is used in Coded Video Sequences (CVS). The sps_reshaper_enabled_flag is set to 0 to specify that the reshaper is not used in CVS.
[0048] The exemplary tile group header / slice header reshaper syntax and semantics are as follows: [Table 4]
[0049] The `tile_group_reshaper_model_present_flag` is set to equal to 1 to indicate that `tile_group_reshaper_model()` exists in the tile group header. The `tile_group_reshaper_model_present_flag` is set to equal to 0 to indicate that `tile_group_reshaper_model()` does not exist in the tile group header. If `tile_group_reshaper_model_present_flag` does not exist, the flag is assumed to be equal to 0. The `tile_group_reshaper_enabled_flag` is set to equal to 1 to indicate that the reshaper is enabled for the current tile group. The `tile_group_reshaper_enabled_flag` is set to 0 to indicate that the reshaper is not enabled for the current tile group. If `tile_group_resharper_enable_flag` does not exist, the flag is assumed to be equal to 0. The `tile_group_reshamer_chroma_residual_scale_flag` is set to equal to 1 to specify that chroma residual scaling is enabled for the current tile group. The `tile_group_reshaper_chroma_residual_scale_flag` is set to 0 to specify that chroma residual scaling is not enabled for the current tile group. If `tile_group_reshaper_chroma_residual_scale_flag` is not present, the flag is assumed to be equal to 0.
[0050] The exemplary tile group header / slice header reshaper model syntax and semantics are as follows: [Table 5]
[0051] `reshape_model_min_bin_idx` specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of `reshape_model_min_bin_idx` can be in the range of 0 to `MaxBinIdx` (including both ends). The value of `MaxBinIdx` may be equal to 15. `reshape_model_delta_max_bin_idx` specifies the maximum allowable bin (or piece) index `MaxBinIdx` minus the maximum bin index used in the reshaper configuration process. The value of `reshape_model_max_bin_idx` is set to be equal to the difference between `MaxBinIdx` and `reshape_model_delta_max_bin_idx`. `reshaper_model_bin_delta_abs_cw_prec_minus1` plus 1 specifies the number of bits used to represent the syntax `reshape_model_bin_delta_abs_CW[i]`. reshape_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value of the i-th bin.
[0052] The variable `reshaper_model_bin_delta_sign_CW_flag[i]` specifies the sign of `reshape_model_bin_delta_abs_CW[i]` as follows: If `reshape_model_bin_delta_sign_CW_flag[i]` is equal to 0, the corresponding variable `RspDeltaCW[i]` is positive. Otherwise (for example, if `reshape_model_bin_delta_sign_CW_flag[i]` is not equal to 0), the corresponding variable `RspDeltaCW[i]` is negative. If `reshape_model_bin_delta_sign_CW_flag[i]` does not exist, the flag is assumed to be equal to 0. The variable `RspDeltaCW[i]` is set to equal to (1 - 2 * `reshape_model_bin_delta_sign_CW[i]`) * `reshape_model_bin_delta_abs_CW[i]`.
[0053] The variable RspCW[i] is derived as follows: The variable OrgCW is set to equal to (1 << BitDepthY) / (MaxBinIdx + 1). If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, then RspCW[i] = OrgCW + RspDeltaCW[i]. Otherwise, RspCW[i] = 0. The value of RspCW[i] is assumed to be in the range of 32 to 2 * OrgCW-1 when the value of BitDepthY is equal to 10. The variable InputPivot[i] with i in the range of 0 to MaxBinIdx + 1 (including both ends) is derived as follows: InputPivot[i] = i * OrgCW. The variables ReshapePivot[i], which have i in the range of 0 to MaxBinIdx + 1 (including both ends), and the variables ScaleCoef[i] and InvScaleCoeff[i], which have i in the range of 0 to MaxBinIdx (including both ends), are derived as follows:
number
[0054] The variable ChromaScaleCoef[i], which has i in the range of 0 to MaxBinIdx (including both endpoints), is derived as follows:
number
[0055] The characteristics of reshaper parameters can be characterized as follows: The size of the set of reshaper parameters contained in the tile_group_reshamer_model() syntax structure is typically around 60 to 100 bits. The reshaper model is typically updated by the encoder about once per second and contains many frames. Furthermore, the parameters of the updated reshaper model are unlikely to be exactly the same as the parameters of previous instances of the reshaper model.
[0056] The aforementioned video coding system has certain problems. Firstly, such a system is configured solely to carry ALF parameters in APS. Furthermore, reshaper / LMCS parameters may be shared by multiple pictures and may contain many variations.
[0057] Disclosed herein are various mechanisms for modifying APS to support improved coding efficiency. In the first example, several types of APS are disclosed. Specifically, an APS of type ALF is called an ALF APS and may include ALF parameters. Furthermore, an APS of type Scaling List is called a Scaling List APS and may include Scaling List parameters. Additionally, an APS of type LMCS is called an LMCS APS and may include LMCS / Reshaper parameters. ALF APS, Scaling List APS, and LMCS APS may each be coded as a separate NAL type and therefore contained within different NAL units. Thus, changes to data (e.g., ALF parameters) in one type of APS do not result in redundant coding of other types of data (e.g., LMCS parameters) that are not changed. Therefore, providing multiple types of APS improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0058] In the second example, each APS contains an APS identifier (ID). Furthermore, each APS type contains a separate value space for the corresponding APS ID. Such value spaces can overlap. Thus, an APS of the first type (e.g., ALF APS) can contain the same APS ID as an APS of the second type (e.g., LMCS APS). This is achieved by identifying each APS by a combination of APS parameter type and APS ID. By allowing each APS type to contain a different value space, the codec does not need to check for ID conflicts across APS types. Furthermore, by allowing value spaces to overlap, the codec can avoid using larger ID values, resulting in bit savings. Thus, using separate overlapping value spaces for different types of APS improves coding efficiency and therefore reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0059] In the third example, the LMCS parameters are included in the LMCS APS. As mentioned above, the LMCS / reshaper parameters may change approximately once per second. A video sequence may display 30 to 60 pictures per second. Therefore, the LMCS parameters do not need to change for 30 to 60 frames. Including the LMCS parameters in the LMCS APS significantly reduces the redundant coding of the LMCS parameters. The slice header and / or the picture header associated with the slice can reference the associated LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters for a slice change. Therefore, using LMCS APS to encode LMCS parameters improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0060] Figure 1 is a flowchart of an exemplary operation method 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. A smaller file size allows the compressed video file to be sent to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is generally very similar to the encoding process, allowing the decoder to consistently reconstruct the video signal.
[0061] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. Alternatively, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may contain both audio and video components. The video component includes a series of image frames that give a visual impression of motion when viewed in sequence. The frames include pixels that, when expressed in terms of light, are called lumens (or lumens samples) as herein, and that, when expressed in terms of color, are called chromens (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0062] In step 103, the video is partitioned into blocks. Partitioning involves subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame may first be divided into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both lumens and chroma samples. A coding tree may be used to divide the CTUs into blocks, and then the blocks may be recursively subdivided until a configuration supporting further coding is achieved. For example, the lumens component of a frame may be subdivided until the individual blocks contain relatively uniform illumination values. Furthermore, the chroma component of a frame may be subdivided until the individual blocks contain relatively uniform color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0063] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, interpretation and / or intrapretation may be used. Interpretation is designed to take advantage of the fact that objects tend to appear in consecutive frames in a common scene. Therefore, blocks depicting objects in a reference frame do not need to be described repeatedly in adjacent frames. Specifically, objects such as tables may remain in the same position across multiple frames. Therefore, a table is described once, and adjacent frames can refer back to the reference frame. Pattern matching mechanisms can be used to match objects across multiple frames. Furthermore, moving objects may be represented across multiple frames, for example, due to the movement of the object or the movement of the camera. As a particular example, a video may show a car moving across the screen across multiple frames. Motion vectors may be used to describe such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of an object in a reference frame. Therefore, interpretation can encode an image block in the current frame as a set of motion vectors indicating the offset from the corresponding block in the reference frame.
[0064] Intra-prediction encodes blocks within a common frame. It leverages the fact that luma and chroma components tend to be concentrated within a frame. For example, green patches in a tree tend to be adjacent to similar green patches. Intra-prediction uses multi-directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional modes indicate that the current block is similar / identical to samples of adjacent blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on adjacent blocks at the edges of the row. Planar mode effectively shows smooth light / color transitions across rows / columns by using a relatively constant slope when changing values. DC mode is used for boundary smoothing and indicates that a block is similar / identical to the mean value related to samples of all adjacent blocks in the angular direction of the directional prediction mode. Thus, intra-predicted blocks can represent image blocks as various relational prediction mode values instead of actual values. Furthermore, the interpretation block can represent the image block as a motion vector value instead of the actual value. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in the residual block. Transformations may be applied to the residual block to further compress the file.
[0065] In step 107, various filtering techniques may be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in the generation of blocky images in the decoder. Furthermore, the block-based prediction scheme may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. The in-loop filtering scheme iteratively applies noise suppression filters, deblocking filters, adaptive loop filters, and sample-adaptive offset (SAO) filters to blocks / frames. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference blocks, making it less likely that artifacts will generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0066] Once the video signal is split, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream contains the data described above, as well as any signaling data desirable to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. The generation of the bitstream is an iterative process. Therefore, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across many frames and blocks. The order shown in Figure 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to a specific order.
[0067] The decoder receives the bitstream and, in step 111, begins the decoding process. Specifically, the decoder uses an entropy decoding scheme that converts the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the partitioning of the frame. The partitioning should match the result of block partitioning in step 103. Entropy coding / decoding, such as that used in step 111, is described below. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible options based on the spatial positioning of values in the input image. Signaling the strict choices may involve using a number of bins. As used herein, a bin is a binary value (e.g., a bit value that may vary depending on the context) treated as a variable. Entropy coding allows the encoder to discard any option that is obviously unfeasible in particular, leaving a set of acceptable options. Each acceptable option is then assigned a code word. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three or four options, etc.). The encoder then encodes the codeword for the selected options. This scheme reduces the size of the codeword, as it is only large enough to uniquely represent a selection from a small subset of acceptable options, as opposed to a codeword uniquely representing a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in a similar manner to the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0068] In step 113, the decoder performs block decoding. Specifically, the decoder uses the inverse transform to generate residual blocks. The decoder then uses the residual blocks and corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. The reconstructed image blocks are then positioned within the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding as described above.
[0069] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to that in step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters may be applied to the frames to remove blocking artifacts. Once the frames are filtered, the video signal may be output to a display for user viewing in step 117.
[0070] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functionality to support implementations of the operation method 100. The codec system 200 is generalized to depict components used in both the encoder and the decoder. The codec system 200 receives and partitions a video signal to obtain a partitioned video signal 201, as discussed with respect to steps 101 and 103 of the operation method 100. The codec system 200 then compresses the partitioned video signal 201 into a coded bitstream when acting as an encoder, as discussed with respect to steps 105, 107, and 109 of the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 of the operation method 100. The codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter-controlled analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context-adaptive binary arithmetic coding (CABAC) component 231. Such components are combined as shown. In Figure 2, black lines show the movement of the data being encoded / decoded, and dashed lines show the movement of the control data that controls the operation of other components. All components of the codec system 200 may reside within the encoder. The decoder may contain a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components will be described below.
[0071] The partitioned video signal 201 is a captured video sequence, which is partitioned into blocks of pixels by a coding tree. The coding tree employs various partitioning modes to subdivide blocks of pixels into smaller blocks. These blocks can then be subdivided into even smaller blocks. Blocks are sometimes called nodes on the coding tree. Larger parent nodes are subdivided into smaller child nodes. The number of times a node is subdivided is called the node / coding tree depth. In some cases, subdivided blocks may be contained within a coding unit (CU). For example, a CU may be a sub-part of a CTU containing a lumen block, a red difference chroma (Cr) block, and a blue difference chroma (Cr) block, along with the corresponding CU syntax instructions. The partitioning modes may include binary trees (BT), triple trees (TT), and quad trees (QT), which are used to partition a node into two, three, or four child nodes of different shapes, depending on the partitioning mode used. The partitioned video signal 201 is transferred to a general coder control component 211, a transformation scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0072] The general coder control component 211 is configured to make decisions related to coding the images of a video sequence into a bitstream, according to application constraints. For example, the general coder control component 211 manages the optimization of bit rate / bitstream size versus reconstruction quality. Such decisions may be made based on memory / bandwidth availability and image resolution requirements. The general coder control component 211 also manages buffer utilization in relation to the transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general coder control component 211 may dynamically increase the complexity of compression to improve resolution and bandwidth utilization, or decrease the complexity of compression to reduce resolution and bandwidth utilization. Thus, the general coder control component 211 controls other components of the codec system 200 to balance bit rate concerns with video signal reconstruction quality. The general coder control component 211 generates control data that controls the operation of other components. The control data is also transferred to the header formatting and CABAC component 231, which is encoded in a bitstream to signal parameters for decoding by the decoder.
[0073] The partitioned video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for interpretation. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform interpretation coding of the received video blocks for one or more blocks within one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes to select an appropriate coding mode for each block of video data, for example.
[0074] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are illustrated separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is the process of generating motion vectors, which estimate the motion of video blocks. The motion vectors may, for example, represent the displacement of a coded object relative to a predicted block. A predicted block is a block that is found to be a precise match to the coded block with respect to pixel difference. Predicted blocks are sometimes also called reference blocks. Such pixel difference can be determined by the sum of absolute difference (SAD), squared difference (SSD), or other difference metrics. HEVC uses several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be split into CTBs, which may then be split into CBs to be included in CUs. A CU may be coded as a prediction unit (PU) containing prediction data and / or a transformation unit (TU) containing transformation residual data of the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc., for the current block / frame and select the reference block, motion vector, etc., with the best rate distortion features. The best rate distortion features balance both the quality of video reconstruction (e.g., amount of data loss due to compression) and coding efficiency (e.g., size of the final encoding).
[0075] In some examples, the codec system 200 may calculate the values of sub-integer pixel positions of the reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate the values of 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation component 221 can perform a motion search on all pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision. The motion estimation component 221 calculates the motion vector of the PU for video blocks in the intercoded slice by comparing the PU position with the predicted block position of the reference picture. The motion estimation component 221 outputs the calculated motion vector for encoding as motion data to the header formatting and CABAC component 231 and outputs the motion to the motion compensation component 219.
[0076] Motion compensation performed by the motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by the motion estimation component 221. In some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the current video block's PU, the motion compensation component 219 can locate the predicted block pointed to by the motion vector. A residual video block is then formed by subtracting the pixel values of the predicted block from the pixel values of the currently coded video block, forming the pixel difference value. Generally, the motion estimation component 221 performs motion estimation for the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predicted block and residual block are transferred to the transformation scaling and quantization component 213.
[0077] The partitioned video signal 201 is also transmitted to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated, but are illustrated separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block for the block in the current frame, instead of the inter-prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames as described above. In particular, the intra-picture estimation component 215 determines the intra-prediction mode to use for encoding the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode from several tested intra-prediction modes to encode the current block. The selected intra-prediction mode is then transmitted to the header formatting and CABAC component 231 for encoding.
[0078] For example, the intra-picture estimation component 215 calculates rate distortion values for various tested intra-prediction modes using rate distortion analysis and selects the intra-prediction mode with the best rate distortion characteristics among the tested modes. Rate distortion analysis generally determines the bit rate (e.g., number of bits) used to generate the coded block, and the amount of distortion (or error) between the original uncoded block encoded to generate the coded block and the coded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various coded blocks and determines which intra-prediction mode exhibits the best rate distortion value for the block. Additionally, the intra-picture estimation component 215 may be configured to code depth blocks of the depth map using a depth modeling mode (DMM) based on rate distortion optimization (RDO).
[0079] The intra-picture prediction component 217, when implemented in an encoder, can generate residual blocks from prediction blocks based on the selected intra-prediction mode determined by the intra-picture prediction component 215, or, when implemented in a decoder, can read residual blocks from the bitstream. The residual blocks contain the difference in values between the prediction blocks and the original blocks, represented as a matrix. The residual blocks are then transferred to the transformation scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 can operate for both lumens and chromens components.
[0080] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the transformation scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transformation coefficients. The quantized transformation coefficients are then transferred to the header formatting and CABAC component 231 and encoded into a bitstream.
[0081] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain, which can then be used as a reference block, for example, to become a predicted block for another current block. The motion estimation component 221 and / or motion compensation component 219 can compute the reference block by converting the residual block back into a corresponding predicted block for use in motion estimation of subsequent blocks / frames. Filters are applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transform. Otherwise, such artifacts could lead to inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0082] The filter-controlled analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, the transformed residual blocks from the scaling and inverse transform component 229 may be combined with the corresponding predicted blocks from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image blocks. The filters can then be applied to the reconstructed image blocks. In some examples, the filters may be applied to residual blocks instead. Like the other components in Figure 2, the filter-controlled analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are drawn separately for conceptual purposes. Filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters to adjust how such filters are applied. The filter-controlled analysis component 227 analyzes the reconstructed reference blocks to determine where such filters should be applied and sets the corresponding parameters. Such data is transferred to the header formatting and CABAC component 231 as filter-controlled data for encoding. The in-loop filter component 225 applies such filters based on filter control data. Filters may include unblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied to the spatial / pixel domain (e.g., in a reconstructed pixel block) or the frequency domain, depending on the example.
[0083] When operating as an encoder, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them to the display. The decoded picture buffer component 223 can be any memory device capable of storing the prediction blocks, residual blocks, and / or reconstructed image blocks.
[0084] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to the decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Furthermore, prediction data, including intra-prediction and motion data, as well as residual data in the form of quantization conversion coefficient data, are all encoded into a bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of coding contexts for various blocks, a representation of the most likely intra-prediction mode, a representation of partition information, and so on. Such data may be encoded using entropy coding. For example, information may be encoded using context-adaptive variable-length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval partitioned entropy (PIPE) coding, or other entropy coding techniques. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.
[0085] Figure 3 is a block diagram illustrating an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 partitions the input video signal, yielding a partitioned video signal 301, which is substantially similar to the partitioned video signal 201. The partitioned video signal 301 is then compressed by the components of the encoder 300 and encoded into a bitstream.
[0086] Specifically, the partitioned video signal 301 is transferred to the intra-picture prediction component 317 for intra-prediction. The intra-picture prediction component 317 may be substantially the same as the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also transferred to the motion compensation component 321 for intra-prediction based on a reference block in the decoded picture buffer component 323. The motion compensation component 321 may be substantially the same as the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transformation and quantization component 313 for transformation and quantization of the residual blocks. The transformation and quantization component 313 may be substantially the same as the transformation scaling and quantization component 213. The transformed and quantized residual blocks and corresponding prediction blocks are transferred to the entropy coding component 331 for coding into a bitstream (along with the associated control data). The entropy coding component 331 may be substantially the same as the header formatting and CABAC component 231.
[0087] The transformed and quantized residual blocks and / or corresponding predicted blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into reference blocks used by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. The in-loop filters in the in-loop filter component 325 are also applied, as example, to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 may be substantially similar to the filter-controlled analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may contain multiple filters, as discussed with respect to the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 and used as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0088] Figure 4 is a block diagram illustrating an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the operation method 100. The decoder 400 receives a bitstream from, for example, the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to the end user.
[0089] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantization conversion coefficients from residual blocks. The quantized conversion coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0090] The reconstructed residual blocks and / or predicted blocks are transferred to the intra-picture prediction component 417 for reconstruction into image blocks based on intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses prediction mode to locate reference blocks in the frame and applies the residual blocks to the result to reconstruct the intra-predicted image blocks. The reconstructed intra-predicted image blocks and / or residual blocks and the corresponding intra-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks and / or predicted blocks, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are transferred to the motion compensation component 421 for interpretation. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or motion compensation component 219. Specifically, the motion compensation component 421 uses motion vectors from a reference block to generate a prediction block and applies a residual block to the result to reconstruct the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into frames via partition information. Such frames may also be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.
[0091] Figure 5 is a schematic diagram illustrating an exemplary bitstream 500 containing multiple types of APS with different types of coding tool parameters. For example, bitstream 500 may be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400. As another example, bitstream 500 may be generated by encoder in step 109 of method 100 for use by decoder in step 111.
[0092] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 511, multiple ALF APS 512, multiple scaling list APS 513, multiple LMCS APS 514, multiple slice headers 515, and image data 520. The SPS 510 contains sequence data common to all pictures in the video sequence contained in the bitstream 500. Such data may include picture size, bit depth, coding tool parameters, bit rate limits, etc. The PPS 511 contains parameters that apply to the entire picture. Thus, each picture in the video sequence may refer to a PPS 511. While each picture refers to a PPS 511, it should be noted that a single PPS 511 may, in some examples, contain data for multiple pictures. For example, multiple similar pictures may be coded according to similar parameters. In such cases, a single PPS 511 may contain data for such similar pictures. PPS511 can indicate the coding tools available for slicing in the corresponding picture, quantization parameters, offset, etc. The slice header 515 contains parameters specific to each slice within the picture. Therefore, in a video sequence, there may be one slice header 515 per slice. The slice header 515 may include slice type information, picture order count (POC), reference picture list, prediction weights, tile entry point, unblocking parameters, etc. Note that the slice header 515 may also be called a tile group header in some contexts.
[0093] An APS is a syntactic structure that contains syntactic elements to be applied to one or more pictures 521 and / or slices 523. In the illustrated example, the APS can be divided into several types. ALF APS512 is an APS of type ALF that includes ALF parameters. An ALF is an adaptive block-based filter that includes a transfer function controlled by variable parameters and uses feedback from a feedback loop to improve the transfer function. Furthermore, an ALF is used to correct coding artifacts (e.g., errors) that result from block-based coding. An adaptive filter is a linear filter that has a transfer function controller by variable parameters that can be controlled by an optimization algorithm such as an RDO process operating in the encoder. Thus, the ALF parameters included in ALF APS512 may include variable parameters selected by the encoder so that the filter removes block-based coding artifacts during decoding in the decoder.
[0094] The scaling list APS513 is an APS of type scaling list that includes scaling list parameters. As described above, the current block is coded according to inter-prediction or intra-prediction, which yields residuals. The residuals are the difference between the lumen and / or chroma values of the block and the values of the corresponding predicted block. A transformation is then applied to the residuals to convert them into transformation coefficients (smaller than the residual values). Encoding high-resolution and / or ultra-high-resolution content may result in increased residual data. A simple transformation process, when applied to such data, may result in significant quantization noise. Therefore, the scaling list parameters included in the scaling list APS513 may include weighting parameters that are applied to scale the transformation matrix and can account for acceptable levels of change in display resolution and / or quantization noise in the resulting decoded video image.
[0095] LMCS APS514 is an APS of type LMCS that includes LMCS parameters, also known as reshaper parameters. The human visual system is less capable of distinguishing differences in color (e.g., chromaticity) than differences in light (e.g., luminance). Therefore, some video systems use a chroma subsampling mechanism to compress video data by reducing the resolution of chroma values without adjusting the corresponding lumens. One concern with such a mechanism is that the interpolation involved may produce interpolated chroma values during decoding that are incompatible with the corresponding lumens in some places. This generates color artifacts in such places, which should be corrected by corresponding filters. This is complicated by the lumens mapping mechanism. Lumen mapping is the process of remapping coded lumens components across the dynamic range of the input lumens signal (e.g., according to a piecewise linear function). This compresses the lumens components. The LMCS algorithm scales the compressed chromens based on lumens mapping to remove artifacts associated with chroma subsampling. Therefore, the LMCS parameters included in LMCS APS514 indicate the chroma scaling used to describe luma mapping. The LMCS parameters are determined by the encoder and can be used by the decoder to filter out artifacts caused by chroma subsampling when luma mapping is used.
[0096] Image data 520 includes video data encoded according to interpretation and / or intrapretation, as well as corresponding transformation and quantization residual data. For example, a video sequence includes multiple pictures 521 coded as image data. One picture 521 is a single frame of the video sequence and is therefore generally displayed as a single unit when the video sequence is displayed. However, partial pictures may be displayed to implement certain techniques such as virtual reality, picture-in-picture, etc. Each of the multiple pictures 521 refers to a PPS 511. The multiple pictures 521 are divided into multiple slices 523. One slice 523 may be defined as a horizontal section of one picture 521. For example, one slice 523 may include a portion of the height of the picture 521 and the full width of the picture 521. In other cases, the picture 521 may be divided into columns and rows, and the slices 523 may be contained in the rectangular portion of the picture 521 produced by such columns and rows. In some systems, slice 523 is subdivided into tiles. In other systems, slice 523 is called a tile group containing tiles. Slice 523 and / or tile groups of tiles refer to slice header 515. Slice 523 is further divided into coding tree units (CTUs). CTUs are further divided into coding blocks based on the coding tree. Coding blocks can then be coded / decoded according to a prediction mechanism.
[0097] Picture 521 and / or slice 523 may directly or indirectly refer to ALF APS512, scaling list APS513, and / or LMCS APS514, which contain the relevant parameters. For example, slice 523 may refer to slice header 515. Furthermore, picture 521 may refer to the corresponding picture header. Slice header 515 and / or picture header may refer to ALF APS512, scaling list APS513, and / or LMCS APS514, which contain the parameters used when coding the relevant slice 523 and / or picture 521. In this way, the decoder can obtain coding tool parameters relevant to slice 523 and / or picture 521 according to the header references related to the corresponding slice 523 and / or picture 521.
[0098] The bitstream 500 is coded into video coding layer (VCL) NAL units 535 and non-VCL NAL units 531. A NAL unit is a coded data unit formed to a size that will be placed as the payload of a single packet for transmission over a network. The VCL NAL unit 535 is a NAL unit that contains coded video data. For example, each VCL NAL unit 535 may contain data, a CTU, and / or one slice 523 and / or a tile group of coding blocks. The non-VCL NAL unit 531 is a NAL unit that contains supporting syntax but does not contain coded video data. For example, the non-VCL NAL unit 531 may contain SPS 510, PPS 511, APS, slice header 515, etc. Thus, the decoder receives the bitstream 500 in discrete VCL NAL units 535 and non-VCL NAL units 531. The access unit is a group of VCL NAL units 535 and / or non-VCL NAL units 531 containing enough data to code a single picture 521.
[0099] In some examples, the ALF APS 512, scaling list APS 513, and LMCS APS 514 are each assigned to a separate non-VCL NAL unit type 531. In such cases, the ALF APS 512, scaling list APS 513, and LMCS APS 514 are contained within the ALF APS NAL unit 532, scaling list APS NAL unit 533, and LMCS APS NAL unit 534, respectively. Thus, the ALF APS NAL unit 532 contains ALF parameters that remain valid until another ALF APS NAL unit 532 is received. Furthermore, the scaling list APS NAL unit 533 contains scaling list parameters that remain valid until another scaling list APS NAL unit 533 is received. Additionally, the LMCS APS NAL unit 534 contains LMCS parameters that remain valid until another LMCS APS NAL unit 534 is received. In this way, it is not necessary to issue a new APS every time the APS parameters change. For example, changing an LMCS parameter results in an additional LMCS APS514, but not an additional ALF APS512 or scaling list APS513. Therefore, by dividing APS into different NAL unit types based on parameter type, redundant signaling of irrelevant parameters is avoided. Thus, dividing APS into different NAL unit types improves coding efficiency and, consequently, reduces the use of processor, memory, and / or network resources in the encoder and decoder.
[0100] Furthermore, slice 523 and / or picture 521 may directly or indirectly reference ALF APS 512, ALF APS NAL unit 532, scaling list APS 513, scaling list APS NAL unit 533, LMCS APS 514, and / or LMCS APS NAL unit 534, which contain coding tool parameters used to code slice 523 and / or picture 521. For example, each APS may include an APS ID 542 and a parameter type 541. The APS ID 542 is a value (e.g., a number) that identifies the corresponding APS. The APS ID 542 may include a predefined number of bits. Thus, the APS ID 542 may be incremented (e.g., by 1) according to a predefined order and may be reset to a minimum value (e.g., 0) when the order reaches the end of a predefined range. The parameter type 541 indicates the type of parameter included in the APS (e.g., ALF, scaling list, and / or LMCS). For example, parameter type 541 may include an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters included in each APS. Thus, parameter type 541 may be used to distinguish between ALF APS512, scaling list APS513, and LMCS APS514. In some examples, ALF APS512, scaling list APS513, and LMCS APS514 may each be uniquely identified by a combination of parameter type 541 and APS ID 542. For example, each APS type may include a separate value space for the corresponding APS ID 542. Thus, each APS type may include an APS ID 542 that increases sequentially based on previous APS of the same type. However, the APS ID 542 of a first APS type may not be related to the APS ID 542 of a previous APS of a different second APS type. Thus, the APS ID 542 of different APS types may include overlapping value spaces.For example, a first type of APS (e.g., ALF APS) may, in some cases, contain the same APS ID 542 as a second type of APS (e.g., LMCS APS). By allowing each APS type to contain a different value space, the codec does not need to check for APS ID 542 conflicts between APS types. Furthermore, by allowing overlapping value spaces, the codec can avoid using larger APS ID 542 values, which results in bit savings. Thus, using separate overlapping value spaces for multiple APS ID 542s of different APS types improves coding efficiency and therefore reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. As described above, the APS ID 542 may span a predefined range. In some examples, the predefined range of the APS ID 542 may vary depending on the APS type indicated by parameter type 541. This may allow different numbers of bits to be assigned to different APS types depending on how often different types of parameters change. For example, the APS ID 542 of ALF APS512 may be in the range of 0 to 7, the APS ID 542 of scaling list APS513 may be in the range of 0 to 7, and the APS ID 542 of LMCS APS514 may be in the range of 0 to 3.
[0101] In another example, the LMCS parameters are included in the LMCS APS514. Some systems include the LMCS parameters in the slice header 515. However, the LMCS / reshaper parameters may change approximately once per second. A video sequence may display 30 to 60 pictures 521 per second. Therefore, the LMCS parameters cannot change for 30 to 60 frames. Including the LMCS parameters in the LMCS APS514 significantly reduces redundant coding of the LMCS parameters. In some examples, the picture headers associated with slice header 515 and / or slice 523 and / or picture 521, respectively, can refer to the associated LMCS APS514. Slice 523 and / or picture 521 then refer to slice header 515 and / or picture header. This allows the decoder to obtain the LMCS parameters for the associated slice 523 and / or picture 521. In this way, LMCS parameters are encoded only when the LMCS parameters for slice 523 and / or picture 521 change. Therefore, using the LMCS APS514 to encode LMCS parameters increases coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Since LMCS is not used for all videos, the SPS510 may include an LMCS enable flag 543. The LMCS enable flag 543 may be set to indicate that LMCS is enabled for the encoded video sequence. Therefore, the decoder may obtain LMCS parameters from the LMCS APS514 based on the LMCS enable flag 543 if it is set (e.g., set to 1). Furthermore, the decoder does not have to attempt to obtain LMCS parameters if the LMCS enable flag 543 is not set (e.g., not set to 0).
[0102] Figure 6 is a schematic diagram illustrating an exemplary mechanism 600 for assigning APS ID 642 to different APS types in different value spaces. For example, mechanism 600 can be applied to bitstream 500 to assign APS ID 542 to ALF APS 512, scaling list APS 513, and / or LMCS APS 514. Furthermore, mechanism 600 may be applied to codec 200, encoder 300, and / or decoder 400 when coding video according to method 100.
[0103] Mechanism 600 assigns APS ID 642 to ALF APS 612, scaling list APS 613, and LMCS APS 614, which may be substantially similar to APS ID 542, ALF APS 512, scaling list APS 513, and LMCS APS 514, respectively. As described above, APS ID 642 may be assigned sequentially to several different value spaces, each value space being specific to an APS type. Furthermore, each value space may span different ranges specific to an APS type. In the example shown, the value space range for APS ID 642 in ALF APS 612 is 0 to 7 (e.g., 3 bits). Furthermore, the value space range for APS ID 642 in scaling list APS 613 is 0 to 7 (e.g., 3 bits). Also, the value space range for APS ID 642 in LMCS APS 611 is 0 to 3 (e.g., 2 bits). When APS ID 642 reaches the end of the value space range, the APS ID 642 of the next APS of the corresponding type returns to the beginning of the range (e.g., 0). If a new APS receives the same APS ID 642 as the previous APS of the same type, the previous APS is no longer active and can no longer be referenced. In this way, the value space range can be expanded to allow more APS of the same type to be actively referenced. Furthermore, the value space range can be narrowed to improve coding efficiency, but such narrowing also reduces the number of APS of the corresponding type that can remain active and available for reference at the same time.
[0104] In the example shown, each ALF APS612, scaling list APS613, and LMCS APS614 is referenced by a combination of APS ID642 and APS type. For example, ALF APS612, LMCS APS614, and scaling list APS613 each receive an APS ID642 of 0. When a new ALF APS612 is received, the APS ID642 is incremented from the value used for the previous ALF APS612. The same sequence applies to scaling list APS613 and LMCS APS614. Thus, each APS ID642 relates to the APS ID642 of a previous APS of the same type. However, the APS ID642 does not relate to the APS ID642 of a previous APS of a different type. In this example, the APS ID642 of ALF APS612 increments from 0 to 7, and then returns to 0 before continuing the increment. Furthermore, the APS ID 642 of scaling list APS613 increments from 0 to 7, then returns to 0 before continuing the increment. Similarly, the APS ID 642 of LMCS APS611 increments from 0 to 3, then returns to 0 before continuing the increment. As illustrated, such value spaces overlap, so different APS of different APS types can share the same APS ID 642 at the same point in the video sequence. It should also be noted that mechanism 600 draws only APS. In the bitstream, the drawn APS are scattered among other VCL and non-VCL NAL units such as SPS, PPS, slice headers, picture headers, and slices.
[0105] Thus, this disclosure includes improvements to the design of APS and several improvements for signaling reshaper / LMCS parameters. APS is designed for signaling information that can be shared across multiple pictures and can include many variations. Reshaper / LMCS parameters are used in the in-adaptive loop reshaper / LMCS video coding tool. The above mechanism can be implemented as follows. To solve the problems enumerated herein, this includes several embodiments that can be used individually and / or in combination.
[0106] The disclosed APS is modified so that multiple APS can be used to carry different types of parameters. Each APS NAL unit is used to carry only one type of parameter. As a result, two APS NAL units are encoded when two types of information are carried for a particular tile group / slice (e.g., one for each type of information). The APS may include an APS parameter type field in the APS syntax. An APS NAL unit may contain only parameters of the type indicated by the APS parameters type field.
[0107] In some examples, different types of APS parameters are indicated by different NAL unit types. For example, two different NAL unit types are used for APS. These two types of APS may be called ALF APS and Reshaper APS, respectively. In another example, the type of tool parameter carried in the APS NAL unit is specified in the NAL unit header. In VVC, the NAL unit header has reserved bits (e.g., 7 bits indicated as nuh_reserved_zero_7bit). In some examples, some of these bits (e.g., 3 of the 7 bits) may be used to specify the APS parameter type field. In some examples, certain types of APS may share the same value space for APS ID. On the other hand, different types of APS use different value spaces for APS ID. Therefore, two APS of different types may coexist and have the same APS ID value at the same moment. Furthermore, a combination of APS ID and APS parameter type may be used to distinguish APS from other APS.
[0108] The APS ID may be included in the tile group header syntax if the corresponding coding tool is enabled for the tile group. Otherwise, the APS ID of the corresponding type may not be included in the tile group header. For example, if ALF is enabled for the tile group, the APS ID of the ALF APS will be included in the tile group header. This can be achieved, for example, by setting the APS parameter type field to indicate the ALF type. Therefore, if ALF is not enabled for the tile group, the APS ID of the ALF APS will not be included in the tile group header. Furthermore, if the reshaper coding tool is enabled for the tile group, the APS ID of the reshaper APS will be included in the tile group header. This can be achieved, for example, by resetting the APS parameter type field to indicate the shaper type. Furthermore, if the reshaper coding tool is not enabled for the tile group, the APS ID of the reshaper APS may not be included in the tile group header.
[0109] In some cases, the presence of APS parameter type information in an APS may be conditional on the use of the coding tool associated with the parameter. If only one coding tool associated with one APS is valid for a single bitstream (e.g., LMCS, ALF, scaling list), then APS parameter type information may not be present and may be inferred instead. For example, if an APS can include parameters for both ALF and reshaper coding tools, but only ALF is valid (e.g., as specified by a flag in the SPS) and reshaper is not valid (e.g., as specified by a flag in the SPS), then the APS parameter type may not need to be signaled and may be inferred to be equal to the ALF parameter.
[0110] In another example, APS parameter type information may be inferred from the APS ID value. For example, a predefined range of APS ID values may be associated with the corresponding APS parameter type. This embodiment may be implemented as follows: Instead of assigning X bits to the APS ID signaling and Y bits to the APS parameter type signaling, X+Y bits may be assigned to the APS ID signaling. Then, different ranges of APS ID values may be specified to indicate different types of APS parameter types. For example, instead of using 5 bits for the APS ID signaling and 3 bits for the APS parameter type signaling, 8 bits may be assigned to the APS ID signaling (e.g., without increasing bit cost). The range of APS ID values from 0 to 63 indicates that the APS contains parameters for ALF, the range of ID values from 64 to 95 indicates that the APS contains parameters for reshaper, and 96 to 255 may be reserved for other parameter types such as scaling lists. In another example, an APS ID value range of 0-31 indicates that the APS contains parameters for ALF, an ID value range of 32-47 indicates that the APS contains parameters for reshaper, and 48-255 may be reserved for other parameter types such as scaling lists. The advantage of this approach is that the APS ID range can be assigned depending on how often the parameters of each tool are changed. For example, ALF parameters may be expected to change more frequently than reshaper parameters. In such cases, a larger APS ID range may be used to indicate that the APS contains ALF parameters.
[0111] In the first embodiment, one or more of the embodiments described above may be implemented as follows: An ALF APS may be defined as an APS having an aps_params_type equal to ALF_APS. A reshaper APS (or LMCS APS) may be defined as an APS having an aps_params_type equal to MAP_APS. An exemplary SPS syntax and semantics are as follows: [Table 6] The sps_reshaper_enabled_flag is set to equal to 1 to specify that the reshaper is used in Coded Video Sequences (CVS). The sps_reshaper_enabled_flag is set to 0 to specify that the reshaper is not used in CVS.
[0112] An example of APS syntax and semantics is as follows: [Table 7]
[0113] aps_params_type specifies the type of APS parameters transported in APS, as defined in the table below. [Table 8]
[0114] The following is an example of tile group header syntax and semantics: [Table 9]
[0115] `tile_group_alf_aps_id` specifies the `adaptation_parameter_set_id` of the ALF APS referenced by the tile group. The TemporalId of an ALF APS NAL unit having an `adaptation_parameter_set_id` equal to `tile_group_alf_aps_id` is assumed to be less than or equal to the TemporalId of the coded tile group NAL unit. If multiple ALF APS with the same `adaptation_parameter_set_id` are referenced by two or more tile groups of the same picture, then multiple ALF APS with the same `adaptation_parameter_set_id` are assumed to have the same content.
[0116] The `tile_group_reshaper_enabled_flag` is set to equal to 1 to indicate that the reshaper is enabled for the current tile group. The `tile_group_reshaper_enabled_flag` is set to 0 to indicate that the reshaper is not enabled for the current tile group. If `tile_group_resharper_enable_flag` does not exist, the flag is assumed to be equal to 0. The `tile_group_reshaper_aps_id` specifies the `adaptation_parameter_set_id` of the reshaper APS referenced by the tile group. The TemporalId of a reshaper APS NAL unit with an `adaptation_parameter_set_id` equal to `tile_group_reshaper_aps_id` is assumed to be less than or equal to the TemporalId of the coded tile group NAL unit. If multiple reshaper APS with the same adaptation_parameter_set_id value are referenced by two or more tile groups of the same picture, then the multiple reshaper APS with the same adaptation_parameter_set_id value shall be considered to have the same content. The tile_group_reshaper_chroma_residual_scale_flag is set to equal to 1 to specify that chroma residual scaling is enabled for the current tile group. The tile_group_reshaper_chroma_residual_scale_flag is set to equal to 0 to specify that chroma residual scaling is not enabled for the current tile group. If the tile_group_reshaper_chroma_residual_scale_flag does not exist, the flag is assumed to be equal to 0.
[0117] An example of reshaper data syntax and semantics is as follows: [Table 10]
[0118] `reshaper_model_min_bin_idx` specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of `reshaper_model_min_bin_idx` is in the range of 0 to MaxBinIdx (including both ends). The value of MaxBinIdx is equal to 15. `reshaper_model_delta_max_bin_idx` specifies the maximum allowable bin (or piece) index MaxBinIdx minus the maximum bin index used in the reshaper configuration process. The value of `reshaper_model_max_bin_idx` is set to equal to MaxBinIdx - `reshaper_model_delta_max_bin_idx`. `resharper_model_bin_delta_abs_cw_prec_minus1 + 1` specifies the number of bits used to represent the syntax element `resharper_model_bin_delta_abs_CW[i]`. `reshaper_model_bin_delta_abs_CW[i]` specifies the absolute delta codeword value of the i-th bin. The `reshaper_model_bin_delta_abs_CW[i]` syntactic element is represented by `reshaper_model_bin_delta_abs_cw_prec_minus1+1 bits`. `resharper_model_bin_delta_sign_CW_flag[i]` specifies the sign of `resharper_model_bin_delta_abs_CW[i]`.
[0119] In a second embodiment, one or more of the embodiments described above may be implemented as follows. Exemplary SPS syntax and semantics are as follows: [Table 11]
[0120] The variables ALFEnabled and ReshaperEnabled are set as follows: ALFEnabled = sps_alf_enabled_flag and ReshaperEnabled = sps_reshaper_enabled_flag.
[0121] An example of APS syntax and semantics is as follows: [Table 12]
[0122] aps_params_type specifies the type of APS parameters that are transported in APS, as shown in the table below. [Table 13]
[0123] If it does not exist, the value of aps_params_type is inferred as follows: If ALFEnabled, aps_params_type is set to equal to 0. Otherwise, aps_params_type is set to equal to 1.
[0124] In a second embodiment, one or more of the embodiments described above may be implemented as follows. Exemplary SPS syntax and semantics are as follows: [Table 14]
[0125] The `adaption_parameter_set_id` provides an identifier for the APS for reference by other syntactic elements. APSs can be shared between pictures and can be different for different tile groups within a picture. The values and descriptions of the variable `APSParamsType` are defined in the table below. [Table 15]
[0126] The following is an example of tile group header semantics: tile_group_alf_aps_id specifies the adaptation_parameter_set_id of the ALF APS referenced by the tile group. The TemporalId of an ALF APS NAL unit having an adaptation_parameter_set_id equal to tile_group_alf_aps_id must be less than or equal to the TemporalId of the coded tile group NAL unit. The value of tile_group_alf_aps_id must be in the range of 0 to 63 (inclusive).
[0127] If multiple ALF APS with the same adaptation_parameter_set_id are referenced by multiple tile groups of the same picture, then multiple ALF APS with the same adaptation_parameter_set_id are considered to have the same content. tile_group_reshaper_aps_id specifies the adaptation_parameter_set_id of the reshaper APS referenced by the tile group. The TemporalId of a reshaper APS NAL unit with an adaptation_parameter_set_id equal to tile_group_reshaper_aps_id is less than or equal to the TemporalId of the coded tile group NAL unit. The value of tile_group_reshaper_aps_id is in the range of 64 to 95 (including both ends). If multiple reshaper APS with the same adaptation_parameter_set_id are referenced by two or more tile groups of the same picture, then multiple reshaper APS with the same adaptation_parameter_set_id are considered to have the same content.
[0128] Figure 7 is a schematic diagram of an exemplary video coding device 700. The video coding device 700 is suitable for implementing the disclosed embodiments / models as described herein. The video coding device 700 includes a transceiver unit (Tx / Rx) 710, which includes a transmitter and / or receiver for communicating data upstream and / or downstream over a network, and a downstream port 720, and / or an upstream port 750. The video coding device 700 also includes a processor 730, which includes a logic unit and / or central processing unit (CPU) for processing data, and a memory 732 for storing data. The video coding device 700 may also include electrical, optical-electrical (OE) components, electrical-optical (EO) components, and / or wireless communication components coupled to the upstream port 750 and / or downstream port 720 for communicating data over an electrical, optical, or wireless communication network. The video coding device 700 may also include input and / or output (I / O) devices 760 for communicating data with a user. The I / O device 760 may include output devices such as a display for showing video data and speakers for outputting audio data. The I / O device 760 may also include input devices such as a keyboard, mouse, and trackball, and / or corresponding interfaces for interacting with such output devices.
[0129] The processor 730 is implemented by hardware and software. The processor 730 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 communicates with downstream port 720, Tx / Rx 710, upstream port 750, and memory 732. The processor 730 includes a coding module 714. The coding module 714 implements embodiments disclosed herein, such as methods 100, 800, and 900, which can use bitstream 500 and / or mechanism 600. The coding module 714 may also implement any other methods / mechanisms described herein. Furthermore, the coding module 714 may implement a codec system 200, an encoder 300, and / or a decoder 400. For example, coding module 714 can encode / decode pictures in a bitstream and encode / decode parameters associated with slices of pictures in multiple APSs. In some examples, different types of parameters may be coded into different types of APSs. Furthermore, different types of APSs may be included in different types of NAL units. Such APS types may include ALF APS, scaling list APS, and / or LMCS APS. Each APS may include an APS ID. APS IDs of different APS types may increment sequentially on different value spaces. Also, slices and / or pictures may refer to corresponding slice headers and / or picture headers. Such headers may then refer to APSs containing the associated coding tools. Such APSs may be uniquely referred to by their APS ID and APS type. Such examples reduce redundant signaling of coding tool parameters and / or reduce bit usage for identifiers.Therefore, the coding module 714 provides additional functionality and / or coding efficiency to the video coding device 700 when coding video data. Thus, the coding module 714 improves the functionality of the video coding device 700 and addresses problems inherent in video coding technology. Furthermore, the coding module 714 performs conversions of the video coding device 700 to different states. Alternatively, the coding module 714 is implemented as instructions stored in memory 732 and executed by the processor 730 (for example, as a computer program product stored on a non-temporary medium).
[0130] Memory 732 includes one or more memory types, such as disks, tape drives, solid-state drives, read-only memory (ROM), random-access memory, flash memory, tri-level associative memory (TCAM), and static random-access memory (SRAM). Memory 732 may also be used as an overflow data storage device, storing the program when such a program is selected for execution, and storing instructions and data read during program execution.
[0131] Figure 8 is a flowchart of an exemplary method 800 for encoding a video sequence into a bitstream such as bitstream 500 by using multiple APS types such as ALF APS512, scaling list APS513, and / or LMCS APS514. Method 800 may be used by an encoder such as a codec system 200, an encoder 300, and / or a video coding device 700 when performing method 100. Method 800 may also assign APS IDs to different types of APS by using different value spaces according to mechanism 600.
[0132] Method 800 may begin when an encoder receives a video sequence containing multiple pictures and determines, for example, based on user input, to encode the video sequence into a bitstream. The video sequence is partitioned into pictures / images / frames for further partitioning prior to encoding. In step 801, the encoder determines chroma mapping (LMCS) parameters with chroma scaling for application to slices. This may include using RDO operations to encode slices of pictures. For example, the encoder may repeatedly encode slices using different coding options / coding tools, decode the coded slices, and filter the decoded slices to increase the output quality of the slices. The encoder can then select an encoding option that provides the best balance between compression and output quality. Once an encoding is selected, the encoder can determine the LMCS parameters (and other optional filter parameters) used to filter the selected encoding. In step 803, the encoder can encode the slices into a bitstream based on the selected encoding.
[0133] In step 805, the LMCS parameters determined in step 801 are encoded into a bitstream in the LMCS APS. Furthermore, data relating to a slice that references the LMCS APS may also be encoded into a bitstream. For example, data relating to a slice may be a slice header and / or a picture header. In one example, the slice header may be encoded into a bitstream. A slice may reference the slice header, which contains data relating to the slice and references the LMCS APS. In another example, the picture header may be encoded into a bitstream. A picture containing a slice may reference a picture header, which contains data relating to the slice and references the LMCS APS. In either case, the header contains enough information to determine a suitable LMCS APS containing the LMCS parameters for the slice.
[0134] In step 807, other filtering parameters may be encoded into other APSs. For example, an ALF APS containing the ALF parameter and a scaling list APS containing the APS parameter may also be encoded into the bitstream. Such APSs may also be referenced by slice headers and / or picture headers.
[0135] As described above, each APS can be uniquely identified by a combination of parameter type and APS ID. Such information can be used by a slice header or picture header to reference the relevant APS. For example, each APS may include an aps_params_type code set to a predefined value indicating the type of parameter contained in the corresponding APS. Furthermore, each APS may include an APS ID selected from a predefined range. The predefined range can be determined based on the parameter type of the corresponding APS. For example, an LMCS APS may have a range of 0 to 3 (2 bits), and an ALF APS and a scaling list APS may have a range of 0 to 7 (3 bits). Such ranges can describe different overlapping value spaces specific to the APS type, as described by mechanism 600. Thus, in steps 805 and 807, both the APS type and the APS ID can be used to reference a particular APS.
[0136] In step 809, the encoder encodes the SPS into a bitstream. The SPS may contain a set of flags indicating that LMCS is enabled for the encoded video sequence, including the picture / slice. The bitstream may then be stored in memory for communication to the decoder, for example, via the transmitter, in step 811. The bitstream contains enough information for the decoder to obtain the LMCS parameters for decoding the encoded slice by including references in the slice header / picture parameter header in the LMCS APS.
[0137] Figure 9 is a flowchart of an exemplary method 900 for decoding a video sequence from a bitstream, such as bitstream 500, by using multiple APS types, such as ALF APS512, scaling list APS513, and / or LMCS APS514. Method 900 may be used by a decoder, such as a codec system 200, a decoder 400, and / or a video coding device 700, when performing method 100. Method 900 may also refer to an APS based on an APS ID assigned according to a mechanism 600, which uses APS IDs assigned according to different value spaces for different types of APS.
[0138] Method 900 may be initiated when the decoder begins receiving a bitstream of coded data representing a video sequence, for example, as a result of Method 800. In step 901, the bitstream is received by the decoder. The bitstream includes pictures divided into slices and an LMCS APS containing LMCS parameters. In some examples, the bitstream may further include an ALF APS containing ALF parameters and a scaling list APS containing APS parameters. The bitstream may also include picture headers and / or slice headers associated with one of the pictures and / or slices, respectively. The bitstream may also include other parameter sets such as SPS, PPS, etc.
[0139] In step 903, the decoder may determine that LMCS APS, ALF APS, and / or scaling list APS are referenced in the data relating to the slice. For example, a slice header or picture header may contain data relating to the slice / picture and may reference one or more APS, including LMCS APS, ALF APS, and / or scaling list APS. Each APS is uniquely identified by a combination of parameter type and APS ID. Therefore, the parameter type and APS ID may be included in the picture header and / or slice header. For example, each APS may include an aps_params_type code set to a predefined value indicating the type of parameter contained in each APS. Furthermore, each APS may include an APS ID selected from a predefined range. For example, APS IDs of APS types may be assigned sequentially across multiple different value spaces over a predefined range. For example, the predefined range may be determined based on the parameter type of the corresponding APS. For example, an LMCS APS may have a range of 0 to 3 (2 bits), and an ALF APS and a scaling list APS may have a range of 0 to 7 (3 bits). Such ranges can describe different overlapping value spaces specific to the APS type, as described by mechanism 600. Therefore, both the APS type and the APS ID may be used in the picture header and / or slice header to refer to a particular APS.
[0140] In step 905, the decoder may decode a slice using LMCS parameters from an LMCS APS based on a reference to the LMCS APS. The decoder may also decode a slice from an ALF APS and / or a scaling list APS using ALF parameters and / or scaling list parameters based on a reference to such an APS in the picture header / slice header. In some examples, the SPS includes a sequence parameter set (SPS) that includes flags set to indicate that the LMCS coding tool is valid for the encoded video sequence containing the slice. In such cases, LMCS parameters from the LMCS APS are retrieved in step 905 to support decoding based on the flags. In step 907, the decoder may transfer the slice for display as part of the decoded video sequence.
[0141] Figure 10 is a schematic diagram of an exemplary system 1000 for coding a video sequence of images in a bitstream, such as bitstream 500, by using multiple APS types such as ALF APS512, scaling list APS513, and / or LMCS APS514. System 1000 may be implemented by an encoder and decoder such as a codec system 200, an encoder 300, a decoder 400, and / or a video coding device 700. Furthermore, system 1000 may be used when implementing methods 100, 800, 900, and / or mechanism 600.
[0142] System 1000 includes a video encoder 1002. The video encoder 1002 includes a determination module 1001 for determining LMCS parameters for application to a slice. The video encoder 1002 further includes an encoding module 1003 for encoding the slice into a bitstream. The encoding module 1003 further encodes LMCS parameters in the LMCS APS into a bitstream. The encoding module 1003 further encodes data relating to the slice that references the LMCS APS into a bitstream. The video encoder 1002 further includes a storage module 1005 for storing the bitstream for communication toward the decoder. The video encoder 1002 further includes a transmission module 1007 for transmitting the bitstream containing the LMCS APS and supporting the decoding of the slice at the decoder. The video encoder 1002 may further be configured to perform any of the steps of Method 800.
[0143] System 1000 also includes a video decoder 1010. The video decoder 1010 includes a receiving module 1011 for receiving a bitstream containing a slice and an LMCS APS containing LMCS parameters. The video decoder 1010 further includes a determination module 1013 for determining whether the LMCS APS is referenced in the data relating to the slice. The video decoder 1010 further includes a decoding module 1015 for decoding the slice using the LMCS parameters from the LMCS APS based on the reference to the LMCS APS. The video decoder 1010 further includes a transfer module 1017 for transferring the slice for display as part of the decoded video sequence. The video decoder 1010 may further be configured to perform any of the steps of method 900.
[0144] The first component is directly joined to the second component when there are no intermediary components other than lines, traces, or other media between the first and second components. The first component is indirectly joined to the second component when there are intermediary components other than lines, traces, or other media between the first and second components. The term "joined" and its variations include both direct and indirect joining. The use of the term "about" means a range including ±10% of the following number unless otherwise specified.
[0145] It should also be understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and that the order of steps in such methods is merely illustrative. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in a manner consistent with various embodiments of the present disclosure.
[0146] While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods can be implemented in many other specific forms without departing from the spirit or scope of this disclosure. These embodiments are illustrative and not limiting, and their intent is not limited to the details given herein. For example, various elements or components may be combined or integrated into other systems, and certain features may be omitted or not implemented.
[0147] Furthermore, the technologies, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined with or integrated with other systems, components, technologies, or methods without departing from the scope of this disclosure. Other examples of modifications, substitutions, and alterations are readily apparent to those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. A method implemented in the decoder, The process involves receiving a bitstream and converting the bitstream into a plurality of syntactic elements by entropy decoding, wherein the plurality of syntactic elements are: The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, The ALF APS includes adaptive loop filter (ALF) parameters associated with the coded slice, and The process involves obtaining the LMCS parameters from the LMCS APS and obtaining the ALF parameters from the ALF APS. This includes obtaining a decoded picture based on the coded slice, Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A method wherein the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
2. The method according to claim 1, wherein each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters included in each APS.
3. The method according to any one of claims 1 to 2, wherein each APS is identified by a combination of parameter type and the APS ID.
4. The method according to any one of claims 1 to 3, wherein the bitstream further includes a sequence parameter set (SPS) which includes a flag set to indicate that the LMCS is valid for an encoded video sequence including the coded slice, and the LMCS parameters from the LMCS APS are obtained based on the flag.
5. The method according to any one of claims 1 to 4, wherein certain types of APS share the same value space for the APS ID, and different types of APS use different value spaces for the APS ID.
6. A method implemented in an encoder, Determine the chroma mapping (LMCS) parameters and adaptive loop filter (ALF) parameters associated with the coded slice containing the current block, and Based on the current block and the predicted block, the residual block is obtained, The process involves performing a transformation and quantization process on the residual block to obtain the quantized coefficients, The process involves performing entropy coding on the quantized coefficients to obtain a bitstream, The LMCS parameters are encoded into the LMCS Adaptive Parameter Set (APS) of the bitstream, This includes encoding the ALF parameters into the ALF APS of the bitstream, Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A method wherein the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
7. The method according to claim 6, wherein each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters included in each APS.
8. The method according to any one of claims 6 to 7, wherein each APS is identified by a combination of parameter type and the APS ID.
9. The method according to any one of claims 6 to 8, further comprising encoding a sequence parameter set (SPS) into the bitstream by the encoder, wherein the SPS includes a flag set to indicate that the LMCS is valid for the encoded video sequence containing the coded slice.
10. The method according to any one of claims 6 to 9, wherein certain types of APS share the same value space for the APS ID, and different types of APS use different value spaces for the APS ID.
11. A video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method according to any one of claims 1 to 10.
12. Non-temporary computer-readable medium containing a computer program used by a video coding device, wherein the computer program includes computer-executable instructions stored on the non-temporary computer-readable medium, which, when executed by a processor, cause the video coding device to perform the method according to any one of claims 1 to 10.
13. A decoder including a processing circuit for carrying out the method according to any one of claims 1 to 5.
14. An encoder including a processing circuit for carrying out the method according to any one of claims 6 to 10.
15. A computer program comprising program code for performing the method described in any one of claims 1 to 10 when executed on a computer or processor.
16. A non-temporary computer-readable medium that, when executed by a computer device, carries program code causing the computer device to perform the method according to any one of claims 1 to 10.
17. A method for transmitting a bitstream, The bitstream is received by at least one receiver, The bitstream is stored in at least one memory, and the bitstream is The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, ALF APS including adaptive loop filter (ALF) parameters associated with the coded slice, The LMCS parameters are obtained from the LMCS APS, and the ALF parameters are obtained from the ALF APS. Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A method wherein the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
18. A device for storing a bitstream, At least one receiver configured to receive the bitstream, The system includes at least one memory configured to store the bitstream, wherein the bitstream is The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, ALF APS including adaptive loop filter (ALF) parameters associated with the coded slice, The LMCS parameters are obtained from the LMCS APS, and the ALF parameters are obtained from the ALF APS. Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A device in which the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
19. A device for transmitting a bitstream, At least one processor configured to acquire the bitstream, The system includes at least one transmitter configured to transmit the bitstream, wherein the bitstream is The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, ALF APS including adaptive loop filter (ALF) parameters associated with the coded slice, The LMCS parameters are obtained from the LMCS APS, and the ALF parameters are obtained from the ALF APS. Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A device in which the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
20. A method for transmitting a bitstream, The bitstream is acquired by at least one processor, The process includes transmitting the bitstream by at least one transmitter, wherein the bitstream is The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, ALF APS including adaptive loop filter (ALF) parameters associated with the coded slice, The LMCS parameters are obtained from the LMCS APS, and the ALF parameters are obtained from the ALF APS. Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A method wherein the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
21. A system for processing a bitstream, comprising a source device, at least one storage medium, and a destination device, The source device is configured to provide a bitstream, The at least one storage medium is configured to store the bitstream, The destination device is used to decode the bitstream. The aforementioned bitstream is The LMCS Adaptive Parameter Set (APS) includes chroma mapping (LMCS) parameters with chroma scaling associated with the coded slice, ALF APS including adaptive loop filter (ALF) parameters associated with the coded slice, The LMCS parameters are obtained from the LMCS APS, and the ALF parameters are obtained from the ALF APS. Each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on the parameter type of the APS. A system in which the APS ID of the ALF APS is in the range of 0 to 7, and the APS ID of the LMCS APS is in the range of 0 to 3.
Citation Information
Patent Citations
Integrated image reshaping and video coding
WO2019006300A1
Video or image coding based on mapping of luma samples and scaling of chroma samples
WO2020262952A1
Scaling list data-based image or video coding
WO2021006630A1