Encoder, decoder, and corresponding methods
By implementing a system of separate Adaptive Parameter Sets for different coding tools within the video coding system, the inefficiencies in signaling coding tool parameters are addressed, resulting in improved coding efficiency and resource utilization.
Patent Information
- Application Number
- JP2025040034
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-21
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-02-26
AI Technical Summary
Existing video coding technologies face challenges in efficiently signaling coding tool parameters, leading to redundant coding and increased resource usage in both encoding and decoding processes.
The implementation of an Adaptive Parameter Set (APS) system that includes separate APS types for different coding tools, such as ALF, scaling list, and LMCS, allows for efficient encoding and decoding by reducing redundant signaling and optimizing resource usage.
This approach significantly reduces redundant coding of parameters, leading to improved coding efficiency, reduced network, memory, and processing resource usage, and enhanced video compression performance.
Smart Images

Figure 2025089316000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to video coding, and more particularly to efficient signaling of coding tool parameters used to compress video data in video coding.
Background Art
[0002] Even the amount of video data required to depict a relatively short video can be substantial, causing difficulties when the data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. Also, the size of the video can be a problem when the video is stored on a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data prior to transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable because network resources are limited and the demand for higher video quality is constantly increasing.
Summary of the Invention
[0003] In one embodiment, the present disclosure is a method implemented in a decoder, the method including: receiving, by a receiver of the decoder, a bitstream including a LMCS adaptation parameter set (APS) including a luma mapping (LMCS) parameter related to chroma scaling for a coded slice; determining, by a processor, that the LMCS APS is referenced in data related to the coded slice; decoding, by the processor, the coded slice using the LMCS parameter from the LMCS APS; and transferring, by the processor, the decoding result for display as part of a decoded video sequence. The APS is used to maintain data related to a plurality of slices over a plurality of pictures. The present disclosure describes improvements related to various APSs. In this example, the LMCS parameter is included in the LMCS APS. The LMCS / rescaler parameter may change about once per second. The video sequence may display 30 to 60 pictures per second. Thus, the LMCS parameter may not change for 30 to 60 frames. Including the LMCS parameter in the LMCS APS instead of a picture-level parameter set significantly reduces (e.g., by 30 to 60 times) the redundant coding of the LMCS parameter. The slice header and / or the picture header related to the slice can reference the related LMCS APS. In this way, the LMCS parameter is coded only when the LMCS parameter for the slice changes. Thus, using the LMCS APS to code the LMCS parameter improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0004] Optionally, in any of the above aspects, another implementation of the aspect is provided that the bitstream further includes a slice header, the coded slice refers to the slice header, the slice header includes data related to the coded slice, and refers to the LMCS APS.
[0005] Optionally, in any of the above aspects, another implementation of the aspect is provided that the bitstream further includes an Adaptive Loop Filter (ALF) APS including ALF parameters and a Scaling List APS including APS parameters.
[0006] Optionally, in any of the above aspects, another implementation of the aspect is provided that each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters included in each APS.
[0007] Optionally, in any of the above aspects, another implementation of the aspect is provided that each APS includes an APS identifier (ID) selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.
[0008] Optionally, in any of the above aspects, another implementation of the aspect is provided that each APS is identified by a combination of parameter type and APS ID.
[0009] Optionally, in any of the above aspects, another implementation of the aspect is provided that the bitstream further includes a Sequence Parameter Set (SPS) including a flag set to indicate that the LMCS is valid for the coded video sequence including the coded slice, and the LMCS parameters from the LMCS APS are obtained based on the flag.
[0010] In one embodiment, the present disclosure is a method implemented in an encoder, comprising: determining, by a processor of the encoder, LMCS parameters for application to a slice; encoding, by the processor, the slice as a coded slice into a bitstream; encoding, by the processor, the LMCS parameters in an LMCS APS into the bitstream; encoding, by the processor, data related to the coded slice that references the LMCS APS into the bitstream; and storing, by a memory coupled to the processor, the bitstream for communication towards a decoder. The APS is used to maintain data related to a plurality of slices over a plurality of pictures. The present disclosure describes improvements related to various APSs. In this example, the LMCS parameters are included in the LMCS APS. The LMCS / rescaler parameters may change about once per second. A video sequence may display 30 to 60 pictures per second. Thus, the LMCS parameters may not need to change for 30 to 60 frames. Including the LMCS parameters in the LMCS APS instead of a picture-level parameter set significantly reduces (e.g., by 30 to 60 times) the redundant coding of the LMCS parameters. A slice header and / or a picture header related to the slice can reference a relevant LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters for the slice change. Thus, using the LMCS APS to encode the LMCS parameters improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0011] Optionally, in any of the above aspects, another implementation of the aspect further includes encoding, by a processor, a slice header into a bitstream, wherein the encoded slice refers to the slice header, the slice header includes data related to the encoded slice, and refers to the LMCS APS.
[0012] Optionally, in any of the above aspects, another implementation of the aspect further includes further encoding, by a processor, an ALF APS including an ALF parameter and a scaling list APS including an APS parameter into a bitstream.
[0013] Optionally, in any of the above aspects, another implementation of the aspect provides that each APS includes an aps_params_type code set to a predefined value indicating the type of parameters included in each APS.
[0014] Optionally, in any of the above aspects, another implementation of the aspect provides that each APS includes an APS ID selected from a predefined range, the predefined range being determined based on the parameter type of each APS.
[0015] Optionally, in any of the above aspects, another implementation of the aspect provides that each APS is identified by a combination of a parameter type and an APS ID.
[0016] Optionally, in any of the above aspects, another implementation of the aspect further includes encoding, by a processor, an SPS into a bitstream, the SPS including a flag set to indicate that the LMCS is valid for an encoded video sequence including the encoded slice.
[0017] In one embodiment, the present disclosure includes a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, and the processor, receiver, memory, and transmitter are configured to execute the method according to any of the above aspects, including a video coding device.
[0018] In one embodiment, the present disclosure includes a non-transitory computer-readable medium including a computer program product used by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium to cause the video coding device to execute the method according to any of the above aspects when executed by a processor.
[0019] In one embodiment, the present disclosure includes receiving means for receiving a bitstream including an LMCS APS including LMCS parameters related to a coded slice, determining means for determining that the LMCS APS is referenced in data related to the coded slice, decoding means for decoding the coded slice using the LMCS parameters from the LMCS APS, and transfer means for transferring the decoding result for display as part of the decoded video sequence, including a decoder.
[0020] Optionally, in any of the above aspects, another implementation of the aspect provides that the decoder is further configured to execute the method according to any of the foregoing aspects.
[0021] In one embodiment, the present disclosure includes a determination means for determining LMCS parameters for application to a slice, an encoding means for encoding the slice into a bitstream as a coded slice, encoding the LMCS parameters into the bitstream in a LMCS Adaptive Parameter Set (APS), and encoding data related to the coded slice that refers to the LMCS APS into the bitstream, and a storage means for storing the bitstream for communication towards a decoder.
[0022] Optionally, in any of the above aspects, another implementation aspect of the aspect provides that the encoder is further configured to execute any of the methods of any of the foregoing aspects.
[0023] For clarity, any one of the foregoing embodiments can be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0024] These and other features will be more clearly understood from the following detailed description, which is to be construed in conjunction with the accompanying drawings and the claims.
Brief Description of the Drawings
[0025] To more fully understand the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.
[0026]
Figure 1
[0027]
Figure 2
[0028]
Figure 3
[0029]
Figure 4
[0030]
Figure 5
[0031]
Figure 6
[0032]
Figure 7
[0033]
Figure 8
[0034]
Figure 9
[0035]
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0036] Initially, exemplary implementations of one or more embodiments are provided below, but it should be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or to exist. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementation forms illustrated and described herein, but may be modified within the scope of the appended claims, along with the full scope of their equivalents.
[0037] The following abbreviations are used herein: Adaptive Loop Filter (ALF), Adaptive Parameter Set (APS), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH), Intra Random Access Point (IRAP), Joint Video Experts Team (JVET), Motion Constrained Tile Set (MCTS), Maximum Transmission Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Low Byte Sequence Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).
[0038] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques can include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with respect to reference samples in adjacent blocks within the same picture. Video blocks in an inter-coded forward prediction (P) or bidirectional prediction (B) slice of a picture may be coded by using spatial prediction with respect to reference samples in adjacent blocks within the same picture or by using temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame and / or an image, and a reference picture may sometimes be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a prediction block representing an image block. Residual data represents the pixel difference between the original image block and the prediction block. Thus, an inter-coded block is coded according to a block of reference samples forming a prediction block and a motion vector pointing to residual data indicating the difference between the coding block and the prediction block. An intra-coded block is coded according to an intra-coding mode and residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain. This results in residual transform coefficients, which may be quantized. The quantized transform coefficients may first be arranged in a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients.Entropy coding can be applied to achieve even more compression. Such video compression techniques are discussed in more detail below.
[0039] To ensure that the encoded video is accurately decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Moving Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or Advanced Video Coding (AVC) also known as ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC) also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multi-View Video Coding (MVC) and Multi-View Video Coding Plus Depth (MVC+D), and Three-Dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multi-View HEVC (MV-HEVC), 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has started developing a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Draft (WD) including JVET-M1001-v5 and JVET-M1002-v1, which provides an algorithm description, an encoder-side description of the VVC WD, and reference software.
[0040] A video sequence is coded by using various coding tools. The encoder selects parameters for the coding tools with the aim of improving compression with minimal quality loss when the video sequence is decoded. The coding tools may be related to different parts of the video at different scopes. For example, some coding tools are related at the video sequence level, some coding tools are related at the picture level, and some coding tools are related at the slice level. An APS may be used to signal information that can be shared by multiple pictures and / or multiple slices across different pictures. Specifically, the APS may carry adaptive loop filter (ALF) parameters. The ALF information may not be suitable for signaling at the sequence level in the sequence parameter set (SPS), at the picture level in the picture parameter set (PPS) or picture header, or at the slice level in the tile group / slice header for various reasons.
[0041] When ALF information is signaled in the SPS, the encoder must generate a new SPS and a new IRAP picture each time the ALF information changes. IRAP pictures significantly reduce coding efficiency. Therefore, placing ALF information in the SPS is particularly problematic in low-latency application environments that do not use frequent IRAP pictures. Also, including ALF information in the SPS may disable out-of-band transmission of the SPS. Out-of-band transmission refers to the transmission of corresponding data in a transport data flow different from the video bitstream (e.g., in the sample description or sample entry of a media file, in a session description protocol (SDP) file, etc.). Signaling ALF information in the PPS may also be problematic for similar reasons. Specifically, including ALF information in the PPS may disable out-of-band transmission of the PPS. Signaling ALF information in the picture header may also be problematic. The picture header may not be used in some cases. Furthermore, ALF information may be applied to multiple pictures. In this way, signaling ALF information in the picture header causes redundant information transmission and thus wastes bandwidth. Signaling ALF information in the tile group / slice header is also problematic because ALF information may be applied to multiple pictures and thus multiple slices / tile groups. Therefore, signaling ALF information in the slice / tile group header causes redundant information transmission and thus wastes bandwidth.
[0042] Based on the above, the APS can be used to signal the ALF parameter. However, the video coding system may use only the APS to signal the ALF parameter. Exemplary APS syntax and semantics are as follows: [Table 1]
[0043] The adaption_parameter_set_id provides an identifier for the APS for reference by other syntax elements. The APS can be shared between pictures and can be different in different tile groups within a picture. The aps_extension_flag is set equal to 0 to specify that the aps_extension_data_flag syntax element does not exist in the APS RBSP syntax structure. The aps_extension_flag is set equal to 1 to specify that the aps_extension_data_flag syntax element exists in the APS RBSP syntax structure. The aps_extension_data_flag may have any value. The presence and value of the aps_extension_data_flag may not affect the decoder's compliance to the profiles specified in VVC. A decoder compliant to VVC may ignore all aps_extension_data_flag syntax elements.
[0044] Exemplary tile group header syntax related to the ALF parameter is as follows.
Table 2
[0045] The tile_group_alf_enabled_flag is set equal to 1 to specify that the adaptive loop filter is enabled and may be applied to the luma (Y), blue chroma (Cb), or red chroma (Cr) color components within a tile group. The tile_group_alf_enabled_flag is set equal to 0 to specify that the adaptive loop filter is disabled for all color components within a tile group. The tile_group_aps_id specifies the adaptation_parameter_set_id of the APS referenced by the tile group. The TemporalId of an APS NAL unit having an adapdation_parameter_set_id equal to the tile_group_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. When multiple APSs with the same value of the adaption_parameter_set_id are referenced by multiple tile groups of the same picture, the multiple APSs with the same value of the adaption_parameter_set_id may contain the same content.
[0046] The reshaper parameter is a parameter used for an adaptive loop in-reshaper video coding tool, also known as luma mapping with chroma scaling (LMCS). Exemplary SPS reshaper syntax and semantics are as follows:
Table 3
[0047] The sps_reshaper_enabled_flag is set equal to 1 to specify that the reshaper is used in the coded video sequence (CVS). The sps_reshaper_enabled_flag is set to 0 to specify that the reshaper is not used in the CVS.
[0048] The exemplary tile group header / slice header reshaper syntax and semantics are as follows: [Table 4]
[0049] The tile_group_reshaper_model_present_flag is set equal to 1 to indicate that the tile_group_reshaper_model() is present in the tile group header. The tile_group_reshaper_model_present_flag is set equal to 0 to indicate that the tile_group_reshaper_model() is not present in the tile group header. If the tile_group_reshaper_model_present_flag is not present, the flag is assumed to be equal to 0. The tile_group_reshaper_enabled_flag is set equal to 1 to indicate that the reshaper is enabled for the current tile group. The tile_group_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not enabled for the current tile group. If the tile_group_resharper_enable_flag is not present, the flag is assumed to be equal to 0. The tile_group_reshamer_chroma_residual_scale_flag is set equal to 1 to indicate that chroma residual scaling is enabled for the current tile group. The tile_group_reshaper_chroma_residual_scale_flag is set to 0 to indicate that chroma residual scaling is not enabled for the current tile group. If the tile_group_reshaper_chroma_residual_scale_flag is not present, the flag is assumed to be equal to 0.
[0050] An exemplary tile group header / slice header reshaper model syntax and semantics are as follows: [Table 5]
[0051] reshape_model_min_bin_idx specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of reshape_model_min_bin_idx may range from 0 to MaxBinIdx (inclusive). The value of MaxBinIdx may be equal to 15. reshape_model_delta_max_bin_idx specifies the result of subtracting the maximum bin index used in the reshaper configuration process from the maximum allowable bin (or piece) index MaxBinIdx. The value of reshape_model_max_bin_idx is set equal to MaxBinIdx minus reshape_model_delta_max_bin_idx. One plus reshaper_model_bin_delta_abs_cw_prec_minus1 specifies the number of bits used in the representation of the syntax reshape_model_bin_delta_abs_CW[i]. reshape_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value of the i-th bin.
[0052] reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshape_model_bin_delta_abs_CW[i] as follows. When reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. Otherwise (for example, reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is a negative value. When reshape_model_bin_delta_sign_CW_flag[i] does not exist, the flag is presumed to be equal to 0. The variable RspDeltaCW[i] is set equal to (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i].
[0053] The variable RspCW[i] is derived as follows. The variable OrgCW is set equal to (1 << BitDepthY) / (MaxBinIdx + 1). When reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, RspCW[i] = OrgCW + RspDeltaCW[i]. Otherwise, RspCW[i] = 0. Assume that the value of RspCW[i] is in the range of 32 to 2 * OrgCW - 1 when the value of BitDepthY is equal to 10. The variable InputPivot[i] with i in the range of 0 to MaxBinIdx + 1 (including both ends) is derived as follows. InputPivot[i] = i * OrgCW. The variables ReshapePivot[i] with i in the range of 0 to MaxBinIdx + 1 (including both ends), and ScaleCoef[i] and InvScaleCoeff[i] with i in the range of 0 to MaxBinIdx (including both ends) are derived as follows:
Number
[0054] The variable ChromaScaleCoef[i] having i in the range of 0 to MaxBinIdx (including both ends) is derived as follows: [Number]
[0055] The characteristics of the reshaper parameters can be characterized as follows. The size of the set of reshaper parameters included in the tile_group_reshamer_model() syntax structure is usually approximately 60 to 100 bits. The reshaper model is usually updated by the encoder about once per second and includes many frames. Furthermore, the parameters of the updated reshaper model are unlikely to be exactly the same as the parameters of the previous instance of the reshaper model.
[0056] The aforementioned video coding system includes specific problems. First, such a system is configured only to carry ALF parameters in the APS. Furthermore, the reshaper / LMCS parameters may be shared by multiple pictures and may include many variations.
[0057] Disclosed herein are various mechanisms for modifying APS to support improved coding efficiency. In a first example, multiple types of APS are disclosed. Specifically, an APS of type ALF is referred to as an ALF APS and can include ALF parameters. Further, an APS of type scaling list is referred to as a scaling list APS and can include scaling list parameters. Additionally, an APS of type LMCS is referred to as an LMCS APS and can include LMCS / rescaler parameters. The ALF APS, the scaling list APS, and the LMCS APS may each be coded as a separate NAL type and thus be included in different NAL units. Thus, changes to data in one type of APS (e.g., ALF parameters) do not result in redundant coding of other types of data that are not changed (e.g., LMCS parameters). Therefore, providing multiple types of APS improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in an encoder and a decoder.
[0058] In the second example, each APS includes an APS identifier (ID). Further, each APS type includes a separate value space for the corresponding APS ID. Such value spaces can overlap. Thus, a first type of APS (e.g., an ALF APS) can include the same APS ID as a second type of APS (e.g., an LMCS APS). This is achieved by identifying each APS by a combination of the APS parameter type and the APS ID. By enabling each APS type to include a different value space, the codec does not need to check for ID conflicts across APS types. Further, by enabling the value spaces to overlap, the codec can avoid using larger ID values, resulting in bit savings. Thus, using separate overlapping value spaces for different types of APSs improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0059] In the third example, LMCS parameters are included in the LMCS APS. As described above, the LMCS / resampler parameters can change about once per second. A video sequence can display 30 to 60 pictures per second. Thus, the LMCS parameters do not need to change for 30 to 60 frames. Including the LMCS parameters in the LMCS APS significantly reduces the redundant coding of the LMCS parameters. The slice header and / or the picture header associated with the slice can reference the associated LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters for the slice change. Thus, using the LMCS APS to encode the LMCS parameters improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder.
[0060] FIG. 1 is a flowchart of an exemplary method of operation 100 for coding a video signal. Specifically, the video signal is encoded by an encoder. The encoding process compresses the video signal by using various mechanisms to reduce the video file size. The smaller file size enables the compressed video file to be sent to the user while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file and reconstructs the original video signal for display to the end user. The decoding process generally closely resembles the encoding process and enables the decoder to consistently reconstruct the video signal.
[0061] In step 101, the video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file can include both an audio component and a video component. The video component includes a series of image frames that give a visual impression of motion when viewed in sequence. The frames are represented from the perspective of light and are referred to herein as the luma component (or luma samples), and also include pixels that are represented from the perspective of color and are referred to as the chroma component (or color samples). In some examples, the frames may also include depth values to support 3D viewing.
[0062] In step 103, the video is partitioned into blocks. Partitioning includes sub - dividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC), also known as H.265 and MPEG - H Part 2, a frame can first be divided into Coding Tree Units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both luma and chroma samples. A coding tree is used to divide the CTU into blocks, and the blocks may be recursively sub - divided until a setting that supports further encoding is achieved. For example, the luma component of a frame can be sub - divided until the individual blocks contain relatively uniform illumination values. Further, the chroma component of a frame can be sub - divided until the individual blocks contain relatively uniform color values. Thus, the partitioning mechanism varies depending on the content of the video frame.
[0063] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects tend to appear in consecutive frames in a common scene. Therefore, blocks depicting objects in a reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a fixed position over a plurality of frames. Therefore, the table can be described once and adjacent frames can refer back to the reference frame. A pattern matching mechanism can be used to match objects over multiple frames. Furthermore, moving objects may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video can show a car moving across the screen over multiple frames. To describe such movement, motion vectors can be used. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of an object in a reference frame. Therefore, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from corresponding blocks in the reference frame.
[0064] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to concentrate in a frame. For example, a green patch in a part of a tree tends to be located adjacent to a similar green patch. Intra prediction uses multi-directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. The direction mode indicates that the current block is similar / same to the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., planar) can be interpolated based on adjacent blocks at the end of the row. The planar mode indicates a smooth transition of light / color across the row / column by using a relatively constant slope when actually changing values. The DC mode is used for boundary smoothing and indicates that the block is similar / same to the average value related to the samples of all adjacent blocks related to the angular direction of the direction prediction mode. Thus, the intra prediction block can represent the image block as various relationship prediction mode values instead of the actual value. Furthermore, the inter prediction block can represent the image block as a motion vector value instead of the actual value. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in the residual block. Transformation can be applied to the residual block to further compress the file.
[0065] In step 107, various filtering techniques can be applied. In HEVC, the filters are applied according to an in-loop filtering scheme. The above-described block-based prediction can result in the generation of a block-shaped image in the decoder. Further, the block-based prediction scheme can encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme repeatedly applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to the blocks / frames. These filters reduce such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters reduce artifacts in the reconstructed reference blocks and reduce the likelihood that the artifacts will generate additional artifacts in subsequent blocks encoded based on the reconstructed reference blocks.
[0066] When the video signal is segmented, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. The bitstream includes the above-described data as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream can be stored in memory for transmission towards the decoder if requested. The bitstream may also be broadcast and / or multicast towards multiple decoders. The generation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 can occur continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to a particular order.
[0067] The decoder receives a bitstream and, in step 111, starts the decoding process. Specifically, the decoder uses an entropy decoding method that converts the bitstream into the corresponding syntax and video data. In step 111, the decoder determines the partition for a frame using the syntax data from the bitstream. The partitioning should match the result of the block partitioning in step 103. The entropy encoding / decoding as used in step 111 will now be described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible options based on the spatial positioning of the values within the input image. Signaling exact options may involve using a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding enables the encoder to discard any option that is clearly infeasible in a particular case, leaving a set of acceptable options. Next, a codeword is assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., one bin for two options, two bins for three to four options, etc.). Next, the encoder encodes the codeword for the selected option. This scheme is desired because the codeword size is reduced since the codeword uniquely indicates a selection from a small subset of acceptable options as opposed to uniquely indicating a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of acceptable options in the same way as the encoder. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.
[0068] In step 113, the decoder performs block decoding. Specifically, the decoder uses inverse transformation to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partitioning. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. The reconstructed image blocks are then positioned within the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding as described above.
[0069] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. When the frame is filtered, the video signal may be output to a display for viewing by the user in step 117.
[0070] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support an implementation of operation method 100. Codec system 200 is generalized to depict components used in both an encoder and a decoder. Codec system 200 receives and partitions a video signal, as discussed with respect to steps 101 and 103 of operation method 100, to obtain a partitioned video signal 201. Codec system 200 then compresses the partitioned video signal 201 into a coded bitstream when acting as an encoder, as discussed with respect to steps 105, 107, and 109 of method 100. When acting as a decoder, codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 in operation method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. Such components are coupled as shown. In FIG. 2, solid lines indicate the movement of data being encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All components of codec system 200 may be present within the encoder. The decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components will be described hereinafter.
[0071] The partitioned video signal 201 is a captured video sequence, which is partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to sub-divide blocks of pixels into smaller blocks of pixels. These blocks can then be further sub-divided into even smaller blocks. A block may sometimes be referred to as a node on the coding tree. A larger parent node is divided into smaller child nodes. The number of times a node is sub-divided is called the depth of the node / coding tree. In some cases, the divided blocks may be included in a coding unit (CU). For example, a CU can be a sub-part of a CTU that includes a luma block, a chroma red difference (Cr) block, and a chroma blue difference (Cb) block, along with the syntax instruction of the corresponding CU. The partitioning mode may include a binary tree (BT), a triple tree (TT), and a quad tree (QT) that are used to partition a node into two, three, or four child nodes of different shapes respectively, depending on the partitioning mode used. The partitioned video signal 201 is transferred to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0072] The general coder control component 211 is configured to make decisions related to coding the images of a video sequence into a bitstream according to the constraints of the application. For example, the general coder control component 211 manages the optimization of the bit rate / bitstream size versus the reconstructed quality. Such decisions can be made based on the availability of memory area / bandwidth and the image resolution requirements. The general coder control component 211 also manages the utilization of the buffer in light of the transmission speed in order to mitigate the problems of buffer underrun and overrun. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general coder control component 211 can dynamically increase the complexity of compression to improve the resolution and improve the use of bandwidth, or can decrease the complexity of compression to reduce the resolution and the use of bandwidth. Thus, the general coder control component 211 controls the other components of the codec system 200 to balance the concern of the bit rate and the quality of the reconstructed video signal. The general coder control component 211 generates control data for controlling the operation of the other components. The control data is also transferred to the header formatting and the CABAC component 231 which are encoded in the bitstream for signaling the parameters for decoding by the decoder.
[0073] The partitioned video signal 201 is also transmitted to the motion estimation component 221 and the motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 can be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can execute a plurality of coding paths to select, for example, an appropriate coding mode for each block of video data.
[0074] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, which estimates the motion of video blocks. The motion vectors may indicate, for example, the displacement of the coded object with respect to the prediction block. The prediction block is a block that is found to exactly match the block to be coded with respect to the pixel difference. The prediction block may sometimes be referred to as a reference block. Such pixel differences can be determined by the sum of the absolute difference (SAD), the sum of squared differences (SSD), or other difference metrics. HEVC uses several coded objects including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU is divided into CTBs, and a CTB can then be divided into CUs for inclusion in a CB. A CU can be coded as a prediction unit (PU) containing prediction data and / or a transform unit (TU) containing the transform residual data of the CU. The motion estimation component 221 generates motion vectors, PUs, and TUs using rate distortion analysis as part of the rate distortion optimization process. For example, the motion estimation component 221 may determine a plurality of reference blocks, a plurality of motion vectors, etc. for the current block / frame, and select the reference block, motion vector, etc. having the best rate distortion characteristics. The best rate distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final coding).
[0075] In some examples, codec system 200 may calculate values of sub-integer pixel positions of reference pictures stored in decoded picture buffer component 223. For example, video codec system 200 may interpolate values at 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of the reference picture. Thus, motion estimation component 221 can perform a motion search for both full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision. Motion estimation component 221 calculates the motion vector of a PU of a video block in an inter-coded slice by comparing the position of the PU with the position of a predicted block of the reference picture. Motion estimation component 221 outputs the motion vector calculated for encoding as motion data to header formatting and CABAC component 231 and outputs the motion to motion compensation component 219.
[0076] Motion compensation performed by motion compensation component 219 may involve fetching or generating a predicted block based on the motion vector determined by motion estimation component 221. Also, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. Upon receiving a motion vector for a PU of the current video block, motion compensation component 219 can identify the position of the predicted block pointed to by the motion vector. By subtracting the pixel values of the predicted block from the pixel values of the currently encoded video block, a residual video block is then formed, forming pixel difference values. Generally, motion estimation component 221 performs motion estimation for the luma component, and motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma component and the luma component. The predicted block and the residual block are transferred to transform scaling and quantization component 213.
[0077] The partitioned video signal 201 is also sent to the intra-picture estimation component 215 and the intra-picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra-picture estimation component 215 and the intra-picture prediction component 217 perform intra prediction of the current block for a block within the current frame instead of the inter prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 described above. In particular, the intra-picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra prediction mode to encode the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then transferred to the header formatting and CABAC component 231 for encoding.
[0078] For example, the intra-picture estimation component 215 calculates rate-distortion values using rate-distortion analysis for various tested intra prediction modes, and selects an intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the bit rate (e.g., number of bits) used to generate an encoded block, and the amount of distortion (or error) between the original unencoded block and the encoded block that is encoded to generate the encoded block. The intra-picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks, and determines which intra prediction mode exhibits the best rate-distortion value for the block. Additionally, the intra-picture estimation component 215 may be configured to code depth blocks of a depth map using a depth modeling mode (DMM) based on rate-distortion optimization (RDO).
[0079] The intra-picture prediction component 217, when implemented in an encoder, can generate a residual block from a prediction block based on the selected intra prediction mode determined by the intra-picture estimation component 215, or, when implemented in a decoder, can read a residual block from a bit stream. The residual block includes the difference in values between the prediction block, represented as a matrix, and the original block. The residual block is then transferred to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 can operate on both the luma component and the chroma component.
[0080] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block that includes residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may transform the residual information from the pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information such that different frequency information is quantized at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix that includes the quantized transform coefficients. The quantized transform coefficients are transferred to the header formatting and CABAC component 231 and encoded into a bitstream.
[0081] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct the residual block in the pixel domain and use it, for example, as a reference block that can later serve as a predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate a reference block by returning the residual block to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. Otherwise, such artifacts can cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.
[0082] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 may be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are depicted separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters to adjust how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data for encoding to the header formatting and CABAC component 231. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, a SAO filter, and an adaptive loop filter. Such a filter may be applied to a spatial / pixel region (e.g., in the reconstructed pixel block) or a frequency region depending on the example.
[0083] When operating as a symbolizer, the filtered and reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them towards the display. The decoded picture buffer component 223 can be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0084] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to a decoder. Specifically, the header formatting and CABAC component 231 generates various headers for encoding control data such as general control data and filter control data. Further, predictive data including intra prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra prediction mode index table (also referred to as a codeword mapping table), the definition of coding contexts for various blocks, the indication of the most likely intra prediction mode, the indication of partition information, etc. Such data may be encoded by using entropy coding. For example, the information may be encoded by using context adaptive variable length coding (CAVLC), CABAC, syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder), or may be archived for later transmission or retrieval.
[0085] FIG. 3 is a block diagram illustrating an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 partitions the input video signal, resulting in a partitioned video signal 301, which is substantially similar to the partitioned video signal 201. The partitioned video signal 301 is then compressed by components of the encoder 300 and encoded into a bitstream.
[0086] Specifically, the partitioned video signal 301 is transferred to the intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also transferred to the motion compensation component 321 for inter prediction based on reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks are transferred to the entropy coding component 331 for coding into the bitstream (along with associated control data). The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231.
[0087] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into the reference blocks used by the motion compensation component 321. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks, depending on the example. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 can include multiple filters, as discussed with respect to the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 and used as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 can be substantially the same as the decoded picture buffer component 223.
[0088] Figure 4 is a block diagram illustrating an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding functions of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the method of operation 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates an output video signal reconstructed based on the bitstream for display to an end user.
[0089] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use the header information to provide a context for interpreting additional data encoded as code words within the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0090] The reconfigured residual block and / or prediction block is transferred to the intra-picture prediction component 417 for the reconstruction of an image block based on an intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to locate a reference block within a frame, applies the residual block to the result, and reconstructs the intra-predicted image block. The reconstructed intra-predicted image block and / or residual block and the corresponding inter prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425, and these components may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses a motion vector from a reference block to generate a prediction block, applies the residual block to the result, and reconstructs the image block. The resulting reconstructed block may also be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that may be reconstructed into a frame via partition information. Such a frame may also be placed in a sequence. This sequence is output to a display as a reconstructed output video signal.
[0091] FIG. 5 is a schematic diagram illustrating an exemplary bitstream 500 including multiple types of APSs including different types of coding tool parameters. For example, the bitstream 500 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400. As another example, the bitstream 500 may be generated by the encoder in step 109 of method 100 for use by the decoder in step 111.
[0092] The bitstream 500 includes a sequence parameter set (SPS) 510, a plurality of picture parameter sets (PPSs) 511, a plurality of ALF APSs 512, a plurality of scaling list APSs 513, a plurality of LMCS APSs 514, a plurality of slice headers 515, and picture data 520. The SPS 510 includes sequence data common to all pictures in the video sequence included in the bitstream 500. Such data can include picture size, bit depth, coding tool parameters, bit rate limit, and the like. The PPS 511 includes parameters applicable to the entire picture. Thus, each picture in the video sequence can refer to the PPS 511. Note that each picture refers to the PPS 511, but a single PPS 511 can, in some examples, include data for a plurality of pictures. For example, a plurality of similar pictures may be coded according to similar parameters. In such a case, a single PPS 511 can include data for such similar pictures. The PPS 511 can indicate coding tools available for slices in the corresponding picture, quantization parameter, offset, and the like. The slice header 515 includes parameters specific to each slice in the picture. Thus, in the video sequence, there can be one slice header 515 per slice. The slice header 515 can include slice type information, picture order count (POC), reference picture list, prediction weight, tile entry point, block release parameters, and the like. Note that the slice header 515 may also be referred to as a tile group header in some contexts.
[0093] The APS is a syntax structure including syntax elements applicable to one or more pictures 521 and / or slices 523. In the illustrated example, the APS can be divided into multiple types. The ALF APS 512 is an APS of type ALF including ALF parameters. The ALF includes a transfer function controlled by variable parameters and is an adaptive block-based filter that uses feedback from a feedback loop to improve the transfer function. Further, the ALF is used to correct coding artifacts (e.g., errors) resulting from block-based coding. The adaptive filter is a linear filter having a transfer function controller by variable parameters that can be controlled by an optimization algorithm such as an RDO process operating in an encoder. Therefore, the ALF parameters included in the ALF APS 512 may include variable parameters selected by the encoder such that the filter removes block-based coding artifacts during decoding in a decoder.
[0094] The scaling list APS 513 is an APS of type scaling list including scaling list parameters. As described above, the current block is coded according to an inter-prediction or intra-prediction that results in a residual. The residual is the difference between the luma value and / or chroma value of the block and the value of the corresponding prediction block. Next, a transform is applied to the residual to transform it into transform coefficients (smaller than the residual value). Encoding high-resolution and / or ultra-high-resolution content may result in increased residual data. A simple transform process may result in significant quantization noise when applied to such data. Therefore, the scaling list parameters included in the scaling list APS 513 may include weighting parameters that are applied to scale the transform matrix and can take into account acceptable levels of change in display resolution and / or quantization noise in the resulting decoded video image.
[0095] The LMCS APS514 is an APS of type LMCS that includes LMCS parameters, which are also known as reshaper parameters. The human visual system has a lower ability to distinguish color differences (e.g., chrominance) than differences in light (e.g., luminance). Therefore, some video systems use a chroma subsampling mechanism to compress video data by reducing the resolution of chroma values without adjusting the corresponding luma values. One concern with such a mechanism is that the associated interpolation may generate interpolated chroma values during decoding that are incompatible with the corresponding luma values in some locations. This generates color artifacts in such locations, and these color artifacts should be corrected by the corresponding filter. This is complicated by the luma mapping mechanism. Luma mapping is the process of remapping the luma component coded across the dynamic range of the input luma signal (e.g., according to a piecewise linear function). This compresses the luma component. The LMCS algorithm scales the compressed chroma values based on luma mapping to remove artifacts related to chroma subsampling. Therefore, the LMCS parameters included in the LMCS APS514 indicate the chroma scaling used to describe the luma mapping. The LMCS parameters are determined by the encoder and can be used by the decoder to remove artifacts caused by chroma subsampling with a filter when luma mapping is used.
[0096] The picture data 520 includes video data encoded according to inter prediction and / or intra prediction, as well as corresponding transform and quantization residual data. For example, a video sequence includes a plurality of pictures 521 coded as picture data. One picture 521 is a single frame of the video sequence, and thus, when the video sequence is displayed, it is generally displayed as a single unit. However, for implementing certain technologies such as virtual reality, picture-in-picture, etc., partial pictures may be displayed. The plurality of pictures 521 each refer to a PPS 511. The plurality of pictures 521 are divided into a plurality of slices 523. One slice 523 can be defined as a horizontal section of one picture 521. For example, one slice 523 can include a part of the height of the picture 521 and the complete width of the picture 521. In other cases, the picture 521 may be divided into columns and rows, and the slice 523 may be included in a rectangular portion of the picture 521 generated by such columns and rows. In some systems, the slice 523 is sub-divided into tiles. In other systems, the slice 523 is called a tile group including tiles. The slice 523 and / or the tile group of tiles refer to a slice header 515. The slice 523 is further divided into coding tree units (CTUs). The CTU is further divided into coding blocks based on a coding tree. The coding block can then be encoded / decoded according to a prediction mechanism.
[0097] Picture 521 and / or slice 523 may directly or indirectly refer to ALF APS512, scaling list APS513, and / or LMCS APS514 that contain related parameters. For example, slice 523 may refer to slice header 515. Further, picture 521 may refer to the corresponding picture header. Slice header 515 and / or picture header may refer to ALF APS512, scaling list APS513, and / or LMCS APS514 that contain parameters used when coding the related slice 523 and / or picture 521. In this way, the decoder can obtain coding tool parameters relevant to slice 523 and / or picture 521 according to the header references related to the corresponding slice 523 and / or picture 521.
[0098] Bitstream 500 is coded into video coding layer (VCL) NAL units 535 and non-VCL NAL units 531. A NAL unit is a coded data unit formed to a size arranged as the payload of a single packet for transmission over a network. VCL NAL unit 535 is a NAL unit that contains coded video data. For example, each VCL NAL unit 535 can contain one slice 523 and / or a tile group of data, CTUs, and / or coding blocks. Non-VCL NAL unit 531 is a NAL unit that contains support syntax but does not contain coded video data. For example, non-VCL NAL unit 531 may include SPS510, PPS511, APS, slice header 515, etc. Therefore, the decoder receives bitstream 500 in discrete VCL NAL units 535 and non-VCL NAL units 531. An access unit is a group of VCL NAL units 535 and / or non-VCL NAL units 531 that contains sufficient data to code a single picture 521.
[0099] In some examples, ALF APS512, Scaling List APS513, and LMCS APS514 are each assigned to a separate non-VCL NAL unit 531 type. In such cases, ALF APS512, Scaling List APS513, and LMCS APS514 are included in an ALF APS NAL unit 532, a Scaling List APS NAL unit 533, and an LMCS APS NAL unit 534, respectively. Thus, the ALF APS NAL unit 532 contains ALF parameters that remain valid until another ALF APS NAL unit 532 is received. Further, the Scaling List APS NAL unit 533 contains scaling list parameters that remain valid until another Scaling List APS NAL unit 533 is received. Additionally, the LMCS APS NAL unit 534 contains LMCS parameters that remain valid until another LMCS APS NAL unit 534 is received. Thus, there is no need to issue a new APS each time an APS parameter changes. For example, a change in the LMCS parameters results in an additional LMCS APS514, but does not result in an additional ALF APS512 or Scaling List APS513. Thus, by separating APSs into different NAL unit types based on parameter type, redundant signaling of unrelated parameters is avoided. Thus, separating APSs into different NAL unit types improves coding efficiency and thus reduces the use of processor, memory, and / or network resources in the encoder and decoder.
[0100] Furthermore, slice 523 and / or picture 521 can directly or indirectly reference ALF APS 512, ALF APS NAL unit 532, scaling list APS 513, scaling list APS NAL unit 533, LMCS APS 514, and / or LMCS APS NAL unit 534 that contain coding tool parameters used to code slice 523 and / or picture 521. For example, each APS may include an APS ID 542 and a parameter type 541. The APS ID 542 is a value (e.g., a number) that identifies the corresponding APS. The APS ID 542 may include a predefined number of bits. Thus, the APS ID 542 can increase (e.g., increment by one) according to a predefined order, and when that order reaches the end of a predefined range, it can be reset to a minimum value (e.g., 0). The parameter type 541 indicates the type of parameter (e.g., ALF, scaling list, and / or LMCS) included in the APS. For example, the parameter type 541 may include an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameter included in each APS. In this way, the parameter type 541 can be used to distinguish ALF APS 512, scaling list APS 513, and LMCS APS 514. In some examples, ALF APS 512, scaling list APS 513, and LMCS APS 514 can each be uniquely identified by a combination of the parameter type 541 and the APS ID 542. For example, each APS type may include a separate value space for the corresponding APS ID 542. Thus, each APS type may include an APS ID 542 that increases in order based on the previous APS of the same type. However, the APS ID 542 of the first APS type may not be related to the APS ID 542 of the previous APS of a different second APS type. In this way, the APS ID 542 of different APS types may include overlapping value spaces.For example, a first type of APS (e.g., ALF APS) can, in some cases, include the same APS ID 542 as a second type of APS (e.g., LMCS APS). By enabling each APS type to include a different value space, the codec does not need to check for conflicts of APS ID 542 between APS types. Further, by enabling the value spaces to overlap, the codec can avoid using larger values of APS ID 542, which results in bit savings. Thus, using separate overlapping value spaces for multiple APS IDs 542 of different APS types improves coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. As described above, the APS ID 542 can span a predefined range. In some examples, the predefined range of the APS ID 542 may vary according to the APS type indicated by parameter type 541. This can potentially allow different numbers of bits to be assigned to different APS types according to the frequency with which different types of parameters generally change. For example, the APS ID 542 of ALF APS 512 may have a range of 0 to 7, the APS ID 542 of scaling list APS 513 may have a range of 0 to 7, and the APS ID 542 of LMCS APS 514 may have a range of 0 to 3.
[0101] In another example, the LMCS parameters are included in the LMCS APS514. Some systems include the LMCS parameters in the slice header 515. However, the LMCS / resampler parameters may change once every about one second. The video sequence may display 30 to 60 pictures 521 per second. Therefore, the LMCS parameters cannot change for 30 to 60 frames. Including the LMCS parameters in the LMCS APS514 significantly reduces the redundant coding of the LMCS parameters. In some examples, the picture headers associated with the slice header 515 and / or the slice 523 and / or the picture 521 can each reference the associated LMCS APS514. The slice 523 and / or the picture 521 then reference the slice header 515 and / or the picture header. This enables the decoder to obtain the LMCS parameters of the associated slice 523 and / or picture 521. In this way, the LMCS parameters are encoded only when the LMCS parameters for the slice 523 and / or picture 521 change. Therefore, using the LMCS APS514 to encode the LMCS parameters increases the coding efficiency and thus reduces the use of network resources, memory resources, and / or processing resources in the encoder and decoder. Since LMCS is not used for all videos, the SPS510 may include an LMCS enable flag 543. The LMCS enable flag 543 can be set to indicate that LMCS is enabled for the encoded video sequence. Therefore, the decoder can obtain the LMCS parameters from the LMCS APS514 based on the LMCS enable flag 543 when the LMCS enable flag 543 is set (e.g., set to 1). Further, the decoder may not need to attempt to obtain the LMCS parameters when the LMCS enable flag 543 is not set (e.g., not set to 0).
[0102] FIG. 6 is a schematic diagram illustrating an exemplary mechanism 600 for assigning APS ID 642 to different APS types on different value spaces. For example, mechanism 600 can be applied to bitstream 500 to assign APS ID 542 to ALF APS 512, scaling list APS 513, and / or LMCS APS 514. Further, mechanism 600 may be applied to codec 200, encoder 300, and / or decoder 400 when coding video according to method 100.
[0103] Mechanism 600 assigns APS ID 642 to ALF APS 612, scaling list APS 613, and LMCS APS 614, which may be substantially the same as APS ID 542, ALF APS 512, scaling list APS 513, and LMCS APS 514, respectively. As described above, APS ID 642 may be sequentially assigned on a plurality of different value spaces where each value space is unique to an APS type. Further, each value space may span a different range that is unique to the APS type. In the example shown, the range of the value space of APS ID 642 for ALF APS 612 is 0 to 7 (e.g., 3 bits). Further, the range of the value space of APS ID 642 for scaling list APS 613 is 0 to 7 (e.g., 3 bits). Also, the range of the value space of APS ID 642 for LMCS APS 611 is 0 to 3 (e.g., 2 bits). When APS ID 642 reaches the end of the range of the value space, the APS ID 642 of the next APS of the corresponding type returns to the beginning of the range (e.g., 0). When a new APS receives the same APS ID 642 as the previous APS of the same type, the previous APS is no longer active and can no longer be referenced. In this way, the range of the value space can be extended to allow more APSs of the same type to be actively referenced. Further, the range of the value space can be reduced to improve coding efficiency, but such a reduction also decreases the number of APSs of the corresponding type that can remain active and available for reference at the same time.
[0104] In the example shown, each ALF APS612, scaling list APS613, and LMCS APS614 is referenced by a combination of an APS ID642 and an APS type. For example, the ALF APS612, LMCS APS614, and scaling list APS613 each receive an APS ID642 of 0. When a new ALF APS612 is received, the APS ID642 is incremented from the value used for the previous ALF APS612. The same sequence applies to the scaling list APS613 and LMCS APS614. Thus, each APS ID642 is related to the APS ID642 of the previous APS of the same type. However, the APS ID642 is not related to the APS ID642 of the previous APS of other types. In this example, the APS ID642 of the ALF APS612 increments from 0 to 7 and then returns to 0 before continuing to increment. Further, the APS ID642 of the scaling list APS613 increments from 0 to 7 and then returns to 0 before continuing to increment. Also, the APS ID642 of the LMCS APS611 increments from 0 to 3 and then returns to 0 before continuing to increment. As shown, such a value space is overlapping since different APSs of different APS types can share the same APS ID642 at the same point within a video sequence. Note also that mechanism 600 only draws APSs. In the bitstream, the drawn APSs are interspersed between other VCL and non-VCL NAL units such as SPS, PPS, slice headers, picture headers, slices, etc.
[0105] Thus, the present disclosure includes improvements to the design of the APS and several improvements for signaling reshaper / LMCS parameters. The APS is designed for signaling information that can be shared across multiple pictures and can include many variations. The reshaper / LMCS parameters are used in the adaptive loop in-reshaper / LMCS video coding tools. The mechanisms can be implemented as follows. To solve the problems listed herein, this includes several aspects that can be used individually and / or in combination.
[0106] The disclosed APS is modified such that multiple APSs can be used to carry different types of parameters. Each APS NAL unit is used to carry only one type of parameter. As a result, two APS NAL units are coded if two types of information are carried for a particular tile group / slice (e.g., one for each type of information). The APS may include an APS parameter type field in the APS syntax. The APS NAL unit may include only the type of parameter indicated by the APS parameters type field.
[0107] In some examples, different types of APS parameters are indicated by different NAL unit types. For example, two different NAL unit types are used for APS. These two types of APS may be referred to as ALF APS and reshaper APS, respectively. In another example, the type of tool parameters carried in the APS NAL unit is specified in the NAL unit header. In VVC, the NAL unit header has reserved bits (e.g., 7 bits indicated as nuh_reserved_zero_7bit). In some examples, some of these bits (e.g., 3 out of 7 bits) may be used to specify the APS parameter type field. In some examples, specific types of APS may share the same value space for the APS ID. On the other hand, different types of APS use different value spaces of the APS ID. Thus, two different types of APS may coexist and have the same value of the APS ID at the same instant. Further, a combination of the APS ID and the APS parameter type may be used to identify an APS from other APSs.
[0108] The APS ID may be included in the tile group header syntax if the corresponding coding tool is valid for the tile group. Otherwise, the corresponding type of APS ID may not be included in the tile group header. For example, if ALF is valid for the tile group, the APS ID of the ALF APS is included in the tile group header. For example, this can be achieved by setting the APS parameter type field to indicate the ALF type. Thus, if ALF is not valid for the tile group, the APS ID of the ALF APS is not included in the tile group header. Further, if the reshaper coding tool is valid for the tile group, the APS ID of the reshaper APS is included in the tile group header. For example, this can be achieved by resetting the APS parameter type field to indicate the reshaper type. Further, if the reshaper coding tool is not valid for the tile group, the APS ID of the reshaper APS may not be included in the tile group header.
[0109] In some cases, the presence of APS parameter type information in the APS can be conditioned by the use of the coding tool associated with the parameter. If only one coding tool related to one APS is valid for one bitstream (e.g., LMCS, ALF, scaling list), the APS parameter type information may not be present and may instead be inferred. For example, the APS can contain parameters for the ALF and reshaper coding tools, but only ALF is valid (e.g., as specified by the flag in the SPS) and the reshaper is not valid (e.g., as specified by the flag in the SPS), the APS parameter type may not be signaled and may be inferred to be equal to the ALF parameter.
[0110] In another example, the APS parameter type information can be inferred from the value of the APS ID. For example, a pre-defined range of APS ID values can be associated with the corresponding APS parameter type. This aspect can be implemented as follows. Instead of allocating X bits for signaling the APS ID and Y bits for signaling the type of the APS parameter, X + Y bits can be allocated for signaling the APS ID. Next, different ranges of the APS ID values can be specified to indicate different types of APS parameters. For example, instead of using 5 bits for signaling the APS ID and 3 bits for signaling the type of the APS parameter, 8 bits can be allocated for signaling the APS ID (e.g., without increasing the bit cost). A range of APS ID values from 0 to 63 indicates that the APS contains parameters for the ALF, a range of ID values from 64 to 95 indicates that the APS contains parameters for the reshaper, and 96 to 255 can be reserved for other parameter types such as scaling lists. In another example, a range of APS ID values from 0 to 31 indicates that the APS contains parameters for the ALF, a range of ID values from 32 to 47 indicates that the APS contains parameters for the reshaper, and 48 to 255 can be reserved for other parameter types such as scaling lists. The advantage of this approach is that the range of the APS ID can be allocated according to the frequency of parameter changes for each tool. For example, it may be expected that the ALF parameters change more frequently than the reshaper parameters. In such a case, a larger APS ID range may be used to indicate that the APS is an APS containing ALF parameters.
[0111] In the first embodiment, one or more of the foregoing aspects may be implemented as follows. The ALF APS may be defined as an APS having an aps_params_type equal to ALF_APS. The reshaper APS (or LMCS APS) may be defined as an APS having an aps_params_type equal to MAP_APS. Exemplary SPS syntax and semantics are as follows.
Table 6
[0112] Exemplary APS syntax and semantics are as follows.
Table 7
[0113] The aps_params_type specifies the type of APS parameters carried in an APS as specified in the following table.
Table 8
[0114] Exemplary tile group header syntax and semantics are as follows.
Table 9
[0115] The tile_group_alf_aps_id specifies the adaptation_parameter_set_id of the ALF APS referred to by the tile group. It is assumed that the TemporalId of the ALF APS NAL unit having an adapdation_parameter_set_id equal to the tile_group_alf_aps_id is less than or equal to the TemporalId of the coded tile group NAL unit. When multiple ALF APSs having the same value of adaptation_parameter_set_id are referred to by two or more tile groups of the same picture, it is assumed that the multiple ALF APSs having the same value of adapation_parameter_set_id have the same content.
[0116] The tile_group_reshaper_enabled_flag is set equal to 1 to indicate that the reshaper is enabled for the current tile group. The tile_group_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not enabled for the current tile group. In the absence of the tile_group_resharper_enable_flag, the flag is assumed to be equal to 0. The tile_group_reshaper_aps_id specifies the adaptation_parameter_set_id of the reshaper APS that the tile group refers to. The TemporalId of the reshaper APS NAL unit with an adapdation_parameter_set_id equal to the tile_group_reshaper_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. When multiple reshaper APSs with the same value of the adapation_parameter_set_id are referred to by two or more tile groups of the same picture, the multiple reshaper APSs with the same value of the adapation_parameter_set_id shall have the same content. The tile_group_reshaper_chroma_residual_scale_flag is set equal to 1 to indicate that chroma residual scaling is enabled for the current tile group. The tile_group_reshaper_chroma_residual_scale_flag is set equal to 0 to indicate that chroma residual scaling is not enabled for the current tile group. In the absence of the tile_group_reshaper_chroma_residual_scale_flag, the flag is assumed to be equal to 0.
[0117] Exemplary reshaper data syntax and semantics are as follows. [Table 10]
[0118] The reshaper_model_min_bin_idx specifies the minimum bin (or piece) index used in the reshaper configuration process. The value of reshaper_model_min_bin_idx shall be in the range from 0 to MaxBinIdx (including both ends). Assume the value of MaxBinIdx is equal to 15. The reshaper_model_delta_max_bin_idx specifies the result of subtracting the maximum bin index used in the reshaper configuration process from the maximum allowable bin (or piece) index MaxBinIdx. The value of reshaper_model_max_bin_idx is set to be equal to MaxBinIdx - reshaper_model_delta_max_bin_idx. The resharper_model_bin_delta_abs_cw_prec_minus1 + 1 specifies the number of bits used for the representation of the syntax element resharper_model_bin_delta_abs_CW[i]. The resharper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value of the i-th bin. The resharper_model_bin_delta_abs_CW[i] syntax element is represented by resharper_model_bin_delta_abs_cw_prec_minus1+1 bits. The resharper_model_bin_delta_sign_CW_flag[i] specifies the sign of resharper_model_bin_delta_abs_CW[i].
[0119] In the second embodiment, one or more of the foregoing aspects may be implemented as follows. Exemplary SPS syntax and semantics are as follows.
Table 11
[0120] The variables ALFEnabled and ReshaperEnabled are set as follows. ALFEnabled = sps_alf_enabled_flag and ReshaperEnabled = sps_reshaper_enabled_flag.
[0121] Exemplary APS syntax and semantics are as follows. [Table 12]
[0122] aps_params_type specifies the type of APS parameters carried in APS as specified in the following table. [Table 13]
[0123] When it does not exist, the value of aps_params_type is estimated as follows. When ALFEnabled, aps_params_type is set equal to 0. Otherwise, aps_params_type is set equal to 1.
[0124] In the second embodiment, one or more of the foregoing aspects may be implemented as follows. Exemplary SPS syntax and semantics are as follows. [Table 14]
[0125] adaption_parameter_set_id provides an identifier for APS for reference by other syntax elements. APS can be shared between pictures and can be different in different tile groups within a picture. The values and descriptions of the variable APSParamsType are defined in the following table. [Table 15]
[0126] Exemplary tile group header semantics are as follows. tile_group_alf_aps_id specifies the adaptation_parameter_set_id of the ALF APS that the tile group refers to. The TemporalId of the ALF APS NAL unit having an adapdation_parameter_set_id equal to tile_group_alf_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. The value of tile_group_alf_aps_id shall be in the range of 0 to 63 (both ends inclusive).
[0127] When multiple ALF APSs having the same value of adaptation_parameter_set_id are referred to by multiple tile groups of the same picture, the multiple ALF APSs having the same value of adapation_parameter_set_id shall have the same content. tile_group_reshaper_aps_id specifies the adaptation_parameter_set_id of the reshaper APS that the tile group refers to. The TemporalId of the reshaper APS NAL unit having an adaptation_parameter_set_id equal to tile_group_reshaper_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. The value of tile_group_reshaper_aps_id shall be in the range of 64 to 95 (both ends inclusive). When multiple reshaper APSs having the same value of adapation_parameter_set_id are referred to by two or more tile groups of the same picture, the multiple reshaper APSs having the same value of adapation_parameter_set_id shall have the same content.
[0128] FIG. 7 is a schematic diagram of an exemplary video coding device 700. The video coding device 700 is suitable for implementing the disclosed embodiments / implementations as described herein. The video coding device 700 includes a transceiver unit (Tx / Rx) 710 that includes a downstream port 720, an upstream port 750, and / or a transmitter and / or receiver for communicating data upstream and / or downstream via a network. The video coding device 700 also includes a processor 730 that includes a logic unit and / or a central processing unit (CPU) for processing data, and a memory 732 for storing data. The video coding device 700 may also include electrical, optical - electrical (OE) components, electro - optical (EO) components, and / or wireless communication components coupled to the upstream port 750 and / or the downstream port 720 for communication of data via an electrical, optical, or wireless communication network. The video coding device 700 may also include an input and / or output (I / O) device 760 for communicating data with a user. The I / O device 760 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The I / O device 760 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.
[0129] Processor 730 is implemented by hardware and software. Processor 730 can be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). Processor 730 communicates with downstream port 720, Tx / Rx 710, upstream port 750, and memory 732. Processor 730 includes coding module 714. Coding module 714 implements the disclosed embodiments described herein, such as methods 100, 800, and 900, which can use bitstream 500 and / or mechanism 600. Coding module 714 can also implement any other method / mechanism described herein. Further, coding module 714 may implement codec system 200, encoder 300, and / or decoder 400. For example, coding module 714 can encode / decode pictures in the bitstream and encode / decode parameters related to slices of pictures in multiple APSs. In some examples, different types of parameters can be coded into different types of APSs. Further, different types of APSs can be included in different types of NAL units. Such APS types can include ALF APS, scaling list APS, and / or LMCS APS. Each APS can include an APS ID. The APS IDs of different APS types can increase sequentially over different value spaces. Also, a slice and / or picture can refer to the corresponding slice header and / or picture header. Such headers can then refer to an APS that includes the relevant coding tools. Such an APS can be uniquely referenced by the APS ID and APS type. Such examples reduce the redundant signaling of coding tool parameters and / or reduce the bit usage for identifiers.Therefore, when coding the video data, the coding module 714 enables the video coding device 700 to have additional functionality and / or coding efficiency. Therefore, the coding module 714 improves the functionality of the video coding device 700 and addresses issues specific to video coding techniques. Further, the coding module 714 performs conversions of the video coding device 700 to different states. Alternatively, the coding module 714 is implemented as instructions stored in the memory 732 and executed by the processor 730 (e.g., as a computer program product stored on a non-transitory medium).
[0130] The memory 732 includes one or more memory types such as a disk, a tape drive, a solid state drive, a read only memory (ROM), a random access memory, a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM), etc. The memory 732 may be used as an overflow data storage device, storing programs when such programs are selected for execution and storing instructions and data read during program execution.
[0131] FIG. 8 is a flowchart of an exemplary method 800 for encoding a video sequence into a bitstream such as bitstream 500 by using multiple APS types such as ALF APS 512, scaling list APS 513, and / or LMCS APS 514. The method 800 may be used by an encoder such as codec system 200, encoder 300, and / or video coding device 700 when performing method 100. The method 800 may also assign APS IDs to different types of APS by using different value spaces according to mechanism 600.
[0132] Method 800 may start when an encoder receives a video sequence including a plurality of pictures and determines to encode the video sequence into a bitstream, for example, based on user input. The video sequence is partitioned into pictures / images / frames for further partitioning prior to encoding. In step 801, the encoder determines the luma mapping with chroma scaling (LMCS) parameters for application to slices. This may include using a rate-distortion (RDO) operation to encode the slices of the picture. For example, the encoder may encode the slices repeatedly multiple times using different coding options / coding tools, decode the coded slices, filter the decoded slices to increase the output quality of the slices, and then the encoder can select the coding option that provides the best balance of compression and output quality. Once the coding is selected, the encoder can determine the LMCS parameters (and any other optional filter parameters) used to filter the selected coding. In step 803, the encoder can encode the slices into a bitstream based on the selected coding.
[0133] In step 805, the LMCS parameters determined in step 801 are encoded into the bitstream in the LMCS APS. Further, data related to the slice referring to the LMCS APS can be encoded into the bitstream. For example, the data related to the slice can be a slice header and / or a picture header. In one example, the slice header may be encoded into the bitstream. A slice can refer to the slice header. And the slice header contains data related to the slice and refers to the LMCS APS. In another example, the picture header may be encoded into the bitstream. A picture containing the slice can refer to the picture header. And the picture header contains data related to the slice and refers to the LMCS APS. In any case, the header contains sufficient information to determine the appropriate LMCS APS including the LMCS parameters for the slice.
[0134] In step 807, other filtering parameters may be encoded into other APSs. For example, the ALF APS including the ALF parameters and the scaling list APS including the APS parameters can also be encoded into the bitstream. Such APSs may also be referred to by the slice header and / or the picture header.
[0135] As described above, each APS can be uniquely identified by a combination of parameter type and APS ID. Such information can be used by a slice header or a picture header to reference relevant APSs. For example, each APS may include an aps_params_type code set to a predefined value indicating the type of parameter included in the corresponding APS. Further, each APS may include an APS ID selected from a predefined range. The predefined range can be determined based on the parameter type of the corresponding APS. For example, the LMCS APS may have a range of 0 to 3 (2 bits), and the ALF APS and the scaling list APS may have a range of 0 to 7 (3 bits). Such ranges can describe different overlapping value spaces specific to the APS type, as described by mechanism 600. Thus, in steps 805 and 807, both the APS type and the APS ID can be used to reference a particular APS.
[0136] In step 809, the encoder encodes the SPS into a bitstream. The SPS may include a set of flags indicating that the LMCS is valid for an encoded video sequence including pictures / slices. The bitstream can then be stored in memory in step 811, for example, for communication towards a decoder via a transmitter. The bitstream includes sufficient information for the decoder to obtain the LMCS parameters for decoding an encoded slice by including a reference in the slice header / picture parameter header to the LMCS APS.
[0137] FIG. 9 is a flowchart of an exemplary method 900 for decoding a video sequence from a bitstream such as bitstream 500 by using a plurality of APS types such as ALF APS 512, scaling list APS 513, and / or LMCS APS 514. Method 900 may be used by a decoder such as codec system 200, decoder 400, and / or video coding device 700 when executing method 100. Also, method 900 may refer to the APS based on the APS ID assigned according to mechanism 600 in which different types of APS are assigned APS IDs according to different value spaces.
[0138] Method 900 may start when the decoder begins receiving a bitstream of coded data representing a video sequence, e.g., as a result of method 800. In step 901, the bitstream is received by the decoder. The bitstream includes pictures sliced into slices and an LMCS APS including LMCS parameters. In some examples, the bitstream may further include an ALF APS including ALF parameters and a scaling list APS including APS parameters. The bitstream may also include a picture header and / or a slice header respectively associated with one of the pictures and / or slices. The bitstream may also include other parameter sets such as SPS, PPS.
[0139] In step 903, the decoder may determine that the LMCS APS, the ALF APS, and / or the scaling list APS are referenced in the data related to the slice. For example, the slice header or the picture header may include data related to the slice / picture and may reference one or more APSs including the LMCS APS, the ALF APS, and / or the scaling list APS. Each APS is uniquely identified by a combination of a parameter type and an APS ID. Therefore, the parameter type and the APS ID may be included in the picture header and / or the slice header. For example, each APS may include an aps_params_type code set to a predefined value indicating the type of the parameters included in each APS. Further, each APS may include an APS ID selected from a predefined range. For example, the APS IDs of the APS types may be sequentially assigned over a plurality of different value spaces within a predefined range. For example, the predefined range may be determined based on the parameter type of the corresponding APS. As a specific example, the LMCS APS may have a range of 0 to 3 (2 bits), and the ALF APS and the scaling list APS may have a range of 0 to 7 (3 bits). Such ranges can describe different overlapping value spaces specific to the APS type, as described by mechanism 600. Therefore, both the APS type and the APS ID can be used by the picture header and / or the slice header to reference a specific APS.
[0140] In step 905, the decoder may decode a slice using the LMCS parameters from the LMCS APS based on the reference to the LMCS APS. The decoder may also decode a slice from the ALF APS and / or the scaling list APS using the ALF parameters and / or the scaling list parameters based on such a reference to the APS in the picture header / slice header. In some examples, the sequence parameter set (SPS) includes a flag set to indicate that the LMCS coding tool is valid for the encoded video sequence including the slice. In such a case, the LMCS parameters from the LMCS APS are obtained to support decoding based on the flag in step 905. In step 907, the decoder can transfer the slice for display as part of the decoded video sequence.
[0141] FIG. 10 is a schematic diagram of an exemplary system 1000 for coding a video sequence of an image in a bitstream such as bitstream 500 by using a plurality of APS types such as ALF APS 512, scaling list APS 513, and / or LMCS APS 514. The system 1000 may be implemented by an encoder and a decoder such as codec system 200, encoder 300, decoder 400, and / or video coding device 700. Further, the system 1000 may be used when implementing methods 100, 800, 900, and / or mechanism 600.
[0142] System 1000 includes a video encoder 1002. The video encoder 1002 includes a determination module 1001 for determining LMCS parameters for application to a slice. The video encoder 1002 further includes an encoding module 1003 for encoding the slice into a bitstream. The encoding module 1003 is further for encoding the LMCS parameters into the bitstream in LMCS APS. The encoding module 1003 is further for encoding data related to the slice that references LMCS APS into the bitstream. The video encoder 1002 further includes a storage module 1005 for storing the bitstream for communication to a decoder. The video encoder 1002 further includes a transmission module 1007 for transmitting the bitstream including LMCS APS and supporting decoding of the slice at the decoder. The video encoder 1002 may further be configured to perform any of the steps of method 800.
[0143] System 1000 also includes a video decoder 1010. The video decoder 1010 includes a reception module 1011 for receiving a bitstream including a slice and LMCS APS including LMCS parameters. The video decoder 1010 further includes a determination module 1013 for determining that the LMCS APS is referenced in data related to the slice. The video decoder 1010 further includes a decoding module 1015 for decoding the slice using the LMCS parameters from the LMCS APS based on the reference to the LMCS APS. The video decoder 1010 further includes a transfer module 1017 for transferring the slice for display as part of the decoded video sequence. The video decoder 1010 may further be configured to perform any of the steps of method 900.
[0144] The first component is directly coupled to the second component when there are no intervening components other than a line, trace, or other medium between the first component and the second component. The first component is indirectly coupled to the second component when there are intervening components other than a line, trace, or other medium between the first component and the second component. The terms "coupled" and its variations include both direct and indirect coupling. The use of the term "about" means a range that includes ±10% of the subsequent number, unless otherwise specified.
[0145] It should also be understood that the steps of the exemplary methods described herein need not necessarily be performed in the order described, and that the order of such method steps is merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in ways consistent with various embodiments of the present disclosure.
[0146] Although several embodiments have been provided in the present disclosure, it will be understood that the disclosed systems and methods may be implemented in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention is not to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0147] Furthermore, the techniques, systems, subsystems, and methods described and illustrated individually or separately in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and modifications are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. 1. A method implemented in a decoder, comprising: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transform operation; and obtaining the LMCS parameters from the LMCS APS, obtaining the ALF parameters from the ALF APS, and obtaining the scaling list parameters from the scaling list APS; and obtaining a decoded picture based on the coded slices, A method, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
2. 2. The method of claim 1, wherein each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters contained in each APS.
3. The method according to any one of claims 1 to 2, wherein each APS is identified by a combination of a parameter type and said APS ID.
4. 4. The method of claim 1, wherein the bitstream further includes a sequence parameter set (SPS) including a flag set to indicate that LMCS is valid for a coded video sequence that includes the coded slice, and the LMCS parameters from the LMCS APS are obtained based on the flag.
5. The method of any one of claims 1 to 4, wherein APSs of a particular type share the same value space for the APS ID, and APSs of different types use different value spaces for the APS ID.
6. 1. A method implemented in an encoder, comprising: determining luma mapping with chroma scaling (LMCS) parameters, adaptive loop filter (ALF) parameters, and scaling list parameters associated with a coded slice that includes the current block; obtaining a residual block based on the current block and the predicted block; performing a transform and quantization process on the residual block to obtain quantized coefficients; encoding the quantized coefficients into a bitstream; and encoding the LMCS parameters into an LMCS adaptation parameter set (APS) of the bitstream; encoding the ALF parameters into an ALF APS of the bitstream; encoding the scaling list parameters in a scaling list APS of the bitstream; A method, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
7. 7. The method of claim 6, wherein each APS includes an APS parameter type (aps_params_type) code set to a predefined value indicating the type of parameters contained in each APS.
8. The method according to any one of claims 6 to 7, wherein each APS is identified by a combination of a parameter type and said APS ID.
9. 9. The method of claim 6, further comprising: encoding, by the encoder, a sequence parameter set (SPS) into the bitstream, the SPS including a flag set to indicate that LMCS is valid for a coded video sequence that includes the coded slice.
10. The method of any one of claims 6 to 9, wherein APSs of a particular type share the same value space for the APS ID, and APSs of different types use different value spaces for the APS ID.
11. A video coding device comprising a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method of any one of claims 1 to 10.
12. A non-transitory computer readable medium comprising a computer program for use by a video coding device, the computer program comprising computer executable instructions stored on the non-transitory computer readable medium such that, when executed by a processor, the computer program causes the video coding device to perform a method according to any one of claims 1 to 10.
13. A decoder comprising processing circuitry for implementing the method according to any one of claims 1 to 5.
14. An encoder comprising processing circuitry for implementing the method according to any one of claims 6 to 10.
15. A computer program comprising a program code for performing the method according to any one of claims 1 to 10, when the computer program is run on a computer or processor.
16. A non-transitory computer readable medium carrying program code which, when executed by a computing device, causes said computing device to perform the method of any one of claims 1 to 10.
17. 1. A method for transmitting a bitstream, comprising: receiving, by at least one receiver, the bitstream; and storing the bitstream in at least one memory, the bitstream comprising: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transformation operation; the LMCS parameters are obtained from the LMCS APS, the ALF parameters are obtained from the ALF APS, and the scaling list parameters are obtained from the scaling list APS; A method, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
18. 1. A device for storing a bitstream, comprising: at least one receiver configured to receive said bitstream; at least one memory configured to store the bitstream, the bitstream comprising: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transformation operation; the LMCS parameters are obtained from the LMCS APS, the ALF parameters are obtained from the ALF APS, and the scaling list parameters are obtained from the scaling list APS; A device, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
19. 1. A device for transmitting a bitstream, comprising: at least one processor configured to obtain the bitstream; and at least one transmitter configured to transmit the bitstream, the bitstream comprising: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transformation operation; the LMCS parameters are obtained from the LMCS APS, the ALF parameters are obtained from the ALF APS, and the scaling list parameters are obtained from the scaling list APS; A device, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
20. 1. A method for transmitting a bitstream, comprising: obtaining, by at least one processor, the bitstream; and transmitting, by at least one transmitter, the bitstream, the bitstream comprising: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transformation operation; the LMCS parameters are obtained from the LMCS APS, the ALF parameters are obtained from the ALF APS, and the scaling list parameters are obtained from the scaling list APS; A method, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
21. 1. A system for processing a bitstream, comprising a source device, at least one storage medium, and a destination device, the system comprising: the source device is configured to provide a bitstream; the at least one storage medium is configured to store the bitstream; the destination device is adapted to decode the bitstream; The bitstream comprises: an LMCS adaptation parameter set (APS) including luma mapping with chroma scaling (LMCS) parameters associated with the coded slice; an adaptive loop filter (ALF APS) containing ALF parameters associated with the coded slice; a scaling list APS including scaling list parameters associated with the coded slice, the scaling list parameters being used for a transformation operation; the LMCS parameters are obtained from the LMCS APS, the ALF parameters are obtained from the ALF APS, and the scaling list parameters are obtained from the scaling list APS; A system, wherein each APS includes an APS identifier (ID) selected from a predefined range, the predefined range being determined based on a parameter type of the APS.
Citation Information
Patent Citations
Integrated image reshaping and video coding
WO2019006300A1
Video or image coding based on mapping of luma samples and scaling of chroma samples
WO2020262952A1
Scaling list data-based image or video coding
WO2021006630A1