Luma mapping and chroma scaling adaptive parameter sets in video coding

Separate adaptation parameter sets for ALF and LMCS parameters in video coding systems address inefficiencies by reducing redundant coding and resource usage, enhancing coding efficiency.

CN120321407APending Publication Date: 2025-07-15HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272022.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-05-21
Filing Date
2020-02-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing video decoding system is inefficient when carrying adaptive loop filter (ALF) parameters, and the redundant decoding of the reshaper/brightness mapping and chromaticity scaling (LMCS) parameters is severe, resulting in waste of network resources, memory resources and processing resources.

Method used

Multiple types of adaptive parameter sets (APS), including ALF APS, scale list APS and LMCS APS, are encoded as separate NAL units, and are identified by a combination of APS ID and parameter types to reduce redundant decoding.

Benefits of technology

It improves the decoding efficiency, reduces the use of network resources, memory resources and processing resources on the encoder and decoder side, and optimizes the video decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321407A_ABST
    Figure CN120321407A_ABST
Patent Text Reader

Abstract

The invention discloses a video coding mechanism. The present invention relates to a method and a system for transmitting and receiving a code stream, the method comprising: receiving a code stream, the code stream comprising a slice and a luma mapping and chroma scaling (LMCS) adaptive parameter set (APS), the LMCS APS comprising an LMCS parameter, and transmitting the code stream to the code stream, the code stream comprising the slice and the luma mapping and chroma scaling (LMCS) adaptive parameter set (APS). The mechanism also includes determining the LMCS APS that is referenced by related data of the slice. The mechanism also includes decoding the slice using LMCS parameters in the LMCS APS based on a reference to the LMCS APS. The mechanism also includes forwarding the slice for display as part of a decoded video sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202080017295.7, the original application date is February 26, 2020, and the entire content of the original application is incorporated herein by reference.

[0002] Cross - reference to related applications

[0003] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 811,358, filed on February 27, 2019, entitled "Adaptation Parameter Set for Video Coding" by Wang Yekui et al., U.S. Provisional Patent Application No. 62 / 816,753, filed on March 11, 2019, entitled "Adaptation Parameter Set for Video Coding" by Wang Yekui et al., and U.S. Provisional Patent Application No. 62 / 850,973, filed on May 21, 2019, entitled "Adaptation Parameter Set for Video Coding" by Wang Yekui et al., the entire disclosures of which are incorporated herein by reference. Technical field

[0004] The present invention generally relates to video coding, and more particularly, to an efficient indication of coding tool parameters for compressing video data in video coding. Background art

[0005] Even relatively short videos require a large amount of video data to describe, which can cause difficulties when the data is to be streamed or otherwise transmitted in a communication network with limited bandwidth capacity. Therefore, video data is usually compressed first and then transmitted through modern telecommunication networks. Also, when storing videos on storage devices, the size of the video can be an issue due to limited memory resources. Video compression devices typically use software and / or hardware on the source side to encode video data and then transmit or store it, thereby reducing the amount of data required to represent digital video images. Then, a video decompression device that decodes the video data receives the compressed data on the destination side. In the context of limited network resources and growing demand for higher video quality, improved compression and decompression techniques are needed that can increase the compression ratio with little impact on image quality. Summary of the invention

[0006] In one embodiment, the present invention includes a method implemented in a decoder, the method comprising: a receiver of the decoder receiving a bitstream, wherein the bitstream includes a luma mapping with chroma scaling (LMCS) adaptation parameter set (APS), and the LMCS APS includes LMCS parameters associated with an encoded slice; a processor determining the LMCS APS referenced (cited) by relevant data of the encoded slice; the processor decoding the encoded slice using the LMCS parameters in the LMCS APS; the processor forwarding the decoding result for display as part of a decoded video sequence. The APS is used to maintain data related to multiple slices on multiple images. The present invention describes various improvements related to the APS. In this example, the LMCS parameters are included in the LMCS APS. The LMCS / resampler parameters can be changed once per second. The video sequence can display 30 to 60 images per second. Therefore, the LMCS parameters may not change within 30 to 60 frames. Including the LMCS parameters in the LMCS APS instead of an image-level parameter set significantly reduces the redundant decoding of the LMCS parameters (e.g., reduced by 30 to 60 times). The slice header and / or the picture header associated with the slice can reference the relevant LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters of the slice change. Therefore, using the LMCS APS to encode the LMCS parameters improves the decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on the encoder side and the decoder side.

[0007] Optionally, according to any of the above aspects, in another implementation of the aspect, the bitstream further includes a slice header, the encoded slice references the slice header, and the slice header contains the data related to the encoded slice and references the LMCS APS.

[0008] Optionally, according to any of the above aspects, in another implementation of the aspect, the bitstream further includes an adaptive loop filter (ALF) APS and a scaling list APS, wherein the ALF APS contains ALF parameters and the scaling list APS contains APS parameters.

[0009] Optionally, according to any one of the above aspects, in another implementation of the aspect, each APS includes an APS parameter type (aps_params_type) code, and the APS parameter type (aps_params_type) code is set to a predefined value, and the predefined value represents the type of parameters included in each APS.

[0010] Optionally, according to any one of the above aspects, in another implementation of the aspect, each APS includes an APS identifier (identifier, ID) selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.

[0011] Optionally, according to any one of the above aspects, in another implementation of the aspect, each APS is identified by a combination of a parameter type and an APS ID.

[0012] Optionally, according to any one of the above aspects, in another implementation of the aspect, the bitstream further includes a sequence parameter set (sequence parameter set, SPS), the sequence parameter set includes a flag, the flag is set to indicate that LMCS has been enabled for the encoded video sequence including the encoded slice, and the LMCS parameters in the LMCS APS are obtained based on the flag.

[0013] In one embodiment, the present invention includes a method implemented in an encoder, the method comprising: a processor of the encoder determining LMCS parameters applied to a slice; the processor encoding the slice into a bitstream as an encoded slice; the processor encoding the LMCS parameters into an LMCS APS in the bitstream; the processor encoding data associated with the encoded slice into the bitstream, wherein the encoded slice refers to (quotes) the LMCS APS; a memory coupled to the processor storing the bitstream for transmission to a decoder. The APS is used to maintain data associated with slices on multiple images. The present invention describes various improvements related to the APS. In this example, the LMCS parameters are included in the LMCS APS. The LMCS / resampler parameters can be changed once per second. A video sequence can display 30 to 60 images per second. Therefore, the LMCS parameters may not change within 30 to 60 frames. Including the LMCS parameters in the LMCS APS instead of a set of picture-level parameters significantly reduces the redundant decoding of the LMCS parameters (e.g., by 30 to 60 times). The slice header and / or picture header associated with the slice can refer to the relevant LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters of the slice change. Therefore, using the LMCS APS to encode the LMCS parameters improves the decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on both the encoder side and the decoder side.

[0014] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further comprises: the processor encoding a slice header into the bitstream, wherein the encoded slice refers to the slice header, and the slice header contains the data associated with the encoded slice and refers to the LMCS APS.

[0015] Optionally, according to any of the above aspects, in another implementation of the aspect, the method further comprises: the processor encoding an ALF APS and a scaling list APS into the bitstream, wherein the ALF APS contains ALF parameters and the scaling list APS contains APS parameters.

[0016] Optionally, according to any of the above aspects, in another implementation of the aspect, each APS includes an aps_params_type code, and the aps_params_type code is set to a predefined value, and the predefined value represents the type of parameters included in each APS.

[0017] Optionally, according to any one of the above aspects, in another implementation of the above aspect, each APS includes an APS ID selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.

[0018] Optionally, according to any one of the above aspects, in another implementation of the above aspect, each APS is identified by a combination of a parameter type and an APS ID.

[0019] Optionally, according to any one of the above aspects, in another implementation of the above aspect, the method further includes: the processor encodes the SPS into the bitstream, where the SPS includes a flag that is set to indicate that LMCS has been enabled for the encoded video sequence including the encoded slice.

[0020] In one embodiment, the present invention includes a video decoding device, the video decoding device including: a processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, where the processor, the receiver, the memory, and the transmitter are configured to perform the method according to any one of the above aspects.

[0021] In one embodiment, the present invention includes a non-transitory computer-readable medium, the non-transitory computer-readable medium including a computer program product for use by a video decoding device, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium, and when the processor executes the computer-executable instructions, causing the video decoding device to perform the method according to any one of the above aspects.

[0022] In one embodiment, the present invention includes a decoder, the decoder including: a receiving module for receiving a bitstream, where the bitstream includes an LMCS APS, and the LMCS APS includes LMCS parameters associated with an encoded slice; a determining module for determining the LMCS APS referred to by relevant data of the encoded slice; a decoding module for decoding the encoded slice using the LMCS parameters in the LMCS APS; and a forwarding module for forwarding the decoding result for display as part of a decoded video sequence.

[0023] Optionally, according to any one of the above aspects, in another implementation of the above aspect, the decoder is further configured to perform the method according to any one of the above aspects.

[0024] In one embodiment, the present invention includes an encoder, the encoder comprising: a determination module configured to determine LMCS parameters applied to a strip; an encoding module configured to: encode the strip into a bitstream as an encoded strip; encode the LMCS parameters into an LMCS adaptation parameter set (APS) in the bitstream; encode data associated with the encoded strip into the bitstream, wherein the encoded strip refers to the LMCS APS; and a storage module configured to store the bitstream for transmission to a decoder.

[0025] Optionally, according to any one of the above aspects, in another implementation of the aspect, the encoder is further configured to perform the method according to any one of the above aspects.

[0026] For clarity, any one of the above embodiments may be combined with any one or more of the other above embodiments to create new embodiments within the scope of the present invention.

[0027] These and other features can be more clearly understood from the following detailed description in conjunction with the accompanying drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more fully understand the present invention, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like components.

[0029] Figure 1 is a flowchart of an exemplary method for decoding a video signal;

[0030] Figure 2 is a schematic diagram of an exemplary encoding and decoding (codec) system for video decoding;

[0031] Figure 3 is a schematic diagram of an exemplary video encoder;

[0032] Figure 4 is a schematic diagram of an exemplary video decoder;

[0033] Figure 5 is a schematic diagram of an exemplary bitstream including multiple types of adaptation parameter sets (APS) including different types of decoding tool parameters;

[0034] Figure 6 is a schematic diagram of an exemplary mechanism for allocating APS identifiers (IDs) for different APS types on different value spaces;

[0035] Figure 7 is a schematic diagram of an exemplary video decoding device;

[0036] Figure 8 is a flowchart of an exemplary method for encoding a video sequence into a bitstream using multiple APS types;

[0037] Figure 9 is a flowchart of an exemplary method for decoding a video sequence from a bitstream using multiple APS types;

[0038] Figure 10 is a schematic diagram of an exemplary system for decoding a video sequence of an image in a bitstream using multiple APS types. Detailed implementation manners

[0039] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or existing. The present invention should in no way be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations illustrated and described herein, but can be modified within the full scope of the appended claims and their equivalents.

[0040] The following abbreviations are used in this document: Adaptive Loop Filter (ALF), Adaptation Parameter Set (APS), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coded Video Sequence (CVS), Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH), Intra-Random Access Point (IRAP), Joint Video Experts Team (JVET), Motion-Constrained Tile Set (MCTS), Maximum Transfer Unit (MTU), Network Abstraction Layer (NAL), Picture Order Count (POC), Raw Byte Sequence Payload (RBSP), Sample Adaptive Offset (SAO), Sequence Parameter Set (SPS), Versatile Video Coding (VVC), and Working Draft (WD).

[0041] Many video compression techniques can be used to reduce the size of video files with minimal data loss. For example, video compression techniques can include performing spatial (e.g., intra-frame) prediction and / or temporal (e.g., inter-frame) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video image or a portion of a video image) can be partitioned into video blocks, which may also be referred to as treeblocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice in an image are coded using spatial prediction relative to reference samples in adjacent blocks in the same image, while video blocks in an inter-coded unidirectional prediction (P) or bidirectional prediction (B) slice in an image can be coded using spatial prediction relative to reference samples in adjacent blocks in the same image or can be coded using temporal prediction relative to reference samples in other reference images. An image can be referred to as a frame, and a reference image can be referred to as a reference frame. Spatial prediction or temporal prediction yields a predicted block representing an image block. Residual data represents the pixel difference between the original image block and the predicted block. Thus, an inter-coded block is coded based on a motion vector and residual data, where the motion vector points to the block of reference samples that make up the predicted block and the residual data represents the difference between the coded block and the predicted block; while an intra-coded block is coded based on an intra-coding mode and residual data. For further compression, the residual data can be transformed from the pixel domain to the transform domain. These result in residual transform coefficients that can be quantized. The quantized transform coefficients are initially arranged in a two-dimensional array. The quantized transform coefficients can be scanned to produce a one-dimensional vector of transform coefficients. Entropy coding can be used to achieve further compression. Such video compression techniques are discussed in more detail below.

[0042] To ensure that an encoded video can be correctly decoded, the video is encoded and decoded according to the corresponding video coding standard. Video coding standards include International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, Advanced Video Coding (AVC) (also known as ITU-T H.264 or ISO / IEC MPEG-4 Part 10), and High Efficiency Video Coding (HEVC) (also known as ITU-T H.265 or MPEG-H Part 2). AVC includes extended versions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), Multiview Video Coding plus Depth (MVC+D), and three dimensional (3D) AVC (3D-AVC). HEVC includes extended versions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The joint video experts team (JVET) of ITU-T and ISO / IEC has started developing a video coding standard called Versatile Video Coding (VVC). VVC is included in the working draft (WD), which includes JVET-M1001-v5 and JVET-M1002-v1, providing algorithm descriptions, an encoding-end description of the VVC WD, and reference software.

[0043] Video sequences are decoded by using various decoding tools. The encoder selects parameters for the decoding tools, aiming to increase compression with minimal quality loss when decoding the video sequence. The decoding tools can be related to different parts of the video in different scopes. For example, some decoding tools are related at the video sequence level, some are related at the picture level, some are related at the slice level, etc. APS can be used to indicate information that can be shared by multiple pictures and / or multiple slices across different pictures. Specifically, APS can carry adaptive loop filter (ALF) parameters. For various reasons, the ALF information may not be suitable for being indicated at the sequence level in the sequence parameter set (SPS), at the picture level in the picture parameter set (PPS) or in the picture header, or at the slice level in the tile group / slice header.

[0044] If the ALF information is indicated in the SPS, whenever the ALF information changes, the encoder must generate a new SPS and a new IRAP picture. IRAP pictures significantly reduce the decoding efficiency. Therefore, for low-latency application scenarios where IRAP pictures are not frequently used, placing the ALF information in the SPS is even more of a problem. Including the ALF information in the SPS may also disable the out-of-band transmission of the SPS. Out-of-band transmission refers to transmitting the corresponding data in a transmission data stream different from the video bitstream (e.g., in the sample description or sample entry of a media file, in a session description protocol (SDP) file, etc.). For similar reasons, indicating the ALF information in the PPS may also be problematic. Specifically, including the ALF information in the PPS may disable the out-of-band transmission of the PPS. Indicating the ALF information in the picture header may also be problematic. In some cases, the picture header may not be used. Additionally, some ALF information can be applied to multiple pictures. Therefore, indicating the ALF information in the picture header will result in redundant information transmission, thus wasting bandwidth. Indicating the ALF information in the tile group / slice header is also problematic because the ALF information can be applied to multiple pictures and thus can also be applied to multiple slices / tile groups. Therefore, indicating the ALF information in the slice / tile group header will result in redundant information transmission, thus wasting bandwidth.

[0045] Based on the above, APS can be used to indicate the ALF parameters. However, a video decoding system can use an APS specifically for indicating the ALF parameters. Examples of the APS syntax and semantics are as follows:

[0046]

[0047]

[0048] The adaptation_parameter_set_id provides an identifier for the APS for reference by other syntax elements. The APS can be shared across images and can be different in different tile groups within an image. The aps_extension_flag is set to 0 to indicate that the aps_extension_data_flag syntax element does not exist in the APS RBSP syntax structure. The aps_extension_flag is set to 1 to indicate that the aps_extension_data_flag syntax element exists in the APS RBSP syntax structure. The aps_extension_data_flag can take any value. The presence and value of the aps_extension_data_flag may not affect the decoder's compliance with the specified profile in VVC. A VVC-compliant decoder may ignore all aps_extension_data_flag syntax elements.

[0049] An example of the tile group header syntax related to the ALF parameters is as follows:

[0050]

[0051] The tile_group_alf_enabled_flag is set to 1 to indicate that the adaptive loop filter is enabled and can be applied to the luma (Y), blue chroma (Cb), or red chroma (Cr) color components in the tile group. The tile_group_alf_enabled_flag is set to 0 to indicate that the adaptive loop filter is disabled for all color components in the tile group. The tile_group_aps_id represents the adaptation_parameter_set_id of the APS referenced by the tile group. The TemporalId of the APS NAL unit with an adaptation_parameter_set_id equal to the tile_group_aps_id should be less than or equal to the TemporalId of the decoded tile group NAL unit. When multiple APSs with the same adaptation_parameter_set_id value are referenced by two or more tile groups of the same image, the multiple APSs with the same adaptation_parameter_set_id value can include the same content.

[0052] The reshaper parameters are parameters for the adaptive in-loop reshaper video coding tool, which is also called luma mapping with chroma scaling (LMCS). An example of the SPS reshaper syntax and semantics is as follows:

[0053]

[0054] sps_reshaper_enabled_flag is set to 1 to indicate that the reshaper is used in a coded video sequence (CVS). sps_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not used in a CVS.

[0055] The following are examples of the block group header / strip header reshaper syntax and semantics:

[0056]

[0057] tile_group_reshaper_model_present_flag is set to 1 to indicate that tile_group_reshaper_model() is present in the tile group header. tile_group_reshaper_model_present_flag is set to 0 to indicate that tile_group_reshaper_model() is not present in the tile group header. When tile_group_reshaper_model_present_flag is not present, the inferred flag is equal to 0. tile_group_reshaper_enabled_flag is set to 1 to indicate that the reshaper is enabled for the current tile group. tile_group_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not enabled for the current tile group. When tile_group_reshaper_enable_flag is not present, the inferred flag is equal to 0. tile_group_reshaper_chroma_residual_scale_flag is set to 1 to indicate that chroma residual scaling is enabled for the current tile group. tile_group_reshaper_chroma_residual_scale_flag is set to 0 to indicate that chroma residual scaling is not enabled for the current tile group. When tile_group_reshaper_chroma_residual_scale_flag is not present, the flag is inferred to be equal to 0.

[0058] Examples of the syntax and semantics of the chunk / group header / band header reshaper model are as follows:

[0059]

[0060]

[0061] reshape_model_min_bin_idx represents the minimum bit (or segment) index to be used during reshaper construction. The value of reshape_model_min_bin_idx can range from 0 to MaxBinIdx (including the end values). The value of MaxBinIdx can be equal to 15. reshape_model_delta_max_bin_idx represents the maximum bit (or segment) index MaxBinIdx minus the maximum bit index to be used during reshaper construction. The value of reshape_model_max_bin_idx is set to MaxBinIdx minus reshape_model_delta_max_bin_idx. reshaper_model_bin_delta_abs_cw_prec_minus1 + 1 represents the number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i]. reshape_model_bin_delta_abs_CW[i] represents the absolute incremental codeword value of the i-th bit.

[0062] reshaper_model_bin_delta_sign_CW_flag[i] represents the sign of reshape_model_bin_delta_abs_CW[i] as follows. If reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is positive. Otherwise (e.g., reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is negative. When reshape_model_bin_delta_sign_CW_flag[i] does not exist, the inferred flag is 0. The variable RspDeltaCW[i] is set to (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i].

[0063] The variable RspCW[i] is derived as follows: The variable OrgCW is set to (1<<BitDepthY) / (MaxBinIdx + 1). If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx, then RspCW[i] = OrgCW + RspDeltaCW[i]. Otherwise, RspCW[i] = 0. If the value of BitDepthY is equal to 10, the value of RspCW[i] should be in the range of 32 to 2*OrgCW - 1. The variable InputPivot[i] is derived as follows, where i is in the range of 0 to MaxBinIdx + 1 (including the end values). InputPivot[i] = i*OrgCW. The variable ReshapePivot[i] (where i is in the range of 0 to MaxBinIdx + 1, including the end values) and the variables ScaleCoef[i] and InvScaleCoeff[i] (where i is in the range of 0 to MaxBinIdx, including the end values) are derived as follows:

[0064]

[0065] The variable ChromaScaleCoef[i] is derived as follows, where i is in the range of 0 to MaxBinIdx (including the end values):

[0066] ChromaResidualScaleLut

[64] = {16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024};

[0067] shiftC = 11

[0068] if (RspCW[i] == 0)

[0069] ChromaScaleCoef[i] = (1 << shiftC)

[0070] Otherwise (RspCW[i] != 0),

[0071] ChromaScaleCoef[i] = ChromaResidualScaleLut[RspCW[i] >> 1]

[0072] The attributes of the resampler parameters can have the following characteristics. The size of the set of resampler parameters included in the tile_group_reshaper_model() syntax structure is typically around 60 to 100 bits. The resampler model is typically updated by the encoder approximately once per second, which includes many frames. In addition, the parameters of the updated resampler model are unlikely to be exactly the same as those of an earlier instance of the resampler model.

[0073] The above video decoding system has certain problems. First, such systems are only used to carry the ALF parameters in the APS. In addition, the resampler / LMCS parameters can be shared by multiple images and can include many variants.

[0074] This disclosure describes various mechanisms for modifying the APS to support increased coding efficiency. In a first example, various types of APS are disclosed. Specifically, an ALF type of APS (referred to as an ALF APS) can contain ALF parameters. In addition, a scaling list type of APS (referred to as a scaling list APS) can contain scaling list parameters. Additionally, an LMCS type of APS (referred to as an LMCS APS) can contain LMCS / resampler parameters. The ALF APS, scaling list APS, and LMCS APS can each be decoded as a separate NAL type and thus be included in different NAL units. In this way, changing the data in one type of APS (e.g., ALF parameters) does not result in redundant decoding of other types of data that have not changed (e.g., LMCS parameters). Therefore, providing multiple types of APS improves the decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on the encoder side and the decoder side.

[0075] In a second example, each APS includes an APS identifier (ID). Additionally, each APS type includes a separate value space for the corresponding APS ID. Such value spaces can overlap. Thus, a first type of APS (e.g., an ALF APS) can include the same APS ID as a second type of APS (e.g., an LMCS APS). This can be achieved by combining the APS parameter type and the APS ID to identify each APS. By allowing each APS type to include a different value space, the codec does not need to check for ID conflicts across APS types. Additionally, by allowing the value spaces to overlap, the codec can avoid using larger ID values, thus saving bits. Therefore, using separate overlapping value spaces for different types of APSs improves decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on both the encoder side and the decoder side.

[0076] In a third example, the LMCS parameters are included in the LMCS APS. As described above, the LMCS / resampler parameters can change once per second. A video sequence can display 30 to 60 images per second. Thus, the LMCS parameters may not change within 30 to 60 frames. Including the LMCS parameters in the LMCS APS significantly reduces the redundant decoding of the LMCS parameters. The strip header and / or picture header associated with the strip can reference the relevant LMCS APS. In this way, the LMCS parameters are encoded only when the LMCS parameters of the strip change. Therefore, using the LMCS APS to encode the LMCS parameters improves decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on both the encoder side and the decoder side.

[0077] Figure 1 A flowchart of an exemplary operation method 100 for decoding a video signal. Specifically, the video signal is encoded on the encoder side. During the encoding process, various mechanisms are employed to compress the video signal in order to reduce the size of the video file. With a smaller file, the compressed video file can be sent to the user while reducing the associated bandwidth overhead. Then, the decoder decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process is typically the inverse of the encoding process so that the video signal reconstructed by the decoder can be consistent with the video signal on the encoder side.

[0078] In step 101, a video signal is input into an encoder. For example, the video signal can be an uncompressed video file stored in a memory. Alternatively, the video file can be captured by a video capture device such as a camera and encoded to support live streaming of the video. The video file can include an audio component and a video component. The video component includes a series of image frames. When these image frames are viewed in sequence, they give a visual effect of motion. These frames include pixels represented by light, referred to herein as luminance components (or luminance samples), and also include pixels represented by color, referred to as chrominance components (or chrominance samples). In some examples, these frames can also include depth values to support three-dimensional viewing.

[0079] In step 103, the video is segmented into blocks. The segmentation includes subdividing the pixels in each frame into square and / or rectangular blocks for compression. For example, in high efficiency video coding (HEVC) (also known as H.265 and MPEG-H Part 2), the frame can first be partitioned into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). These CTUs include luminance samples and chrominance samples. The coding tree can be used to partition the CTUs into blocks, and then these blocks are repeatedly subdivided until a configuration that supports further coding is obtained. For example, the luminance component of the frame can be subdivided until each block includes relatively uniform luminance values. Additionally, the chrominance component of the frame can be subdivided until each block includes relatively uniform chrominance values. Thus, the segmentation mechanism varies depending on the content of the video frame.

[0080] In step 105, various compression mechanisms are employed to compress the image blocks obtained from the segmentation in step 103. For example, inter-frame prediction and / or intra-frame prediction can be used. Inter-frame prediction is designed to take advantage of the fact that objects in a scene often appear in consecutive frames. In this way, blocks that describe objects in a reference frame do not need to be repeatedly described in adjacent frames. Specifically, an object (e.g., a table) can remain in a fixed position in multiple frames. Therefore, the table is described once, and adjacent frames can refer back to the reference frame. A pattern matching mechanism can be used to match objects across multiple frames. Additionally, due to reasons such as object movement or camera movement, moving objects can be represented across multiple frames. In a specific example, a video can show a car moving across the screen over multiple frames. Motion vectors can be used to describe such movement. A motion vector is a two-dimensional vector that provides the offset of the coordinates of an object in one frame to the coordinates of that object in a reference frame. Thus, an image block in the current frame can be encoded as a set of motion vectors through inter-frame prediction, representing the offset between the image block in the current frame and the corresponding block in the reference frame.

[0081] Encode blocks in a common frame by intra prediction. Intra prediction exploits the fact that luminance and chrominance components tend to cluster in a frame. For example, a patch of green in a part of a tree tends to be adjacent to several similar patches of green. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a direct current (DC) mode. These directional modes indicate that the samples of the current block are similar / same as the samples of adjacent blocks in the corresponding direction. The planar mode indicates that a series of blocks in a row / column (e.g., a plane) can be interpolated based on adjacent blocks on the edge of that row. The planar mode actually represents a smooth transition of light / color across rows / columns by adopting a relatively constant slope of varying values. The DC mode is used for boundary smoothing and indicates that the block is similar / same as the average of the samples of all adjacent blocks that are related to the angular directions of the directional prediction modes. Thus, an intra prediction block can represent an image block as various relationship prediction mode values rather than as actual values. Additionally, an inter prediction block can represent an image block as motion vector values rather than as actual values. In either case, the prediction block may not accurately represent the image block in some cases. All differences are stored in a residual block. A transform can be applied to the residual block to further compress the file.

[0082] In step 107, various filtering techniques can be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in a blocky image at the decoder side. Additionally, the block-based prediction scheme can encode a block and then reconstruct the encoded block for subsequent use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and a sample adaptive offset (SAO) filter to blocks / frames. These filters reduce block artifacts so that the encoded file can be accurately reconstructed. Additionally, these filters reduce artifacts in the reconstructed reference blocks so that artifacts are less likely to generate other artifacts in subsequent blocks encoded based on the reconstructed reference blocks.

[0083] Once the video signal is segmented, compressed, and filtered, in step 109, the resulting data is encoded into a bitstream. The bitstream includes the data described above and any indication data needed to support proper video signal reconstruction at the decoder side. For example, this data can include segmentation data, prediction data, residual blocks, and various flags that provide decoding instructions to the decoder. The bitstream can be stored in a memory for sending to the decoder upon request. The bitstream can also be broadcast and / or multicast to multiple decoders. Creating the bitstream is an iterative process. Thus, step 101, step 103, step 105, step 107, and step 109 can be performed continuously and / or simultaneously over multiple frames and blocks. Figure 1The order shown is presented for clarity and ease of description and is not intended to limit the video decoding process to a particular order.

[0084] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax data and video data. In step 111, the decoder uses the syntax data in the bitstream to determine the partitioned parts of the frame. The partitioning should match the result of the block partitioning in step 103. The entropy encoding / decoding used in step 111 is described below. The encoder makes many choices during the compression process. For example, the encoder selects a block partitioning scheme from several possible choices based on the spatial placement of values in one or more input images. Indicating the exact choice may require a large number of bits. As used herein, a "bit" is a binary value that serves as a variable (e.g., a bit value that may vary depending on the content). Entropy encoding enables the encoder to discard any options that are clearly not suitable for a particular situation, leaving a set of available options. Then, a codeword is assigned to each available option. The length of the codeword depends on the number of available options (e.g., one bit corresponds to two options, two bits correspond to three or four options, etc.). Then, the encoder encodes the codewords of the selected options. This scheme shrinks the codewords because the codewords are as large as expected, uniquely indicating a selection from a small subset of allowable options rather than uniquely indicating a selection from a potentially large set of all possible options. Then, the decoder decodes this selection by determining the set of allowable options in a manner similar to the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the choice made by the encoder.

[0085] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct image blocks according to the partitioning. The prediction blocks can include intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. Then, the reconstructed image blocks are placed in the frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax for step 113 can also be indicated in the bitstream by the entropy encoding described above.

[0086] In step 115, the frame of the reconstructed video signal is filtered in a manner similar to step 107 on the encoder side. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter can be applied to the frame to remove blocking artifacts. Once the frame is filtered, in step 117, the video signal can be output to a display for viewing by the end user.

[0087] Figure 2Schematic diagram of an exemplary encoding and decoding (codec) system 200 for video decoding. Specifically, the codec system 200 provides functions to support the implementation of the operation method 100. The codec system 200 is generally used to describe components used on both the encoder and decoder sides. The codec system 200 receives a video signal and segments the video signal to obtain a segmented video signal 201, as described in connection with steps 101 and 103 in the operation method 100. Then, when acting as an encoder, the codec system 200 compresses the segmented video signal 201 into an encoded bitstream, as described in connection with steps 105, 107, and 109 in method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described in connection with steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general decoder control component 211, a transform scaling and quantization component 213, an intra prediction component 215, an intra prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded image buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In Figure 2 it, the black lines represent the movement of data to be encoded / decoded, and the dashed lines represent the movement of control data that controls the operation of other components. All components in the codec system 200 can be present in the encoder. The decoder can include a subset of the components in the codec system 200. For example, the decoder can include an intra prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded image buffer component 223. These components are described below.

[0088] The split video signal 201 is a captured video sequence that has been divided into pixel blocks by an encoding tree. The encoding tree uses various partitioning modes to break down the pixel blocks into smaller pixel blocks. These blocks can then be further broken down into smaller blocks. These blocks can be referred to as nodes on the encoding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / encoding tree. In some cases, the partitioned blocks are included in a coding unit (CU). For example, a CU can be a sub-part of a CTU, including a luminance block, one or more chrominance red difference (Cr) blocks, one or more chrominance blue difference (Cb) blocks, and the corresponding syntax instructions for the CU. The partitioning modes can include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to divide a node into two, three, or four child nodes of different shapes respectively according to the adopted partitioning mode. The split video signal 201 is forwarded to a general decoder control component 211, a transform scaling and quantization component 213, an intra prediction component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0089] The general decoder control component 211 is used to make decisions related to encoding images in a video sequence into a bitstream according to application constraints. For example, the general decoder control component 211 manages the optimization of the bitrate / bitstream size with respect to the reconstructed quality. These decisions can be made according to the storage space / bandwidth availability and the image resolution request. The general decoder control component 211 also manages the buffer utilization according to the transmission speed to alleviate buffer underflow and overflow problems. To solve these problems, the general decoder control component 211 manages the partitioning, prediction, and filtering performed by other components. For example, the general decoder control component 211 can dynamically increase the compression complexity to increase the resolution and bandwidth utilization, or reduce the compression complexity to reduce the resolution and bandwidth utilization. Therefore, the general decoder control component 211 controls other components in the codec system 200 to balance the video signal reconstruction quality and the bitrate issue. The general decoder control component 211 generates control data, which is used to control the operations of other components. The control data is also forwarded to a header format and CABAC component 231 for encoding into the bitstream to indicate the parameters used by the decoder for decoding.

[0090] The split video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter-frame prediction. The frames or strips of the split video signal 201 can be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-frame predictive coding on the received video blocks with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform multiple coding rounds to select a suitable coding mode for each video data block, and so on.

[0091] The motion estimation component 221 and the motion compensation component 219 can be highly integrated, but are described separately for conceptual purposes. Motion estimation performed by the motion estimation component 221 is a process of generating motion vectors, where these motion vectors are used to estimate the motion of video blocks. For example, a motion vector can represent the displacement of a decoded object relative to a predicted block. A predicted block is a block that is found to highly match the block to be coded in terms of pixel differences. A predicted block can also be referred to as a reference block. Such pixel differences can be determined by the sum of absolute difference (SAD), the sum of squared difference (SSD), or other difference metrics. HEVC employs several decoded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into multiple CTBs, and then the CTBs can be divided into CUs to be included in CUs. A CU can be coded as a prediction unit (PU) including prediction data and / or a transform unit (TU) including the transform residual data of the CU. The motion estimation component 221 uses rate-distortion analysis as part of the rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame, and can select the reference blocks, motion vectors, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance the quality of video reconstruction (e.g., the amount of data loss caused by compression) and coding efficiency (e.g., the final coded size).

[0092] In some examples, the codec system 200 may calculate values of sub-integer pixel positions of a reference image stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, the motion estimation component 221 may perform motion search relative to integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates the motion vectors of PUs of video blocks in the inter-coded slice by comparing the positions of the PUs with the positions of the predicted blocks of the reference image. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and CABAC component 231 for encoding and to the motion compensation component 219 as motion data.

[0093] The motion compensation performed by the motion compensation component 219 may involve obtaining or generating a predicted block according to the motion vectors determined by the motion estimation component 221. Similarly, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. After receiving the motion vectors of the PUs of the current video block, the motion compensation component 219 may locate the predicted block pointed to by the motion vectors. Then, the pixel values of the predicted block are subtracted from the pixel values of the current video block being decoded to obtain a pixel difference, thus forming a residual video block. Generally, the motion estimation component 221 performs motion estimation relative to the luminance component, while the motion compensation component 219 uses the motion vectors calculated according to the luminance component for both the chrominance component and the luminance component. The predicted block and the residual block are forwarded to the transform scaling and quantization component 213.

[0094] The segmented video signal 201 is also sent to the intra-estimation component 215 and the intra-prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra-estimation component 215 and the intra-prediction component 217 may be highly integrated, but are described separately for conceptual purposes. The intra-estimation component 215 and the intra-prediction component 217 perform intra-prediction on the current block relative to each block in the current frame to replace the inter-prediction performed between frames by the motion estimation component 221 and the motion compensation component 219 as described above. Specifically, the intra-estimation component 215 determines the intra-prediction mode to encode the current block. In some examples, the intra-estimation component 215 selects a suitable intra-prediction mode from multiple tested intra-prediction modes to encode the current block. Then, the selected intra-prediction mode is forwarded to the header format and CABAC component 231 for encoding.

[0095] For example, the intra prediction component 215 performs rate-distortion analysis on various tested intra prediction modes to calculate rate-distortion values, and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to produce the encoded block, and determines the bit rate (e.g., number of bits) used to produce the encoded block. The intra prediction component 215 calculates a ratio based on the distortion and rate of various encoded blocks to determine the intra prediction mode that exhibits the best rate-distortion value for the block. Additionally, the intra prediction component 215 can be used to decode depth blocks of a depth image using a depth modeling mode (DMM) according to rate-distortion optimization (RDO).

[0096] When implemented on the encoder, the intra prediction component 217 can generate a residual block from a prediction block according to the selected intra prediction mode determined by the intra prediction component 215, or when implemented on the decoder, it can read the residual block from the bitstream. The residual block includes the difference between the prediction block and the original block, represented as a matrix. Then, the residual block is forwarded to the transform scaling and quantization component 213. The intra prediction component 215 and the intra prediction component 217 can operate on the luminance component and the chrominance component.

[0097] The transform scaling and quantization component 213 is used to further compress the residual block. The transform scaling and quantization component 213 applies a transform such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform to the residual block, thereby generating a video block including residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used. The above-mentioned transforms can convert the residual information from the pixel domain to the transform domain (e.g., frequency domain). The transform scaling and quantization component 213 is also used to scale the transform residual information according to frequency, etc. This scaling involves applying a scaling factor to the residual information in order to quantify different frequency information at different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also used to quantize the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The quantization level can be modified by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 can then scan the matrix including the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the bitstream.

[0098] The scaling and inverse transform component 229 performs operations opposite to those of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, inverse transform, and / or inverse quantization to reconstruct the residual block in the pixel domain, e.g., for subsequent use as a reference block. This reference block can become the prediction block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block back to the corresponding prediction block for motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to reduce artifacts generated during scaling, quantization, and transformation. These artifacts may make the prediction inaccurate (and generate additional artifacts) when predicting subsequent blocks.

[0099] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding prediction block from the intra prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, a filter can be applied to the reconstructed image block. In some examples, the filter can instead be applied to the residual block. Similar to Figure 2 other components in

[0100] the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and can be implemented together, but are described separately for conceptual purposes. The filters applied to the reconstructed reference block are applied to specific spatial regions and these filters include multiple parameters to adjust the way these filters are used. The filter control analysis component 227 analyzes the reconstructed reference block to determine where these filters are needed and sets the corresponding parameters. This data is forwarded as filter control data to the header format and CABAC component 231 for encoding. The in-loop filter component 225 applies these filters according to the filter control data. These filters can include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. These filters can be applied in the spatial / pixel domain (e.g., for reconstructed pixel blocks) or in the frequency domain according to examples.

[0101] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode this data into the coded bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generate various headers to encode control data (such as general control data and filter control data). In addition, prediction data (including intra-prediction data and motion data) and residual data in the form of quantized transform coefficient data are both encoded into the bitstream. The final bitstream includes all the information required by the decoder to reconstruct the original segmented video signal 201. This information may also include an intra-prediction mode index table (also referred to as a codeword mapping table), the definition of the coding context of various blocks, an indication of the most likely intra-prediction mode, an indication of the segmentation information, etc. This data can be encoded using entropy coding. For example, this information can be encoded using context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. After entropy coding, the coded bitstream can be sent to another device (e.g., a video decoder) or archived for subsequent transmission or retrieval.

[0102] Figure 3 FIG. is a block diagram of an exemplary video encoder 300. The video encoder 300 can be used to implement the encoding function of the codec system 200 and / or perform the steps 101, 103, 105, 107, and / or 109 in the operation method 100. The encoder 300 segments the input video signal to obtain a segmented video signal 301 that is substantially similar to the segmented video signal 201. Then, the segmented video signal 301 is compressed by the components in the encoder 300 and encoded into the bitstream.

[0103] Specifically, the segmented video signal 301 is forwarded to the intra prediction component 317 for intra prediction. The intra prediction component 317 can be substantially similar to the intra estimation component 215 and the intra prediction component 217. The segmented video signal 301 is also forwarded to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 can be substantially similar to the motion estimation component 221 and the motion compensation component 219. The predicted blocks and residual blocks from the intra prediction component 317 and the motion compensation component 321 are forwarded to the transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 can be substantially similar to the transform scaling and quantization component 213. The transform quantized residual blocks and the corresponding predicted blocks (along with relevant control data) are forwarded to the entropy coding component 331 for encoding into the bitstream. The entropy coding component 331 can be substantially similar to the header format and CABAC component 231.

[0104] The transform quantized residual blocks and / or the corresponding predicted blocks are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 can be substantially similar to the scaling and inverse transform component 229. According to an example, the in-loop filter in the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks. The in-loop filter component 325 can be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 can include multiple filters as described in connection with the in-loop filter component 225. Then, the filtered blocks are stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 can be substantially similar to the decoded picture buffer component 223.

[0105] Figure 4 It is a block diagram of an exemplary video decoder 400. The video decoder 400 can be used to implement the decoding function of the codec system 200 and / or execute steps 111, 113, 115, and / or 117 in the operation method 100. The decoder 400 receives the bitstream from the encoder 300 etc. and generates a reconstructed output video signal based on the bitstream for display to the end user.

[0106] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is used to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE decoding, or other entropy decoding techniques. For example, the entropy decoding component 433 can use the header information to provide context for parsing additional data encoded as codewords in the bitstream. The decoded information includes any information required to decode the video signal, such as general control data, filter control data, segmentation information, motion data, prediction data, and quantized transform coefficients in the residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 to reconstruct the residual blocks. The inverse transform and quantization component 429 can be similar to the inverse transform and quantization component 329.

[0107] The reconstructed residual blocks and / or prediction blocks are forwarded to the intra prediction component 417 to be reconstructed as image blocks according to the intra prediction operation. The intra prediction component 417 can be similar to the intra estimation component 215 and the intra prediction component 217. Specifically, the intra prediction component 417 uses the prediction mode to locate the reference blocks in the frame and adds the residual blocks to the above results to reconstruct the intra prediction image blocks. The reconstructed intra prediction image blocks and / or residual blocks and the corresponding inter prediction data are forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 and the in-loop filter component 425 can be substantially similar to the decoded image buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image blocks, residual blocks, and / or prediction blocks. This information is stored in the decoded image buffer component 423. The reconstructed image blocks from the decoded image buffer component 423 are forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 can be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses the motion vectors of the reference blocks to generate the prediction blocks and applies the residual blocks to the above results to reconstruct the image blocks. The resulting reconstructed blocks can also be forwarded to the decoded image buffer component 423 through the in-loop filter component 425. The decoded image buffer component 423 continues to store other reconstructed image blocks. These reconstructed image blocks can be reconstructed into frames through the segmentation information. These frames can also be placed in a sequence. This sequence is output as the reconstructed output video signal to the display.

[0108] Figure 5 It is a schematic diagram of an exemplary bitstream 500 that includes multiple types of APSs including parameters of different types of decoding tools. For example, the bitstream 500 can be generated by the codec system 200 and / or the encoder 300 and decoded by the codec system 200 and / or the decoder 400. In another example, in step 109 of method 100, the bitstream 500 can be generated by the encoder and used by the decoder in step 111.

[0109] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPSs) 511, multiple ALF APSs 512, multiple scaling list APSs 513, multiple LMCS APSs 514, multiple slice headers 515, and picture data 520. The SPS 510 contains sequence data common to all pictures in the video sequence included in the bitstream 500. This data may include picture size, bit depth, decoding tool parameters, bitrate limits, etc. The PPS 511 contains parameters applied to the entire picture. Thus, each picture in the video sequence can refer to the PPS 511. It should be noted that although each picture refers to the PPS 511, in some examples, a single PPS 511 may contain data for multiple pictures. For example, multiple similar pictures can be decoded according to similar parameters. In this case, a single PPS 511 can contain data for such similar pictures. The PPS 511 can represent decoding tools, quantization parameters, offsets, etc. available for slices in the corresponding picture. The slice header 515 includes parameters specific to each slice in the picture. Thus, each slice in the video sequence can have a slice header 515. The slice header 515 can contain slice type information, picture order count (POC), reference picture list, prediction weights, block entry points, deblocking parameters, etc. It should be noted that in some contexts, the slice header 515 can also be referred to as a block group header.

[0110] An APS is a syntax structure that contains syntax elements applied to one or more pictures 521 and / or slices 523. In the example shown, the APS can be divided into multiple types. The ALF APS 512 is an ALF-type APS that includes ALF parameters. ALF is a block-based adaptive filter that includes a transfer function controlled by variable parameters and uses feedback from a feedback loop to correct the transfer function. In addition, ALF is used to correct decoding artifacts (e.g., errors) that occur due to block-based decoding. An adaptive filter is a linear filter that has a transfer function controlled by variable parameters, and the variable parameters can be controlled by an optimization algorithm, such as the RDO process running on the encoder side. Thus, the ALF parameters included in the ALF APS 512 can include variable parameters selected by the encoder to enable the filter to remove block-based decoding artifacts during decoding on the decoder side.

[0111] The Scaling List APS 513 is a Scaling List type of APS that includes Scaling List parameters. As described above, the current block is decoded according to inter - prediction or intra - prediction that generates a residual. The residual is the difference between the luminance and / or chrominance values of the block and the corresponding values of the predicted block. Then, a transform is applied to the residual to convert the residual into transform coefficients (which are smaller than the residual values). Encoding high - definition and / or ultra - high - definition content may result in an increase in residual data. When applied to such data, a simple transform process may generate a large amount of quantization noise. Therefore, the Scaling List parameters included in the Scaling List APS 513 may include weighting parameters that can be used to scale the transform matrix to account for changes in the display resolution in the resulting decoded video image and / or an acceptable level of quantization noise.

[0112] The LMCS APS 514 is an LMCS type of APS that includes LMCS parameters, which are also referred to as reshaper parameters. The human visual system has a lower ability to distinguish color differences (e.g., chrominance) than to distinguish light differences (e.g., luminance). Therefore, some video systems use a chroma subsampling mechanism to compress video data by reducing the resolution of chrominance values without adjusting the corresponding luminance values. One problem with this mechanism is that the associated interpolation may generate interpolated chrominance values during decoding that are incompatible with the corresponding luminance values at some positions. This will create color artifacts at these positions, which should be corrected by corresponding filters. The luminance mapping mechanism complicates this. Luminance mapping is the process of remapping the decoded luminance component within the dynamic range of the input luminance signal (e.g., according to a piece - wise linear function). This compresses the luminance component. The LMCS algorithm scales the compressed chrominance values according to the luminance mapping to remove the artifacts associated with chroma subsampling. Therefore, the LMCS parameters included in the LMCS APS 514 represent chroma scaling for considering luminance mapping. The LMCS parameters are determined by the encoder, and when luminance mapping is used, the decoder can use the LMCS parameters to filter out the artifacts caused by chroma subsampling.

[0113] The image data 520 includes video data encoded according to inter-frame prediction and / or intra-frame prediction, as well as corresponding transform and quantization residual data. For example, the video sequence includes a plurality of images 521 decoded as image data. The images 521 are individual frames of the video sequence and are typically displayed as a single unit when the video sequence is displayed. However, a portion of an image may be displayed to implement certain techniques such as virtual reality, picture-in-picture, etc. The images 521 each refer to the PPS 511. The images 521 are divided into strips 523, which may be defined as horizontal portions of the images 521. For example, a strip 523 may include a portion of the height of the image 521 and the full width of the image 521. In other cases, the image 521 may be divided into columns and rows, and a strip 523 may include a rectangular portion of the image 521 created by these columns and rows. In some systems, the strips 523 are subdivided into blocks. In other systems, the strips 523 are referred to as groups of blocks that contain blocks. The strips 523 and / or the groups of blocks refer to the strip headers 515. The strips 523 are further divided into coding tree units (CTUs). The CTUs are further divided into coding blocks according to a coding tree. The coding blocks can then be encoded / decoded according to a prediction mechanism.

[0114] The images 521 and / or the strips 523 may directly or indirectly refer to the ALF APS 512, the scaling list APS 513, and / or the LMCS APS 514 that contain relevant parameters. For example, a strip 523 may refer to the strip header 515. In addition, an image 521 may refer to a corresponding image header. The strip header 515 and / or the image header may refer to the ALF APS 512, the scaling list APS 513, and / or the LMCS APS 514 that contain the parameters used when decoding the relevant strip 523 and / or image 521. In this way, the decoder can obtain the decoding tool parameters related to the strip 523 and / or image 521 according to the header references related to the corresponding strip 523 and / or image 521.

[0115] The bitstream 500 is decoded into video coding layer (VCL) NAL units 535 and non-VCL NAL units 531. A NAL unit is a decoded data unit, and its size can be set to the payload of a single data packet for transmission over a network. The VCL NAL unit 535 is a NAL unit containing decoded video data. For example, each VCL NAL unit 535 may contain a slice 523, and / or a partition group of data, CTUs, and / or coded blocks. The non-VCL NAL unit 531 is a NAL unit containing syntax support but not decoded video data. For example, the non-VCL NAL unit 531 may contain SPS 510, PPS 511, APS, slice header 515, etc. Thus, the decoder receives the bitstream 500 in discrete VCL NAL units 535 and non-VCL NAL units 531. An access unit is a set of VCL NAL units 535 and / or non-VCL NAL units 531, including data sufficient to decode a single picture 521.

[0116] In some examples, the ALF APS 512, the scaling list APS 513, and the LMCS APS 514 are each assigned to a separate non-VCL NAL unit 531 type. In this case, the ALF APS 512, the scaling list APS 513, and the LMCS APS 514 are included in the ALF APS NAL unit 532, the scaling list APS NAL unit 533, and the LMCS APS NAL unit 534, respectively. Thus, the ALF APS NAL unit 532 contains ALF parameters that remain valid until another ALF APS NAL unit 532 is received. Further, the scaling list APS NAL unit 533 contains scaling list parameters that remain valid until another scaling list APS NAL unit 533 is received. Additionally, the LMCS APS NAL unit 534 contains LMCS parameters that remain valid until another LMCS APS NAL unit 534 is received. In this way, it is not necessary to send a new APS every time the APS parameters change. For example, a change in the LMCS parameters results in an additional LMCS APS 514, but not an additional ALF APS 512 or scaling list APS 513. Therefore, by dividing the APS into different NAL unit types according to the parameter type, redundant indication of irrelevant parameters is avoided. Thus, dividing the APS into different NAL unit types improves the decoding efficiency, thereby reducing the use of processor resources, memory resources, and / or network resources on the encoder side and the decoder side.

[0117] In addition, stripe 523 and / or image 521 may directly or indirectly refer to ALF APS 512, ALF APS NAL unit 532, scaling list APS 513, scaling list APS NAL unit 533, LMCS APS 514, and / or LMCS APS NAL unit 534 that contain decoding tool parameters for decoding stripe 523 and / or image 521. For example, each APS may contain an APS ID 542 and a parameter type 541. The APS ID 542 is a value (e.g., a number) that identifies the corresponding APS. The APS ID 542 may contain a predefined number of bits. Thus, the APS ID 542 may increment (e.g., increment by 1) according to a predefined sequence, and once the sequence reaches the end of the predefined range, the APS ID 542 may be reset to the minimum value (e.g., 0). The parameter type 541 indicates the type of parameter contained in the APS (e.g., ALF, scaling list, and / or LMCS). For example, the parameter type 541 may include an APS parameter type (aps_params_type) code that is set to a predefined value that represents the type of parameter included in each APS. Thus, the parameter type 541 may be used to distinguish between ALF APS 512, scaling list APS 513, and LMCS APS 514. In some examples, ALF APS 512, scaling list APS 513, and LMCS APS 514 may each be uniquely identified by combining the parameter type 541 and the APS ID 542. For example, each APS type may include a separate value space for the corresponding APS ID 542. Thus, each APS type may include an APS ID 542 that increments sequentially from the previous APS of the same type. However, the APS ID 542 of the first APS type may not be related to the APS ID 542 of the previous APS of a different second APS type. Thus, the APS ID 542 of different APS types may include overlapping value spaces. For example, in some cases, an APS of the first type (e.g., ALF APS) may include the same APS ID 542 as an APS of the second type (e.g., LMCS APS). By allowing each APS type to include a different value space, the codec does not need to check for APS ID 542 conflicts across APS types. In addition, by allowing the value spaces to overlap, the codec can avoid using larger APS ID 542 values, thereby saving bits. Thus, using separate overlapping value spaces for the APS ID 542 of different APS types improves coding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on the encoder side and the decoder side.As described above, the APS ID 542 can be extended within a predefined range. In some examples, the predefined range of the APS ID 542 can vary according to the APS type indicated by the parameter type 541. This can enable different numbers of bits to be allocated to different APS types according to the change frequency of different types of parameters under normal circumstances. For example, the range of the APS ID 542 for the ALF APS 512 can be from 0 to 7, the range of the APS ID 542 for the scaling list APS 513 can be from 0 to 7, and the range of the APS ID 542 for the LMCS APS 514 can be from 0 to 3.

[0118] In another example, the LMCS parameters are included in the LMCS APS 514. Some systems include the LMCS parameters in the strip header 515. However, the LMCS / resampler parameters can change once per second. The video sequence can display 30 to 60 images 521 per second. Therefore, the LMCS parameters may not change within 30 to 60 frames. Including the LMCS parameters in the LMCS APS 514 significantly reduces the redundant decoding of the LMCS parameters. In some examples, the strip header 515 and / or the picture header associated with the strip 523 and / or the picture 521, respectively, can refer to the relevant LMCS APS 514. Then, the strip 523 and / or the picture 521 refer to the strip header 515 and / or the picture header. This allows the decoder to obtain the LMCS parameters of the relevant strip 523 and / or picture 521. In this way, the LMCS parameters are encoded only when the LMCS parameters of the strip 523 and / or the picture 521 change. Therefore, using the LMCS APS 514 to encode the LMCS parameters improves the decoding efficiency, thereby reducing the use of network resources, memory resources, and / or processing resources on the encoder side and the decoder side. Since LMCS is not used for all videos, the SPS 510 can include an LMCS enable flag 543. The LMCS enable flag 543 can be set to indicate that LMCS has been enabled for the encoded video sequence. Therefore, when the LMCS enable flag 543 is set (e.g., to 1), the decoder can obtain the LMCS parameters from the LMCS APS 514 according to the LMCS enable flag 543. In addition, when the LMCS enable flag 543 is not set (e.g., to 0), the decoder can refrain from attempting to obtain the LMCS parameters.

[0119] Figure 6FIG. 600 is a schematic illustration of an exemplary mechanism 600 for allocating APS ID 642 to different APS types on different value spaces. For example, mechanism 600 may be applied to bitstream 500 to allocate APS ID 542 to ALF APS 512, scaling list APS 513, and / or LMCS APS 514. Additionally, when decoding video according to method 100, mechanism 600 may be applied to codec 200, encoder 300, and / or decoder 400.

[0120] Mechanism 600 allocates APS ID 642 to ALF APS 612, scaling list APS 613, and LMCS APS 614, where APS ID 642, ALF APS 612, scaling list APS 613, and LMCS APS 614 may be substantially similar to APS ID 542, ALF APS 512, scaling list APS 513, and LMCS APS 514, respectively. As described above, APS ID 642 may be sequentially allocated on multiple different value spaces, where each value space is specific to an APS type. Additionally, each value space may extend over a different range specific to the APS type. In the example shown, the value space for APS ID 642 of ALF APS 612 ranges from 0 to 7 (e.g., 3 bits). Additionally, the value space for APS ID 642 of scaling list APS 613 ranges from 0 to 7 (e.g., 3 bits). Additionally, the value space for APS ID 642 of LMCS APS 611 ranges from 0 to 3 (e.g., 2 bits). When APS ID 642 reaches the end of the value space range, the APS ID 642 of the next APS of the corresponding type returns to the start of the range (e.g., 0). When a new APS receives the same APS ID 642 as the previous APS of the same type, the previous APS is no longer active and can no longer be referenced. Thus, the range of the value space may be extended to allow more APSs of a type to remain active and be referenced. The range of the value space may be decreased to improve decoding efficiency, but such a decrease also reduces the number of APSs of the corresponding type that can remain active and be available for reference simultaneously.

[0121] In the example shown, each of the ALF APS 612, the Scaling List APS 613, and the LMCS APS 614 is referenced by combining the APS ID 642 and the APS type. For example, the ALF APS 612, the LMCS APS 614, and the Scaling List APS 613 each receive an APS ID 642 of 0. When a new ALF APS 612 is received, the APS ID 642 is incremented from the value used for the previous ALF APS 612. The same sequence applies to the Scaling List APS 613 and the LMCS APS 614. Thus, each APS ID 642 is related to the APS ID 642 of the previous APS of the same type. However, the APS ID 642 is not related to the APS ID 642 of the previous APS of other types. In this example, the APS ID 642 of the ALF APS 612 increments from 0 to 7 and then returns to 0 before continuing to increment. Additionally, the APS ID 642 of the Scaling List APS 613 increments from 0 to 7 and then returns to 0 before continuing to increment. Additionally, the APS ID 642 of the LMCS APS 611 increments from 0 to 3 and then returns to 0 before continuing to increment. As shown, these value spaces overlap because different APSs of different APS types can share the same APS ID 642 at the same point in the video sequence. It should be noted that the mechanism 600 only describes the APS. In the bitstream, the described APSs will be scattered among other VCL and non-VCL NAL units such as SPS, PPS, slice headers, picture headers, slices, etc.

[0122] Accordingly, the present invention includes improvements to the APS design and some improvements to the indication of the resampler / LMCS parameters. The APS is designed to indicate information that can be shared by multiple pictures and can include many variants. The resampler / LMCS parameters are for the adaptive in-loop resampler / LMCS video decoding tool. The above mechanism can be implemented as follows. To solve the problems listed herein, several aspects that can be used alone and / or in combination are included here.

[0123] Modify the disclosed APS such that multiple APSs can be used to carry different types of parameters. Each APS NAL unit is only used to carry one type of parameter. Thus, when two types of information (e.g., one for each type of information) are to be carried for a particular block group / slice, two APS NAL units are encoded. The APS can include an APS parameter type field in the APS syntax. Only the parameters of the type indicated by the APS parameter type field can be included in the APS NAL unit.

[0124] In some examples, different types of APS parameters are represented by different NAL unit types. For example, two different NAL unit types are used for APS. These two types of APS can be referred to as ALF APS and resampler APS respectively. In another example, the tool parameter type carried in the APS NAL unit is specified in the NAL unit header. In VVC, the NAL unit header has reserved bits (e.g., 7 bits represented as nuh_reserved_zero_7bits). In some examples, some of these bits (e.g., 3 out of 7 bits) can be used to specify the APS parameter type field. In some examples, a specific type of APS can share the same value space for the APS ID. At the same time, different types of APS use different value spaces for the APS ID. Therefore, two APS of different types can coexist and have the same APS ID value at the same time. In addition, the combination of the APS ID and the APS parameter type can be used to identify an APS among multiple other APSs.

[0125] When the corresponding decoding tool is enabled for a tile group, the APS ID can be included in the tile group header syntax. Otherwise, the APS ID of the corresponding type may not be included in the tile group header. For example, when ALF is enabled for a tile group, the APS ID of the ALF APS is included in the tile group header. For example, this can be achieved by setting the APS parameter type field to indicate the ALF type. Therefore, when ALF is not enabled for a tile group, the APS ID of the ALF APS is not included in the tile group header. In addition, when the resampler decoding tool is enabled for a tile group, the APS ID of the resampler APS is included in the tile group header. For example, this can be achieved by setting the APS parameter type field to indicate the resampler type. In addition, when the resampler decoding tool is not enabled for a tile group, the APS ID of the resampler APS may not be included in the tile group header.

[0126] In some examples, the presence of APS parameter type information in the APS can be regulated by using the decoding tool related to the parameter. When only one APS-related decoding tool (e.g., LMCS, ALF, or scaling list) is enabled for the bitstream, the APS parameter type information may not be present but can be inferred. For example, when the APS can contain the parameters of both the ALF and resampler decoding tools, but only ALF is enabled (e.g., as specified by a flag in the SPS) and the resampler is not enabled (e.g., as specified by a flag in the SPS), the APS parameter type may not be indicated and can be inferred to be equal to the ALF parameter.

[0127] In another example, the APS parameter type information can be inferred from the APS ID value. For example, a predefined range of APS ID values can be associated with corresponding APS parameter types. This can be achieved as follows. Instead of allocating X bits to indicate the APS ID and Y bits to indicate the APS parameter type, X + Y bits can be allocated to indicate the APS ID. Then, different ranges of APS ID values can be specified to represent different types of APS parameter types. For example, instead of using 5 bits to indicate the APS ID and 3 bits to indicate the APS parameter type, 8 bits can be allocated to indicate the APS ID (e.g., without increasing the bit cost). An APS ID value range of 0 to 63 indicates that the APS contains parameters for the ALF, an ID value range of 64 to 95 indicates that the APS contains parameters for the reshaper, and for other parameter types (such as the scaling list), the value range of 96 to 255 can be reserved. In another example, an APS ID value range of 0 to 31 indicates that the APS contains parameters for the ALF, an ID value range of 32 to 47 indicates that the APS contains parameters for the reshaper, and for other parameter types (such as the scaling list), the value range of 48 to 255 can be reserved. The advantage of this method is that the APS ID range can be allocated according to the frequency of parameter changes for each tool. For example, it can be expected that the ALF parameters change more frequently than the reshaper parameters. In this case, a larger APS ID range can be used to indicate that the APS is an APS containing ALF parameters.

[0128] In a first example, one or more of the above aspects can be implemented as follows. An ALF APS can be defined as an APS where aps_params_type is equal to ALF_APS. A reshaper APS (or LMCS APS) can be defined as an APS where aps_params_type is equal to MAP_APS. Examples of SPS syntax and semantics are as follows.

[0129]

[0130] The sps_reshaper_enabled_flag is set to 1 to indicate that the reshaper is used for decoding a coded video sequence (CVS). The sps_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not used for the CVS.

[0131] Examples of APS syntax and semantics are as follows.

[0132]

[0133]

[0134] The aps_params_type represents the type of APS parameters carried in APS, as shown in the following table.

[0135] Table 1: APS parameter type code and type of APS parameter

[0136] aps_params_type Name of aps_params_type Type of APS parameters 0 ALF_APS ALF parameters 1 MAP_APS In-loop mapping (i.e., resampler) parameters 2..7 Reserved Reserved

[0137] Examples of the syntax and semantics of the tile group header are as follows.

[0138]

[0139]

[0140] The tile_group_alf_aps_id represents the adaptation_parameter_set_id of the ALF APS referred to by the tile group. The TemporalId of the ALF APS NAL unit with the adaptation_parameter_set_id equal to the tile_group_alf_aps_id should be less than or equal to the TemporalId of the coded tile group NAL unit. When multiple ALF APSs with the same adaptation_parameter_set_id value are referred to by two or more tile groups of the same picture, the multiple ALF APSs with the same adaptation_parameter_set_id value should have the same content.

[0141] The tile_group_reshaper_enabled_flag is set to 1 to indicate that the reshaper is enabled for the current tile group. The tile_group_reshaper_enabled_flag is set to 0 to indicate that the reshaper is not enabled for the current tile group. When the tile_group_reshaper_enable_flag does not exist, the inferred flag is equal to 0. The tile_group_reshaper_aps_id represents the adaptation_parameter_set_id of the reshaper APS referenced by the tile group. The TemporalId of the reshaper APS NAL unit with an adaptation_parameter_set_id equal to tile_group_reshaper_aps_id should be less than or equal to the TemporalId of the coded tile group NAL unit. When multiple reshaper APSs with the same adaptation_parameter_set_id value are referenced by two or more tile groups of the same picture, the multiple reshaper APSs with the same adaptation_parameter_set_id value should have the same content. The tile_group_reshaper_chroma_residual_scale_flag is set to 1 to indicate that chroma residual scaling is enabled for the current tile group. The tile_group_reshaper_chroma_residual_scale_flag is set to 0 to indicate that chroma residual scaling is not enabled for the current tile group. When the tile_group_reshaper_chroma_residual_scale_flag does not exist, the inferred flag is equal to 0.

[0142] Examples of reshaper data syntax and semantics are as follows.

[0143]

[0144] The reshaper_model_min_bin_idx represents the minimum bin (or segment) index to be used during the reshaper construction. The value of reshaper_model_min_bin_idx should be in the range from 0 to MaxBinIdx (including the end values). The value of MaxBinIdx should be equal to 15. The reshaper_model_delta_max_bin_idx represents the maximum bin (or segment) index MaxBinIdx minus the maximum bin index to be used during the reshaper construction. The value of reshaper_model_max_bin_idx is set to MaxBinIdx – reshaper_model_delta_max_bin_idx. The reshaper_model_bin_delta_abs_cw_prec_minus1 + 1 represents the number of bits used to represent the syntax element reshaper_model_bin_delta_abs_CW[i]. The reshaper_model_bin_delta_abs_CW[i] represents the absolute delta codeword value of the i-th bin. The syntax element reshaper_model_bin_delta_abs_CW[i] is represented by reshaper_model_bin_delta_abs_cw_prec_minus1 + 1 bits. The reshaper_model_bin_delta_sign_CW_flag[i] represents the sign of reshaper_model_bin_delta_abs_CW[i].

[0145] In a second example, one or more of the above aspects may be implemented as follows. Examples of SPS syntax and semantics are as follows.

[0146]

[0147]

[0148] The variables ALFEnabled and ReshaperEnabled are set as follows. ALFEnabled = sps_alf_enabled_flag and ReshaperEnabled = sps_reshaper_enabled_flag.

[0149] Examples of APS syntax and semantics are as follows.

[0150]

[0151] aps_params_type represents the type of APS parameters carried in APS, as shown in the following table.

[0152] Table 2: APS Parameter Type Codes and Types of APS Parameters

[0153] aps_params_type Name of aps_params_type Type of APS parameters 0 ALF_APS ALF parameters 1 MAP_APS In-loop mapping (i.e., resampler) parameters 2..7 Reserved Reserved

[0154] When it does not exist, the value of aps_params_type is inferred as follows. If ALFEnabled, then aps_params_type is set to 0. Otherwise, aps_params_type is set to 1.

[0155] In a second example, one or more of the above aspects may be implemented as follows. Examples of SPS syntax and semantics are as follows.

[0156]

[0157] adaptation_parameter_set_id provides an identifier for APS for other syntax elements to refer to. APS can be shared across images and can be different in different tile groupings within an image. The values and descriptions of the variable APSParamsType are defined in the following table.

[0158] Table 3: APS Parameter Type Codes and Types of APS Parameters

[0159] APS ID Range ApsParamsType Type of APS parameters 0~63 0:ALF_APS ALF parameters 64~95 1:MAP_APS In-loop mapping (i.e., resampler) parameters 96..255 Reserved Reserved

[0160] Examples of tile group header semantics are as follows. tile_group_alf_aps_id represents the adaptation_parameter_set_id of the ALF APS referred to by the tile group. The TemporalId of the ALF APS NAL unit whose adaptation_parameter_set_id is equal to tile_group_alf_aps_id should be less than or equal to the TemporalId of the coded tile group NAL unit. The value of tile_group_alf_aps_id should be in the range of 0 to 63 (including the end values).

[0161] When multiple ALF APSs with the same adaptation_parameter_set_id value are referenced by two or more tile groups of the same picture, the multiple ALF APSs with the same adaptation_parameter_set_id value shall have the same content. The tile_group_reshaper_aps_id represents the adaptation_parameter_set_id of the reshaper APS referenced by the tile group. The TemporalId of the reshaper APS NAL unit with the adaptation_parameter_set_id equal to the tile_group_reshaper_aps_id shall be less than or equal to the TemporalId of the coded tile group NAL unit. The value of the tile_group_reshaper_aps_id shall be in the range of 64 to 95 (including the end values). When multiple reshaper APSs with the same adaptation_parameter_set_id value are referenced by two or more tile groups of the same picture, the multiple reshaper APSs with the same adaptation_parameter_set_id value shall have the same content.

[0162] Figure 7 FIG. is a schematic diagram of an exemplary video decoding device 700. The video decoding device 700 is suitable for implementing the disclosed examples / embodiments described herein. The video decoding device 700 includes a downlink port 720, an uplink port 750, and / or a transceiver (Tx / Rx) 710, where the transceiver includes a transmitter and / or a receiver for transmitting data uplink and / or downlink over a network. The video decoding device 700 further includes: a processor 730, including a logic unit and / or a central processing unit (CPU) for processing data and a memory 732 for storing data. The video decoding device 700 may further include electronic components, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components coupled to the uplink port 750 and / or the downlink port 720 for data communication over electrical, optical, or wireless communication networks. The video decoding device 700 may further include an input and / or output (I / O) device 760 for sending data to and receiving data from a user. The I / O device 760 may include an output device, such as for displaying video data

[0163] A display for presenting data, a speaker for outputting audio data, etc. The I / O device 760 may also include input devices such as a keyboard, a mouse, a trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0164] The processor 730 is implemented by hardware and software. The processor 730 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 communicates with the downstream port 720, Tx / Rx 710, the upstream port 750, and the memory 732. The processor 730 includes a decoding module 714. The decoding module 714 implements the embodiments disclosed herein, such as method 100, method 800, and method 900, which may employ the bitstream 500 and / or the mechanism 600. The decoding module 714 may also implement any other method / mechanism described herein. In addition, the decoding module 714 may implement the codec system 200, the encoder 300, and / or the decoder 400. For example, the decoding module 714 may encode / decode images in the bitstream and encode / decode parameters associated with the stripes of the images in the multiple APSs. In some examples, different types of parameters may be encoded into different types of APSs. In addition, different types of APSs may be included in different NAL unit types. Such APS types may include ALF APS, scaling list APS, and / or LMCS APS. Each APS may include an APS ID. The APS IDs of different APS types increase sequentially over different value spaces. In addition, the stripe and / or the image may refer to the corresponding stripe header and / or image header. Then, such headers may refer to the APSs containing the relevant decoding tools. Such APSs may be uniquely referred to by the APS ID and the APS type. Such examples reduce the redundant indication of decoding tool parameters and / or reduce the bit usage of the identifiers. Therefore, the decoding module 714 enables the video decoding device 700 to provide other functions and / or decoding efficiency when decoding video data. Therefore, the decoding module 714 improves the function of the video decoding device 700 and solves problems specific to the field of video coding. In addition, the decoding module 714 may transform the video decoding device 700 into different states. Alternatively, the decoding module 714 may be implemented as instructions stored in the memory 732 and executed by the processor 730 (e.g., a computer program product stored on a non-transitory medium).

[0165] Memory 732 includes one or more types of memory, such as magnetic disks, tape drives, solid state drives, read only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. Memory 732 can be used as an overflow data storage device to store these programs and the instructions and data read during program execution when a program is selected for execution.

[0166] Figure 8 FIG. 800 is a flowchart of an exemplary method for encoding a video sequence into a bitstream 500, etc., using multiple APS types such as ALF APS 512, scaled list APS 513, and / or LMCS APS 514. An encoder (e.g., codec system 200, encoder 300, and / or video decoding device 700) may use method 800 when performing method 100. In method 800, APS IDs may also be assigned to different types of APS according to mechanism 600 by using different value spaces.

[0167] Method 800 may begin with: the encoder receiving a video sequence including multiple images and determining, for example, based on user input, to encode the video sequence into a bitstream. The video sequence is segmented into pictures / frames before encoding for further segmentation. In step 801, the encoder determines luma mapping with chroma scaling (LMCS) parameters applied to a slice. This may include performing an RDO operation to encode a slice of an image. For example, the encoder may iteratively encode a slice multiple times using different decoding options / decoding tools, decode the encoded slice, and filter the decoded slice to improve the output quality of the slice. The encoder may then select an encoding option to achieve an optimal balance between compression and output quality. Once the encoding is selected, the encoder may determine the LMCS parameters (and any other filter parameters) for filtering the selected encoding. In step 803, the encoder may encode the slice into the bitstream based on the selected encoding.

[0168] In step 805, the LMCS parameters determined in step 801 are encoded into the LMCS APS in the bitstream. Additionally, data related to the strip may also be encoded into the bitstream, where the strip refers to the LMCS APS. For example, the data related to the strip may be a strip header and / or an image header. In one example, the strip header may be encoded into the bitstream. The strip may refer to the strip header. Then, the strip header contains the data related to the strip and refers to the LMCS APS. In another example, the image header may be encoded into the bitstream. The image containing the strip may refer to the image header. Then, the image header contains the data related to the strip and refers to the LMCS APS. In either case, the header contains sufficient information to determine the appropriate LMCS APS that contains the LMCS parameters for the strip.

[0169] In step 807, other filtering parameters may be encoded into other APSs. For example, the ALF APS containing the ALF parameters and the scaling list APS containing the APS parameters may also be encoded into the bitstream. The strip header and / or the image header may also refer to such APSs.

[0170] As described above, each APS may be uniquely identified by a combination of a parameter type and an APS ID. Such information may be used for the strip header or the image header to refer to the relevant APS. For example, each APS may include an aps_params_type code that is set to a predefined value representing the type of parameters included in the corresponding APS. Additionally, each APS may include an APS ID selected from a predefined range. The predefined range may be determined based on the parameter type of the corresponding APS. For example, the range of the LMCS APS may be from 0 to 3 (e.g., 2 bits), and the ranges of the ALF APS and the scaling list APS may be from 0 to 7 (e.g., 3 bits). Such ranges may describe different overlapping value spaces specific to the APS type as described in mechanism 600. Therefore, in step 805 and step 807, both the APS type and the APS ID may be used to refer to a specific APS.

[0171] In step 809, the encoder encodes the SPS into the bitstream. The SPS may include a flag that is set to indicate that LMCS has been enabled for the encoded video sequence including the image / strip. Then, in step 811, the bitstream may be stored in a memory for sending to a decoder, e.g., via a transmitter. By referring to the LMCS APS in the strip header / image parameter header, the bitstream includes sufficient information for the decoder to obtain the LMCS parameters to decode the encoded strip.

[0172] Figure 9 It is a flowchart of an exemplary method 900 for decoding a video sequence from a bitstream such as bitstream 500 using multiple APS types such as ALF APS 512, scaled list APS 513, and / or LMCS APS 514. A decoder (e.g., codec system 200, decoder 400, and / or video coding device 700) may use method 900 when executing method 100. In method 900, an APS may also be referenced based on an APS ID, where the APS ID is assigned according to mechanism 600, and different types of APSs use APS IDs assigned according to different value spaces.

[0173] Method 900 may begin with the decoder starting to receive a bitstream representing decoded data of a video sequence, for example, as a result of method 800. In step 901, the decoder receives the bitstream. The bitstream includes images segmented into slices and an LMCS APS containing LMCS parameters. In some examples, the bitstream may also include an ALF APS and a scaled list APS, where the ALF APS contains ALF parameters and the scaled list APS contains APS parameters. The bitstream may also include an image header and / or a slice header respectively associated with an image and one of the slices. The bitstream may also include other parameter sets, such as SPS, PPS, etc.

[0174] In step 903, the decoder may determine the LMCS APS, ALF APS, and / or Scaling List APS referred to by the relevant data of the strip. For example, the strip header or picture header may contain data related to the strip / picture and may refer to one or more APSs, and the one or more APSs include LMCS APS, ALF APS, and / or Scaling List APS. Each APS may be uniquely identified by a combination of a parameter type and an APS ID. Therefore, the parameter type and the APS ID may be included in the picture header and / or strip header. For example, each APS may include an aps_params_type code, and the aps_params_type code is set to a predefined value, and the predefined value represents the type of parameters included in each APS. In addition, each APS may include an APS ID selected from a predefined range. For example, the APS IDs of the APS types may be sequentially allocated in multiple different value spaces within the predefined range. For example, the predefined range may be determined based on the parameter type of the corresponding APS. In a specific example, the range of the LMCS APS may be from 0 to 3 (e.g., 2 bits), and the ranges of the ALF APS and the Scaling List APS may be from 0 to 7 (e.g., 3 bits). Such ranges may describe different overlapping value spaces specific to the APS type as described in mechanism 600. Therefore, both the APS type and the APS ID can be used for the picture header and / or strip header to refer to a specific APS.

[0175] In step 905, the decoder may decode the strip using the LMCS parameters in the LMCS APS based on the reference to the LMCS APS. The decoder may also decode the strip using the ALF parameters and / or Scaling List parameters in the ALF APS and / or Scaling List APS based on the reference to the APS in the picture header / strip header. In some examples, the SPS includes a sequence parameter set (SPS), and the sequence parameter set includes a flag that is set to indicate that the LMCS decoding tool has been enabled for the encoded video sequence including the strip. In this case, the LMCS parameters in the LMCS APS are obtained to support decoding in step 905 based on the flag. In step 907, the decoder may forward the strip for display as part of the decoded video sequence.

[0176] Figure 10FIG. is a schematic diagram of an exemplary system 1000 for decoding a video sequence of images in a bitstream, such as bitstream 500, using multiple APS types, such as ALF APS 512, Scaled List APS 513, and / or LMCS APS 514. The system 1000 may be implemented by an encoder and a decoder (e.g., codec system 200, encoder 300, decoder 400, and / or video decoding device 700). In addition, the system 1000 may be used when implementing method 100, method 800, method 900, and / or mechanism 600.

[0177] The system 1000 includes a video encoder 1002. The video encoder 1002 includes a determination module 1001 for determining LMCS parameters applied to a strip. The video encoder 1002 further includes an encoding module 1003 for encoding the strip into a bitstream. The encoding module 1003 is further configured to encode the LMCS parameters into the LMCS APS in the bitstream. The encoding module 1003 is further configured to encode data related to the strip into the bitstream, wherein the strip refers to the LMCS APS. The video encoder 1002 further includes a storage module 1005 for storing the bitstream for transmission to a decoder. The video encoder 1002 further includes a transmission module 1007 for transmitting the bitstream including the LMCS APS to support decoding of the strip on the decoder side. The video encoder 1002 may also be used to perform any of the steps in the method 800.

[0178] The system 1000 further includes a video decoder 1010. The video decoder 1010 includes a receiving module 1011 for receiving a bitstream, wherein the bitstream includes a strip and an LMCS APS, and the LMCS APS includes LMCS parameters. The video decoder 1010 further includes a determination module 1013 for determining the LMCS APS referred to by the relevant data of the strip. The video decoder 1010 further includes a decoding module 1015 for decoding the strip using the LMCS parameters in the LMCS APS based on the reference to the LMCS APS. The video decoder 1010 further includes a forwarding module 1017 for forwarding the strip for display as part of a decoded video sequence. The video decoder 1010 may also be used to perform any of the steps in the method 900.

[0179] When there is no intermediate component between the first component and the second component other than a wire, a trace, or other medium, the first component and the second component are directly coupled. When there is an intermediate component between the first component and the second component in addition to a wire, a trace, or other medium, the first component and the second component are indirectly coupled. The term "coupled" and its variants include direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range including ±10% of the subsequent number.

[0180] It should also be understood that the steps of the exemplary methods described herein do not necessarily need to be executed in the order described, and the order of these method steps should be understood as merely exemplary. Similarly, in methods consistent with various embodiments of the present invention, such methods may include other steps, and certain steps may be omitted or combined.

[0181] Although the present invention provides several embodiments, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present invention. The examples of the present invention will be regarded as illustrative rather than restrictive, and the present invention is not limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0182] In addition, without departing from the scope of the present invention, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques, or methods. Other variations, substitutions, and change examples can be determined by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. A method implemented in a decoder, characterized in that, The method includes: Receiving a bitstream, wherein the bitstream includes a luminance mapping and chrominance scaling (LMCS) adaptive parameter set (APS) and an adaptive loop filter (ALF) APS, the LMCS APS includes LMCS parameters associated with an encoded slice, and the ALF APS includes ALF parameters associated with the encoded slice; Obtaining the LMCS parameters from the LMCS APS and obtaining the ALF parameters from the ALF APS; Obtaining a decoded image according to the encoded slice; Wherein each APS includes an APS identifier ID selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.

2. The method according to claim 1, wherein Each APS includes an APS parameter type (aps_params_type) code, and the APS parameter type (aps_params_type) code is set to a predefined value, and the predefined value represents the type of parameters included in each APS.

3. The method according to claim 1 or 2, characterized in that Each APS is identified by a combination of a parameter type and the APS ID.

4. The method according to any one of claims 1 to 3, characterized in that The bitstream further includes a sequence parameter set (SPS), the SPS includes a flag, and the flag is set to indicate that the encoded video sequence including the encoded slice enables LMCS, and the LMCS parameters in the LMCS APS are obtained based on the flag.

5. The method according to any one of claims 1 to 4, characterized in that, Specific types of APSs share the same value space of the APS ID, and different types of APSs use different value spaces of the APS ID.

6. A method implemented in an encoder, characterized in that, The method includes: Determining the luminance mapping and chrominance scaling (LMCS) parameters and the adaptive loop filter (ALF) parameters applied to a slice; Encoding the slice into the bitstream as an encoded slice; Encoding the LMCS parameters into the LMCS adaptive parameter set (APS) in the bitstream and encoding the ALF parameters into the ALF APS in the bitstream; Wherein each APS includes an APS identifier ID selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.

7. The method according to claim 6, wherein Each APS includes an APS parameter type (aps_params_type) code, and the APS parameter type (aps_params_type) code is set to a predefined value, and the predefined value represents the type of parameters included in each APS.

8. The method according to claim 6 or 7, characterized in that, Each APS is identified by a combination of a parameter type and the APS ID.

9. The method according to any one of claims 6 to 8, characterized in that, The method further includes: encoding a sequence parameter set (SPS) into the bitstream, wherein the SPS includes a flag, and the flag is set to indicate that the encoded video sequence including the encoded slice enables LMCS.

10. The method according to any one of claims 6 to 9, characterized in that Specific types of APSs share the same value space of the APS ID, and different types of APSs use different value spaces of the APS ID.

11. A video decoding device, characterized in that, The video decoding device includes: A processor, a receiver coupled to the processor, a memory coupled to the processor, and a transmitter coupled to the processor, wherein the processor, the receiver, the memory, and the transmitter are configured to perform the method according to any one of claims 1 to 10.

12. A non-transitory computer-readable medium, characterized in that, The non-transitory computer-readable medium includes a computer program product for use by a video decoding device, the computer program product including computer-executable instructions stored in the non-transitory computer-readable medium, which, when executed by a processor, cause the video decoding device to perform the method according to any one of claims 1 to 10.

13. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores a bitstream, the bitstream including: A luminance mapping and chrominance scaling LMCS adaptive parameter set APS and an adaptive loop filter ALF APS, the LMCS APS including LMCS parameters associated with an encoded slice, and the ALF APS including ALF parameters associated with the encoded slice; wherein each APS includes an APS identifier ID selected from a predefined range, and the predefined range is determined based on the parameter type of each APS.

14. The computer-readable storage medium according to claim 13, wherein Each APS includes an APS parameter type (aps_params_type) code, the APS parameter type (aps_params_type) code being set to a predefined value that indicates the type of parameters included in each APS.

15. The computer-readable storage medium according to claim 13 or 14, characterized in that, Each APS is identified by a combination of a parameter type and the APS ID.

16. The computer-readable storage medium according to any one of claims 13 to 15, characterized in that, The bitstream further includes a sequence parameter set SPS, the SPS including a flag that is set to indicate that the encoded video sequence including the encoded slice enables LMCS, and the LMCS parameters in the LMCS APS are obtained based on the flag.

17. The computer-readable storage medium according to any one of claims 13 to 16, wherein APSs of a specific type share the same value space of the APS ID, and APSs of different types use different value spaces of the APS ID.

18. A computer program product, characterized in that, Includes computer-executable instructions stored in a non-transitory computer storage medium, which, when executed by a processor, cause a computer to perform the method according to any one of claims 1 to 10.