Coefficient coding in transform skip mode

By introducing coefficient encoding and decoding technology in transform skip mode into video encoding and decoding, and by optimizing the conversion of video blocks and bitstreams using context index offset and adaptive parameter set, the inefficiency problem in existing technologies is solved, and more efficient video data compression and decoding is achieved.

CN115428463BActive Publication Date: 2026-05-15DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2021-04-01
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing video codec standards suffer from inefficiency and redundant signaling when handling the conversion between video blocks and bitstreams, especially in transform skip mode, where it is difficult to effectively utilize context information for encoding and decoding.

Method used

By introducing coefficient encoding and decoding techniques in transform skip mode, and utilizing the regularization of context index offset and symbol flags, the conversion process between video blocks and bitstreams is optimized, and adaptive parameter sets and auxiliary enhancement information are introduced to improve encoding and decoding efficiency.

Benefits of technology

It improves the efficiency and flexibility of video encoding and decoding, reduces redundant signaling, and enhances the compression rate and decoding quality of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115428463B_ABST
    Figure CN115428463B_ABST
Patent Text Reader

Abstract

Methods and apparatus for video processing are disclosed. The processing can include video encoding, video decoding, or video transcoding. An example method includes performing a conversion between a current block of a video and a bitstream of the video, wherein the bitstream conforms to a rule that specifies a context index offset is used for including a first sign flag of a first coefficient in the bitstream, wherein the rule specifies that a value of the context index offset is based on whether a first coding mode is applied to the current block in the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is based on International Patent Application No. PCT / CN2021 / 084869, filed on April 1, 2021, which claims priority and interest in International Patent Application No. PCT / CN2020 / 082983, filed on April 2, 2020. All of the aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This paper discloses techniques that can be used by video encoders and decoders to process the codec representation of video using control information useful for decoding the codec representation.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a current block of video and a bitstream of video, wherein the bitstream conforms to a rule specifying a context index offset for including a first symbol flag of a first coefficient in the bitstream, wherein the rule specifies that the value of the context index offset is based on whether a first codec mode is applied to the current block in the bitstream.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between a current block of video and a bitstream of video, wherein the bitstream conforms to a rule that specifies the symbol flag of the current block and is included in the bitstream using either a context mode or a bypass mode based on the number of remaining context codec bits.

[0008] In another example, a video processing method is disclosed. This method includes performing a conversion between a current block of video and a bitstream of video, wherein the bitstream conforms to a rule specifying a context index offset for including symbolic flags of the current block in the bitstream, and the rule specifies that the context index offset is determined based on information from the current block.

[0009] In another example, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more video layers and a bitstream of the video according to a rule, wherein the rule specifies the use of a plurality of adaptive parameter set network abstraction layer units (APNs) of the video, wherein each APN has a corresponding adaptive parameter type value, wherein each APN is associated with a corresponding video layer identifier, wherein each APN is a prefix unit or a suffix unit, and wherein the rule specifies that, in response to the plurality of APNs sharing the same adaptive parameter type value, the adaptive parameter set identifier values ​​of the plurality of APNs belong to the same identifier space.

[0010] In another example, a video processing method is disclosed. The method includes performing a conversion between a current block of video and a bitstream of video, wherein the bitstream conforms to a rule specifying that a first auxiliary enhancement message with specific characteristics is not allowed to: (1) be repeated within a stripe unit in the bitstream in response to a second auxiliary enhancement message with specific characteristics being included in a stripe unit, or (2) be updated within a stripe unit in the bitstream in response to the first auxiliary enhancement message, wherein the stripe unit comprises a set of network abstraction layer units consecutively in decoding order, and wherein the set of network abstraction layers comprises a single codec stripe and one or more non-video codec layer network abstraction layer units associated with that single codec stripe.

[0011] In another example, a video processing method is disclosed. The method includes performing a conversion between a current block of video and a bitstream of video, wherein the bitstream conforms to a rule specifying that in response to: (1) a stripe unit comprising a second non-video codec layer network abstraction layer unit having the same characteristics as a first non-video codec layer network abstraction layer unit, and (2) the first non-video codec layer network abstraction layer unit having a network abstraction layer unit type other than prefix auxiliary enhancement information or suffix auxiliary enhancement information, repetition of the first non-video codec layer network abstraction layer unit is not allowed in the stripe unit of the bitstream.

[0012] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising multiple video layers and a video bitstream, based on rules specifying which adaptive parameter sets from multiple adaptive parameter sets are not allowed to be shared across multiple video layers.

[0013] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising one or more video layers and a video bitstream, according to a rule, wherein the bitstream includes one or more adaptive loop filter adaptive parameter sets, and wherein the rule specifies whether the one or more adaptive loop filter adaptive parameter sets are allowed to be updated within a picture cell.

[0014] In another example, a video processing method is disclosed. This method includes performing a conversion between a video comprising one or more codec layer video sequences and a video bitstream, according to a rule, wherein the bitstream includes an adaptive loop filter adaptive parameter set, and wherein the rule specifies that, in response to the adaptive loop filter adaptive parameter set having one or more specific characteristics, the adaptive loop filter adaptive parameter set is not allowed to be shared across one or more codec layer video sequences.

[0015] In another example, a video processing method is disclosed. This method includes performing a conversion between a current block of video and a codec representation of the video. The codec representation conforms to a format rule specifying a range of context index offsets for encoding and decoding symbolic flags of the current block in the codec representation, wherein a range function of the codec mode is used to represent the current block in the codec representation.

[0016] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a current block of video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that the symbol flags of the current block are encoded and decoded in the codec representation using either context codec bits or a bypass mode, depending on the number of remaining context codec bits.

[0017] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a current block of video and a codec representation of the video; wherein the conversion uses a block-based incremental pulse codec modulation (BDPCM) mode, where the codec representation conforms to a format rule specifying that symbolic flags from BDPCM are context-coded in the codec representation, such that the context index offset used for encoding and decoding the symbolic flags is a function of the codec conditions of the current block.

[0018] In another example, a different video processing method is disclosed. The method includes: performing a conversion between a current block of video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that, at most once, a stripe unit in the codec representation corresponding to a set of network abstraction layer units arranged in sequential decoding order and containing a single codec stripe is allowed to include at least a portion of auxiliary enhancement information (SEI).

[0019] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a current block of video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that a stripe unit in the codec representation corresponding to a set of network abstraction layer units arranged in a sequential decoding order and containing a single codec stripe comprises one or more Video Codec Layer Network Abstraction Layer (VCL NAL) units, wherein the format rule further specifies a first type of unit that is allowed to repeat in a stripe unit and a second type of unit that is not allowed to repeat in a stripe unit.

[0020] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a video comprising one or more video layers and a codec representation of the video, according to a rule; wherein the codec representation includes one or more adaptive parameter sets (APS); and wherein the rule specifies the applicability of some of the one or more APSs to the conversion of the one or more video layers.

[0021] In another example, a different video processing method is disclosed. This method includes: performing a conversion between a video comprising one or more video layers and a codec representation of the video, according to rules; wherein the codec representation is arranged in one or more Network Abstraction Layer (NAL) units; wherein the codec representation includes one or more adaptive parameter sets for controlling characteristics of the conversion.

[0022] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0023] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0024] In yet another example, a computer-readable medium storing code is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0025] These and other features will be described in this article. Attached Figure Description

[0026] Figure 1 This is a block diagram of an example video processing system;

[0027] Figure 2 This is a block diagram of a video processing device;

[0028] Figure 3 This is a flowchart of an example method for video processing;

[0029] Figure 4This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure;

[0030] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure;

[0031] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure;

[0032] Figure 7 An example of the shape of an adaptive loop filter (ALF) is shown (chroma: 5×5 rhombus, luminance: 7×7 rhombus);

[0033] Figure 8 Examples of ALF and CC-ALF plots are shown; and

[0034] Figures 9 to 17 This is a flowchart of an example method for video processing. Detailed Implementation

[0035] Chapter headings are used in this document for ease of understanding, not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding, not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.

[0036] 1. Introduction

[0037] This patent document relates to video codec technology. Specifically, it relates to coefficient encoding and decoding in transform skip mode and the repetition and updating of non-VCL data units in video encoding and decoding. It can be applied to existing video codec standards (such as HEVC) or upcoming standards (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs.

[0038] 2. Abbreviation

[0039] ALF Adaptive Loop Filter

[0040] APS Adaptive Parameter Set

[0041] AU Access Unit

[0042] AUD Access Unit Separator

[0043] AVC Advanced Video Codec

[0044] CLVS codec layer video sequence

[0045] CPB image buffer

[0046] CRA Fully Random Access

[0047] CTU (Codec Tree Unit)

[0048] CVS codec video sequence

[0049] DCI decoding capability information

[0050] DPB Decoding Image Buffer

[0051] DU decoding unit

[0052] End of EOB bitstream

[0053] End of EOS sequence

[0054] GDR gradually decoded and refreshed

[0055] HEVC High-Efficiency Video Encoding and Decoding

[0056] HRD Assumption Reference Decoder

[0057] IDR Instant Decoding and Refresh

[0058] JEM Joint Exploration Model

[0059] LMCS Luminance Mapping and Chroma Scaling

[0060] MCTS Motion Restraint Piece Set

[0061] NAL Network Abstraction Layer

[0062] OLS Output Layer Set

[0063] PH image header

[0064] PPS Image Parameter Set

[0065] PTL levels, tiers, and grades

[0066] PU Image Unit

[0067] RADL Random Access Decoding Pre-amplifier (Image)

[0068] RAP Random Access Point

[0069] RASL random access skips preprocessor (image)

[0070] RBSP raw byte sequence payload

[0071] RPL Reference Image List

[0072] SAO Sample Adaptive Offset

[0073] SEI Assist Enhancement Information

[0074] SPS Sequence Parameter Set

[0075] STSA Step-by-Step Temporal Sublayer Access

[0076] SVC Scalable Video Codec

[0077] VCL (Video Codec Layer)

[0078] VPS Video Parameter Set

[0079] VTM VVC Test Model

[0080] VUI Video Availability Information

[0081] VVC Multi-Functional Video Encoding and Decoding

[0082] 3. Introduction to Video Encoding and Decoding

[0083] Video codec standards are primarily developed from well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 video standards. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, which uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new codec standards is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and the VTM test model are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.

[0084] 3.1 Concepts and definitions most relevant to this disclosure in VVC

[0085] Access Unit (AU): A collection of PUs belonging to different layers and containing codec images associated with the same time output from the DPB.

[0086] Adaptive Loop Filter (ALF): A filtering process that is applied as part of the decoding process and controlled by parameters transmitted in the APS.

[0087] ALF APS: APS that controls the ALF process.

[0088] Adaptive Parameter Set (APS): Contains a syntax structure applicable to zero or more stripes, as determined by zero or more syntax elements found in the strip header.

[0089] Associated non-VCL NAL units: Non-VCL NAL units of VCL NAL units (when present), where VCLNAL units are associated VCL NAL units of non-VCL NAL units.

[0090] The associated VCL NAL unit: the preceding VCL NAL unit in the decoding order if nal_unit_type is equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, or a non-VCL NAL unit in the range UNSPEC_30..UNSPEC_31; otherwise, the next VCL NAL unit in the decoding order.

[0091] Decoding Unit (DU): If DecodingUnitHrdFlag equals 0, it is an AU; otherwise, it is a subset of the AU, consisting of one or more VCL NAL units in the AU and associated non-VCL NAL units.

[0092] Layer: A collection of all VCL NAL units with a specific value of nuh_layer_id and their associated non-VCL NAL units.

[0093] LMCS APS: APS that controls the LMCS process.

[0094] Luminance Mapping and Chroma Scaling (LMCS): A process applied as part of the decoding process to map luminance samples to specific values ​​and to apply scaling operations to the values ​​of chrominance samples.

[0095] Image Unit (PU): A collection of NAL units that are related to each other according to specified classification rules, are consecutive in decoding order, and contain exactly one encoded / decoded image.

[0096] Scaling list: A list that associates each frequency index with a scaling factor in the scaling process.

[0097] Scaling List APS: An APS with syntax elements for constructing a scaling list.

[0098] Video codec layer (VCL) NAL unit: A collective term for codec strip NAL units and subsets of NAL units, which have a reserved value of nal_unit_type and are classified as VCL NAL units in this specification.

[0099] 3.2 Coefficient Encoding and Decoding in Transform Skip Mode

[0100] In the current VVC draft, several modifications are proposed to the coefficient coding and decoding in Transform Skip (TS) mode compared to non-TS coefficient coding and decoding, in order to adapt the residual coding and decoding to the statistical and signal characteristics of the transform skip level.

[0101] The latest text related to this section in JVET-Q2001-vE is shown below.

[0102] 7.3.10.11 Residual Encoding / Decoding Syntax

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117] 3.2.1 Context Modeling and Context Index Offset Derivation of the Symbol Flag coeff_sign_flag

[0118] Table 51 – Association between ctxIdx and each syntax element of initializationType during initialization.

[0119]

[0120] Table 125 – Specification of initValue and shiftIdx of ctxInc for coeff_sign_flag

[0121]

[0122] Table 131 – Assigning ctxInc to syntax elements with context encoding / decoding bits

[0123]

[0124] 9.3.4.2.10 Derivation of ctxInc of the syntax element coeff_sign_flag for transforming skip modes.

[0125] The inputs to this process are the color component index cIdx, the brightness position (x0, y0) of the top-left sample relative to the top-left sample of the current image for the current transform block, and the current coefficient scan position (xC, yC).

[0126] The output of this process is the variable ctxInc.

[0127] The variables leftSign and aboveSign are derived as follows:

[0128] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1595)

[0129] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1596)

[0130] The variable ctxInc is derived as follows:

[0131] – If leftSign equals 0 and aboveSign equals 0, or if leftSign equals -aboveSign, then the following applies:

[0132] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?0:3) (1597)

[0133] Otherwise, if leftSign is greater than or equal to 0 and aboveSign is greater than or equal to 0, the following applies:

[0134] ctxInc=(BdpcmFlag[x0][y0][cIdx]?1:4) (1598)

[0135] –Otherwise, the following applies:

[0136] ctxInc=(BdpcmFlag[x0][y0][cIdx]?2:5) (1599)

[0137] 3.3 Block-based Quantization Residual Domain DPCM (BDPCM)

[0138] In JVET-M0413, Quantization Residual Block Differential Pulse Coding and Decoding Modulation (BDPCM) was proposed and adopted by the VVC draft to efficiently encode and decode screen content.

[0139] The prediction direction used in QR-BDPCM can be either vertical or horizontal prediction mode. Similar to intra-frame prediction, intra-frame prediction is performed on the entire block by copying samples along the prediction direction (horizontal or vertical prediction). The residual is quantized, and the increment between the quantized residual and its predicted value (horizontal or vertical) is encoded and decoded. This can be described as follows: For a block of size M (rows) × N (columns), let r i,j ,0≤i≤M-1,0≤j≤N-1 is the prediction residual after performing intra-frame prediction using unfiltered samples from the upper or left block boundaries, either horizontally (copying left neighbor pixel values ​​line-by-line across the prediction block) or vertically (copying top neighbor lines to each line in the prediction block). Let Q(r) i,j ), 0≤i≤M-1, 0≤j≤N-1 represent residuals r i,j The quantized version is then used, where the residual is the difference between the original block and the predicted block values. The block DPCM is then applied to the quantized residual samples to produce a result with element-wise... Modified M×N array When the vertical BDPCM is signaled:

[0140]

[0141] For horizontal prediction, similar rules apply, and the residual quantization samples are obtained by the following formula.

[0142]

[0143] Residual Quantization Samples It is transmitted to the decoder.

[0144] On the decoder side, the above calculation is inverted to produce Q(r). i,j ), 0≤i≤M-1, 0≤j≤N-1. For the vertical prediction case,

[0145]

[0146] Regarding the horizontal situation

[0147]

[0148] Inverse quantization residual Q -1 (Q(r i,j The values ​​are added to the intra-block prediction values ​​to produce reconstructed sample values.

[0149] The main benefit of this approach is that inverse DPCM can be completed during the coefficient analysis run, simply by adding the predicted values ​​when analyzing the coefficients, or it can be performed after analysis.

[0150] 3.4 Scalable Video Codec (SVC) in Typical and VVC Contexts

[0151] Scalable video codec (SVC, sometimes also called scalability in video codec) refers to video codec using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can act as a BL, while the top layer can act as an EL. Intermediate layers can act as either an EL or an RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below the intermediate layer (such as a base layer or any intermediate enhancement layer) and simultaneously act as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0152] In SVC, parameters used by the encoder or decoder are grouped into parameter sets based on the codec level at which they can be utilized (e.g., video level, sequence level, picture level, stripe level, etc.). For example, parameters available for one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters available for one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters utilized for one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and additional parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.

[0153] Because of VVC's support for Reference Picture Resampling (RPR), it's possible to design support for bitstreams containing multiple layers (e.g., two layers in VVC with SD and HD resolutions) without requiring any additional signal processing level codecs, as the upsampling needed for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires higher-level syntax changes (compared to no scalability support). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions to AVC and HEVC), VVC's scalability is designed to be as friendly as possible to single-layer decoder designs. Decoding capabilities for multi-layer bitstreams are specified as if there were only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams do not require many changes to be able to decode multi-layer bitstreams. Compared to the multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, the IRAP AU is required to contain an image of each layer present in CVS.

[0154] A VVC bitstream can consist of one or more Output Layer Sets (OLS). An OLS is a set of one or more layers designated as output layers. Output layers are the layers that are output after decoding.

[0155] 3.5 Parameter Set

[0156] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. AVC, HEVC, and VVC all support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but it is included in the latest VVC draft text.

[0157] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving fault tolerance.

[0158] A VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.

[0159] APS is introduced to carry such image-level or strip-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.

[0160] In VVC, APS is used to carry parameters for ALF, LMCS, and scaling list parameters.

[0161] 3.6 NAL Unit Type in VVC, and NAL Unit Header Syntax and Semantics

[0162] In the latest VVC text (in JVET-Q2001-vE / v15), the NAL unit header syntax and semantics are as follows.

[0163] 7.3.1.2 NAL Unit Header Syntax

[0164]

[0165] 7.4.2.2 NAL Unit Header Semantics

[0166] forbidden_zero_bit should be equal to 0.

[0167] The nuh_reserved_zero_bit should be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified in the future by ITU-T|ISO / IEC. The decoder should ignore (i.e., remove and discard) NAL cells where nuh_reserved_zero_bit is equal to 1.

[0168] The nuh_layer_id specifies the identifier of the layer to which a VCL NAL element belongs, or the identifier of the layer to which a non-VCL NAL element applies. The value of nuh_layer_id should be in the range of 0 to 55 (inclusive). Other values ​​of nuh_layer_id are reserved for future use by ITU-T|ISO / IEC.

[0169] The value of nuh_layer_id should be the same for all VCL NAL units of the codec image. The nuh_layer_id value of the codec image or PU is the nuh_layer_id value of the VCL NAL unit of the codec image or PU.

[0170] The nuh_layer_id values ​​for AUD, PH, EOS, and FD NAL cells are constrained as follows:

[0171] - If nal_unit_type equals AUD_NUT, then nuh_layer_id should equal vps_layer_id[0].

[0172] Otherwise, when nal_unit_type is equal to PH_NUT, EOS_NUT, or FD_NUT, nuh_layer_id should be equal to the nuh_layer_id of the associated VCL NAL unit. Note 1 – The value of nuh_layer_id for DCI, VPS, and EOB NAL units is not restricted.

[0173] The value of nal_unit_type should be the same for all images in CVSS AU.

[0174] nal_unit_type specifies the NAL unit type, that is, the type of RBSP data structure contained in the NAL unit as specified in Table 5.

[0175] NAL units with nal_unit_type within the range of UNSPEC_28..UNSPEC_31 (unspecified semantics) (inclusive of UNSPEC_28 and UNSPEC_31) should not affect the decoding process specified in this specification.

[0176] Note 2 – NAL unit types within the range of UNSPEC_28..UNSPEC_31 may be used as determined by the application. The decoding process for these values ​​of nal_unit_type is not specified in this specification. Because different applications may use these NAL unit types for different purposes, special care must be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the contents of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values ​​may only be suitable for use in contexts where “conflicts” (i.e., different definitions of the meaning of NAL unit contents with the same nal_unit_type value) are not important, impossible, or managed (e.g., in environments defined or managed in control applications or transport specifications, or distributed through control bitstreams).

[0177] For purposes other than determining the amount of data in the DU of the bitstream (as specified in Appendix C), the decoder should ignore (remove and discard) the contents of all NAL units using the reserved value of nal_unit_type.

[0178] Note 3 – This requirement allows for future definitions of compatible extensions to this specification.

[0179] Table 5 – NAL Unit Type Codes and NAL Unit Type Classification

[0180]

[0181]

[0182] Note 4 – A fully random access (CRA) picture can have an associated RASL or RADL picture present in the bitstream.

[0183] Note 5 – An Instant Decode Refresh (IDR) picture with a nal_unit_type equal to IDR_N_LP does not have an associated preceding picture present in the bitstream. An IDR picture with a nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture present in the bitstream.

[0184] The value of nal_unit_type should be the same for all VCL NAL units of a subpicture. A subpicture is defined as having the same NAL unit type as the VCL NAL units of the subpicture.

[0185] For any given image's VCL NAL unit, the following applies:

[0186] - If mixed_nalu_types_in_pic_flag equals 0, then the value of nal_unit_type should be the same for all VCL NAL units of the picture, and the picture or PU is called to have the same NAL unit type as the NAL unit of the codec stripe of the picture or PU.

[0187] - Otherwise (mixed_nalu_types_in_pic_flag equals 1), the picture should have at least two subpicks, and the VCL NAL units of the picture should have exactly two different nal_unit_type values, as follows: the VCL NAL units of at least one subpick of the picture should have a specific value of nal_unit_type equal to STSA_NUT, RADL_NUT, RASL_NUT, IDR_W_RADL, IDR_N_LP or CRA_NUT, while the VCL NAL units of the other subpicks in the picture should have different specific values ​​of nal_unit_type equal to TRAIL_NUT, RADL_NUT or RASL_NUT.

[0188] For single-layer bitstreams, the following constraints apply:

[0189] - Each image in the bitstream, except for the first image in the decoding order, is considered to be associated with a previous IRAP image in the decoding order.

[0190] - When the image is a preceding image to an IRAP image, it should be a RADL or RASL image.

[0191] - When the image is a trailing image of an IRAP image, it should not be a RADL or RASL image.

[0192] - There should be no RASL images associated with IDR images in the bitstream.

[0193] - There should be no RADL image associated with an IDR image that has a nal_unit_type equal to IDR_N_LP in the bitstream.

[0194] Note 6 – Random access can be performed at the location of the IRAP PU (and the IRAP picture and all subsequent non-RASL pictures can be correctly decoded in the decoding order) by discarding all PUs preceding the IRAP PU, provided that each parameter set is available when it is referenced (either in the bitstream or by an external means not specified in this specification).

[0195] - Any image that precedes the IRAP image in the decoding order should precede the IRAP image in the output order, and should precede any RADL image associated with the IRAP image in the output order.

[0196] - Any RASL images associated with a CRA image should precede any RADL images associated with a CRA image in the output order.

[0197] - Any RASL images associated with a CRA image should follow any IRAP images that precede the CRA image in the output order.

[0198] - If field_seq_flag equals 0, and the current image is a preceding image associated with an IRAP image, then it should precede all non-preceding images associated with the same IRAP image in decoding order. Otherwise, assuming picA and picB are the first and last preceding images associated with the IRAP image in decoding order, there should be at most one non-preceding image preceding picA in decoding order, and there should be no non-preceding images between picA and picB in decoding order.

[0199] nuh_temporal_id_plus1 minus 1 specifies the time-domain identifier of the NAL cell.

[0200] The value of nuh_temporal_id_plus1 should not be equal to 0.

[0201] The variable TemporalId is derived as follows:

[0202] TemporalId=nuh_temporal_id_plus1-1 (36)

[0203] When nal_unit_type is in the range from IDR_W_RADL to RSV_IRAP_12 (inclusive), TemporalId should be equal to 0.

[0204] When nal_unit_type equals STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equals 1, TemporalId should not be equal to 0.

[0205] The TemporalId value should be the same for all VCL NAL units of the AU. The TemporalId value of the codec image, PU, ​​or AU is the TemporalId value of the VCL NAL unit of the codec image, PU, ​​or AU. The TemporalId value of the sublayer representation is the maximum value of the TemporalId of all VCL NAL units in the sublayer representation.

[0206] The TemporalId value of non-VCL NAL units is constrained as follows:

[0207] - If nal_unit_type is equal to DCI_NUT, VPS_NUT, or SPS_NUT, then TemporalId should be equal to 0, and the TemporalId of the AU containing the NAL unit should be equal to 0.

[0208] Otherwise, if nal_unit_type is equal to PH_NUT, then TemporalId should be equal to the TemporalId of the PU containing the NAL unit.

[0209] Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, then TemporalId should be equal to 0.

[0210] Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, then TemporalId should be equal to the TemporalId of the AU containing the NAL unit.

[0211] Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT or SUFFIX_APS_NUT, TemporalId should be greater than or equal to the TemporalId of the PU containing the NAL unit.

[0212] Note 7 – When the NAL unit is a non-VCL NAL unit, the TemporalId value is equal to the minimum TemporalId value among all AUs applicable to the non-VCL NAL unit. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, the TemporalId can be greater than or equal to the TemporalId of the AU, because all PPS and APS can be included at the beginning of the bitstream (e.g., when they are transmitted out of band and the receiver places them at the beginning of the bitstream), where the first codec picture has a TemporalId equal to 0.

[0213] 3.7 Adaptive Loop Filter (ALF)

[0214] Two diamond filter shapes (such as) Figure 7 (As shown) is used for block-based ALF. 7×7 rhombuses are applied to the luma component, and 5×5 rhombuses are applied to the chroma component. One of up to 25 filters is selected for each 4×4 block based on the orientation and activity of the local gradient. Each 4×4 block in the image is categorized according to orientation and activity. Before filtering each 4×4 block, simple geometric transformations, such as rotation or diagonal and vertical flips, can be applied to the filter coefficients based on the gradient values ​​calculated for that block. This is equivalent to applying these transformations to samples in the filter's support region. The idea is to make different blocks more similar by aligning the orientation of the different blocks to which the ALF is applied. Block-based categorization is not applied to the chroma component.

[0215] ALF filter parameters are signaled in the Adaptive Parameter Set (APS). Within an APS, a set of up to 25 luma filter coefficients and clipping value indices, and a set of up to 8 chroma filter coefficients and clipping value indices can be signaled. To reduce bit overhead, filter coefficients from different categories of luma components can be merged. In the picture or strip header, up to 7 APS IDs can be signaled to specify the luma filter set used for the current picture or strip. The filtering process is further controlled at the CTB level. The luma CTB can select a filter set from 16 fixed filter sets and the filter sets signaled in the APS. For the chroma component, the APS ID is signaled in the picture or strip header to indicate the chroma filter set used for the current picture or strip. At the CTB level, if there is more than one chroma filter set in the APS, the filter index is signaled for each chroma CTB. When ALF is enabled for a CTB, for each sample within the CTB, a diamond filter with weights provided by signaling is executed, where a clipping operation is applied to clip the difference between neighboring samples and the current sample. The clipping operation introduces non-linearity to make ALF more efficient by reducing the influence of neighboring sample values ​​that differ too much from the current sample value.

[0216] The Cross-Component Adaptive Loop Filter (CC-ALF) can further enhance each chroma component on top of the previously described ALF. The goal of CC-ALF is to refine each chroma component using luminance sample values. This is achieved by applying a diamond-shaped high-pass linear filter and then using the output of this filtering operation for chroma refinement. Figure 8 A system-level diagram of the CC-ALF process for other loop filters is provided. For example... Figure 8 As shown, CC-ALF uses the same input as the luminance ALF to avoid additional steps in the entire loop filtering process.

[0217] 3.8 Luminance Mapping and Chroma Scaling (LMCS)

[0218] Unlike other loop filters (i.e., deblocking, SAO, and ALF) that typically apply filtering to the current sample using information from its spatial neighbors to reduce encoding / decoding artifacts, Luminosity Mapping and Chromaticity Scaling (LMCS) modifies the input signal before encoding by redistributing codewords across the entire dynamic range to improve compression efficiency. LMCS has two main components: (a) loop mapping of the luminosity component based on an adaptive piecewise linear model, and (b) luminosity-dependent chroma residual scaling for the chroma component. Luminosity mapping utilizes the forward mapping function FwdMap and the corresponding inverse mapping function InvMap. The FwdMap function is signaled using a piecewise linear model with 16 equal segments. The InvMap function is not signaled but is derived from the FwdMap function. The luminosity mapping model is signaled in the APS. Up to four LMCS APSs can be used in an encoded / decoded video sequence. When LMCS is enabled for an image, the APS ID is signaled in the image header to identify the APS carrying the luminosity mapping parameters. When LMCS is enabled for a stripe, the InvMap function is applied to all reconstructed luma blocks to convert samples back to the original domain. For inter-frame codec blocks, an additional mapping process is required, which applies the FwdMap function to map the luma prediction blocks in the original domain to the mapped domain after the normal compensation process. Chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. When luma mapping is enabled, signaling notifies an additional flag indicating whether luma-dependent chroma residual scaling is enabled. The chroma residual scaling factor depends on the average of the reconstructed neighboring luma samples at the top and / or left of the current CU. Once the scaling factor is determined, forward scaling is applied to the intra-frame and inter-frame prediction residuals during the coding phase, and inverse scaling is applied to the reconstructed residuals.

[0219] 4 Examples of technical problems solved by the disclosed technical solutions

[0220] The existing design in the latest VVC documentation (in JVET-Q2001-vE / v15) has the following issues:

[0221] 1) Although coefficient encoding and decoding in JVET-N0280 can achieve the encoding and decoding benefits of screen content encoding and decoding, coefficient encoding and decoding and TS mode may still have some drawbacks.

[0222] a) For example, when BDPCM is disabled for a sequence, the context used for encoding and decoding symbol flags (as described in Section 3.2.1) may be non-contiguous. Such a design would make context switching difficult because the memory addresses associated with non-contiguous context indices would also be non-contiguous.

[0223] b) It is unclear whether to use bypass encoding / decoding or context encoding / decoding for symbolic flags in the following cases:

[0224] – The number of remaining allowed context codec bits (represented by RemCcbs) is equal to 0.

[0225] – The current block is encoded and decoded in TS mode.

[0226] –slice_ts_residual_coding_disabled_flag is false (false).

[0227] 2) The repetition of most SEI messages is limited to a maximum of 4 times within the PU, and the repetition of decoded unit information SEI messages is limited to a maximum of 4 times within the DU. However, it is still possible for the same SEI information to be repeated multiple times between two VCL NAL units, which is meaningless.

[0228] 3) Allowing the repetition of non-VCL NAL units other than SEI NAL units between two VCL NAL units, however, this is meaningless.

[0229] 4) APS NAL units can be shared across layers. However, for ALF, filters are highly dependent on reconstruction samples, and different layers can use different QPs even with the same video content. Therefore, in most cases, inheriting ALF filters from images generated for different layers does not provide any negligible codec gain benefit. Therefore, sharing APS NAL units across layers should be prohibited to simplify ALF APS operation. The same applies to other types of APS (i.e., LMCSAPS and scaled list APS).

[0230] 5) The following constraints exist:

[0231] All APS NAL cells within a PU that have specific values ​​for adaptation_parameter_set_id and aps_params_type should have the same content, regardless of whether they are prefix APS NAL cells or suffix APS NAL cells.

[0232] This constraint disallows updates to the contents of APS NAL units within the PU. However, when an image is divided into multiple stripes, it makes sense to allow the encoder to compute the ALF parameters for each strip individually. When the number of stripes is large, for example in 360° video applications, it may be necessary to update the ALF APS NAL units (i.e., change the values ​​of the syntax elements of the APS NAL units, in addition to the NAL unit type, APS ID, and APS type fields).

[0233] 6) All APS NAL elements with a specified value of aps_params_type share the same value space of adaptation_parameter_set_id, regardless of the nuh_layer_id value. However, on the other hand, all APSNAL elements within a PU with specified values ​​of both adaptation_parameter_set_id and aps_params_type should have the same content, regardless of whether they are prefix APS NAL elements or suffix APSNAL elements. Therefore, it makes sense for APS elements of the same type with different NAL element types to also share the same value space of adaptation_parameter_set_id.

[0234] 7) Allows APS NAL units (with specific values ​​for nal_unit_type, adaptation_parameter_set_id, and aps_params_type) to be shared across PUs and even across CLVSs. However, good codec gain is not expected to allow specific APS NAL units to be shared across CLVSs.

[0235] 5. List of technical solutions and embodiments

[0236] To address the above and other issues, the following summarized methods are presented. These items should be considered as examples for explaining general concepts, and not interpreted in a narrow way. Furthermore, these items can be applied individually or combined in any way.

[0237] Solution to Problem 1:

[0238] Let BdpcmFlag be the current BDPCM flag, leftSign be the sign of the left neighbor coefficient, and aboveSign be the sign of the upper neighbor coefficient.

[0239] Let condition M be leftSign equal to 0 and aboveSign equal to 0, or leftSign equal to -aboveSign.

[0240] 1. It is proposed that if the current block is encoded or decoded in mode X, the allowed context index offset for encoding or decoding the symbol flag can be within a first range [N0, N1], otherwise within [N2, N3].

[0241] a. In one example, mode X can be a luminance and / or chrominance BDPCM mode.

[0242] b. In one example, N0, N1, N2, and N3 can be 0, 2, 3, and 5 respectively.

[0243] c. In one example, when the condition M is false, the above example can be applied.

[0244] d. In one example, when the transform skip is applied to a block, the above example can be applied.

[0245] e. N0, N1, N2, and N3 can be determined based on one or more of the following:

[0246] i. Indications signaled in SPS / VPS / PPS / picture header / strip header / slice group header / LCU row / LCU group / LCU / CU

[0247] ii. Block dimensions of the current block and / or its neighboring blocks

[0248] iii. Block shapes of the current block and / or its neighboring blocks

[0249] iv. Indications of color format (such as 4:2:0, 4:4:4)

[0250] v. Whether to use a separate codec tree structure or a dual codec tree structure

[0251] vi. Strip type and / or picture type

[0252] vii. Number of color components

[0253] 2. It is proposed that when the number of remaining allowed context coding binary bits (represented by RemCcbs) is less than N, the sign flag is coded in bypass mode

[0254] a. In one example, when RemCcbs < N, the sign flag is coded in bypass mode.

[0255] i. Alternatively, in one example, when RemCcbs >= N, the sign flag is coded in context mode.

[0256] b. In one example, when RemCcbs is equal to N, the sign flag is coded in bypass mode.

[0257] i. Alternatively, in one example, when RemCcbs > N, the sign flag is coded in bypass mode.

[0258] c. In one example, N can be set to be equal to 4.

[0259] i. Alternatively, in one example, N can be set to be equal to 0.

[0260] d. In one example, N is an integer and can be based on

[0261] i. Signaling notification instructions in SPS / VPS / PPS / Image header / Strip header / Piece group header / LCU line / LCU group / LCU / CU

[0262] ii. Block dimensions of the current block and / or its neighboring blocks

[0263] iii. The block shape of the current block and / or its neighboring blocks

[0264] iv. Instructions for color format (such as 4:2:0, 4:4:4)

[0265] v. Use a single codec tree structure or a dual codec tree structure

[0266] vi. Strip type and / or image type

[0267] vii. Number of color components

[0268] e. The above examples can be applied to transform blocks and / or transform skip blocks that include or exclude BDPCM codec blocks.

[0269] 3. The context index offset used for encoding and decoding symbolic flags can be determined based on one or more of the following:

[0270] a. Signaling notification instructions in SPS / VPS / PPS / Image Header / Strip Header / Piece Group Header / LCU Line / LCU Group / LCU / CU

[0271] b. Block dimensions of the current block and / or its neighboring blocks

[0272] c. Block shape of the current block and / or its neighboring blocks

[0273] i. In one example, the context value of the symbol flag can be separate for different block shapes.

[0274] d. Prediction modes of neighboring blocks of the current block (intra-frame / inter-frame)

[0275] i. In one example, for intra-blocks and inter-blocks, the context value of the symbol flag can be separate.

[0276] e. Indication of the BDPCM mode of the current block's neighboring blocks

[0277] f. Indication of color format (such as 4:2:0, 4:4:4)

[0278] g. Whether to use a single codec tree structure or a dual codec tree structure.

[0279] h. Strip type and / or image type

[0280] i. Number of color components

[0281] i. In one example, for the luminance color component and the chrominance color component, the context value of the symbol can be separate.

[0282] Solutions to problems 2 through 5:

[0283] The term "strip cell" is defined as follows:

[0284] Stripe Unit (SU): A set of NAL units that are consecutive in the decoding order and contain exactly one codec stripe and all its associated non-VCL NAL units.

[0285] 4. To solve problem 2, it is possible to prevent one or more of all types of SEI messages from being repeated within the SU.

[0286] a. In one example, any SEI message with a specific payloadType value that has specific content is not allowed to be repeated within the SU, and the constraint is specified as follows: the number of equivalent sei_payload() syntax structures with any specific value of payloadType within the SU should not be greater than 1.

[0287] b. Alternatively, any SEI message with a specific payloadType value (regardless of its content) is not allowed to be repeated within SU.

[0288] i. In one example, the constraint is specified as follows: the number of sei_payload() syntax structures with any specific value of payloadType within SU should not be greater than 1.

[0289] c. Alternatively, any SEI message with a specific payloadType is not allowed to be updated within SU.

[0290] i. In one example, the constraint is specified as follows: All sei_payload() syntax constructs within SU that have a specific value for payloadType should have the same content.

[0291] 5. To address issue 3, for one or more non-VCL NAL types other than PREFIX_SEI_NUT and SUFFIX_SEI_NUT, duplication of non-VCL NAL units within SU may be disallowed.

[0292] a. In one example, one or more of the following constraints are specified:

[0293] i. In one example, the constraint is specified as follows: the number of DCI NAL cells within the SU should not be greater than 1.

[0294] ii. In one example, the constraint is specified as follows: the number of VPS NAL units within the SU that have a specific value of vps_video_parameter_set_id should not be greater than 1.

[0295] iii. In one example, the constraint is specified as follows: the number of SPS NAL cells within the SU with a specific value of sps_seq_parameter_set_id should not be greater than 1.

[0296] iv. In one example, the constraint is specified as follows: the number of PPS NAL cells within the SU with a specific value of pps_pic_parameter_set_id should not be greater than 1.

[0297] v. In one example, the constraint is specified as follows: The number of APS NAL cells within the SU with specific values ​​for adaptation_parameter_set_id and aps_params_type should not be greater than 1.

[0298] vi. In one example, the constraint is specified as follows: the number of APS NAL units within the SU that have specific values ​​for nal_unit_type, adaptation_parameter_set_id, and aps_params_type should not be greater than 1.

[0299] b. Alternatively or additionally, DCI NAL units are not allowed to be repeated within CLVS or CVS.

[0300] c. Alternatively or additionally, VPS NAL units with a specific value of vps_video_parameter_set_id are not allowed to be repeated within CLVS or CVS.

[0301] d. Alternatively or additionally, SPSNAL cells with a specific value of sps_seq_parameter_set_id are not allowed to be repeated within CLVS.

[0302] 6. To address issue 4, it can be specified that all types of APS should not be shared across layers.

[0303] a. Alternatively, it can be specified that ALF APS should not be shared across layers.

[0304] 7. To solve problem 5, you can specify that ALF APS can be updated within the PU.

[0305] a. Alternatively, it can be specified that no adaptive parameter set is allowed to be updated within the image cell.

[0306] 8. To address issue 6, it is possible to specify that all APS NAL cells with a specific value of aps_params_type share the same value space of adaptation_parameter_set_id, regardless of the nuh_layer_id value or whether they are prefix APS NAL cells or suffix APS NAL cells.

[0307] a. Alternatively, LMCS APS NAL elements can be specified to share the same value space of adaptation_parameter_set_id, regardless of the nuh_layer_id value and whether they are prefix APS NAL elements or suffix APS NAL elements.

[0308] b. Alternatively, ALF APS NAL cells can be specified to share the same value space of adaptation_parameter_set_id, regardless of the nuh_layer_id value and whether they are prefix APS NAL cells or suffix APS NAL cells.

[0309] 9. To address issue 6, it can be specified that APS NAL units (with specific values ​​for nal_unit_type, adaptation_parameter_set_id, and aps_params_type) should not be shared across CLVS.

[0310] a. In one example, the APS NAL unit referenced by the specified VCL NAL unit vclNalUnitA should not be a non-VCL NAL unit associated with a VCL NAL unit in a CLVS different from the CLVS containing vclNalUnitA.

[0311] b. Alternatively, the APS NAL unit referenced by the specified VCL NAL unit vclNalUnitA should not be a non-VCLNAL unit associated with an IRAP image that is different from the IRAP image associated with vclNalUnitA.

[0312] c. Alternatively, the APS NAL unit referenced by the specified VCL NAL unit vclNalUnitA should not be a non-VCL NAL unit associated with an IRAP or GDR image that is different from the IRAP or GDR image associated with vclNalUnitA.

[0313] d. Alternatively, it can be specified that APS should not be shared across CVS.

[0314] i. Alternatively, it can be specified that APS with a specific APS type (e.g., ALF, LMCS, SCALING) should not be shared across CVS.

[0315] 6 Examples

[0316] These embodiments are based on JVET-Q2001-vE. Deleted text is marked with double braces (e.g., [[]]), where the deleted text is enclosed within the double braces. Added text is marked with... mark.

[0317] 6.1 Example #1

[0318]

[0319] The inputs to this process are the color component index cIdx, the brightness position (x0, y0) of the top-left sample relative to the top-left sample of the current transform block in the current image, and the current coefficient scan position (xC, yC).

[0320] The output of this process is the variable ctxInc.

[0321] The variables leftSign and aboveSign are derived as follows:

[0322] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1595)

[0323] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1596)

[0324] The variable ctxInc is derived as follows:

[0325] – If leftSign equals 0 and aboveSign equals 0, or if leftSign equals

[0326] The -aboveSign option applies to the following:

[0327] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?0:3) (1597)

[0328] Otherwise, if leftSign is greater than or equal to 0, and aboveSign is greater than or equal to 0,

[0329] The following applies:

[0330] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?1:4) (1598)

[0331] –Otherwise, the following applies:

[0332] ctxInc=(BdpcmFlag[x0][y0][cIdx]==0?2:5) (1599)

[0333] 6.2 Example #2

[0334]

[0335] The inputs to this process are the color component index cIdx, the brightness position (x0, y0) of the top-left sample relative to the top-left sample of the current transform block in the current image, and the current coefficient scan position (xC, yC).

[0336] The output of this process is the variable ctxInc.

[0337] The variables leftSign and aboveSign are derived as follows:

[0338] leftSign=(xC==0)? 0:CoeffSignLevel[xC-1][yC] (1595)

[0339] aboveSign=(yC==0)? 0:CoeffSignLevel[xC][yC-1] (1596)

[0340] The variable ctxInc is derived as follows:

[0341] – If leftSign equals 0 and aboveSign equals 0, or if leftSign equals -aboveSign, then the following applies:

[0342] ctxInc=[[(BdpcmFlag[x0][y0][cIdx]==0?]]0[[:3)]] (1597)

[0343] Otherwise, if leftSign is greater than or equal to 0 and aboveSign is greater than or equal to 0, the following applies:

[0344] ctxInc=[[(BdpcmFlag[x0][y0][cIdx]?]]1[[:4)]] (1598)

[0345] –Otherwise, the following applies:

[0346] ctxInc=[[(BdpcmFlag[x0][y0][cIdx]?]]2[[:5)]] (1599)

[0347] – If BdpcmFlag[x0][y0][cIdx] equals 1, then the following applies:

[0348] ctxInc = ctxInc + 3

[0349] 6.3 Example #3

[0350] [xS][yS] Specifies the following for the sub-block at position (xS, yS) within the current transform block, where the sub-block is an array of transform coefficients:

[0351] – If sb_coded_flag[xS][yS] equals 0, then the 16 transform coefficient levels of the sub-block at position (xS,yS) are inferred to be equal to 0.

[0352] – Otherwise (sb_coded_flag[xS][yS] equals 1), the following applies:

[0353] – If (xS, yS) equals (0, 0) and (LastSignificantCoeffX, LastSignificantCoeffY) does not equal (0, 0), then for the sub-block at position (xS, yS), there exists [

[16] ]. At least one of the sig_coeff_flag syntax elements.

[0354] Otherwise, the sub-block at position (xS, yS) [

[16] ] At least one of the transformation coefficient levels has a non-zero value.

[0355] When sb_coded_flag[xS][yS] does not exist, it is inferred to be equal to 1.

[0356] 6.4 Example #4

[0357] This embodiment pertains to items 5 through 7.

[0358] 7.4.3.5 Adaptive Parameter Set Semantics

[0359] Each APS RBSP should be available for the decoding process before being referenced, and it is included in at least one AU of the TemporalId of the NAL unit of the codec strip that references it, or provided by external means.

[0360] The PU contains a specific value for adaptation_parameter_set_id and All APS NAL cells with a specific value of aps_params_type should have the same content, regardless of whether they are prefix APS NAL cells or suffix APS NAL cells.

[0361] The adaptation_parameter_set_id provides an identifier for the APS for reference by other syntax elements.

[0362] When aps_params_type is equal to ALF_APS or SCALING_APS, the value of adaptation_parameter_set_id should be in the range of 0 to 7 (inclusive).

[0363] When aps_params_type equals LMCS_APS, the value of adaptation_parameter_set_id should be in the range of 0 to 3 (inclusive).

[0364] Let apsLayerId be the value of nuh_layer_id for a specific APS NAL cell, and let vclLayerId be the value of nuh_layer_id for a specific VCL NAL cell. A specific VCL NAL cell should not reference a specific APS NAL cell unless apsLayerId is less than or equal to vclLayerId and the layer whose nuh_layer_id is equal to apsLayerId is included in at least one OLS that includes the layer whose nuh_layer_id is equal to vclLayerId.

[0365] Specify the type of APS parameters carried in the APS, as specified in Table 6.

[0366] 1. Table 6 – APS Parameter Type Codes and APS Parameter Types

[0367] aps_params_type The name of aps_params_type Types of APS parameters 0 ALF_APS ALF parameters 1 LMCS_APS LMCS parameters 2 SCALING_APS Scaling list parameters 3..7 reserve reserve

[0368] All APS NAL cells with a specific value of aps_params_type share the same value space for adaptation_parameter_set_id, regardless of nuh_layer_id. how APS NAL cells with different values ​​of aps_params_type use a separate value space for adaptation_parameter_set_id.

[0369] Note 1 – APS NAL unit (with) Specific values ​​for adaptation_parameter_set_id and aps_params_type can be shared across images, and different stripes within an image can reference different ALF APSs.

[0370] Note 2 – The suffix APS NAL unit associated with a specific VCL NAL unit (which precedes the suffix APS NAL unit in the decoding order) is not used by the specific VCL NAL unit, but by the VCL NAL unit that follows the suffix APS NAL unit in the decoding order.

[0371] A value of 0 indicates that the `aps_extension_data_flag` syntax element does not exist in the APS RBSP syntax structure. A value of 1 indicates that the `aps_extension_data_flag` syntax element exists in the APS RBSP syntax structure.

[0372] It can have any value. Its presence and value do not affect the grade specified in this version of the decoder that conforms to this specification. Decoders conforming to this version of the specification should ignore all aps_extension_data_flag syntax elements.

[0373] 6.5 Examples of Encoding and Decoding Symbols

[0374] In one embodiment, bypass encoding / decoding is used when RemCcbs is less than 0.

[0375] Table 131 – Assigning ctxInc to syntax elements with context encoding / decoding bits

[0376]

[0377] Table 131 – Assigning ctxInc to syntax elements with context encoding / decoding bits

[0378]

[0379] Alternatively, in one embodiment, when the current block is encoded / decoded in transform skip mode and RemCcbs 0 and ! slice_ts_residual_coding_disabled_flag, use context encoding / decoding.

[0380] Table 131 – Assigning ctxInc to syntax elements with context encoding / decoding bits

[0381]

[0382] Alternatively, in one embodiment, bypass encoding / decoding is used when RemCcbs is less than 4.

[0383] Table 131 – Assigning ctxInc to syntax elements with context encoding / decoding bits

[0384]

[0385] Figure 1 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0386] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described herein. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection as indicated by component 1906. The bitstream (or codec) representation of the video received at input 1902, whether stored or communicated, can be used by component 1908 to generate pixel values ​​or transmit as displayable video to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.

[0387] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0388] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple) 3602 can be configured to implement one or more methods described in this document. The memories (multiple) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in a hardware circuit system.

[0389] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0390] like Figure 4As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, and this source device 110 may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110, and this target device 120 may be referred to as a video decoding device.

[0391] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0392] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0393] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0394] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with an external display device.

[0395] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or additional standards.

[0396] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 4 The video encoder 114 in the system 100 shown.

[0397] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0398] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0399] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0400] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretive purposes, in Figure 5 The example is represented separately.

[0401] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0402] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).

[0403] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0404] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip.

[0405] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0406] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0407] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.

[0408] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0409] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0410] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0411] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling Notification.

[0412] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0413] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0414] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0415] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0416] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0417] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.

[0418] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0419] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0420] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of that tool or mode. That is, when a video processing tool or mode is enabled based on that decision or determination, the conversion from video blocks to video bitstream (or bitstream representation) will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from video bitstream to video blocks will be performed using the video processing tool or mode enabled based on that decision or determination.

[0421] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 4 The video decoder 114 in the system 100 shown.

[0422] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0423] exist Figure 6In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The encoding process described is the opposite of the decoding process.

[0424] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and from the entropy-coded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.

[0425] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0426] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0427] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0428] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0429] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on the display device.

[0430] The following is a list of preferred technical solutions for some embodiments.

[0431] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0432] 1. A video processing method (e.g., Figure 3 The method 3000 shown includes performing a conversion between the current block of the video and the codec representation of the video (3002); wherein the codec representation conforms to a format rule specifying a context index offset within a range for encoding and decoding symbolic flags of the current block in the codec representation, wherein a range function of the codec mode is used to represent the current block in the codec representation.

[0433] 2. According to the method of technical solution 1, when the encoding / decoding mode is a specific mode, the rule specifies a range of [N0, N1], otherwise it is [N2, N3].

[0434] 3. The method according to technical solution 2, wherein the specific mode is a chroma or luminance block differential pulse encoding / decoding modulation mode.

[0435] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0436] 4. A video processing method, comprising: performing a conversion between a current block of a video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that the symbol flag of the current block is encoded and decoded in the codec representation using either a context codec bit or a bypass mode based on the number of remaining context codec bits.

[0437] 5. The method according to technical solution 4, wherein the bypass mode is used for encoding and decoding if and only if the number of remaining context encoding / decoding bits is less than N, where N is a positive integer.

[0438] 6. The method according to technical solution 5, wherein the bypass mode is used for encoding and decoding if and only if the number of remaining context encoding / decoding bits is equal to N, where N is a positive integer.

[0439] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0440] 7. A video processing method, comprising: performing a conversion between a current block of a video and a codec representation of the video; wherein the conversion uses a block-based incremental pulse codec modulation (BDPCM) mode, wherein the codec representation conforms to a format rule specifying that symbolic flags from the BDPCM are context-coded in the codec representation such that a context index offset used for encoding and decoding the symbolic flags is a function of the encoding and decoding conditions of the current block.

[0441] 8. The method according to technical solution 7, wherein the encoding / decoding conditions correspond to the syntax elements included in the encoding / decoding representation in the sequence parameter set, video parameter set, picture parameter set, picture header, strip header, slice group header, logical codec unit level, LCU group, or codec unit level.

[0442] 9. The method according to technical solution 7, wherein the encoding / decoding conditions correspond to the block dimension of the current block and / or neighboring blocks.

[0443] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 4).

[0444] 10. A video processing method, comprising: performing a conversion between a current block of video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that at most once, a stripe unit in the codec representation corresponding to a set of network abstraction layer units in sequential decoding order and containing a single codec stripe includes at least a portion of auxiliary enhancement information (SEI).

[0445] 11. The method according to technical solution 10, wherein a portion of the SEI corresponds to the entire SEI.

[0446] 12. The method according to technical solution 10, wherein a portion of the SEI corresponds to a specific type of SEI information field.

[0447] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 5).

[0448] 13. A video processing method, comprising: performing a conversion between a current block of video and a codec representation of the video; wherein the codec representation conforms to a format rule specifying that a stripe unit in the codec representation corresponding to a set of network abstraction layer units in a sequential decoding order and containing a single codec stripe includes one or more Video Codec Layer Network Abstraction Layer (VCL NAL) units, wherein the format rule further specifies a first type of unit that is allowed to be repeated in a stripe unit and a second type of unit that is not allowed to be repeated in a stripe unit.

[0449] 14. The method according to technical solution 13, wherein the second type of unit includes units having a range of types.

[0450] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 6 and 7).

[0451] 15. A video processing method comprising: performing a conversion between a video comprising one or more video layers and a codec representation of the video according to a rule; wherein the codec representation includes one or more adaptive parameter sets (APS); and wherein the rule specifies the applicability of some of the one or more APS to the conversion of the one or more video layers.

[0452] 16. The method according to technical solution 15, wherein the rule specifies disabling all sharing across one or more video layers in one or more APSs.

[0453] The following technical solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., items 8 and 9).

[0454] 17. A video processing method comprising: performing a conversion between a video comprising one or more video layers and a codec representation of the video according to rules; wherein the codec representation is arranged in one or more network abstraction layer (NAL) units; wherein the codec representation includes one or more adaptive parameter sets for controlling characteristics of the conversion.

[0455] 18. The method according to technical solution 17, wherein the rule specifies that all adaptive parameter sets of a particular type also share the same spatial values ​​of the corresponding identifier values.

[0456] 19. The method according to technical solution 17, wherein the rule specifies that one or more video layers are not allowed to share a specific adaptive parameter set network abstraction layer unit.

[0457] 20. The method according to any one of technical solutions 1 to 19, wherein the conversion includes encoding the video into a codec representation.

[0458] 21. The method according to any one of technical solutions 1 to 19, wherein the conversion includes decoding the codec representation to generate pixel values ​​of the video.

[0459] 22. A video decoding apparatus, comprising a processor configured to implement one or more of the methods according to claims 1 to 21.

[0460] 23. A video encoding apparatus, comprising a processor configured to implement one or more of the methods according to claims 1 to 21.

[0461] 24. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 21.

[0462] 25. A method, apparatus or system described in this document.

[0463] Figure 9 This is a flowchart of an example method 900 for video processing. Operation 902 includes performing a conversion between the current block of the video and the bitstream of the video, wherein the bitstream conforms to a rule for including a first symbol flag of a first coefficient in the bitstream by a specified context index offset, wherein the rule specifies that the value of the context index offset is based on whether a first codec mode is applied to the current block in the bitstream.

[0464] In some embodiments of method 900, the rule specifies that, without applying the first codec mode, the value of the context index offset is within a first range from N0 to N1, the first range including N0 and N1, and the rule specifies that, with applying the first codec mode, the value of the context index offset is within a second range from N2 to N3, the second range including N2 and N3. In some embodiments of method 900, in the first codec mode, pulse codec modulation is used to represent the difference between the quantization residual and the prediction of the quantization residual in the bitstream. In some embodiments of method 900, N0, N1, N2, and N3 are 0, 2, 3, and 5, respectively. In some embodiments of method 900, in response to the first neighboring symbol value and the second neighboring symbol value satisfying a first condition, the value of the context index offset is equal to 0 or 3, and the first condition is: (1) the first neighboring symbol value is equal to 0 and the second neighboring symbol value is equal to 0, or (2) the first neighboring symbol value is equal to the negative of the second neighboring symbol value, and wherein in response to the horizontal coordinate of the first coefficient being 0, the first neighboring symbol value is equal to 0, and in response to the horizontal coordinate of the first coefficient being not 0, the first neighboring symbol value is equal to the symbol value of the left neighboring coefficient, and in response to the vertical coordinate of the first coefficient being 0, the second neighboring symbol value is equal to 0, and in response to the vertical coordinate of the first coefficient being not 0, the second neighboring symbol value is equal to the symbol value of the upper neighboring coefficient.

[0465] In some embodiments of method 900, in response to the first codec mode not being applied to the current block, the value of the context index offset is equal to 0, and in response to the first codec mode being applied to the current block, the value of the context index offset is equal to 3. In some embodiments of method 900, in response to the first neighboring symbol value and the second neighboring symbol value not satisfying a first condition, and in response to the first neighboring symbol value and the second neighboring symbol value satisfying a second condition, the value of the context index offset is equal to 1 or 4, and the second condition is: (1) the first neighboring symbol value is greater than or equal to 0, and the second neighboring symbol value is greater than or equal to 0. In some embodiments of method 900, in response to the first codec mode not being applied to the current block, the value of the context index offset is equal to 1, and in response to the first codec mode being applied to the current block, the value of the context index offset is equal to 4. In some embodiments of method 900, in response to the first neighboring symbol value and the second neighboring symbol value not satisfying the first condition and the second condition, the value of the context index offset is equal to 2 or 5.

[0466] In some embodiments of method 900, the context index offset is equal to 2 in response to the first codec mode not being applied to the current block, and equal to 5 in response to the first codec mode being applied to the current block. In some embodiments of method 900, a transform skip operation is applied to the current block. In some embodiments of method 900, the rule specifies that N0, N1, N2, and N3 are determined based on any one or more of the following: (1) the sequence parameter set, video parameter set, picture parameter set, picture header, strip header, slice header, maximum codec unit line, maximum codec unit group, maximum codec unit, or indication included in the codec unit; (2) the block dimension of the current block and / or the block dimension of the current block's neighboring blocks; (3) the block shape of the current block and / or the block shape of the current block's neighboring blocks; (4) the indication of the video's color format; (5) whether a single codec tree structure or a dual codec tree structure is applied to the current block; (6) the strip type and / or picture type to which the current block belongs; and (7) the number of color components of the current block.

[0467] Figure 10 This is a flowchart of an example method 1000 for video processing. Operation 1002 includes performing a conversion between the current block of the video and the bitstream of the video, wherein the bitstream conforms to the rule that the symbol flag of the specified current block is included in the bitstream using either a context mode or a bypass mode based on the number of remaining context codec bits.

[0468] In some embodiments of method 1000, the rule specifies that, in response to the number of remaining context codec bits being less than N, a bypass mode is used to include a symbol flag in the bitstream, where N is an integer. In some embodiments of method 1000, the rule specifies that, in response to the number of remaining context codec bits being greater than or equal to N, a context mode is used to include a symbol flag in the bitstream, where N is an integer. In some embodiments of method 1000, the rule specifies that, in response to the number of remaining context codec bits being equal to N, a bypass mode is used to include a symbol flag in the bitstream, where N is an integer. In some embodiments of method 1000, the rule specifies that, in response to the number of remaining context codec bits being greater than N, a bypass mode is used to include a symbol flag in the bitstream, where N is an integer. In some embodiments of method 1000, N equals 4. In some embodiments of method 1000, N equals 0.

[0469] In some embodiments of method 1000, N is based on the following integers: (1) sequence parameter set, video parameter set, picture parameter set, picture header, strip header, slice header, maximum codec unit row, maximum codec unit group, maximum codec unit, or an indication included in a codec unit; (2) the block dimension of the current block and / or the block dimension of the current block's neighboring blocks; (3) the block shape of the current block and / or the block shape of the current block's neighboring blocks; (4) an indication of the video's color format; (5) whether a single codec tree structure or a dual codec tree structure is applied to the current block; (6) the strip type and / or picture type to which the current block belongs; and (7) the number of color components in the current block. In some embodiments of method 1000, the rule specifies that the current block is a transform block or a transform skip block. In some embodiments of method 1000, the rule specifies that a block differential pulse codec modulation mode is applied to the current block. In some embodiments of method 1000, the rule specifies that a block differential pulse codec modulation mode is not applied to the current block.

[0470] Figure 11 This is a flowchart of an example method 1100 for video processing. Operation 1102 includes performing a conversion between the current block of the video and the bitstream of the video, wherein the bitstream conforms to a rule specifying a context index offset for including the symbol flags of the current block in the bitstream, and wherein the rule specifies that the context index offset is determined based on information about the current block.

[0471] In some embodiments of method 1100, the information includes a sequence parameter set, a video parameter set, a picture parameter set, a picture header, a stripe header, a slice header, a maximum codec unit line, a maximum codec unit group, a maximum codec unit, or an indication within a codec unit. In some embodiments of method 1100, the information includes the block dimension of the current block and / or the block dimensions of neighboring blocks of the current block. In some embodiments of method 1100, the information includes the block shape of the current block and / or the block shapes of neighboring blocks of the current block. In some embodiments of method 1100, the context value of the symbol flag is separate for different block shapes. In some embodiments of method 1100, the information includes one or more prediction modes of one or more neighboring blocks of the current block. In some embodiments of method 1100, the one or more prediction modes include intra-frame prediction modes. In some embodiments of method 1100, the one or more prediction modes include inter-frame prediction modes.

[0472] In some embodiments of method 1100, the context value of the symbol flag is separate for one or more inter-frame codec blocks and for one or more intra-frame codec blocks. In some embodiments of method 1100, the information includes an indication of the block-based delta pulse codec modulation mode of the neighboring blocks of the current block. In some embodiments of method 1100, the information includes an indication of the color format of the video. In some embodiments of method 1100, the information includes whether a single codec tree structure or a dual codec tree structure is applied to the current block. In some embodiments of method 1100, the information includes the stripe type and / or picture type to which the current block belongs. In some embodiments of method 1100, the information includes the number of color components of the current block. In some embodiments of method 1100, the context value of the symbol flag is separate for the luma color component of the current block and for the chroma color component of the current block. In some embodiments of methods(s) 900-1100, the current block is represented in the bitstream using a binary delta pulse codec modulation mode.

[0473] Figure 12 This is a flowchart of an example method 1200 for video processing. Operation 1202 includes performing a conversion between a video and a video bitstream comprising one or more video layers according to a rule, wherein the rule specifies the use of multiple adaptive parameter set network abstraction layer units (APNs) for the video, each APN having a corresponding adaptive parameter type value, each APN being associated with a corresponding video layer identifier, each APN being a prefix unit or a suffix unit, and the rule specifies that, in response to multiple APNs sharing the same adaptive parameter type value, the adaptive parameter set identifier values ​​of the multiple APNs belong to the same identifier space.

[0474] In some embodiments of method 1200, the rule specifies that multiple adaptive parameter set network abstraction layer units (APNs) share the same adaptive parameter type values, regardless of the multiple video layer identifiers of the multiple APNs. In some embodiments of method 1200, the rule specifies that multiple APNs share the same adaptive parameter type values, regardless of whether the multiple APNs are prefix units or suffix units. In some embodiments of method 1200, the multiple APNs are multiple first-mode APNs, wherein in the first mode, mapped luminance prediction samples are derived based on a linear model and further used to derive luminance reconstruction samples, and chroma residual values ​​are scaled based on the luminance reconstruction samples. In some embodiments of method 1200, the multiple APNs are multiple second-mode APNs, wherein in the second mode, luminance reconstruction samples are filtered using a classification operation that generates filter indices based on the difference between luminance reconstruction samples in different directions, and chroma reconstruction samples are filtered without a classification operation.

[0475] Figure 13 This is a flowchart of an example method 1300 for video processing. Operation 1302 includes performing a conversion between the current block of video and the bitstream of video, wherein the bitstream conforms to a rule that specifies that a first auxiliary enhancement message with specific characteristics is not allowed to be repeated within a slice unit in the bitstream in response to a second auxiliary enhancement message with specific characteristics being included in a slice unit, or (2) updated within a slice unit in the bitstream in response to a first auxiliary enhancement message being included in a slice unit, wherein the slice unit comprises a set of network abstraction layer units that are consecutive in decoding order, and the set of network abstraction layers comprises a single codec slice and one or more non-video codec layer network abstraction layer units associated with the single codec slice.

[0476] In some embodiments of method 1300, the rule specifies that, in response to a first auxiliary enhancement message having a value of the same specific payload type as a second auxiliary enhancement message included in a stripe unit, the first auxiliary enhancement message is not allowed to be repeated within the stripe unit. In some embodiments of method 1300, the rule specifies that the number of equivalent syntax structures of a specific payload type within a stripe unit is not greater than one. In some embodiments of method 1300, the rule specifies that, in response to a first auxiliary enhancement message having a specific payload type, the first auxiliary enhancement message is not allowed to be repeated within a stripe unit. In some embodiments of method 1300, the rule specifies that the number of syntax structures of a specific payload type within a stripe unit is not greater than one.

[0477] In some embodiments of method 1300, the rule specifies that, in response to a first auxiliary enhancement message having a specific payload type, the first auxiliary enhancement message is not allowed to be updated within a stripe cell. In some embodiments of method 1300, the rule specifies that one or more syntax structures of a specific payload type within a stripe cell have the same content.

[0478] Figure 14 This is a flowchart of an example method 1400 for video processing. Operation 1402 includes performing a conversion between the current block of video and the bitstream of video, wherein the bitstream conforms to a rule specifying that in response to: (1) the stripe units in the bitstream include a second non-video codec layer network abstraction layer unit having the same characteristics as the first non-video codec layer network abstraction layer unit, and (2) the first non-video codec layer network abstraction layer unit has a network abstraction layer unit type other than prefix auxiliary enhancement information or suffix auxiliary enhancement information, the stripe units in the bitstream are not allowed to be repeated by the first non-video codec layer network abstraction layer unit.

[0479] In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a decoding capability information network abstraction layer unit, and the rule specifies that the number of decoding capability information network abstraction layer units within a stripe unit is no greater than 1. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a video parameter set network abstraction layer unit with a specific identifier, and the rule specifies that the number of video parameter set network abstraction layer units with a specific identifier within a stripe unit is no greater than 1. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a sequence parameter set network abstraction layer unit with a specific identifier, and the rule specifies that the number of sequence parameter set network abstraction layer units with a specific identifier within a stripe unit is no greater than 1. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a picture parameter set network abstraction layer unit with a specific identifier, and the rule specifies that the number of picture parameter set network abstraction layer units with a specific identifier within a stripe unit should not be greater than 1.

[0480] In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is an adaptive parameter set network abstraction layer unit with a specific identifier and a specific parameter type, and the rule specifies that the number of adaptive parameter set network abstraction layer units with a specific identifier and a specific parameter type within a stripe unit is no greater than 1. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is an adaptive parameter set network abstraction layer unit with a specific type, specific identifier, and specific parameter type of a network abstraction layer unit, and the rule specifies that the number of adaptive parameter set network abstraction layer units with a specific type, specific identifier, and specific parameter type of a network abstraction layer unit within a stripe unit is no greater than 1. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a decoding capability information network abstraction layer unit, and the rule specifies that, in response to the first non-video codec layer network abstraction layer unit being a decoding capability information network abstraction layer unit, the first non-video codec layer network abstraction layer unit is not allowed to be repeated within a codec layer video sequence or a codec video sequence.

[0481] In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a video parameter set network abstraction layer unit with a specific identifier, and the rule specifies that, in response to the first non-video codec layer network abstraction layer unit being a video parameter set network abstraction layer unit with a specific identifier, the first non-video codec layer network abstraction layer unit is not allowed to be repeated within a codec layer video sequence or a codec video sequence. In some embodiments of method 1400, the second non-video codec layer network abstraction layer unit is a sequence parameter set network abstraction layer unit with a specific identifier, and the rule specifies that, in response to the first non-video codec layer network abstraction layer unit being a sequence parameter set network abstraction layer unit with a specific identifier, the first non-video codec layer network abstraction layer unit is not allowed to be repeated within a codec layer video sequence or a codec video sequence.

[0482] Figure 15 This is a flowchart of example method 1500 for video processing. Operation 1502 includes performing a conversion between video and video bitstreams comprising multiple video layers according to a rule, wherein the rule specifies which adaptive parameter sets from multiple adaptive parameter sets are not allowed to be shared across multiple video layers.

[0483] In some embodiments of method 1500, the rule specifies that multiple adaptive parameter sets cannot be shared. In some embodiments of method 1500, the rule specifies that multiple adaptive parameter sets of an adaptive loop filter type cannot be shared.

[0484] Figure 16This is a flowchart of example method 1600 for video processing. Operation 1602 includes performing a conversion between video and a video bitstream comprising one or more video layers according to a rule, wherein the bitstream comprises one or more adaptive loop filter adaptive parameter sets, and wherein the rule specifies whether one or more adaptive loop filter adaptive parameter sets are allowed to be updated within a picture cell.

[0485] In some embodiments of method 1600, the rule specifies that one or more adaptive loop filter adaptive parameter sets are allowed to be updated within a picture cell. In some embodiments of method 1600, the rule specifies that no adaptive parameter set is allowed to be updated within a picture cell.

[0486] Figure 17 This is a flowchart of example method 1700 for video processing. Operation 1702 includes performing a conversion between video and a bitstream of video comprising one or more codec layer video sequences, according to a rule, wherein the bitstream comprises an adaptive loop filter adaptive parameter set, and wherein the rule specifies that, in response to the adaptive loop filter adaptive parameter set having one or more specific characteristics, the adaptive loop filter adaptive parameter set is not allowed to be shared across one or more codec layer video sequences.

[0487] In some embodiments of method 1700, one or more specific features include an adaptive loop filter adaptive parameter set having a specific network abstraction layer (NALB) unit type. In some embodiments of method 1700, one or more specific features include an adaptive loop filter adaptive parameter set having a specific identifier. In some embodiments of method 1700, one or more specific features include an adaptive loop filter adaptive parameter set having a specific parameter type. In some embodiments of method 1700, the rule specifies that the adaptive loop filter adaptive parameter set referenced by a video codec layer NLB unit in a first codec layer video sequence is not associated with a non-video codec layer NLB unit in a second codec layer video sequence, and the first codec layer video sequence is different from the second codec layer video sequence. In some embodiments of method 1700, the rule specifies that the adaptive loop filter adaptive parameter set referenced by a video codec layer NLB unit associated with a first intra-frame random access picture is not associated with a non-video codec layer NLB unit associated with a second intra-frame random access picture, and the first intra-frame random access picture is different from the second intra-frame random access picture.

[0488] In some embodiments of method 1700, the rule specifies that the adaptive loop filter adaptive parameter set referenced by the video codec layer network abstraction layer unit associated with the first intra-frame random access picture or the first progressively decoded refresh picture is not associated with the non-video codec layer network abstraction layer unit associated with the second intra-frame random access picture or the second progressively decoded refresh picture, and the first intra-frame random access picture or the first progressively decoded refresh picture is different from the second intra-frame random access picture or the second progressively decoded refresh picture. In some embodiments of method 1700, the rule specifies that the adaptive loop filter adaptive parameter set is included in one or more adaptive parameter sets of a video that are not shared across one or more codec video sequences. In some embodiments of method 1700, the rule specifies that one or more adaptive parameter sets belonging to a specific type are not allowed to be shared across one or more codec video sequences. In some embodiments of method 1700, the specific type includes the adaptive loop filter adaptive parameter set. In some embodiments of method 1700, the specific type includes the luma mapping and chroma scaling adaptive parameter set. In some embodiments of method 1700, the specific type includes the scaling adaptive parameter set.

[0489] In some embodiments of methods 900-1700, performing the conversion includes encoding video into a bitstream. In some embodiments of methods 900-1700, performing the conversion includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 900-1700, performing the conversion includes decoding video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to perform the operations described for methods 900-1700.

[0490] In some embodiments, a video encoding apparatus includes a processor configured to perform operations described for methods 900-1700(x). In some embodiments, a computer program product storing computer instructions that, when executed by a processor, cause the processor to perform operations described for methods 900-1700(x). In some embodiments, a non-transitory computer-readable storage medium storing a bitstream generated according to the operations described for methods 900-1700(x). In some embodiments, a non-transitory computer-readable storage medium storing instructions that cause a processor to perform operations described for methods 900-1700(x). In some embodiments, a method for generating a bitstream includes: generating a bitstream of video according to the operations described for methods 900-1700(x), and storing the bitstream on a computer-readable program medium. In some embodiments, a method, an apparatus, or a bitstream generated according to the methods or systems disclosed herein.

[0491] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to juxtaposed positions or bits propagated at different positions in the bitstream defined by the syntax. For example, a macroblock can be encoded based on the error residual value after transformation and encoding, and can also use bits in the header and other fields of the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above technical solutions. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0492] The disclosed and other technical solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of a data processing apparatus. The computer-readable medium can be a combination of a machine-readable storage device, a machine-readable storage substrate, a storage device, a substance that influences machine-readable propagated signals, or one or more such combinations. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0493] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed and executed on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0494] The processing and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0495] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as one or more of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as intra-frame hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0496] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of features of specific embodiments of a particular technology. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features from a combination of claims may be removed from the combination, and a combination of claims may refer to a sub-combination or a variation of a sub-combination.

[0497] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence shown to perform such operations, or all the described operations, in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments of this patent document should not be construed as requiring such separation in all embodiments.

[0498] Only some implementations and examples are described; other implementations, enhancements, and variations can be made based on the content described and illustrated in this patent document.

Claims

1. A video processing method, comprising: Perform the conversion between the current block of the video and the bitstream of the video. The bitstream conforms to a first rule that specifies a context increment for including a first symbol flag of a first coefficient in the bitstream. The first rule specifies that the value of the context increment is based on whether the first codec mode is applied to the current block. The bitstream conforms to a second rule specifying the use of multiple adaptive parameter set network abstraction layer units for the video. Each adaptive parameter set network abstraction layer unit has a corresponding adaptive parameter set type value. Each adaptive parameter set network abstraction layer unit is associated with a corresponding video layer identifier. In this context, each unit in the adaptive parameter set network abstraction layer is either a prefix unit or a suffix unit, and The second rule stipulates that, in response to the plurality of adaptive parameter set network abstraction layer units having specific adaptive parameter set type values, regardless of the plurality of video layer identifiers of the plurality of adaptive parameter set network abstraction layer units, and regardless of whether the plurality of adaptive parameter set network abstraction layer units are the prefix unit or the suffix unit, the adaptive parameter set identifier values ​​of the plurality of adaptive parameter set network abstraction layer units belong to the same adaptive parameter set identifier value space.

2. The method according to claim 1, in, Without applying the first codec mode, the first rule specifies that the value of the context increment is within a first range from N0 to N1, wherein the first range includes N0 and N1, and When the first codec mode is applied, the first rule specifies that the value of the context increment is in a second range from N2 to N3, wherein the second range includes N2 and N3.

3. The method according to claim 2, wherein, The first encoding / decoding mode is a block-based incremental pulse encoding / decoding modulation / decoding mode.

4. The method according to claim 2, wherein, N0, N1, N2 and N3 are 0, 2, 3 and 5 respectively.

5. The method according to claim 1, in, In response to the first neighboring symbol value and the second neighboring symbol value satisfying a first condition, the value of the context increment is equal to 0 or 3, and the first condition is: (1) The first neighboring symbol value is equal to 0, and the second neighboring symbol value is equal to 0, or (2) The first neighboring symbol value is equal to the negative of the second neighboring symbol value; Wherein, in response to the horizontal coordinate of the current coefficient scan position being 0, the first neighboring sign value is equal to 0, and in response to the horizontal coordinate of the current coefficient scan position not being 0, the first neighboring sign value is equal to the sign value of the left neighboring coefficient, and Wherein, in response to the vertical coordinate of the current coefficient scan position being 0, the second neighboring sign value is equal to 0, and in response to the vertical coordinate of the current coefficient scan position not being 0, the second neighboring sign value is equal to the sign value of the upper neighboring coefficient.

6. The method according to claim 5, wherein, In response to the first codec mode not being applied to the current block, the value of the context increment is equal to 0, and in response to the first codec mode being applied to the current block, the value of the context increment is equal to 3.

7. The method according to claim 5, wherein, In response to the first neighboring symbol value and the second neighboring symbol value not satisfying the first condition, and in response to the first neighboring symbol value and the second neighboring symbol value satisfying the second condition, the value of the context increment is equal to 1 or 4, and the second condition is: (1) The first neighboring symbol value is greater than or equal to 0, and the second neighboring symbol value is greater than or equal to 0.

8. The method according to claim 7, wherein, In response to the first codec mode not being applied to the current block, the value of the context increment is equal to 1, and in response to the first codec mode being applied to the current block, the value of the context increment is equal to 4.

9. The method according to claim 7, wherein, In response to the first neighboring symbol value and the second neighboring symbol value not satisfying the first condition and the second condition, the value of the context increment is equal to 2 or 5.

10. The method according to claim 9, wherein, In response to the first codec mode not being applied to the current block, the value of the context increment is equal to 2, and in response to the first codec mode being applied to the current block, the value of the context increment is equal to 5.

11. The method according to claim 1, in, A transform skip operation is applied to the current block.

12. The method according to claim 2, wherein, The first rule specifies that N0, N1, N2, and N3 are determined based on any one or more of the following: (1) Sequence parameter set, video parameter set, picture parameter set, picture header, strip header, slice header, maximum codec unit line, maximum codec unit group, maximum codec unit or indications included in the codec unit, (2) The block dimension of the current block and / or the block dimensions of the neighboring blocks of the current block, (3) The block shape of the current block and / or the block shape of the neighboring blocks of the current block, (4) Indication of the color format of the video, (5) Whether a single codec tree structure or a dual codec tree structure is applied to the current block, (6) The strip type and / or image type to which the current block belongs, and (7) The number of color components of the current block.

13. The method according to claim 1, in, The bitstream conforms to the third rule included in the bitstream, which uses either the context mode or the bypass mode based on the number of remaining context codec bits, according to the first symbol flag specifying the current block.

14. The method according to claim 13, wherein, The third rule specifies that, in response to the number of remaining context codec bits being less than N, the bypass mode is used to include the first symbol flag in the bitstream, where N is an integer.

15. The method according to claim 13, wherein, The third rule specifies that, in response to the number of remaining context codec bits being greater than or equal to N, the context mode is used to include the first symbol flag in the bitstream, where N is an integer.

16. The method according to claim 13, wherein, The third rule specifies that, in response to the number of remaining context codec bits being equal to N, the bypass mode is used to include the first symbol flag in the bitstream, where N is an integer.

17. The method according to claim 13, wherein, The third rule specifies that, in response to the number of remaining context codec bits being greater than N, the bypass mode is used to include the first symbol flag in the bitstream, where N is an integer.

18. The method according to claim 14, wherein, N equals 4.

19. The method of claim 14, wherein, N equals 0.

20. The method of claim 14, wherein, N is based on the following integers: (1) Sequence parameter set, video parameter set, picture parameter set, picture header, strip header, slice header, maximum codec unit line, maximum codec unit group, maximum codec unit or indications included in the codec unit, (2) The block dimension of the current block and / or the block dimensions of the neighboring blocks of the current block, (3) The block shape of the current block and / or the block shape of the neighboring blocks of the current block, (4) Indication of the color format of the video, (5) Whether a single codec tree structure or a dual codec tree structure is applied to the current block, (6) The strip type and / or image type to which the current block belongs, and (7) The number of color components of the current block.

21. The method according to claim 13, wherein, The third rule specifies that the current block is a transform block or a transform skip block.

22. The method according to claim 13, wherein, The third rule specifies that the block differential pulse coding / decoding modulation mode is applied to the current block.

23. The method according to claim 13, wherein, The third rule specifies that the block differential pulse coding / decoding modulation mode should not be applied to the current block.

24. The method according to claim 1, wherein, The current block is represented in the bitstream using a binary incremental pulse encoding / decoding modulation mode.

25. The method according to any one of claims 1 to 24, wherein, Performing the conversion includes encoding the video into the bitstream.

26. The method according to any one of claims 1 to 24, wherein, Performing the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

27. The method according to any one of claims 1 to 24, wherein, Performing the conversion includes decoding the video from the bitstream.

28. A video decoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 24 and 27.

29. A video encoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 26.

30. A non-transitory computer-readable storage medium having a computer program and a bit stream stored thereon, wherein the computer program, when executed by a processor, implements the method of any one of claims 1 to 25 to generate the bit stream.

31. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to perform the method according to any one of claims 1 to 27.

32. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform the conversion between the current block of the video and the bitstream of the video. in, The bitstream conforms to a first rule that specifies a context increment for including a first symbol flag of a first coefficient in the bitstream. The first rule specifies that the value of the context increment is based on whether the first codec mode is applied to the current block. The bitstream conforms to a second rule specifying the use of multiple adaptive parameter set network abstraction layer units for the video. Each adaptive parameter set network abstraction layer unit has a corresponding adaptive parameter set type value. Each adaptive parameter set network abstraction layer unit is associated with a corresponding video layer identifier. In this context, each unit in the adaptive parameter set network abstraction layer is either a prefix unit or a suffix unit, and The second rule stipulates that, in response to the plurality of adaptive parameter set network abstraction layer units having specific adaptive parameter set type values, regardless of the plurality of video layer identifiers of the plurality of adaptive parameter set network abstraction layer units, and regardless of whether the plurality of adaptive parameter set network abstraction layer units are the prefix unit or the suffix unit, the adaptive parameter set identifier values ​​of the plurality of adaptive parameter set network abstraction layer units belong to the same adaptive parameter set identifier value space.