Context derivation for arithmetic coding of transform coefficients generated by non-separable transforms

The method addresses inefficient context modeling for non-separable transforms by using neighboring coefficients from different scan lines to derive context models, enhancing coding efficiency and compression performance in video encoding.

JP2025537058APending Publication Date: 2025-11-14TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025515553
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2024-04-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing video encoding methods struggle with efficient context modeling for transform coefficients generated by non-separable transforms, particularly in scenarios where neighboring coefficients are not on the same scan line, leading to suboptimal coding efficiency.

Method used

A video decoding method that determines a context model for transform coefficient levels based on neighboring coefficients positioned on different scan lines, using a non-separable transform like LFNST or NSPT, to improve coding efficiency by deriving context models from a broader range of neighboring coefficients.

Benefits of technology

Enhances coding efficiency by allowing parallel parsing of syntax elements related to transform coefficient levels without interruption, improving the overall compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025537058000001_ABST
    Figure 2025537058000001_ABST
Patent Text Reader

Abstract

A video bitstream including a current transform block (TB) in a current picture is received. A context model is determined for a syntax element associated with a transform coefficient level of a first coefficient group (CG) in the current TB based on a transform coefficient level of at least one first neighboring CG of the first CG. The first CG is positioned on a first scan line. The at least one first neighboring CG is positioned on a second scan line scanned before the first scan line. The context model is a probability model for a non-separable transform. The first CG is reconstructed according to the determined transform coefficient level based on the determined context model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. patent application Ser. No. 18 / 620.945, filed March 28, 2024, entitled "CONTEXT DERIVATION FOR ARITHMETIC CODING OF TRANSFORM COEFFICIENTS GENERATED BY NON-SEPARABLE TRANSFORMS," which in turn claims priority to U.S. provisional application Ser. No. 63 / 460.877, filed April 20, 2023, entitled "Context Derivation for Arithmetic Coding of Transform Coefficients Generated by Non-Separable Transforms," ​​the entire disclosure of which is incorporated herein by reference.

[0002] [Technical field] This disclosure generally describes embodiments related to video encoding. [Background technology]

[0003] The background art provided herein is intended to represent the present disclosure as a whole, and to the extent that the work described in the background art section is not currently the work of the undersigned inventors, and aspects of the description that are not otherwise considered prior art at the time of filing, are not expressly or implicitly admitted as prior art to the present disclosure.

[0004] Image / video compression allows image / video data to be transmitted between different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques compress video based on spatial and temporal redundancy. In some examples, video codecs compress images based on spatial redundancy using a technique called intra-frame prediction. For example, intra-frame prediction predicts samples using reference data of the current picture being reconstructed. In other examples, video codecs compress images based on temporal redundancy using a technique called inter-frame prediction. For example, inter-frame prediction uses motion compensation to predict samples in the current picture based on previously reconstructed pictures. Motion compensation is indicated by motion vectors (MVs). Summary of the Invention [Problem to be solved by the invention]

[0005] Aspects of the present disclosure include video encoding / decoding methods and apparatus. In some examples, the video decoding apparatus includes a processing circuit system. [Means for solving the problem]

[0006] According to one aspect of the present disclosure, there is provided a video decoding method executed in a video decoder. The method includes receiving a video bitstream including a current transform block (TB) in a current picture, determining a context model for a syntax element associated with a transform coefficient level of a first coefficient group (CG) in the current TB based on transform coefficient levels of at least one first neighboring CG of the first CG, the first CG being positioned on a first scan line, and the at least one first neighboring CG being positioned on a second scan line scanned before the first scan line, the context model being a probability model for a non-separable transform, and reconstructing the first CG according to the determined transform coefficient levels based on the determined context model.

[0007] In the example, the at least one first nearby CG includes a first nearby CG positioned on a second scanning line and another first nearby CG positioned on a third scanning line scanned before the first scanning line and the second scanning line.

[0008] In the example, at least one first neighboring CG is positioned farther from the top left sample position of the current TB than the first CG.

[0009] In the example, the first and second scan lines are parallel diagonal lines.

[0010] In the example, the number of at least one first neighboring CG is determined based on the position of the first CG in the current TB.

[0011] In the example, the number of the at least one first neighboring CG is in the range of 1-15.

[0012] In the example, the syntax element relates to the absolute value of the transform coefficient level of the first CG.

[0013] According to one aspect, the non-separable transform is one of a low frequency non-separable transform (LFNST) mode and a non-separable linear transform (NSPT) mode.

[0014] In the example, a context model of a syntax element associated with a transform coefficient level of each CG among a plurality of CGs positioned in a first scan line is determined based on the transform coefficient levels of at least one first neighboring CG of the first CG.

[0015] In one example, the context model for a syntax element associated with a transform coefficient level of a second CG is determined based on a transform coefficient level of at least one second neighboring CG of the second CG on a first scan line, where the at least one second neighboring CG includes a second neighboring CG located on the second scan line and another second neighboring CG different from the at least one first neighboring CG.

[0016] According to another aspect of the present disclosure, there is provided an apparatus, the apparatus including a processing circuitry system, the processing circuitry system being arranged to perform any one of the described video decoding / encoding methods.

[0017] Aspects of the present disclosure further provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a video decoding / encoding method. [Brief explanation of the drawings]

[0018] Other features, properties and advantages of the disclosed subject matter will become more apparent based on the following detailed description and drawings. [Figure 1] 1 is a schematic drawing of an exemplary block diagram of a communication system (100). [Figure 2] 1 is a schematic drawing of an exemplary block diagram of a decoder. [Figure 3] 1 is a schematic drawing of an example block diagram of an encoder; [Figure 4] Low-Frequency Non-Separable Transform [Figure 5] 1 is an exemplary template for selecting a context model for coefficient coding. [Figure 6] 1 is an exemplary template for selecting a context model for coefficient coding. [Figure 7] 1 is a schematic diagram of a reverse diagonal scan. [Figure 8A] 1 is an exemplary process for selecting a probability model based on a scanline. [Figure 8B] 1 is an exemplary process for selecting a probability model based on a scanline. [Figure 8C] 1 is an exemplary process for selecting a probability model based on a scanline. [Figure 8D] 1 is an exemplary process for selecting a probability model based on a scanline. [Figure 8E]1 is an exemplary process for selecting a probability model based on a scanline. [Figure 8F] 1 is an exemplary process for selecting a probability model based on a scanline. [Figure 9A] 1 is an exemplary process for selecting a probability model based on position in a scanline. [Figure 9B] 1 is an exemplary process for selecting a probability model based on position in a scanline. [Figure 9C] 1 is an exemplary process for selecting a probability model based on position in a scanline. [Figure 10] 1 is a flowchart that outlines a decoding process of some embodiments according to the present disclosure. [Figure 11] 1 is a flowchart that outlines an encoding process of some embodiments according to the present disclosure. [Figure 12] 1 is a schematic diagram of an exemplary computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] 1 is a block diagram of a video processing system 100 according to some examples. The video processing system 100 illustrates the application of the disclosed subject matter, a video encoder and a video decoder, in a streaming transmission environment. The disclosed subject matter is equally applicable to other applications that support video, such as video conferencing, digital TV, streaming transmission services, and the storage of compressed video on digital media, including Compact Discs (CDs), Digital Versatile Discs (DVDs), memory sticks, etc.

[0020] The video processing system 100 includes a capture subsystem 113, which includes a video source 101, e.g., a digital camera, that produces an uncompressed video picture stream 102. In the illustrated example, the video picture stream 102 includes samples captured by the digital camera. The video picture stream 102 is depicted as a bold line to emphasize its higher data content than the encoded video data 104 (or encoded video bitstream), and the video picture stream 102 is processed by electronics 120, which includes a video encoder 103, coupled to the video source 101. The video encoder 103 may include hardware, software, or a combination thereof to implement or perform aspects of the disclosed subject matter, as described in more detail below. The encoded video data 104 (or encoded video bitstream) is rendered as thin lines to emphasize its lower data volume than the video picture stream 102, and the encoded video data 104 (or encoded video bitstream) is stored on a streaming transport server 105 for subsequent use. One or more streaming transport client subsystems, such as the client subsystems 106 and 108 of FIG. 1, access the streaming transport server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 includes a video decoder 110, such as in an electronic device 130. The video decoder 110 decodes the incoming copy 107 of the encoded video data to create an outgoing stream of video pictures 111 that can be displayed on a display 112 (e.g., a display screen) or other display device (not rendered). In some streaming transmission systems, the coded video data (104), (107) and (109) (eg, video bitstream) are coded according to some video coding / compression standard.Examples of these standards include the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T) H.265 Recommendation. In the examples, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter is applicable in the context of VVC.

[0021] Here, the electronic devices 120 and 130 may include other components (not shown). For example, the electronic device 120 may include a video decoder (not shown), and the electronic device 130 may include a video encoder (not shown).

[0022] 2 is an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in electronic equipment (230). The electronic equipment (230) may include a receiver (231) (e.g., a receiving circuit system). The video decoder (210) may replace the video decoder (110) in the example of FIG. 1.

[0023] The receiver (231) can receive, for example, one or more coded video sequences contained in a bitstream to be decoded by the video decoder (210). In some embodiments, the receiver receives one coded video sequence at a time, with the decoding of each coded video sequence being independent of the decoding of other coded video sequences. The receiver receives coded video sequences from a channel (201), which is a hardware / software link to a storage device where the coded video data is stored. The receiver (231) can receive coded video data and other data, such as coded audio data and / or auxiliary data streams, which can be forwarded to respective using entities (not depicted). The receiver (231) can separate the coded video sequences from other data. To prevent network jitter, a buffer memory (215) is coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory 215 is part of the video decoder 210. In other embodiments, the buffer memory is located external to the video decoder 210 (not shown). In other applications, there is a buffer memory (not shown) external to the video decoder 210, for example, to prevent network jitter, and there is another buffer memory 215 internal to the video decoder 210, for example, to handle playback timing. If the receiver 231 is receiving data from a store-and-forward device or synchronous network with sufficient bandwidth and controllability, then the buffer memory 215 may not be required, or may be small. To maximize use of packet networks such as the Internet, a buffer memory 215 may be required, which may be large and advantageously of adaptive size, and at least a portion of which may be implemented in an operating system or similar device (not shown) external to the video decoder 210.

[0024] The video decoder 210 includes a parser 220 to reconstruct codes 221 based on the encoded video sequence. These code categories include information that governs the operation of the video decoder 210 and information that controls a display device, such as the display device 212 (e.g., a display screen), which, as shown in FIG. 2, is not a component of the electronic device 230 but can be coupled to it. Control information for the display device is expressed in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser 220 performs parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence is based on a video coding technology or standard and follows various principles, including variable-length coding, Huffman coding, and arithmetic coding with or without context sensitivity. The parser (220) extracts a subgroup parameter set for at least one subgroup of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group, including a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) also extracts information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0025] The parser (220) performs an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to produce a code (221).

[0026] Depending on the type of coded video picture, or part of it (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks), and other factors, the reconstruction of the code (221) may involve several different units. Which units are involved, and in what form, is controlled by subgroup control information parsed by the parser (220) from the coded video sequence. For clarity, the flow of such subgroup control information between the parser (220) and the following units is not shown.

[0027] Beyond the functional blocks noted above, the video decoder (210) is conceptually subdivided into a number of functional units, as described below. In an actual implementation operating within business constraints, many of these units will interact closely and be at least partially integrated with one another. However, for purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.

[0028] The first unit is the scalar / inverse transform unit (251), which receives quantized transform coefficients as codes (221) and control information from the parser (220), including the transform used, the size of the block, the quantization factor, the quantization scaling matrix, etc. The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to the aggregation device (255).

[0029] In some cases, the output samples of the scaler / inverse transform unit (251) belong to intra-coded blocks. Intra-coded blocks are blocks that do not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information is provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block similar in size and shape to the block being reconstructed using surrounding reconstructed information extracted from the current picture buffer unit (258). The current picture buffer unit (258) buffers, for example, a partially reconstructed and / or fully reconstructed current picture. In some cases, the aggregation unit (255) adds the predictive information generated by the intra-frame prediction unit (252) based on each sample to the output sample information provided by the scaler / inverse transform unit (251).

[0030] In other cases, the output samples of the scalar / inverse transform unit (251) belong to an inter-frame coded block and potentially to a motion compensated block. In this case, the motion compensated prediction unit (253) accesses the reference picture memory (257) to extract prediction samples. After performing motion compensation on the extracted samples based on the code (221) belonging to the block, the aggregation device (255) adds these samples to the output of the scalar / inverse transform unit (251) (called residual samples or residual signal in this case) to generate output sample information. The addresses in the reference picture memory (257) from which the motion compensated prediction unit (253) extracts prediction samples may be controlled by a motion vector, which is provided to the motion compensated prediction unit (253) in the form of a code (221), which may have, for example, an X component, a Y component, and a reference picture component. When using sub-sample accurate motion vectors, motion compensation may further include, for example, interpolation of sample values ​​extracted from the reference picture memory (257), a motion vector prediction mechanism, etc.

[0031] The output samples of the aggregation device (255) are used in various loop filtering techniques in a loop filter unit (256). Video compression techniques include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called coded video bitstream) and available to the loop filter unit (256) as codes (221) from the parser (220). Video compression can be responsive to meta-information obtained during decoding of a coded picture or previous part of the coded video sequence (in decoding order), and to previously reconstructed and loop-filtered sample values.

[0032] The output of the loop filter unit (256) may be a stream of samples that are output to a display device (212) and stored in a reference picture memory (257) for use in subsequent inter-frame picture prediction.

[0033] As soon as some coded pictures are fully reconstructed, they can be used as reference pictures for subsequent prediction. For example, as soon as the coded picture corresponding to the current picture is fully reconstructed and that coded picture is marked as a reference picture (e.g., using the parser (220)), the current picture buffer (258) becomes part of the reference picture memory (257), and a new current picture buffer unit is re-allocated before the reconstruction of the subsequent coded picture.

[0034] The video decoder 210 performs decoding operations based on a given video compression technology or standard, such as the ITU-T H.265 recommendation. The encoded video sequence conforms to the syntax specified by the video compression technology or standard in use, in the sense that it conforms to both the syntax of the video compression technology or standard and the configuration file recorded in the video compression technology or standard. Specifically, the configuration file selects some tools from all available tools in the video compression technology or standard as tools available only in that configuration file. For compliance, the complexity of the encoded video sequence is required to be within a range limited by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the level limits are further limited by a Hypothetical Reference Decoder (HRD) specification and metadata managed by an HRD buffer device signaled in the encoded video sequence.

[0035] In an embodiment, the receiver (231) receives additional (redundant) data along with the encoded video. The additional data is included as part of the encoded video sequence. The video decoder (210) uses the additional data to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0036] 3 is an exemplary block diagram of a video encoder (303). The video encoder (303) is included in electronic equipment (320). The electronic equipment (320) includes a transmitter (340) (e.g., a transmission circuit system). The video encoder (303) is used in place of the video encoder (103) in the example of FIG. 1.

[0037] The video encoder (303) receives video samples from a video source (301) (not part of the electronics (320) in the example of FIG. 3), which captures video images for encoding by the video encoder (303). In other examples, the video source (301) is part of the electronics (320).

[0038] The video source (301) provides a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303). The digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any suitable color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) is a storage device that stores previously prepared video. In a video conferencing system, the video source (301) is a camera that captures local image information as a video sequence. The video data is provided as a number of single pictures that, when viewed in sequence, impart motion. The pictures themselves are organized as a spatial pixel array, with each pixel containing one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses primarily on samples.

[0039] According to an embodiment, the video encoder (303) encodes pictures of a source video sequence in real time or under any other required time constraints and compresses them into an encoded video sequence (343). Enforcing the appropriate encoding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units described below. For simplicity, the coupling is not depicted. Parameters set by the controller (350) include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) placement, maximum motion vector search range, etc. The controller (350) is configured to include other appropriate functions associated with the video encoder (303) that are optimized for a particular system design.

[0040] In some embodiments, the video encoder (303) is configured to operate in a coding loop. For simplicity, the illustrative coding loop includes a source encoder (330) (e.g., generating a code, e.g., a code stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the code to generate sample data in a manner similar to that of a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the code stream produces bit-accurate results independent of the decoder's location (local or remote), the contents of the reference picture memory (334) are accurate between the local encoder and the remote encoder. In other words, the prediction portion of the encoder "considers" as reference picture samples sample values ​​that are exactly the same as the sample values ​​the decoder "saw" when predicting during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example due to channel errors) is also used in several related techniques.

[0041] The operation of the "local" decoder (333) may be similar to the operation of the "remote" decoder, e.g., the video decoder (210), described in detail above in conjunction with Figure 2. Also, referring to Figure 2, because the codes are available and the codes encoded / decoded into the encoded video sequence by the entropy coder (345) and parser (220) are lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), cannot be fully implemented in the local decoder (333).

[0042] In some embodiments, the decoder techniques present in the decoder, with the exception of analysis / entropy decoding, are the same or essentially the same functionally as those present in the corresponding encoder. Therefore, the focus of the disclosed subject matter is on decoder operation. Because the encoder techniques are not fully described, the description of the encoder techniques may be simplified. More detailed descriptions are provided in several sections below.

[0043] During operation, in some instances, the source encoder (330) performs motion-compensated predictive encoding, predictively encoding an input picture with reference to one or more previously encoded pictures from the video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes differences between pixel blocks of the input picture and pixel blocks of reference pictures that can be selected as predictive references for the input picture.

[0044] The local video decoder (333) decodes the coded video data of pictures that can be designated as reference pictures based on the codes created by the source encoder (330). Preferably, the operation of the coding engine (332) can be a lossy process. If the coded video data can be decoded in a video decoder (not shown in FIG. 3), the reconstructed video sequence is generally a copy of the source video sequence, with some error. The local video decoder (333) replicates the decoding process that the video decoder can perform on the reference pictures and stores the reconstructed reference pictures in the reference picture memory (334). In this manner, the video encoder (303) locally stores a copy of the reconstructed reference picture, which is obtained from the remote video decoder and has common content (no transmission errors) with the reconstructed reference picture.

[0045] The predictor (335) performs a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) searches the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as the reference picture's motion vectors and block shape, as suitable prediction references for the new picture. The predictor (335) finds suitable prediction references by operating pixel block by pixel block based on the sample blocks. In some cases, the input picture has prediction references obtained from multiple reference pictures stored in the reference picture memory (334), for example, as determined by the search results obtained by the predictor (335).

[0046] The controller (350) manages the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0047] The outputs of all the functional units mentioned above are entropy coded in the entropy coder (345), which converts the codes generated by the various functional units into an encoded video sequence by using lossless compression based on techniques such as Huffman coding, variable length coding, and arithmetic coding.

[0048] The transmitter (340) buffers the encoded video sequence produced by the entropy encoder (345) and prepares it for transmission over a communication channel (360), which may be a hardware or software link to a storage device where the encoded video data is stored. The transmitter (340) merges the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0049] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) assigns each encoded picture a coded picture type, which may affect the coding technique applicable to the corresponding picture. For example, pictures are typically assigned to one of the following picture types:

[0050] Intraframe pictures (I-pictures) can be coded and decoded without using other pictures in the sequence as prediction sources. Some video codecs allow different types of intraframe pictures, such as Independent Decoder Refresh ("IDR") pictures.

[0051] Predictive pictures (P pictures) can be coded and decoded using intra-frame or inter-frame prediction, which predicts the sample values ​​of each block by means of motion vectors and reference indices.

[0052] Bidirectionally predictive pictures (B-pictures) can be coded and decoded using intra-frame or inter-frame prediction, which predicts the sample values ​​of each block using two motion vectors and reference indices. Similarly, multi-predictive pictures can reconstruct a single block using two or more reference pictures and associated metadata.

[0053] Generally, a source picture is spatially subdivided into a plurality of sample blocks (e.g., each block is a block of 4*4, 8*8, 4*8, or 16*16 samples) and coded block by block. These blocks are predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding picture applied to the block. For example, blocks of an I-picture may be non-predictively coded, or predictively coded (spatial prediction or intra-frame prediction) with reference to previously coded blocks of the same picture. Pixel blocks of a P-picture may be predictively coded through spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded through spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0054] The video encoder (303) performs encoding operations based on a predetermined video encoding technique or standard, such as the ITU-T H.265 recommendation. During this operation, the video encoder (303) can perform various compression operations, including predictive coding operations based on temporal and spatial redundancy in the input video sequence. Thus, the encoded video data conforms to the syntax specified by the video encoding technique or standard used.

[0055] In an embodiment, the transmitter (340) transmits additional data along with the encoded video. The source encoder (330) includes such data as part of the encoded video sequence. The additional data includes temporal / spatial / SNR enhancement layers, other forms of redundant data, such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0056] Video is captured as multiple source pictures (video pictures) in a time sequence. Intra-frame picture prediction (commonly abbreviated to intra-frame prediction) exploits spatial correlations within a given picture, while inter-frame picture prediction exploits correlations (temporal or otherwise) between pictures. In an example, a particular picture in encoding / decoding, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture is coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension marking the reference picture when multiple reference pictures are used.

[0057] In some embodiments, bidirectional prediction techniques can be applied to inter-frame picture prediction. Based on the bidirectional prediction technique, two reference pictures, for example, a first reference picture and a second reference picture, both of which are before the current picture in the video according to decoding order (but may be past and future, respectively, according to display order), are used. A block in the current picture is coded using a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block is predicted using a combination of the first reference block and the second reference block.

[0058] Also, merge mode techniques can be used in inter-frame picture prediction to improve coding efficiency.

[0059] According to some embodiments of the present disclosure, prediction, e.g., inter-frame picture prediction and intra-frame picture prediction, is performed block-based. For example, based on the High Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, e.g., 64*64 pixels, 32*32 pixels, or 16*16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU is recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64*64 pixel CTU is divided into one 64*64 pixel CU, four 32*32 pixel CUs, or sixteen 16*16 pixel CUs. In the example, each CU is analyzed to determine the prediction type of the CU, e.g., inter-frame prediction type or intra-frame prediction type. A CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations are performed in coding (encoding / decoding) using prediction blocks as units. A luma prediction block is an example of a prediction block, and the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8*8 pixels, 16*16 pixels, 8*16 pixels, 16*8 pixels, etc.

[0060] Here, the video encoders (103), (303) and video decoders (110), (210) are implemented using any suitable technology. In some embodiments, the video encoders (103), (303) and video decoders (110), (210) are implemented by one or more integrated circuits. In other embodiments, the video encoders (103), (303) and video decoders (110), (210) are implemented by one or more processors executing software instructions.

[0061] The present disclosure includes related aspects of context derivation for arithmetic coding of transform coefficients generated by a non-separable transform. For example, a context model is derived for a syntax element associated with a transform coefficient level of a coefficient group (CG) in a transform block (TB) based on at least one neighboring CG of the CG, where the at least one neighboring CG is not located on the scan line (e.g., a diagonal scan line) on which the CG is located.

[0062] A low-frequency non-separable transform (LFNST) is applied between the forward primary transform and quantization (e.g., in an encoder) and between the inverse quantization and inverse primary transform (e.g., in a decoder), as shown in FIG. 4. For example, as shown in FIG. 4, an LFNST (400) is used as a forward LFNST between the forward primary transform (402) and quantization (404), and an inverse LFNST between the inverse primary transform (406) and inverse quantization (408). In the LFNST, a 4*4 non-separable transform or an 8*8 non-separable transform is applied based on the size of the block. For example, a 4*4 LFNST is applied to small blocks (e.g., min(width, height)<8) and an 8*8 LFNST is applied to large blocks (e.g., min(width, height)>4).

[0063] For the application of a non-separable transform, e.g., LFNST, the input is exemplarily written as follows: To apply a 4*4 LFNST, provide a 4*4 input block X, e.g., in equation (1).

number

number

number

number

number

number

[0064] In context modeling for coefficient coding, the selection of a probability model for a syntax element associated with the absolute value of a transform coefficient level depends on the absolute level (or absolute value of the transform coefficient level) in a local neighborhood domain, or the value of a partially reconstructed absolute level. See Figure 5 for a template for context modeling. As shown in Figure 5, a probability model (e.g., context model) for coding a syntax element associated with a transform coefficient level at a current scanning position (or current CG) (502) is derived based on the absolute values ​​of the transform coefficient levels at neighboring positions (or neighboring CGs), e.g., neighboring CG (504), CG (506), CG (508), CG (510), and CG (512).

[0065] As shown in FIG. 6, in the modification of context modeling, the coding efficiency is improved as follows: instead of the template of FIG. 5, the previous five coefficients (e.g., coefficients of the previous five scanning positions) (603) to (607) in the coding order are used to model the context of the syntax element related to the absolute value of the transform coefficient level (or quantized transform coefficient) of the current scanning position (602).

[0066] Syntax elements related to absolute values ​​of transform coefficient levels are parsed according to the reverse diagonal scan order of Figure 7. Context modeling using the template of Figure 5 improves throughput by parsing all syntax elements related to absolute transform coefficient levels along the same diagonal in parallel. However, using the template of Figure 6, parallelization (or parallel parsing) may not be possible. Therefore, seamless parallel parsing may require more effective context modeling to capture coding efficiency.

[0067] The present disclosure provides a context modeling method. In the provided context modeling, the previous N transform coefficients (or transform coefficients of the previous N CGs) in the coding order (or scan line) can be used to model the context of a syntax element related to the absolute value of the transform coefficient level of the current CG. The previous N transform coefficients do not belong to the same diagonal scan line where the current CG is positioned. The provided context modeling can improve coding efficiency when parsed in parallel without interruption.

[0068] According to one embodiment, the previous N transform coefficients in coding order that do not belong to the same diagonal scan line are used to model the context of syntax elements related to the absolute value of the transform coefficient level.

[0069] In an embodiment, a current block (or TB) in a current picture is divided into multiple CGs. Each CG includes multiple samples, for example, 4*4 samples. To encode the current block, for example, a reverse diagonal scanning encoding order as shown in FIG. 7 is used. To encode the first CG in the first scanline of the encoding order, the transform coefficients of N neighboring CGs of the first CG can be used to derive a context model for a syntax element (e.g., sb_coded_flag) related to the absolute values ​​of the transform coefficient levels (or quantized transform coefficients) of the first CG. The N neighboring CGs are located in one or more scanlines different from the first scanline of the encoding order.

[0070] According to one embodiment, for a coefficient group (CG) having 16 samples, the template for selecting a probability model for LFNST / NSPT (Non-Separable Primary Transform, NSPT) is related to a scan line, and may refer to, but is not limited to, FIGS. 8A-8F.

[0071] In an embodiment, a reverse diagonal scanning coding order, for example, is applied to the transform block (TB) (800) of Figures 8A to 8F. The TB (800) is divided into multiple CGs, for example, 16 CGs in Figure 8A. Each CG includes multiple samples, for example, 16 samples. Based on the reverse diagonal scanning, as shown in Figure 8F, the CG (801) located in the lower right corner of the TB (800) is first coded. For example, entropy coding is performed on the transform coefficient levels of the CG (801). Based on the reverse diagonal scanning, CGs (802) and (803) located on a scanning line, for example, the diagonal scanning line (804), are further coded. In the example, the context model of a syntax element, such as sb_coded_flag, associated with the absolute value of the transform coefficient level (or quantized transform coefficient) of the current scanning position, such as CG(802) or CG(803), is derived based on the transform coefficient level of CG(801).

[0072] The encoding process continues until the process shown in FIG. 8E. As shown in FIG. 8E, CGs (805) to (807) in a scan line (or first scan line) (808) are encoded in the encoding order. In this example, a context model for a syntax element related to the absolute value of the transform coefficient level of one of CGs (805) to (807) in the current scan position (or current CG), e.g., the first scan line (808), is derived based on the transform coefficient levels of neighboring CGs of the current CG, e.g., CGs (801) to (803). Still referring to FIG. 8E, the neighboring CGs include CGs (e.g., (802) to (803)) that are scanned before the scan line (e.g., (808)) in which the current CG (e.g., (805)) is positioned and are positioned in a different scan line (or second scan line) (804) from the current scan line. In the example, the first scan line (808) and the second scan line (804) are parallel to the diagonal of the TB (800). In the example, the neighboring CGs, e.g., CGs (801) to (803), are positioned farther from the top left sample position of the TB (800) than the current CG, e.g., CG (805).

[0073] The encoding process continues until the end of FIG. 8D. As shown in FIG. 8D, CG (809) to CG (812) in a scan line (or first scan line) (813) are encoded in the encoding order. In this example, a context model for a syntax element related to the absolute value of the transform coefficient level of one of the current scan position (or current CG, or first CG), for example, CG (809) to CG (812), is derived based on the transform coefficient levels of neighboring CGs (or first neighboring CGs) of the current CG, for example, CG (802) to CG (803) and CG (805) to CG (807). Still referring to FIG. 8D, the neighboring CGs are positioned on one or more scan lines, for example, the second scan line (804) and the third scan line (808), that are scanned before the first scan line (813) in which the current CG (e.g., (809)) is positioned and that are different from the first scan line (813). In the example, scan lines 804, 808, and 813 are parallel to the diagonals of TB 800.

[0074] Referring to Figure 8C, for example, each of CGs (814) through (816) along scan line (817) is coded based on neighboring CGs (809) through (812) and CG (805). In Figure 8B, for example, each of CGs (818) through (819) along scan line (820) is coded based on neighboring CGs (809) through (810) and CGs (814) through (816). In Figure 8A, for example, the current scanning position (or current CG) (821) is coded based on neighboring CGs (814) through (816) and CGs (818) through (819). Thus, as shown in Figures 8A through 8F, each CG in the current scan line is coded based on neighboring CGs of the same group that are located on one or more scan lines different from the current scan line.

[0075] According to one embodiment, for a coefficient group (CG) having 16 samples, the template for selecting a probability model for LFNST / NSPT is related to the scanline and position (within the scanline), as shown in Figures 9A, 9B, and 9C.

[0076] 9A, in an embodiment, TB (900) includes multiple CGs, for example, CG (902) to CG (904) positioned on scan line (908). When CG (902) is the current scan position (or first CG), a context model for a syntax element related to the absolute value of the transform coefficient level of the current scan position (or current CG) (902) is derived based on the transform coefficient levels of neighboring CGs (or at least one first neighboring CG) of the current CG, for example, CG (910) to CG (912) on scan line (919) and CG (914) to CG (915) on scan line (920). If CG (903) is the current scanning position (or second CG), the context model of the syntax element related to the absolute value of the transform coefficient level of the current scanning position (or current CG) (903) is derived based on the transform coefficient levels of neighboring CGs (or at least one second neighboring CG) of the current CG, for example, CG (910) to CG (913) in scanning line (919) and CG (915) in scanning line (920). If CG (904) is the current scanning position, the context model of the syntax element related to the absolute value of the transform coefficient level of the current scanning position (or current CG) (904) is derived based on the transform coefficient levels of neighboring CGs of the current CG, for example, CG (911) to CG (913) in scanning line (919) and CG (915) to CG (916) in scanning line (920). Thus, each CG in the same scanning line is coded based on a different group of neighboring CGs, which depends on the position of the current scanning position in the scanning line. In some embodiments, a first CG (e.g., 902) and a second CG (e.g., 903) in the same scan line (e.g., 908) are coded based on one or more identical neighboring CGs (e.g., 911 and 912). In some embodiments, the neighboring CG group of the first CG (e.g., 902) includes at least one CG (e.g., 914) that is not included in the neighboring CG group of the second CG (e.g., 903).

[0077] According to one embodiment, the value of N is position-related: it depends on the position (x, y) of the transform coefficient currently being analyzed, where (x, y) corresponds to the Euclidean coordinates of the coefficient relative to the top-left sample position.

[0078] In the example shown in Figures 8D, 8E, and 8F, CG(802) to CG(803) are coded based on one neighbor, CG(801), CG(805) to CG(807) are coded based on three neighbors, CG(801) to CG(803), and CG(809) to CG(812) are coded based on five neighbors, CG(805) to CG(807) and CG(802) to CG(803). CG(802) to CG(803) are positioned farther from the top-left sample position of TB(800) than CG(805) to CG(807). CG(805) to CG(807) are positioned farther from the top-left sample position of TB(800) than CG(809) to CG(812).

[0079] According to one embodiment, the value of N may include, but is not limited to, elements of the set {1, 2, 3, 4, 5, . . . , 15}.

[0080] For example, CG(802) to CG(803) are coded based on one neighboring CG, such as CG(801). CG(805) to CG(807) are coded based on three neighboring CGs, such as CG(801) to CG(803). CG(809) to CG(812) are coded based on five neighboring CGs, such as CG(805) to CG(807) and CG(802) to CG(803).

[0081] 10 is a flowchart outlining a process (1000) according to an embodiment of the present disclosure. The process (1000) is used in a video decoder. In various embodiments, the process (1000) is performed by a processing circuitry system, such as a processing circuitry system that performs the functions of the video decoder (110), a processing circuitry system that performs the functions of the video decoder (210), etc. In some embodiments, the process (1000) is implemented by software instructions, such that the processing circuitry system executes the software instructions to perform the process (1000). The process begins at (S1001) and continues through (S1010).

[0082] At (S1010), a video bitstream including a current transform block (TB) in a current picture is received.

[0083] In step (S1020), a context model is determined for a syntax element associated with a transform coefficient level of a first coefficient group (CG) based on a transform coefficient level of at least one first neighboring CG of the first CG in the current TB. The first CG is located on a first scan line. The at least one first neighboring CG is located on a second scan line scanned before the first scan line. The context model is a probability model for a non-separable transform.

[0084] According to one aspect, the previous N transform coefficients in coding order that do not belong to the same diagonal scanline are used to model the context of syntax elements related to the absolute value of a transform coefficient level. In an example, as shown in Figures 8A-8F, the template for selecting a probability model for LFNST / NSPT is related to a scanline. In an example, for a coefficient group (CG) having 16 samples, the template for selecting a probability model for LFNST / NSPT is related to a scanline and a position (within a scanline). See Figures 9A, 9B, and 9C for an example.

[0085] In (S1030), the first CG is reconstructed according to the determined transform coefficient levels based on the determined context model.

[0086] In the example, the at least one first nearby CG includes a first nearby CG positioned on a second scanning line and another first nearby CG positioned on a third scanning line scanned before the first scanning line and the second scanning line.

[0087] In the example, at least one first neighboring CG is positioned farther from the top left sample position of the current TB than the first CG.

[0088] In the example, the first and second scan lines are parallel diagonal lines.

[0089] In the example, the number of at least one first neighboring CG is determined based on the position of the first CG in the current TB.

[0090] In the example, the number of the at least one first neighboring CG is in the range of 1-15.

[0091] In the example, the syntax element relates to the absolute value of the transform coefficient level of the first CG.

[0092] According to one aspect, the non-separable transform is one of a low frequency non-separable transform (LFNST) mode and a non-separable linear transform (NSPT) mode.

[0093] In the example, a context model of a syntax element associated with a transform coefficient level of each CG among a plurality of CGs positioned in a first scan line is determined based on the transform coefficient levels of at least one first neighboring CG of the first CG.

[0094] In one example, a context model is determined for a syntax element associated with a transform coefficient level of a second CG based on a transform coefficient level of at least one second neighboring CG of the second CG on a first scan line, where the at least one second neighboring CG includes a second neighboring CG located on the second scan line and another second neighboring CG different from the at least one first neighboring CG.

[0095] Then, the process is carried out up to (S1099) and ends.

[0096] Process 1000 may be adjusted as appropriate. Steps in process 1000 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0097] 11 is a flowchart outlining a process (1100) according to an embodiment of the present disclosure. The process (1100) is used in a video encoder. In various embodiments, the process (1100) is performed by a processing circuitry system, which may be, for example, a processing circuitry system that performs the functions of the video encoder (103), a processing circuitry system that performs the functions of the video encoder (303), etc. In some embodiments, the process (1100) is implemented by software instructions, such that the processing circuitry system executes the software instructions to perform the process (1100). The process begins at (S1101) and continues through (S1110).

[0098] In (S1110), the coding order of the current transform block (TB) in the current picture is determined.

[0099] In step (S1120), a context model is determined for a syntax element associated with a transform coefficient level of a first CG based on a transform coefficient level of at least one first neighboring CG of the first CG in the current TB, the first CG being located on a first scan line in the coding order, and the at least one first neighboring CG being located on at least a second scan line scanned before the first scan line in the coding order.

[0100] According to one aspect, the context of a syntax element related to the absolute value of a transform coefficient level is modeled using the previous N transform coefficients that do not belong to the same diagonal scanline in coding order. In an example, as shown in Figures 8A-8F, the template for selecting a probability model for LFNST / NSPT is related to the scanline. In an example, for a coefficient group (CG) having 16 samples, the template for selecting a probability model for LFNST / NSPT is related to the scanline and position (within the scanline). See Figures 9A, 9B, and 9C for an example.

[0101] In step S1130, syntax elements associated with the transform coefficient level of the first CG are coded based on the determined context model.

[0102] Then, the process is carried out up to (S1199) and ends.

[0103] The process 1100 may be adjusted as appropriate. Steps in the process 1100 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0104] The techniques described above may be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, Figure 12 illustrates a computer system (1200) for implementing some embodiments of the disclosed subject matter.

[0105] Any suitable machine code or computer language may be used to encode computer software, which may be assembled, compiled, linked, or otherwise processed to create code containing instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or the like, or may be executed by interpretation, microcode execution, or the like.

[0106] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0107] 12 is exemplary in nature and does not limit the scope or functionality of the computer software of the embodiments implementing the present disclosure, nor should the arrangement of components be construed as having any dependency or requirement regarding any one component or combination of components shown in the exemplary embodiment of computer system 1200.

[0108] The computer system 1200 includes several human-machine interface input devices that can respond to inputs made by one or more human users, for example, through tactile input (e.g., clicks, slides, data glove movements), audio input (e.g., voice, taps), visual input (e.g., posture), and olfactory input (undrawn). The human-machine interface devices also capture several media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographed images obtained from still image capture devices), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).

[0109] The input human-machine interface devices include one or more (only one of each is shown) of a keyboard (1201), a mouse (1202), a touchpad (1203), a touchscreen (1210), a data glove (not shown), a joystick (1205), a microphone (1206), a scanner (1207), and an imaging device (1208).

[0110] The computer system 1200 further includes several human-machine interface output devices that stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices include haptic output devices (e.g., haptic feedback via a touchscreen (1210), data gloves (not shown), or joystick (1205), although there may be haptic feedback devices that are not used as input devices), audio output devices (e.g., speakers (1209), head-mounted headphones (not shown)), visual output devices (e.g., screens (1210), including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, and organic light-emitting diode (OLED) screens, each with or without touchscreen input capability and each with or without haptic feedback capability. Some of the screens can provide two-dimensional visual output or three or more dimensions, for example in the form of stereoscopic output, including virtual reality glasses (not shown), holographic displays, and smoke generators (not shown)), and printers (not shown).

[0111] The computer system (1200) may further include human-accessible storage devices and associated media, such as optical media such as CD / DVD ROM (Read Only Memory) / RW (1220) with media such as CD / DVD (1221), thumb drives (1222), removable hard drives or solid state drives (1223), conventional magnetic media such as magnetic tape and floppy disks (not shown), dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD (Programmable Logic Device, PLD) based devices such as security dongles (not shown), etc.

[0112] As will be appreciated by those skilled in the art, the term "computer-readable medium" as used in conjunction with the subject matter of the present disclosure does not include transmission media, carriers, or other transitory signals.

[0113] The computer system 1200 further includes an interface 1254 to one or more communication networks 1255. The network may be, for example, a wireless network, a wired network, or an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular or industrial network, a real-time, delay-tolerant network, or the like. Examples of networks include local networks such as Ethernet, wireless local area networks (LANs), cellular networks including Global System for Mobile Communications (GSM), 3G (Third Generation, 3G), 4G (Fourth Generation, 4G), 5G (Fifth Generation, 5G), and LTE (Long Term Evolution, LTE), TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and automotive and industrial networks including Controller Area Network (CAN) buses. Some networks are typically coupled to external network interface adapters on some general-purpose data ports or peripheral buses 1249 (e.g., Universal Serial Bus (USB) ports on the computer system 1200), while other networks are typically integrated into the core of the computer system 1200 by being coupled to the system bus (e.g., via an Ethernet interface to a personal computer (PC) computer system or a cellular network interface to a smartphone computer system). Through any of these networks, the computer system 1200 can communicate with other entities. Such communications can be one-way receive (e.g., broadcast TV), one-way transmit (e.g., a CAN bus to some CAN bus devices), or bidirectional, for example, to other computer systems using local-area digital networks or wide-area digital networks.Several protocols and protocol stacks may be used in each of these networks and network interfaces described above.

[0114] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be coupled to the core 1240 of the computer system 1200.

[0115] The cores (1240) may include one or more central processing units (CPUs) (1241), graphics processing units (GPUs) (1242), dedicated programmable processing units (1243) in the form of field programmable gate arrays (FPGAs), hardware accelerators (1244) for certain tasks, graphics adapters (1250), etc. These devices, along with read-only memory (ROM) (1245), random access memory (1246), and internal mass storage devices such as internal non-user-accessible hard disk drives or solid-state drives (SSDs) (1247), are connected via a system bus (1248). In some computer systems, expansion is achieved by adding additional CPUs, GPUs, etc., by accessing the system bus (1248) in the form of one or more physical plugs. Peripherals may be directly connected to the core system bus 1248 or may be connected to the core system bus 1248 via a peripheral bus 1249. In the example, a screen 1210 is connected to a graphics adapter 1250. Peripheral bus architectures include Peripheral Component Interconnect / Interface (PCI), USB, etc.

[0116] The CPU (1241), GPU (1242), FPGA (1243), and accelerator (1244) can execute several instructions, which, when combined, constitute the computer code. The computer code is stored in ROM (1245) or RAM (1246). Temporary data may be stored in RAM (1246), while persistent data is stored, for example, in an internal mass storage device (1247). A cache memory can be used to achieve rapid storage and retrieval from any of the memory devices, and the cache memory is closely associated with one or more of the CPU (1241), GPU (1242), mass storage device (1247), ROM (1245), RAM (1246), etc.

[0117] The computer-readable medium contains computer code for performing various computer-implemented operations, either media and computer code specially designed and constructed for the purposes of the present disclosure, or of the type well-known and available to those skilled in the art of computer software.

[0118] As an illustrative, non-limiting example, the computer system (1200) having the architecture, and in particular the core (1240), can provide the provided functionality by the processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with the user-accessible mass storage devices mentioned above, as well as some storage devices with the non-transitory core (1240), such as the core's internal mass storage device (1247) or ROM (1245). Software embodying the present disclosure in various embodiments is stored in such devices and executed by the core (1240). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software causes the core (1240), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform specific processes or portions of specific processes described herein, including defining data structures stored in RAM (1246) and modifying such data structures through software-defined operations. Additionally, or alternatively, a computer system may provide functionality represented by hardwired or otherwise implemented circuitry (e.g., accelerator 1244) that operates in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Where appropriate, references to software may include logic, and references to logic may include software. Where appropriate, references to computer-readable media include circuitry (e.g., integrated circuits (ICs)) that store software for execution, circuitry that implements logic for execution, or both. The present disclosure includes any appropriate combination of hardware and software.

[0119] The use of "at least one of" or "one of" in this disclosure is intended to include any one of the listed elements or combinations thereof. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B, or (A and B). The use of "one of" does not exclude any combination of the listed elements, where appropriate, e.g., where the elements are not mutually exclusive.

[0120] While this disclosure has described several exemplary embodiments, there are alterations, substitutions, and various alternative equivalent solutions that fall within the scope of this disclosure. Thus, many systems and methods that, although not explicitly shown or described herein, will occur to those skilled in the art for carrying out the principles of this disclosure, are within the spirit and scope of this disclosure.

Claims

1. 1. A video decoding method, the method comprising: receiving a video bitstream including a current transform block (TB) in a current picture; determining a context model for a syntax element associated with a transform coefficient level of a first coefficient group (CG) based on transform coefficient levels of at least one first neighboring CG of the first CG in the current table, the first CG being located on a first scan line and the at least one first neighboring CG being located on a second scan line scanned before the first scan line, the context model being a probability model for a non-separable transform; reconstructing the first CG in response to the transform coefficient levels determined based on the determined context model.

2. 2. The method of claim 1, wherein the at least one first neighboring CG includes a first neighboring CG positioned on the second scanning line and another first neighboring CG positioned on a third scanning line scanned before the first scanning line and the second scanning line.

3. The method of claim 1 , wherein the at least one first neighboring CG is positioned further away from the top-left sample position of the current TB than the first CG.

4. 2. The method of claim 1, wherein the first scan line and the second scan line are parallel diagonal lines.

5. The method of claim 1 , further comprising determining the number of the at least one first neighboring CG based on a position of the first CG in the current TB.

6. The method of claim 1 , wherein the number of said at least one first neighboring CG is in the range of 1 to 15.

7. The method of claim 1 , wherein the syntax elements relate to absolute values ​​of the transform coefficient levels of the first CG.

8. The method of claim 1 , wherein the non-separable transform is one of a low frequency non-separable transform (LFNST) mode and a non-separable linear transform (NSPT) mode.

9. 2. The method of claim 1, further comprising determining a context model for syntax elements associated with a transform coefficient level of each CG among a plurality of CGs positioned on the first scan line based on the transform coefficient levels of the at least one first neighboring CG of the first CG.

10. 2. The method of claim 1, further comprising determining a context model for a syntax element associated with a transform coefficient level of the second CG based on a transform coefficient level of at least one second neighboring CG of the second CG on the first scanning line, wherein the at least one second neighboring CG includes a second neighboring CG positioned on the second scanning line and another second neighboring CG different from the at least one first neighboring CG.

11. 1. An apparatus, comprising:

11. Apparatus comprising a processing circuitry system, the processing circuitry system being configured to perform a method according to any one of claims 1 to 10.

12. A program causing a computer to carry out the method according to any one of claims 1 to 10.

13. 1. A video encoding method, the method comprising: determining a coding order for a current transform block (TB) in a current picture; determining a context model for a syntax element associated with a transform coefficient level of a first coefficient group (CG) based on transform coefficient levels of at least one first neighboring CG of the first CG in the current table, the first CG being located in a first scan line in the coding order, the at least one first neighboring CG being located in a second scan line scanned before the first scan line in the coding order, and the context model being a probability model for a non-separable transform; and encoding syntax elements associated with the transform coefficient levels of the first CG based on the determined context model.

Citation Information

Patent Citations

  • Video decoding and encoding method, apparatus and computer program

    JP2022171910A

  • Video processing method, device and storage medium

    JP2022532517A

  • Context modeling of reduced secondary transforms in video

    US20220295099A1