Symbol prediction of transform coefficients

By optimizing the number of symbol predictions and adaptive region settings based on coding conditions and parameters in video coding, the problem of low symbol prediction efficiency in existing technologies is solved, and a more efficient video coding and decoding process is achieved.

CN121128161APending Publication Date: 2025-12-12INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480026426.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-21
Filing Date
2024-04-16
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing video coding techniques suffer from inefficiency in predicting the sign of transform coefficients. In particular, existing methods fail to effectively utilize coding conditions and parameters during the sign prediction process, resulting in poor coding complexity and compression efficiency.

Method used

By determining the number of symbol predictions based on coding conditions and parameters, using signals to indicate whether the symbols of the transform coefficients match the predicted symbols, performing inverse transform and decoding, and combining adaptive symbol prediction region settings and improved symbol selection methods, the symbol prediction process is optimized.

Benefits of technology

It improves the compression efficiency and coding complexity of video coding, enhances the accuracy and efficiency of symbol prediction, and reduces the computational burden of the coding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128161A_ABST
    Figure CN121128161A_ABST
Patent Text Reader

Abstract

To improve or simplify symbol prediction of transform coefficients, several aspects are presented. For example, symbols to be predicted may be predicted based on weighted qIdx values by ranking transform coefficients having the same qIdx value. Alternatively, instead of qIdx-based ordering, we may place saliency coefficients with absolute level values greater than a threshold at the beginning to perform symbol prediction in a transform block. In another example, prediction residuals may be generated using a model that is more complex than simple linear prediction. For prediction symbols, the CABAC context derivation may also be modified. In addition, the maximum number of prediction symbols may be adapted to block size, maximum symbol prediction region, energy, transform type, QP, block prediction mode of TB, or other parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This embodiment generally relates to a method and apparatus for predicting residual symbol prediction in video encoding and decoding. Background Technology

[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to utilize spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to leverage intra-frame or inter-frame image correlations. Then, the differences between the original block and the predicted block (usually represented as prediction error or prediction residual) are transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the Invention

[0003] According to an embodiment, a video decoding method is proposed, the method comprising: determining a limit on the number of symbols to be predicted for a block based on coding conditions and parameters; obtaining a signal indicating whether the symbol of a transform coefficient matches the predicted symbol of the transform coefficient; obtaining the predicted symbol of the transform coefficient in response to the transform coefficient belonging to a subset of transform coefficients of the block, wherein the symbol of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit; obtaining the symbol of the transform coefficient based on the predicted symbol and the signal; performing an inverse transform on a transform coefficient block corresponding to the block, the transform coefficient block including the transform coefficient; and decoding the block based on the predicted block corresponding to the block and the transform coefficient block.

[0004] According to an embodiment, a video coding method is proposed, the method comprising: determining a constraint on the number of symbols to be predicted for a block based on coding conditions and parameters; obtaining symbols of transform coefficients; obtaining predicted symbols of the transform coefficients in response to the transform coefficients belonging to a subset of transform coefficients of the block, wherein the symbol of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the constraint; obtaining a signal indicating whether the predicted symbols match the symbols of the transform coefficients; and encoding the signal.

[0005] According to another embodiment, an apparatus is proposed, comprising at least one memory and one or more processors, wherein the one or more processors are configured to: determine a limit on the number of symbols to be predicted for a block based on encoding conditions and parameters; obtain a signal indicating whether the symbol of a transform coefficient matches the predicted symbol of the transform coefficient; obtain the predicted symbol of the transform coefficient in response to the transform coefficient belonging to a subset of transform coefficients of the block, wherein the symbol of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit; obtain the symbol of the transform coefficient based on the predicted symbol and the signal; perform an inverse transform on a block of transform coefficients corresponding to the block, the block of transform coefficients including the transform coefficients; and decode the block based on the predicted block corresponding to the block and the block of transform coefficients.

[0006] According to another embodiment, an apparatus is proposed, comprising at least one memory and one or more processors, wherein the one or more processors are configured to: determine a limit on the number of symbols to be predicted for a block based on encoding conditions and parameters; obtain symbols of transform coefficients; obtain predicted symbols of the transform coefficients in response to the transform coefficients belonging to a subset of transform coefficients of the block, wherein the symbol of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit; obtain a signal indicating whether the predicted symbol matches the symbol of the transform coefficient; and encode the signal.

[0007] One or more embodiments also provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to perform an encoding or decoding method according to any embodiment described herein. One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for encoding or decoding video according to the methods described herein.

[0008] One or more embodiments also provide a computer-readable storage medium storing video data generated according to the method described above. One or more embodiments also provide methods and apparatus for transmitting or receiving video data generated according to the methods described herein. Attached Figure Description

[0009] Figure 1 The diagram illustrates a block diagram of a system in which various aspects of this embodiment can be implemented.

[0010] Figure 2 A block diagram illustrating an embodiment of a video encoder is shown.

[0011] Figure 3 A block diagram illustrating an embodiment of a video decoder is shown.

[0012] Figure 4 The illustration shows how to obtain the prediction residual using neighboring samples.

[0013] Figure 5 The diagram illustrates DQ (Dependency Quantization).

[0014] Figure 6 The illustration shows a method for symbol prediction at the encoder according to an embodiment.

[0015] Figure 7 The illustration shows a method for symbol prediction at the decoder according to an embodiment.

[0016] Figure 8 The illustration shows a method for predicting the maximum number of symbols based on block size adaptation, according to an embodiment. Detailed Implementation

[0017] Figure 1 The diagram illustrates an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100 may be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0018] System 100 includes at least one processor 110 configured to execute instructions loaded thereon for implementing various aspects, such as those described in this application. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 140 may include internal storage devices, attached storage devices, and / or network-accessible storage devices.

[0019] System 100 includes an encoder / decoder module 130, which is configured to process data, for example, to provide encoded or decoded video, and the encoder / decoder module 130 may include its own processor and memory. Encoder / decoder module 130 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100, or may be incorporated within processor 110 as a combination of hardware and software known to those skilled in the art.

[0020] Program code to be loaded onto processor 110 or encoder / decoder 130 to execute the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0021] In several embodiments, the memory within processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 110 or encoder / decoder module 130) is used for one or more of these functions. External memory may be memory 120 and / or storage device 140, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, fast external volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, HEVC, or VVC.

[0022] Inputs can be provided to the components of system 100 through various input devices as indicated in box 105. Such input devices include, but are not limited to: (i) an RF section that receives, for example, RF signals transmitted over the air by a broadcasting company, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0023] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or limiting a signal band to a certain band); (ii) down-converting the selected signal; (iii) further limiting the band to a narrower band to select, for example, a signal band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF section may include tuners that perform various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the aforementioned (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0024] Additionally, USB and / or HDMI terminals may include corresponding interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, in a separate input processing IC or in processor 110. Similarly, as needed, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, processor 110 and encoder / decoder 130, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on the output device.

[0025] Various components of system 100 can be provided within an integrated housing. Within the integrated housing, various components can be interconnected and transmit data between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit board.

[0026] System 100 includes a communication interface 150, which is capable of communicating with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card, and the communication channel 190 may be implemented, for example, in a wired and / or wireless medium.

[0027] In various embodiments, a Wi-Fi network such as IEEE 802.11 is used to stream data to system 100. The Wi-Fi signals in these embodiments are received via a communication channel 190 and a communication interface 150 adapted for Wi-Fi communication. The communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 100, delivering data via an HDMI connection to input box 105. Still other embodiments use an RF connection to input box 105 to provide streaming data to system 100.

[0028] System 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. In various examples of embodiments, the other peripheral devices 185 include one or more of the following: a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 100. In various embodiments, control signals are transmitted between system 100 and the display 165, speaker 175, or other peripheral devices 185 using signaling (such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention). Output devices can be communicatively coupled to system 100 via dedicated connections through corresponding interfaces 160, 170, and 180. Alternatively, output devices can be connected to system 100 via communication interface 150 using communication channel 190. The display 165 and speaker 175 can be integrated into a single unit within an electronic device (e.g., a television set) along with other components of system 100. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0029] For example, if the RF portion of input 105 is part of a separate set-top box, then display 165 and speaker 175 may alternatively be separated from one or more other components. In various embodiments where display 165 and speaker 175 are external components, output signals may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0030] Figure 2 The illustration shows an example video encoder 200, such as a VVC (Variety Video Coding) encoder. Figure 2 The diagram may also show an encoder that improves upon the VVC standard or an encoder that uses a technology similar to VVC.

[0031] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "encoded" and "coded," and the terms "image," "picture," and "frame" are used interchangeably. Generally (but not necessarily), the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0032] Before being encoded, the video sequence may undergo pre-encoding processing (201), such as applying color transformations to the input color image (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input image components to obtain a more resilient signal distribution to compression (e.g., using histogram equalization with one of the color components). Metadata may be associated with pre-processing and appended to the bitstream.

[0033] In encoder 200, the image is encoded by encoder elements as described below. The image to be encoded is partitioned (202) and processed into units, such as CUs. Each unit is encoded using, for example, an intra-frame mode or an inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use for encoding the unit, and indicates the intra / inter-frame decision by, for example, a prediction mode flag. After prediction, prediction enhancement (285) is applied to the prediction block. The prediction residual is calculated, for example, by subtracting (210) the prediction block from the original image block.

[0034] Then, the predicted residual is transformed (225) and quantized (230). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (245) to output a bitstream. The encoder can skip the transform and directly quantize the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., directly encode the residual without applying the transform or quantization process.

[0035] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct the image blocks. An in-loop filter (265) is applied to the reconstructed image to perform, for example, deblocking / SAO (Sample Adaptive Shift) filtering to reduce coding artifacts. The filtered image is stored in a reference image buffer (280).

[0036] Figure 3 A block diagram of an example video decoder 300 is shown. In decoder 300, the bitstream is decoded by decoder elements, as described below. Video decoder 300 typically performs operations similar to... Figure 2 The encoder 200 is the reciprocal of the encoding path described in the document. The encoder 200 also typically performs video decoding as part of the encoded video data.

[0037] Specifically, the decoder's input includes a video bitstream, which can be generated by the video encoder 200. First, entropy decoding (330) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Image partitioning information indicates how the image should be partitioned. Therefore, the decoder can partition the image based on the decoded image partitioning information (335). The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (355) to reconstruct the image block. The prediction block (370) can be obtained from intra-frame prediction (360) or motion-compensated prediction (i.e., inter-frame prediction) (375). After prediction, prediction enhancement (390) is applied to the prediction block. An in-loop filter (365) is applied to the reconstructed image. The filtered image is stored in a reference image buffer (380).

[0038] The decoded image can be further processed by post-decoding (385), such as inverse color transformation (e.g., transformation from YCbCr 4:2:0 to RGB 4:4:4) or performing the inverse of the remapping process performed in pre-encoding (201). Post-decoding can use metadata derived in pre-encoding and sent as a signal in the bitstream.

[0039] Sign prediction of transform coefficients in recent video encoding and decoding In the residual coding process of the Exploratory Coding Model (ECM), coefficient symbol prediction is an encoding tool used to compress some symbols of the transform coefficients using a prediction system. The predicted symbols are no longer EP (equal probability) encoded in the bitstream (entropy encoding in bypass mode), but are instead replaced by a "residual" (signaled using the associated CABAC context), which indicates whether the symbol prediction was correct.

[0040] Examples of such methods are described in “Description of SDR, HDR and 360° video coding technology proposal by Qualcomm and Technicolor – low and high complexity versions” (JVET-J0021) proposed by Y.-W. Chen et al. in San Diego, USA in April 2018, and “Residual Coefficient Sign Prediction” (JVET-D0031) proposed by A. Alshin et al. in Chengdu, China in October 2016.

[0041] The symbol prediction method operates by performing multiple inverse transforms on the transform coefficients of a coded block. For each inverse transform, the sign of a non-zero transform coefficient is either set to negative or positive. The combination of symbols that minimizes the cost function is selected as the symbol predictor to predict the sign of the transform coefficients for the current block. For an example illustrating this idea, suppose the current block contains two non-zero coefficients; then there are four possible symbol combinations: (+, +), (+, -), (-, +), and (-, -). For all four combinations, the cost function is computed, and the combination with the minimum cost is selected as the symbol predictor. In this example, the cost function is computed as a measure of discontinuity across the block boundary. To derive the optimal symbol prediction hypothesis from all possible combinations, a cost function needs to be defined. In ECM, the cost function is defined as a measure of discontinuity across the block boundary. It is measured for all hypotheses, and the hypothesis with the minimum cost is selected as the predictor for the coefficient symbols. The predicted value of the current symbol is derived from this hypothesis. If the prediction matches the truth value of the symbol, a "0" is sent as the symbol residual; otherwise, a "1" is sent. Both the encoder and decoder need to generate reconstructed samples (also known as hypotheses) at the upper and left boundaries of the TB (transform block) based on different symbol combinations of the selected coefficients.

[0042] like Figure 4 As shown, for each predicted pixel in the first left column of the current block P 0,y Perform a simple linear prediction using the two reconstructed pixels on the left to obtain its prediction residual. This prediction residual and the residual assumptions r 0,y The absolute difference between them is added to the cost of the assumption.

[0043] For pixels in the top row of the current block, a similar process occurs, thus converting each prediction residual... and residual assumptions r x,0 Sum the absolute differences.

[0044] The cost function can be mathematically modeled using the following equation: (1) in P This is the prediction signal for the current block. R It is about rebuilding the neighborhood, and r It's the residual assumption. (Term) Each block can be computed only once, and only the residual assumption is subtracted.

[0045] Order-based symbol selection To control complexity, there is a limit to the number of symbols to be predicted for a TB. In ECM-2.0, the maximum number of symbols to be predicted (maxNumPredSigns) is signaled to the decoder via the SPS (Sequence Parameter Set). The allowed values ​​for maxNumPredSigns are 0 to 8 (inclusive). In CTC (Common Test Condition), the number of symbols to be predicted (NumSignPred) is set to 8, which is also the value assigned to maxNumPredSigns. Only the symbols in the top-left 4×4 block of the TB are predicted. The remaining symbols are encoded by EP. If the number of symbols in the top-left 4×4 block of the TB exceeds the maximum limit, the first maxNumPredSigns symbols (in raster scan order) are predicted.

[0046] However, the raster scan order may not be optimal for all TBs. Generally, the sign of large transform coefficients is relatively easy to predict, as sign errors in large transform coefficients have a relatively high impact on the reconstructed block. Based on this observation, some sorting-based methods have been proposed.

[0047] Sort based on the absolute value of qIdx JVET-X0120 (see Yan Ye et al., “AHG12: On sign prediction”, JVET-X0120, October 2021, at the 24th JVET conference call) and JVET-Y0141 (see Jie Chen et al., “EE2-4.3 related: More combined test results for sign prediction”, January 2022, at the 25th JVET conference call) proposes to adaptively select the sign to be predicted based on the absolute value of qIdx (qIdx is the level of transform coefficients after compensating for the effects of multiple quantizers in DQ).

[0048] Since ECM-2.0 (also in VVC), due to the use of two quantizers, the coefficient level cannot accurately reflect the magnitude of the transform coefficients. For the same level value, due to the two quantizers used in DQ ( Figure 5 The dequantized transform coefficients (Q0 and Q1 shown) may be different. Assume the following two cases: Case 1: Level = 2, quantizer Q0, dequantized transform coefficients = 4∆ k Case 2: Level = 2, quantizer Q1, dequantized transform coefficients = 3∆ k In both of the above cases, the level is equal to 2; however, the dequantized transform coefficients in case 1 (i.e., 4∆) k ) is greater than case 2 (i.e., 3∆) k Based on this observation, the sorting is performed based on the qIdx value (dequantized transform coefficients = qIdx ×∆ k The level qIdx depends on the DQ (Dependency Quantization) state and can be calculated as follows: (2) In the proposed method, after decoding the absolute values ​​of the levels, the transform coefficients in the TB are sorted based on their qIdx values. The transform coefficients with the highest qIdx values ​​are placed at the beginning of the sorted TBs. The first maxNumPredSigns symbols in the sorted TBs are predicted, and the remaining symbols are encoded by EP. This proposed method is adopted in ECM-4.0.

[0049] Ranking based on the energy influence of reconstructed boundary samples JVET-X0150 (see Xiaoyu Xiu et al., “AHG12: Enhanced sign prediction”, JVET-X0150, presented at the 24th JVET conference call in October 2021) proposes selecting signs based on their impact on the quality of a TB of reconstructed boundary samples. Specifically, to select transform coefficients for sign prediction, the following cost function is applied, which measures the energy induced by a transform coefficient on the reconstructed boundary samples: (3) in C i,j Represents the coordinates in TB ( i, j The transformation coefficient level at ) and T i,j This represents the L-shaped template corresponding to the transform coefficients, which is formed by the upper and left boundaries of TB. N The system consists of several reconstructed samples. Based on the cost function described above, the encoder / decoder will select the sign of the coefficient with the maximum cost of maxNumPredSigns for prediction.

[0050] Adaptive symbol prediction region setting In ECM-2.0, only symbols in the top-left 4×4 block of the TB are predicted. JVET-X0120 proposes extending the symbol prediction area of ​​the TB to a maximum of 32×32 blocks. Specifically, in the proposed method, symbols in the top-left M×N block are predicted. The values ​​of M and N are calculated as follows: (4) The width and height refer to the width and height of the transform block.

[0051] Building upon the method proposed in JVET-X0120, JVET-Y0141 proposes that the maximum region for symbol prediction is not always set to 32×32. Instead, the encoder sets the maximum region based on configuration, sequence class, and quantization parameters (QP), and this region is signaled in the SPS. This proposed method is adopted in ECM-4.0.

[0052] Hypothesis generation Using these n (up to maxNumPredSigns) selected transformation coefficients, the following procedure is performed. 2 n The simplified boundary reconstruction is performed once for each unique combination of the signs of the n coefficients.

[0053] Only the leftmost and topmost pixels of the block are recreated by adding the result from the inverse transform to the block prediction. Additionally, to avoid multiple inverse transforms, existing symbol prediction in ECM-2.0 applies a template-based hypothesis generation scheme, where each template is a set of reconstructed boundary samples of a TB. The templates are pre-computed at both the encoder and decoder based on the inverse transform of some unit coefficient matrices, each generated by setting one specific coefficient to 1 while keeping all other coefficients 0. In this way, when predicting n symbols in a block, only n+1 inverse transform operations are performed: 1. A single inverse transform operation on the dequantized transform coefficients, where the values ​​of all predicted symbols are set to positive. Once added to the prediction of the current block, this corresponds to the boundary reconstruction of the first hypothesis; 2. For each of the n coefficients whose sign is predicted, an inverse transform operation is performed on an originally empty block containing the corresponding dequantized (and positive) coefficient as its only non-zero element. The leftmost and topmost boundary values ​​are preserved in a so-called L-shaped template. T In the middle, for use during subsequent reconstruction.

[0054] Boundary reconstruction for subsequent hypotheses begins by employing a properly preserved reconstruction of the previous hypothesis, requiring only a single predicted sign to change from "positive" to "negative" in order to construct the desired current hypothesis. This sign change is then approximated by doubling the template corresponding to the predicted sign and subtracting it from the hypothesis boundary. After computational cost, the boundary reconstruction is preserved if it is known that it will be reused to construct subsequent hypotheses. Note that these approximations are used only during the sign prediction process, not during the final reconstruction.

[0055] Table 1: Shows the save / restore and template application for n = 3 (3 symbols, 8 entries). Symbols are transmitted, parsed, and reconstructed using signals. When predicting residuals using signals to send specific symbols, eight CABAC contexts are used. The CABAC context to use is determined by whether the dequantized coefficients in the raster order are luma / chroma, intra / inter, and below / above a threshold (i.e., a threshold of 2).

[0056] The decoder resolves the transform coefficients, the signs of some transform coefficients, and the sign residuals of others as part of its resolution process. The signs and sign residuals are resolved at the end of the TU, and at this point, the decoder knows the absolute values ​​of all coefficients. Knowledge of "correct" or "incorrect" predictions is stored only as part of the TU data of the resolved block. The true signs of their associated coefficients are unknown at this point.

[0057] Subsequently, during reconstruction, the decoder performs operations similar to those of the encoder described above. First, sort-based symbol selection is performed. Based on the absolute value of qIdx, the coefficient with the highest qIdx value is placed at the beginning of the sorted TB. The stored CABAC-encoded "residual" symbols are applied to the first maxNumPredSigns coefficients in the sorted TB. Then, the true symbols to be applied to coefficients whose symbols have been predicted are determined by XORing the following: 1. Predicted value of the symbol; 2. "Correct" or "incorrect" data stored in the TU during bitstream parsing.

[0058] Symbol prediction applied to LFNST blocks In ECM-2.0, symbol prediction is enabled only for TB, where only the main transform, including the DCT-2 and MTS transform kernels, is applied. For TB applying the Low Frequency Inseparable Transform (LFNST), symbol prediction is always skipped.

[0059] JVET-X0150 and JVET-Y0141 propose applying symbolic prediction to LFNST blocks. Additionally, to achieve a better gain / complexity tradeoff, a maximum of 4 coefficients (i.e., M = 4) are allowed to be predicted for an LFNST TB.

[0060] This document aims to improve or simplify coefficient sign prediction. It proposes several modifications, including: 1) Symbol selection: a. The coefficients can be further ordered based on the weighted qIdx value (qIdx is the transformed coefficient level after compensating for the effects of multiple quantizers in DQ). b. Instead of sorting, place the significance coefficients with absolute level values ​​greater than a predefined threshold Th (e.g., 1) at the beginning to perform sign prediction in TB; 2) Generate prediction residuals, which have a more complex model than simple linear prediction; 3) Improved CABAC context derivation for prediction symbols; 4) Adaptive maximum number of prediction symbols set based on block size, maximum symbol prediction area, energy, transform type, block prediction mode (QP) and / or TB, or other parameters; 5) If LMCS is used, then the scan order, coefficient value range, cropping, and chroma scaling are encoded / decoded.

[0061] Symbol selection In the current ECM-8.0, the symbols to be predicted are selected based on the absolute value of qIdx (qIdx is the transform coefficient level after compensating for the effects of multiple quantizers in DQ). However, if multiple transform coefficients have the same absolute qIdx value, they may not contribute equally to the reconstructed boundary samples within the L-shaped template. Therefore, in one embodiment, we propose that the ranking of coefficients with the same qIdx value can be further based on weighted qIdx values, where the weights reflect the energy impact of the coefficients on the samples in the L-shaped template. For example, weights w It can be the sum of the absolutely reconstructed samples in the L-shaped template corresponding to the coefficients: (5) Weight w Alternatively, it can be measured using other methods, such as the coordinates of the transformation coefficients. i, j The Euclidean distance between the top left corner of TB (0, 0) and TB (0, 0), or the diagonal scan order index (scanIdx) of the transform coefficients in TB, etc.

[0062] In a variant of this embodiment, the sorting can be performed based on the weighted qIdx values of all transform coefficients within the allowed symbol prediction region of the TB, and the cost function can be calculated as follows: (6) where qIdx i, j represents the transform coefficient level at the coordinates ( i, j ) in the TB after compensating for the effects of multiple quantizers in DQ.

[0063] By observing that symbol errors of large transform coefficients have a relatively high impact on the reconstructed block, it has been found that sorting-based symbol selection can enhance the coding efficiency by making the prediction of the symbols of these large coefficients relatively easy. However, this sorting also involves a large amount of computation. To achieve a better gain / complexity trade-off, in one embodiment, we propose to place significant / non-zero coefficients whose absolute level values are higher than a predefined threshold Th (e.g., the value of this threshold Th can be set to 1) at the beginning of the buffer that collects n coefficients for symbol prediction in the TB. If the number of coefficients n in this buffer reaches the maximum limit (n = maxNumPredSigns), the symbol selection can be terminated, and these maxNumPredSigns symbols can be predicted; if the number of coefficients n is less than the maximum limit (n < maxNumPredSigns), other significant coefficients whose absolute level values are lower than the threshold Th are continuously placed in this collection buffer, and the symbols of the first maxNumPredSigns transform coefficients in the buffer are predicted.

[0064] In some examples, the value of this threshold Th can be predefined and fixed for all sequences, or can be signaled in the Sequence Parameter Set (SPS), View Parameter Set (VPS), Picture Parameter Set (PPS), or picture header. Alternatively, the value of this threshold Th can depend on the block size (e.g., the width and / or height of the current block), QP, color component, block prediction mode (intra or inter coding), transform type / core, slice type, sequence category, and configuration.

[0065] Prediction residual generation The prior art prediction residual described above is generated by using simple linear prediction, and this simple linear prediction can be further improved.

[0066] For each predicted pixel at the first left column of the current block P 0,yPerform a simple linear prediction using the two reconstructed pixels on the left to obtain its prediction residual. The predicted residuals and residual assumptions are then compared. r 0,y The absolute difference between the two is added to the cost of the hypothesis. Generating predicted residuals that are close to the actual residuals is crucial, as this can significantly affect the cost of the hypothesis. In one embodiment, we propose using a linear prediction estimated using the least mean square (LMS) method. The hypothesis is that for the predicted pixels... P 0,y Using the two reconstructed pixels on the left to obtain their prediction residuals, we can estimate the prediction residuals using linear prediction, as follows: (7) in a 1 and a 2 is the scaling factor. b It's the offset. For example, a 1. a 2 and b The value is estimated using the least mean square method, which minimizes the leftmost... and the very top The actual value of the reconstructed pixel is compared with the value at the left side of the boundary ( and ) / Above ( and The sum of the squared differences between the predicted values ​​of two reconstructed pixels.

[0067] In a variant of this embodiment, the linear prediction used to estimate the prediction residual may be without offset. b Of, such as: (8) Instead of using a simple linear regression with two adjacent reconstructed pixels, another variation of this embodiment can use a multiple linear regression (MLR) with more than two adjacent reconstructed pixels, and formulate it as follows: (9) In another variant of this embodiment, a polynomial model is proposed to estimate the prediction residuals, such as the conventional filter-based polynomial model used in the convolutional cross-component model (CCCM).

[0068] Context Export The existing CABAC context derivation technique described above for symbol prediction residuals is determined by whether the absolute level of the dequantized coefficients in the raster order is less than / greater than 2 at the decoder parsing process. However, the first maxNumPredSigns dequantized coefficients in the raster order are not always the coefficients used to perform symbol prediction.

[0069] In one embodiment, we propose performing order-based symbol selection during the resolution process before decoding the symbols, thus knowing which symbols to predict, and for each predicted symbol, deriving context to resolve the symbol prediction residuals based on the associated dequantized coefficient values. In a variant of this embodiment, for CABAC context derivation, the dependency on dequantized coefficient values ​​can be eliminated.

[0070] In another embodiment, we propose employing alternative derivation rules to select the CABAC context, such as based on the energy of the TB, specifically counting the number of transform coefficients whose absolute level in the TB is greater than a predefined threshold (e.g., 1 or 2). The CABAC context derivation for the prediction symbol in the TB is determined by whether the counted number is below or above another threshold (i.e., half the number of significant transform coefficients within the TB).

[0071] Adaptive maximum number setting for prediction symbols To limit complexity, there is a limit to the number of symbols to be predicted for a TB. In ECM-8.0, the maximum number of symbols to be predicted (maxNumPredSigns) is signaled to the decoder via SPS. The allowed values ​​for maxNumPredSigns are 0 to 8 (inclusive). The coding parameter NumSignPred assigned to symbol prediction is fixed at 8 under CTC, which is also the value assigned to maxNumPredSigns. Additionally, to achieve a better gain / complexity tradeoff, by using LFNST, a maximum of 4 coefficient symbols will be allowed to be predicted for a TB.

[0072] When the maximum number of symbols to be predicted (maxNumPredSigns) is set to 8, this means that in the worst case, up to 2^36 symbols may need to be generated and computed. 8 This involves several assumptions and associated costs, which require significant computation. Because a smaller maximum number of prediction symbols results in shorter processing time and lower computational complexity. In one embodiment, we propose setting an adaptive maximum number of prediction symbols based on certain conditions / parameters.

[0073] In a variant of this embodiment, the maximum number of prediction symbols used for TB can be based on the block size. Typically, this is achieved if the block size (e.g., the width and / or height of the current block) meets a specific threshold (e.g., less than, greater than, equal to, etc.). Th (i.e., 8), then for this TB, a maximum of K coefficient symbols will be predicted; otherwise, for this TB, a maximum of μK (e.g., μ=2) coefficient symbols will be predicted. For example, when maxNumPredSigns is set to 8, if the width or height of the current block is less than 8, then for this TB, a maximum of 4 coefficient symbols will be allowed to be predicted; otherwise, a maximum of 8 coefficient symbols will be predicted.

[0074] In one embodiment, the maximum number of prediction symbols used for the TB increases with increasing block size in terms of improving compression efficiency. In another embodiment, if the goal is to reduce complexity, the maximum number of prediction symbols used for the TB decreases with increasing block size.

[0075] In some examples, this threshold Th The value can be predefined and fixed for all sequences, or it can be variable and sent as a signal in the SPS, VPS, PPS, or image header. Alternatively, this threshold... Th The value can depend on the block size (e.g., the width and / or height of the current block), QP, color components, block prediction mode (intra-frame or inter-frame coding), transform type / core, slice type, sequence class, and configuration.

[0076] In some examples, coefficient sign prediction can also be disabled if the block size is very small and / or very large.

[0077] In another variation of this embodiment, the maximum number of prediction symbols for the TB can be based on a specific size satisfying (e.g., less than, greater than, equal to, etc.) the maximum symbol prediction region for the TB. For example, if the maximum region for symbol prediction for the TB is less than 32×32, then for the TB, the maximum number of predicted coefficient symbols will be reduced by a maximum value (i.e., a maximum reduction of 4 when maxNumPredSigns is set to 8).

[0078] In another variation of this embodiment, the maximum number of predicted symbols for the TB can be based on the energy of the TB, for example, based on the number of transform coefficients whose absolute level in the TB is greater than a predefined threshold (i.e., 1 or 2). For the TB, the maximum number of predicted reduced coefficient symbols (e.g., 4 when maxNumPredSigns is set to 8) will be predicted, depending on whether the number counted is below / above another threshold (i.e., half the number of significant coefficients within the TB).

[0079] In another variant of this embodiment, the maximum number of prediction symbols used for the TB can be based on the transform type or transform core. For example, in VVC and ECM, blocks can use different horizontal / vertical transforms, known as multiple transform selection (MTS). It can also be encoded using sub-block transform (SBT). The maximum number of prediction symbols used for the TB can be selected depending on whether DCT-II / DCT-VIII / DST-VII / MTS / SBT is used for that TB.

[0080] In another variation of this embodiment, the maximum number of prediction symbols used for TB can be based on QP, color components, block prediction mode (intra-frame or inter-frame coding), slice type, sequence category, and configuration. These parameters / conditions can be combined to determine the maximum number of prediction symbols set.

[0081] other In one embodiment, we propose using a diagonal / zigzag order instead of a raster order during the parsing process at the decoder and the encoding process at the encoder to improve encoding efficiency.

[0082] In another embodiment, we propose using the internal or input bit depth instead of the current SIGN_PRED_SHIFT = 8 to represent the range of coefficient values ​​to improve coding efficiency.

[0083] In another embodiment, we propose pruning the prediction residual or prediction reconstruction based on the internal or input bit depth, such as... or .

[0084] When using Luminance Mapping and Chroma Scaling (LMCS), in one embodiment, we propose performing chroma scaling first on the previous reconstruction hypothesis and on each template corresponding to the predicted symbol, rather than directly performing chroma scaling on the current reconstruction hypothesis. For example, the current reconstruction hypothesis in the prior art is obtained using the following formula: And in our proposed method, it is obtained using the following formula: .

[0085] Figure 6 The illustration depicts a method for symbol prediction at the encoder according to an embodiment. In step 610, a maximum number of predicted symbols is adapted based on some conditions / parameters. In step 620, symbols to be predicted are selected. In step 630, symbols are predicted, and symbol prediction residuals are generated. In step 640, a CABAC context is derived for the symbol prediction residuals. In step 650, the symbol prediction residuals are encoded.

[0086] Figure 7The illustration depicts a method for symbol prediction at the decoder according to an embodiment. In step 710, a maximum number of predicted symbols is adapted based on some conditions / parameters. In step 720, symbols to be predicted are selected. In step 730, the symbols are predicted. In step 740, a CABAC context is derived for the symbol prediction residuals. In step 750, the symbol prediction residuals are decoded. In step 760, the symbols are reconstructed. Further details are described below. Figure 6 and Figure 7 Some steps of the method shown.

[0087] Figure 8 The illustration shows a method for adapting the maximum number of prediction symbols based on the block size according to an embodiment. The width or height of the current TB is compared with a specific threshold (8 in this case) (810). If the width or height of the current TB is less than 8, half the value of maxNumPredSigns is allowed for the maximum number of prediction symbols used for that TB (820); otherwise, the normal maxNumPredSigns is used (830).

[0088] 1. Adaptive maximum number of prediction symbols (610, 710) based on certain conditions / parameters: - Block Size: If the block size (e.g., the width and / or height of the current block) meets (e.g., less than, greater than, equal to, etc.) a specific threshold Th (i.e., 8), then for that TB, the maximum number of prediction symbols allowed to be reduced. Specifically, when the width or height of the current TB is less than 8, the maximum number of prediction symbols used for that TB is allowed to be half the value of maxNumPredSigns (4 when maxNumPredSigns is set to 8). In some examples, the value of the threshold Th can be predefined and fixed for all sequences, or it can be signaled in the SPS, VPS, PPS, or image header. Alternatively, the value of the threshold Th can depend on the block size, QP, color components, block prediction mode (intra- or inter-frame coding), transform type / core, slice type, sequence category, and configuration. - Maximum symbol prediction region of a TB: The maximum number of prediction symbols that can be reduced for a given TB when it meets (e.g., less than, greater than, equal to, etc.) a specific size (i.e., 32×32). - TB Energy: Counts the number of transform coefficients whose absolute level in the TB is greater than a predefined threshold (e.g., 1 or 2). If the counted number is below / above another threshold (i.e., half the number of significant transform coefficients within the TB), then the maximum number of prediction symbols allowed to be reduced for that TB. - Transformation type / core: Enables the maximum number of prediction symbols used for the TB to be adapted to whether DCT-II / DCT-VIII / DST-VII / MTS / SBT is used for the TB. - Other factors, such as QP, color classification, block prediction mode (intra-frame or inter-frame coding), slice type, sequence category, and configuration. - These parameters / conditions can be used individually or combined to determine the maximum number of predictive symbols to set.

[0089] 2. Select the symbols to be predicted (620, 720): - In addition to the ranking based on the absolute value of qIdx (qIdx is the level of transform coefficients after compensating for the effects of multiple quantizers in DQ) used to select coefficients for sign prediction, the ranking can be further based on weighted qIdx values. The weights can be derived from the following: ■ The effect of the transform coefficients on the energy of samples (the first left column and the first top row of the current block) in the L-shaped template, specifically the sum of the absolute reconstructed samples in the L-shaped template corresponding to the transform coefficients, or ■ The Euclidean distance between the coordinates of the transform coefficients and the top-left corner of the TB, or the diagonal scan order index of the transform coefficients in the TB. Consider using weighted qIdx ■ Only when multiple transformation coefficients have the same absolute qIdx value, or ■ All transform coefficients within the allowed symbol prediction region of TB Instead of sorting, the significance coefficients with absolute level values ​​greater than a predefined threshold Th (i.e., 1) are placed at the beginning to perform symbolic prediction in TB; • In some examples, the value of the threshold Th can be predefined and fixed for all sequences, or it can be signaled in the SPS, VPS, PPS, or picture header. Alternatively, the value of the threshold Th can depend on the block size, QP, color components, block prediction mode (intra-frame or inter-frame coding), transform type / core, slice type, sequence category, and configuration.

[0090] 3. Predicting residual generation (630): - Linear forecasting using the least mean square (LMS) method The linear parameters (multiple scaling factors and multiple offsets) are estimated using the least mean square method, which minimizes the sum of the squared differences between the actual value of the leftmost / topmost reconstructed pixel and the predicted values ​​of the two reconstructed pixels to its left / above. - Linear prediction using multiple linear regression (MLR) estimation, which utilizes more than two adjacent reconstructed pixels. - Use a polynomial model, such as the traditional filter-based polynomial model used in the Convolutional Cross Component Model (CCCM).

[0091] 4. CABAC context derivation for predicting symbols (640, 740): - Before encoding or decoding symbols, order-based symbol selection is performed, thus determining which symbols to predict. For each predicted symbol, it can derive context to resolve the symbol residuals based on the associated dequantized coefficient values, or... - Other derived rules: The energy of the TB is specifically counted as the number of coefficients whose absolute level in the TB is greater than a predefined threshold (i.e., 1 or 2). The CABAC context used for the prediction symbol in the TB is based on whether the counted number is below / above another threshold (i.e., half the number of significant transformation coefficients within the TB).

[0092] 5. Other: - Use diagonal / zigzag order instead of raster order during parsing at the decoder and encoding at the encoder. - Use the internal or input bit depth, instead of the current SIGN_PRED_SHIFT=8, to represent the range of coefficient values. - Pruning of prediction residuals or prediction reconstruction based on internal or input bit depth - When using Luminance Mapping and Chromaticity Scaling (LMCS), chroma scaling is first performed on the previous reconstruction hypothesis and each template corresponding to the predicted symbol, rather than performing chroma scaling directly on the current reconstruction hypothesis.

[0093] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first," "second," etc., can be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of these terms does not imply a modified ordering of operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0094] The various methods and other aspects described in this application can be used to modify the module, for example. Figure 2 and Figure 3 The reconstruction modules (255, 355) of the video encoder 200 and decoder 300 are shown. Furthermore, this aspect is not limited to ECM and VVC, and can be applied to, for example, other standards and recommendations, as well as any extensions of such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application may be used individually or in combination.

[0095] Various numerical values ​​are used in this application. Specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0096] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process, such as performing a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more procedures typically performed by a decoder, such as entropy coding, inverse quantization, inverse transform, and differential decoding. Whether the phrase “decoding process” is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.

[0097] Various implementations involve encoding. In a manner similar to the discussion of "decoding" above, "encoding" as used in this application can include, for example, performing all or part of a process on an input video sequence to produce an encoded bitstream.

[0098] Note that the grammatical elements used in this article are descriptive terms. Therefore, the use of other grammatical element names is not excluded.

[0099] The implementations and aspects described herein can be implemented, for example, as methods or processes, apparatuses, software programs, data streams, or signals. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the features under discussion can be implemented in other forms (e.g., apparatuses or programs). Apparatuses can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in apparatuses, such as processors, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0100] References to "an embodiment" or "an implementation" or "an implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in one embodiment" or "in one implementation" or "in one implementation" appearing in various places throughout this application, and any other variations, do not necessarily all refer to the same embodiment.

[0101] Additionally, this application may relate to "determining" various information segments. Determining information may include one or more of the following: for example, estimation information, calculation information, prediction information, or information retrieved from memory.

[0102] Furthermore, this application may involve "accessing" various information segments. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0103] Additionally, this application may relate to "receiving" various segments of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., from memory). Further, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0104] It should be understood that the use of any of the following “ / ”, “and / or”, and “…at least one of…”—for example, in the cases of “A / B”, “A and / or B”, and “at least one of A and B”—is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As further examples, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, this wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). Those skilled in the art and related fields will appreciate that this can be extended to as many items as possible listed.

[0105] Similarly, as used herein, among other things, the phrase “signal” indicates something to the corresponding decoder. For example, in some embodiments, the decoder signals a quantization matrix used for dequantization. In this way, in one embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters to the decoder (explicit signaling) so that the decoder can use the same specific parameters. Conversely, if the decoder already has the specific parameters as well as other parameters, signaling can be used without transmission (implicit signaling) to allow only the decoder to know and select the specific parameters. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be done in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc., are used to signal information to the corresponding decoder. Although the verb form of the phrase “signal” is referred to above, the word “signal” can also be used as a noun herein.

[0106] It will be apparent to those skilled in the art that implementations can generate various signals formatted to carry information that can, for example, be stored or transmitted. For example, the information may include instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. For example, formatting may include encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal can be transmitted via various wired or wireless links. The signal may be stored on a processor-readable medium.

Claims

1. A video decoding method, the method comprising: The limits on the number of symbols to be predicted for a block are determined based on the coding conditions and parameters; A signal is obtained indicating whether the sign of the transform coefficient matches the predicted sign of the transform coefficient; In response to the transform coefficients belonging to a subset of transform coefficients of the block, the predicted sign of the transform coefficients is obtained, wherein the sign of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit. Based on the predicted symbol and the signal, the symbol of the transform coefficient is obtained; An inverse transform is performed on the transform coefficient block corresponding to the block, the transform coefficient block including the transform coefficients; and The block is decoded based on the prediction block corresponding to the block and the transform coefficient block.

2. A video encoding method, the method comprising: The limits on the number of symbols to be predicted for a block are determined based on the coding conditions and parameters; Obtain the sign of the transformation coefficients; In response to the transform coefficients belonging to a subset of transform coefficients of the block, a predicted sign of the transform coefficients is obtained, wherein the sign of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit. Obtain a signal to indicate whether the predicted symbol matches the symbol of the transform coefficient; as well as The signal is encoded.

3. The method according to claim 1 or 2, wherein the limitation is based on the block size of the block.

4. The method of claim 3, wherein the limitation increases with the increase of the block size.

5. The method of claim 3, wherein the limitation increases as the block size decreases.

6. The method according to any one of claims 3 to 5, wherein the limitation is set to a first value and a second value respectively in response to the block size being greater than and less than a third value.

7. The method of claim 6, wherein the second value is twice the first value.

8. The method of claim 6, wherein the third value depends on at least one of block size, color component, quantization parameter, prediction mode, transform type or core, slice type, sequence category, and configuration.

9. The method of claim 1, wherein the limitation is based on the energy of the block.

10. The method of claim 1, wherein the limitation is based on at least one of the block's transform type or transform core, color components, prediction mode, transform type or core, slice type, sequence category, and configuration.

11. An apparatus comprising at least one memory and one or more processors, wherein the one or more processors are configured to: The limits on the number of symbols to be predicted for a block are determined based on the coding conditions and parameters; A signal is obtained indicating whether the sign of the transform coefficient matches the predicted sign of the transform coefficient; In response to the transform coefficients belonging to a subset of transform coefficients of the block, the predicted sign of the transform coefficients is obtained, wherein the sign of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit. Based on the predicted symbol and the signal, the symbol of the transform coefficient is obtained; An inverse transform is performed on the transform coefficient block corresponding to the block, the transform coefficient block including the transform coefficients; and The block is decoded based on the prediction block corresponding to the block and the transform coefficient block.

12. An apparatus comprising at least one memory and one or more processors, wherein the one or more processors are configured to: The limits on the number of symbols to be predicted for a block are determined based on the coding conditions and parameters; Obtain the sign of the transformation coefficients; In response to the transform coefficients belonging to a subset of transform coefficients of the block, a predicted sign of the transform coefficients is obtained, wherein the sign of each transform coefficient in the subset of transform coefficients is to be predicted, and wherein the number of transform coefficients in the subset of transform coefficients is within the limit. Obtain a signal to indicate whether the predicted symbol matches the symbol of the transform coefficient; as well as The signal is encoded.

13. The apparatus of claim 11 or 12, wherein the limitation is based on the block size of the block.

14. The apparatus of claim 13, wherein the limitation increases with the increase of the block size.

15. The apparatus of claim 13, wherein the limitation increases as the block size decreases.

16. The apparatus according to any one of claims 13 to 15, wherein the limitation is set to a first value and a second value respectively in response to the block size being greater than and less than a third value.

17. The apparatus of claim 16, wherein the second value is twice the first value.

18. The apparatus of claim 16, wherein the third value depends on at least one of block size, color components, quantization parameters, prediction mode, transform type or core, slice type, sequence category, and configuration.

19. The apparatus of claim 11, wherein the limitation is based on the energy of the block.

20. The apparatus of claim 11, wherein the limitation is based on at least one of the block's transform type or transform core, color components, prediction mode, transform type or core, slice type, sequence category, and configuration.

21. A signal comprising a bit stream, the bit stream being formed by performing the method according to any one of claims 2 to 10.

22. A computer-readable storage medium having instructions stored thereon for encoding or decoding video according to any one of claims 1 to 10.