Encoding method, decoding method, encoder, decoder, code stream, and storage medium
By indicating the identification information of inseparable transformation at the transform block level, the problem of inseparable transformation in the prior art is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- PCT/CN2024/072599
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-17
AI Technical Summary
The existing inseparable transformation operations are unreasonable in some scenarios, resulting in a reduced encoding efficiency.
By introducing identification information at the transform block level to indicate whether an indivisible transformation is used, the adjacent information of the current transform block is determined based on the information, thereby selecting suitable adjacent information for encoding or decoding.
Improves the encoding and decoding performance, ensures the accurate determination of transform block coefficients, and improves the encoding efficiency.
Smart Images

Figure CN2024072599_17072025_PF_FP_ABST
Abstract
Description
Coding and decoding method, codec, code stream and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on January 9, 2024, with application number PCT / CN2024 / 071312 and application name “Codec Method, Codec, Bit Stream and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of video coding and decoding technology, and in particular to a coding and decoding method, a codec, a bit stream, and a storage medium. Background Art
[0003] Transform operations can be used to remove redundant information from residual data, improving video compression performance. Non-separable transforms are an important type of transform. However, current encoding and decoding methods for non-separable transforms are sometimes illogical, potentially degrading encoding and decoding performance.
[0004] Summary of the Invention
[0005] The present application provides a coding and decoding method, a codec, a bit stream, and a storage medium. The following introduces various aspects of the present application.
[0006] In a first aspect, a decoding method is provided, which is applied to a decoder, and the decoding method includes: determining first information, where the first information is used to indicate whether a current transform block uses an inseparable transform; determining adjacent information of the current transform block based on the first information; and determining coefficients of the current transform block based on the adjacent information.
[0007] In a second aspect, a coding method is provided, which is applied to an encoder, and the coding method includes: determining first information, where the first information is used to indicate whether the current transformation block uses an inseparable transformation; determining adjacent information of the current transformation block based on the first information; and encoding the coefficients of the current transformation block based on the adjacent information.
[0008] According to a third aspect, a decoder is provided, comprising: a first determination unit configured to determine first information, wherein the first information is used to indicate whether a current transform block uses an inseparable transform; a second determination unit configured to determine adjacent information of the current transform block based on the first information; and a third determination unit configured to determine coefficients of the current transform block based on the adjacent information.
[0009] In a fourth aspect, a decoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the first aspect when running the computer program.
[0010] In the fifth aspect, an encoder is provided, including: a first determination unit, configured to determine first information, wherein the first information is used to indicate whether the current transform block uses an inseparable transform; a second determination unit, configured to determine the adjacent information of the current transform block based on the first information; and an encoding unit, used to encode the coefficients of the current transform block based on the adjacent information.
[0011] In a sixth aspect, an encoder is provided, comprising: a memory for storing a computer program; and a processor for executing the method of the second aspect when running the computer program.
[0012] In a seventh aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method of the first aspect or the second aspect is implemented.
[0013] In an eighth aspect, a computer program product is provided, comprising a computer program, which implements the method of the first aspect or the second aspect when the computer program is executed.
[0014] In a ninth aspect, a non-volatile computer-readable storage medium for storing a bit stream is provided, wherein the bit stream is generated by an encoding method of an encoder, or the bit stream is decoded by a decoding method of a decoder, wherein the decoding method is the method described in the first aspect and the encoding method is the method described in the second aspect.
[0015] In a tenth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed, the method described in the first aspect or the second aspect is implemented.
[0016] According to an eleventh aspect, a code stream is provided, including a code stream generated according to the method described in the second aspect.
[0017] The embodiment of the present application introduces transform block-level identification information (i.e., the first information mentioned above) to indicate whether to use an inseparable transform. Compared with the coding block-level identification information in the related art, the transform block-level identification information helps to select appropriate adjacent information, thereby helping to improve the encoding and decoding performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] FIG1 is a structural diagram illustrating an example of a video encoder to which an embodiment of the present application may be applied.
[0019] FIG2 is a diagram showing an example structure of a video decoder to which an embodiment of the present application can be applied.
[0020] FIG3 is an example diagram of the LFNST transformation and inverse transformation process.
[0021] FIG4 is an example diagram of two types of adjacent information.
[0022] FIG5 is a flowchart of a decoding method provided in an embodiment of the present application.
[0023] FIG6 is a flow chart of the encoding method provided in an embodiment of the present application.
[0024] FIG7 is a schematic flowchart of determining neighbor information provided in an embodiment of the present application.
[0025] FIG8 is a schematic flowchart of another method for determining neighbor information provided in an embodiment of the present application.
[0026] FIG9 is a schematic flowchart of another method for determining neighbor information provided in an embodiment of the present application.
[0027] FIG10 is a schematic flowchart of another method for determining neighbor information provided in an embodiment of the present application.
[0028] FIG11 is a schematic diagram of determining a gradient histogram provided in an embodiment of the present application.
[0029] FIG12 is a schematic diagram of the structure of a decoder provided in one embodiment of the present application.
[0030] FIG13 is a schematic structural diagram of a decoder provided in another embodiment of the present application.
[0031] FIG14 is a schematic diagram of the structure of an encoder provided in one embodiment of the present application.
[0032] FIG15 is a schematic diagram of the structure of an encoder provided in another embodiment of the present application. DETAILED DESCRIPTION
[0033] The technical solution in this application will be described below with reference to the accompanying drawings.
[0034] FIG1 is a schematic block diagram of a video encoder according to an embodiment of the present application.
[0035] It should be understood that the video encoder 100 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless compression or mathematically lossless compression.
[0036] The video encoder 100 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents brightness (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) for describing color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0037] For example, the video encoder 100 reads video data and, for each image in the video data, divides the image into a number of coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," a "largest coding unit" (LCU) or a "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chrominance sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU may be further divided into a number of coding units (CUs) for encoding. A CU may be a rectangular block or a square block. A CU may correspond to a prediction unit (PU) and a transform unit (TU).
[0038] In some embodiments, as shown in FIG1 , the video encoder 100 may include a prediction module 110, a residual module 120, a transform / quantization module 130, an inverse transform / quantization module 140, a reconstruction module 150, a loop filter module 160, a decoded image buffer 170, and an entropy coding module 180. It should be noted that the video encoder 100 may include more, fewer, or different functional components.
[0039] Optionally, in this application, the current block may be referred to as the current coding unit (CU). The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstruction block. Due to the need for parallel processing, an image may be divided into slices. Slices in the same image may be processed in parallel, meaning that there is no data dependency between them. The term "frame" is commonly used, and it can generally be understood that a frame is an image. The term "frame" herein may also be replaced by "image" or "slice," etc.
[0040] In some embodiments, the prediction module 110 includes an inter-frame prediction module 111 and an intra-frame prediction module 112. Because there is a strong correlation between adjacent pixels in a video image, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent images in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent images, thereby improving coding efficiency.
[0041] The inter-frame prediction module 111 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation. It can refer to image information of different images. Inter-frame prediction uses motion information to find a reference block from the reference image and generates a prediction block based on the reference block to eliminate temporal redundancy. Motion information includes the reference image list where the reference image is located, the reference image index, and the motion vector. The motion vector can be an integer pixel or a sub-pixel. If the motion vector is a sub-pixel, it is necessary to use interpolation filtering in the reference image to generate the required sub-pixel block. Here, the integer pixel or sub-pixel block in the reference image found based on the motion vector is called a reference block. Some technologies will directly use the reference block as the prediction block, while other technologies will further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0042] The intra-frame prediction module 112 only refers to information of the same image to predict pixel information within the current image block to eliminate spatial redundancy.
[0043] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. High-efficiency video coding (HEVC) uses planar, direct current (DC), and 33 angular modes for a total of 35 intra-frame prediction modes. Versatile video coding (VVC) uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.
[0044] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of the development of high-definition and ultra-high-definition digital videos.
[0045] Residual module 120 may generate a residual block for a CU based on the pixel block of the CU and the prediction block of the CU. For example, residual module 120 may generate a residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the CU.
[0046] The transform / quantization module 130 may quantize the transform coefficients. The transform / quantization module 130 may quantize the transform coefficients associated with the CU based on a quantization parameter (QP) value associated with the CU. The video encoder 100 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0047] The inverse transform / quantization module 140 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0048] Reconstruction module 150 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction module 110 to generate a reconstructed image block associated with the CU. By reconstructing each sample block of the CU in this manner, video encoder 100 can reconstruct the pixel blocks of the CU.
[0049] The loop filter module 160 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent pixel encoding. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.
[0050] In some embodiments, the loop filtering module 160 includes a deblocking filtering module and a sample adaptive offset / adaptive loop filtering (SAO / ALF) module, wherein the deblocking filtering module is used to remove blocking effects, and the SAO / ALF module is used to remove ringing effects.
[0051] The decoded image buffer 170 may store the reconstructed pixel blocks. The inter prediction module 111 may use a reference image containing the reconstructed pixel blocks to perform inter prediction on PUs of other images. In addition, the intra prediction module 112 may use the reconstructed pixel blocks in the decoded image buffer 170 to perform intra prediction on other PUs in the same image as the CU.
[0052] The entropy encoding module 180 may receive the quantized transform coefficients from the transform / quantization module 130. The entropy encoding module 180 may perform one or more entropy encoding operations on the quantized transform coefficients to generate entropy-encoded data.
[0053] FIG2 is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0054] 2 , video decoder 200 includes an entropy decoding module 210, a prediction module 220, an inverse quantization / transformation module 230, a reconstruction module 240, a loop filter module 250, and a decoded image buffer 260. It should be noted that video decoder 200 may include more, fewer, or different functional components.
[0055] The video decoder 200 may receive a bitstream. The entropy decoding module 210 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding module 210 may parse the entropy-encoded syntax elements in the bitstream. The prediction module 220, the inverse quantization / transformation module 230, the reconstruction module 240, and the loop filter module 250 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.
[0056] In some embodiments, the prediction module 220 includes an intra-frame prediction module 222 and an inter-frame prediction module 221 .
[0057] The intra prediction module 222 may perform intra prediction to generate a prediction block for a PU. The intra prediction module 222 may use an intra prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. The intra prediction module 222 may also determine the intra prediction mode for the PU based on one or more syntax elements parsed from the codestream.
[0058] The inter-frame prediction module 221 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. Furthermore, if a PU is encoded using inter-frame prediction, the entropy decoding module 210 may parse the motion information of the PU. The inter-frame prediction module 221 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction module 221 may generate a prediction block for the PU based on the one or more reference blocks of the PU.
[0059] The inverse quantization / transform module 230 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform module 230 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0060] After inverse quantizing the transform coefficients, inverse quantization / transform module 230 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
[0061] Reconstruction module 240 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, reconstruction module 240 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
[0062] The loop filtering module 250 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.
[0063] The video decoder 200 may store the reconstructed image of the CU in the decoded image buffer 260. The video decoder 200 may use the reconstructed image in the decoded image buffer 260 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0064] The basic process of video encoding and decoding is as follows: At the encoder end, an image is divided into blocks. For the current block, the prediction module 110 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual module 120 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block can also be referred to as residual information. This residual block undergoes transformation and quantization by the transform / quantization module 130, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization module 130 can be referred to as a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization module 130 can be referred to as a frequency residual block or a frequency-domain residual block. The entropy coding module 180 receives the quantized transform coefficients output by the transform / quantization module 130 and performs entropy coding on these quantized transform coefficients to output a bitstream. For example, the entropy coding module 180 can eliminate character redundancy based on the target context model and probability information of the binary bitstream.
[0065] At the decoding end, the entropy decoding module 210 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction module 220 uses intra-frame prediction or inter-frame prediction on the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation module 230 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction module 240 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering module 250 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference image for inter-frame prediction of subsequent images.
[0066] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.
[0067] It is understandable that the "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text. The "transformation" and "inverse transformation" in the embodiments of the present application correspond to two opposite processes. For example, if the "transformation" converts the numerical values in the spatial domain to the coefficients in the frequency domain, then the "inverse transformation" converts the coefficients in the frequency domain to the numerical values in the spatial domain. If the standard only stipulates decoding, then the "transformation" in the standard text is the decoding part, which refers to the "inverse transformation" in this article. The "inverse transformation" of the transform coefficients at the decoding end may also be referred to as "transformation" in the standard text.
[0068] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0069] The preceding text describes in detail the codec framework provided by the embodiments of this application. This application relates to transformation technology. This transformation technology can be applied to the transformation / inverse transformation module in the codec framework. The following describes the transformation technology involved in this application.
[0070] Dual tree partition (DT) partition
[0071] As can be seen from the foregoing, the encoder and decoder in the embodiments of the present application can be applied to image data in luminance and chrominance (YCbCr, YUV) formats. Generally, there is a significant difference in the image details contained in the luminance component and the chrominance component. The luminance component has a large amount of detailed information, while the chrominance component carries less information and appears relatively flat. Therefore, the luminance component is often more suitable for being divided into smaller blocks to express more details. Conversely, the chrominance component does not need to be finely divided in most cases.
[0072] The Versatile Video Coding (VCC) standard uses a partitioning technique, known as DT partitioning, that allows for different partitioning of luma and chroma components. DT partitioning is applied to intra-coded frames. For intra-coded frames, luma and chroma components can be partitioned differently. Starting from each coding tree unit, the luma component uses the luma partitioning tree, while the chroma components use the chroma partitioning tree. Since the luma and chroma components have their own coding tree units, they also have their own coding units.
[0073] For inter-frame coded frames or coded frames that do not apply the DT partitioning method, the luminance component and the chrominance component do not use different partitioning trees, and a coding unit includes a luminance block and a chrominance block.
[0074] KL transform (Karhunen-Loève transform)
[0075] The name of the KL transform comes from Kari Karhunen and Michel Loève. The KL transform is an orthogonal transformation of the input vector X so that the output vector can remove the correlation of the data.
[0076] The KL transform is a transform based on statistical properties. It is the best transform in terms of mean square error (MSE). Therefore, it plays an important role in data compression technology.
[0077] While the KL transform offers the best performance in terms of MSE, it requires prior knowledge of the input signal and requires complex mathematical operations, such as calculation of covariance and eigenvectors. Consequently, it's not widely used in engineering practice. However, the KL transform is theoretically optimal, so when searching for less-than-optimal but more practical transformation methods, it can provide a performance metric for evaluating transformation performance.
[0078] For example, in image processing, the KL transform concentrates the image's energy, helping to compress it. However, the KL transform is input-dependent, requiring a different transformation mechanism for each input image, making it impractical in practical applications.
[0079] Inseparable transformation
[0080] An inseparable transform refers to a two-dimensional transform that cannot be separated into two one-dimensional transforms in the horizontal and vertical directions. In this application, inseparable transforms can include low-frequency non-separable secondary transforms (LFNST) and non-separable primary transforms (NSPT). These two transform methods are introduced below.
[0081] LFNST is a transform based on the KL transform. Due to complexity considerations, in VCC, LFNST operates on the transform coefficients after the base transform, so it is called a secondary transform. LFNST can be applied to the low-frequency coefficients after the base transform to remove redundant information in the low-frequency coefficients. The base transform can be, for example, a DCT-II transform. This transform method can also be called a DCT-II + LFNST transform. In some implementations, the base transform can also be called a primary transform.
[0082] In VCC, LFNST can have two different sizes of transform kernels. One is a transform kernel with 64 coefficients input and 16 coefficients output, and the other is a transform kernel with 16 coefficients input and 8 coefficients output. Transform kernels of different sizes can be applied to transform blocks of different sizes. For example, a transform kernel with 64 coefficients input is applicable to transform blocks of size greater than or equal to 8×8, while a transform kernel with 16 coefficients input is applicable to transform blocks of other sizes (such as 4×4). A transform kernel with 64 coefficients input is also called a large-size transform kernel, and a transform kernel with 16 coefficients input is also called a small-size transform kernel.
[0083] The following describes the transformation and inverse transformation process of LFNST in conjunction with Figure 3.
[0084] At the encoding end, the residual information can be transformed using forward basis transform and LFNST. For example, the residual information can be subjected to a forward basis transform to obtain the first coefficient, and then the first coefficient can be transformed using LFNST to obtain the transformed coefficient. When using LFNST, the 64 coefficients of the 8×8 area in the upper left corner can be used as the input of the large-size transform kernel, or the 16 coefficients of the 4×4 area in the upper left corner can be used as the input of the small-size transform kernel. The other areas are defaulted to all-zero coefficients. After processing by the transform kernel, the transformed coefficients are obtained, and then the transformed coefficients are quantized to obtain the coefficients of the transform block. The transformed coefficients are encoded into the bitstream (or bitstream). Similarly, at the decoding end, the coefficients of the transform block can be parsed from the bitstream, the coefficients of the transform block can be inversely quantized to obtain the transformed coefficients, and then the transformed coefficients can be subjected to inverse LFNST and inverse basis transform to obtain the residual information. When performing inverse LFNST, the 16 positions of the upper left 4×4 area may include the output of the large-size inverse transform kernel, or the first 8 positions in the upper left scan order may include the output of the small-size inverse transform kernel, and the coefficients at other positions are defaulted to 0.
[0085] In VCC, the LFNST transform kernel selection is not only related to the size of the transform block, but also to the intra-frame prediction mode used during the transform block prediction and the LFNST index parsed from the bitstream. In implementation, the LFNST transform coefficients are stored in a multidimensional array g_lfnstMxN[][][][], where the first dimension is the index related to the intra-frame prediction mode, the second dimension is the transform kernel index parsed from the bitstream, and the third and fourth dimensions are the transform coefficients.
[0086] Table 1 shows the correspondence between intra prediction modes and related indexes.
[0087] Table 1
[0088] In Table 1, IntraPredMode is the intra prediction mode, and Tr.set index is the index related to the intra prediction mode. When the intra prediction mode is less than 0, the relevant index is 1; when the intra prediction mode is greater than or equal to 0 and less than or equal to 1, the relevant index is 0; when the intra prediction mode is greater than or equal to 2 and less than or equal to 12, the relevant index is 1; when the intra prediction mode is greater than or equal to 13 and less than or equal to 23, the relevant index is 2; when the intra prediction mode is greater than or equal to 24 and less than or equal to 44, the relevant index is 3; when the intra prediction mode is greater than or equal to 45 and less than or equal to 55, the relevant index is 2; when the intra prediction mode is greater than or equal to 56 and less than or equal to 80, the relevant index is 1; when the intra prediction mode is greater than or equal to 81 and less than or equal to 83, the relevant index is 0.
[0089] In the enhanced compression model (ECM), some enhancement techniques of LFNST and NSPT are adopted.
[0090] Compared to the LFNST technology in VCC, the enhanced LFNST technology includes a wider range of transform kernel sizes and finer kernel grouping. Different kernel sizes can be used for transform blocks greater than or equal to 16×16, greater than or equal to 8×8 and less than 16×16, and greater than or equal to 4×4 and less than 8×8. Kernel groups can be divided into 35 groups based on intra-frame prediction modes and into three groups based on kernel indices parsed from the bitstream.
[0091] NSPT technology shares similar transformation principles with LFNST. NSPT replaces the traditional DCT-II base transform and LFNST secondary transform. Instead, it directly uses group-based NSPT base transform coefficients for the residual transform and inverse transform. In ECM-11.0, block sizes that can use NSPT technology include 4x4, 4x8, 4x16, 4x32, 8x4, 8x8, 8x16, 8x32, 16x4, 16x8, 32x4, and 32x8. NSPT and LFNST share the same group classification and number of input and output coefficients. For block sizes that can use NSPT, NSPT transform coefficients can be directly retrieved using the parsed LFNST sign and index, replacing the traditional DCT-II + LFNST transform.
[0092] Whether LFNST or NSPT is used can be indicated by an LFNST flag. The LFNST flag can be represented by, for example, cu.lfnstIdx or lfnstIdx. The LFNST flag can be a coding unit level flag.
[0093] Taking LFNST as an example, if the LFNST flag indicates that LFNST is used, the transform coefficients are obtained after the basic transform and LFNST. If the LFNST flag indicates that LFNST is not used, the transform coefficients are obtained only after the basic transform. On the encoding side, if LFNST is used, the transform and quantization module can first perform a basic transform on the residual information, and then perform LFNST to obtain the transform coefficients. After quantization, the transform coefficients are represented using the corresponding flag and encoded into the bitstream. On the decoding side, the decoder can parse the bitstream to obtain the LFNST flag and the corresponding flag, first construct the quantized coefficients, and then perform inverse quantization to obtain the transform coefficients. If the LFNST flag indicates that LFNST is used, the decoder can first perform an LFNST inverse transform on the transform coefficients, and then perform a basic inverse transform to obtain the reconstructed residual information. If the LFNST flag indicates that LFNST is not used, the decoder can perform a basic inverse transform on the transform coefficients to obtain the reconstructed residual information.
[0094] Taking NSPT as an example, if the LFNST flag indicates that NSPT is used, the transform coefficients are obtained after NSPT. If the LFNST flag indicates that NSPT is not used, the transform coefficients are obtained after basic transform. At the encoding end, if NSPT is used, the transform and quantization module can first perform NSPT on the residual information to obtain the transform coefficients. After quantization, the transform coefficients are represented by the corresponding flag and encoded into the bitstream. At the decoding end, the decoder can parse the LFNST flag and the corresponding flag from the bitstream, first construct the quantized coefficients, and the quantized coefficients are inverse quantized to obtain the transform coefficients. If the LFNST flag indicates that NSPT is used, the decoder can perform NSPT inverse transform on the transform coefficients to obtain reconstructed residual information. If the LFNST flag indicates that NSPT is not used, the decoder can perform a basic inverse transform on the transform coefficients to obtain reconstructed residual information.
[0095] As can be seen from the previous text, for intra-coded frames, the luminance component and the chrominance component can use different partitioning trees. In this case, the luminance transform block can use its own LFNST identifier to indicate whether the luminance transform block uses an inseparable transform, and the chrominance transform block can also use its own LFNST identifier to indicate whether the chrominance transform block uses an inseparable transform. For inter-coded frames or coded frames that do not apply the DT partitioning method, since a coding unit includes both luminance blocks and chrominance blocks, in this case, the LFNST identifier only controls whether the luminance block uses an inseparable transform, while the chrominance block does not use an inseparable transform, but only uses the basic transform or the basic inverse transform. For example, if the LFNST identifier indicates that the coding unit uses an inseparable transform, it only means that the luminance block corresponding to the coding unit uses an inseparable transform, while the chrominance block corresponding to the coding unit does not use an inseparable transform.
[0096] Inseparable transforms are important transform operations. However, in some scenarios, encoding and decoding methods for these operations are sometimes inappropriate, potentially degrading encoding and decoding performance. The following example illustrates this issue.
[0097] In ECM-11.0, transform block coefficients can be represented by target parameters. Target parameters represent the absolute value of a transform block coefficient through accumulation and combination. Target parameters can include, for example, one or more of the following: sig_coeff_flag, abs_level_gtx_flag, par_level_flag, abs_remainder, and dec_abs_level.
[0098] sig_coeff_flag is a flag for context model-based coding and decoding, which is used to indicate whether the coded coefficient at the current position of the transform block is zero. If the coefficient of the current transform block is 0, sig_coeff_flag=1.
[0099] abs_level_gtx_flag is a context model-based encoding and decoding flag that indicates whether the absolute value of the coefficient encoded and decoded at the current position of the transform block is greater than x. x is a positive number. In ECM-x 11.0, x can be 1 or 3. For example, if x = 1, if the coefficient of the transform block is greater than 1, then abs_level_gt1_flag = 1. For example, if N = 3, if the coefficient of the transform block is greater than 3, then abs_level_gt3_flag = 1.
[0100] par_level_flag is a flag for context model-based coding and decoding, used to indicate whether the coefficient coded and decoded at the current position of the transform block is odd or even. If the coefficient of the transform block is odd, par_level_flag = 1.
[0101] abs_remainder is a remainder syntax element, bypass-coded using Golombre coding, and is used to represent the remainder of the absolute value of the coefficient encoded or decoded at the current position of the current transform block. For example, if the absolute value of the level of the transform block coefficient is greater than the second value, the fourth parameter may be used to represent the absolute value of the remaining quantized transform block coefficient.
[0102] dec_abs_level is the absolute value of the current coefficient, and bypass coding is performed using the Golombirai code. During the encoding and decoding process of the current block coefficient, when the number of identifiers encoded based on the context model exceeds the threshold, the remaining coefficients will no longer be represented by sig_coeff_flag, abs_level_gtx_flag, par_level_flag and abs_remainder, but will directly use the Golombirai code to encode and decode the absolute value of the coefficient.
[0103] At the encoding end, the encoder can transform and quantize the residual information to obtain the coefficients of the transform block, and further encode the target parameters according to the coefficients of the transform block, so that the encoded target parameters can represent the coefficients of the transform block.
[0104] At the decoding end, the decoder can obtain target parameters from the bitstream, determine the coefficients of the transform block based on the values of the target parameters, and dequantize and inverse transform the coefficients of the transform block to obtain residual information.
[0105] When encoding and decoding target parameters, context model index and / or Rice parameters are required. The context model index can be used to select the context model, and the Rice parameters can be the Rice parameters used when using Columbus coding when encoding and decoding abs_remainder and dec_abs_level.
[0106] The context model index and / or Rice parameter can be determined based on the adjacent information of the transform block. Which specific adjacent information is used is related to whether the transform block uses an inseparable transform. If the transform block uses an inseparable transform, the adjacent information is the adjacent information of the transform block in the scanning order; if the transform block does not use an inseparable transform, the adjacent information is the adjacent information of the spatial position of the transform block. For example, if the transform block uses an inseparable transform, the context model index and / or Rice parameter used for encoding and decoding the above parameters can be determined based on the adjacent information in the scanning order. If the transform block does not use an inseparable transform, the context model index and / or Rice parameter used for encoding and decoding the above parameters can be determined based on the adjacent information of the spatial position.
[0107] The embodiment of the present application does not specifically limit the amount of adjacent information. For example, the adjacent information may include information about five adjacent positions. Of course, the adjacent information may also include information about other numbers of adjacent positions.
[0108] The above-mentioned adjacent information may refer to information on a position (or transform block) adjacent to the current transform block, and the information on the adjacent position may include a decoded value or a partially decoded value on the adjacent position. When determining the context model index and / or Rice parameter, the context model index and / or Rice parameter may be determined based on the decoded value or the partially decoded value on the adjacent position. For example, the decoded value or the partially decoded value on the adjacent position may be accumulated, and based on the accumulated value, the context model index and / or Rice parameter may be determined.
[0109] Below, in combination with Figure 4, taking the adjacent information including information of five adjacent positions as an example, the method of determining the context model index and / or Rice parameter is introduced.
[0110] Referring to Figure 4, the left side of Figure 4 shows five adjacent positions in the spatial domain. If the transform block does not use an inseparable transform, the decoded values or partial decoded values at these five positions can be accumulated, and the accumulated values can be used as a basis to derive the context model index and / or Rice parameters. The right side of Figure 4 shows five adjacent positions in the scanning order. If the transform block uses an inseparable transform, the decoded values or partial decoded values at these five positions can be accumulated, and the accumulated values can be used as a basis to derive the context model index and / or Rice parameters.
[0111] In some scenarios (such as ECM-11.0), the type of neighboring information used by the current transform block is determined based on the LFNST flag (cu.lfnstIdx). If the LFNST flag indicates that an inseparable transform is used, the neighboring information used by the current transform block is the neighboring information in the scan order; if the LFNST flag indicates that an inseparable transform is not used, the neighboring information used by the current transform block is the neighboring information in the spatial domain. This judgment method is applicable to using different partitioning trees for luminance and chrominance, but there will be some problems if different partitioning trees are not used for luminance and chrominance. The following analyzes this problem.
[0112] Since the LFNST identifier is a coding unit-level identifier, the luminance block and the chrominance block share a coding unit. For inter-frame coded frames or coded frames that do not apply DT division, if the LFNST identifier indicates that an inseparable transform is used, both the luminance block and the chrominance block will use the adjacent information in the scanning order as the adjacent information. However, as mentioned above, for inter-frame coded frames or coded frames that do not apply DT division, if the LFNST identifier indicates that an inseparable transform is used, the LFNST identifier only indicates that an inseparable transform is used for the luminance block, and the chrominance block does not use an inseparable transform. In theory, the chrominance block should use the adjacent information of the spatial position as the adjacent information. Therefore, if the LFNST identifier is used as a condition for determining which adjacent information to use, the chrominance block will use unreasonable information to determine the coefficients of the transform block, thereby reducing the transform performance.
[0113] Similarly, at the encoding end, the chroma block will use unreasonable information to encode the coefficients of the transform block, thereby reducing the transform performance.
[0114] The above is merely an example to illustrate an inappropriate method of determining transform block coefficients, and the embodiments of the present application are not limited thereto.
[0115] In response to the above problem, an embodiment of the present application provides a coding method, including: determining first information, where the first information is used to indicate whether the current transform block uses an inseparable transform; determining adjacent information of the current transform block based on the first information; and determining coefficients of the current transform block based on the adjacent information.
[0116] In addition, an embodiment of the present application also provides a decoding method, including: determining first information, where the first information is used to indicate whether the current transform block uses an inseparable transform; determining adjacent information of the current transform block based on the first information; and encoding the coefficients of the current transform block based on the adjacent information.
[0117] The embodiment of the present application introduces transform block-level identification information (i.e., the first information mentioned above) to indicate whether to use an inseparable transform. Compared with the coding block-level identification information in the related art, the transform block-level identification information helps to select appropriate adjacent information, thereby helping to improve the encoding and decoding performance.
[0118] The following first describes the decoding method of the embodiment of the present application in detail with examples.
[0119] Figure 5 is a flowchart of a decoding method provided by an embodiment of the present application. The method of Figure 5 can be applied to a decoder.
[0120] Referring to FIG. 5 , in step S510, first information is determined, where the first information indicates whether a non-separable transform is used for the current transform block. The current transform block may refer to the current transform block to be decoded. In some implementations, the current transform block is a luma block. In other implementations, the current transform block may also be a chroma block. The current transform block may be an inter-frame prediction block or an intra-frame prediction block.
[0121] The non-separable transform may refer to the LFNST or NSPT described above.
[0122] There are multiple ways to determine the first information, and the following will introduce the ways to determine the first information in detail.
[0123] 5 , in step S520 , the neighboring information of the current transform block is determined based on the first information. The neighboring information may be, for example, the neighboring information of the spatial position of the current transform block or the neighboring information of the current transform block in the scanning order.
[0124] In some implementations, if the first information indicates that the current transform block uses a non-separable transform, the neighboring information of the current transform block is neighboring information of the current transform block in a scanning order. In other implementations, if the first information indicates that the current transform block does not use a non-separable transform, the neighboring information of the current transform block is neighboring information of the spatial position of the current transform block.
[0125] Continuing to refer to FIG. 5 , in step S530 , coefficients of the current transform block are determined based on the neighboring information.
[0126] The embodiment of the present application determines the coefficients of the transform block according to appropriate adjacent information, thereby obtaining appropriate transform block coefficients, which is beneficial to improving transform performance.
[0127] The embodiment of the present application does not specifically limit the determination method of step S530. In some implementations, the coefficients of the current transform block can be directly obtained based on the neighboring information. In other implementations, the context model index and / or Rice parameters can be determined based on the neighboring information; and then the coefficients of the current transform block can be determined based on the context model index and / or Rice parameters.
[0128] The coefficients of the current transform block can be represented by target parameters. These target parameters can be parsed from the bitstream and can be superimposed and / or combined to represent the coefficients of the current transform block. The target parameters may include one or more of the following parameters: a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter. These five parameters can be superimposed and combined to represent the coefficients of the current transform block. These five parameters are explained below.
[0129] The first parameter may be used to indicate whether the coefficients of the current transform block are zero. The first parameter may also be referred to as a non-zero flag. The first parameter may be represented by sig_coeff_flag (of course, the first parameter may also be represented by any other letters and / or numbers). If the coefficients of the current transform block are zero, the value of the first parameter may be 1; if the coefficients of the current transform block are non-zero, the value of the first parameter may be 0.
[0130] The second parameter can be used to indicate whether the absolute value of the coefficients of the current transform block is greater than the first value. The second parameter can be represented by abs_level_gtx_flag (of course, the second parameter can also be represented by any other letters and / or numbers), and the first value can be x. The value of x can be 1, 2, 3, etc. If the absolute value of the coefficients of the current transform block is greater than the first value, the value of the second parameter can be 1; if the absolute value of the coefficients of the current transform block is not greater than the first value, the value of the second parameter can be 0.
[0131] The third parameter may be used to indicate the parity of the coefficients of the current transform block. The third parameter may be represented by par_level_flag (of course, the third parameter may also be represented by any other letters and / or numbers). If the coefficients of the current transform block are odd, the value of the third parameter may be 1; if the coefficients of the current transform block are even, the value of the third parameter may be 0.
[0132] The fourth parameter may be used to represent the remainder of the coefficients of the current transform block. For example, if the absolute value of the level of the transform block coefficient is greater than the second value, the fourth parameter may be used to represent the remaining transform block coefficients. The fourth parameter may be represented by abs_remainder (of course, the fourth parameter may also be represented by any other letters and / or numbers).
[0133] The fifth parameter may be used to represent the absolute value of the coefficient of the current transform block. When the context model-based identifier used for encoding and decoding the current block exceeds a threshold, the absolute value of the coefficient may be represented by directly encoding and decoding the fifth parameter. The fifth parameter may be represented by dec_abs_level (of course, the fifth parameter may also be represented by any other letters and / or numbers).
[0134] Determining the coefficients of the current transform block according to the context model index and / or Rice parameter may refer to decoding the target parameters in the code stream according to the context model index and / or Rice parameter, and determining the coefficients of the current transform block according to the target parameters. After obtaining the target parameters, the target parameters may be superimposed and / or combined to obtain the coefficients of the transform block.
[0135] The following describes a method for determining the first information.
[0136] The first information may be determined based on one or more of the following information: the second information; whether the current transform block uses a transform skip mode; whether different partitioning trees are used for the luma component and the chroma component; the type of the current transform block; and the type of the current coding block. The above information is explained below.
[0137] The second information is used to indicate whether a non-separable transform is used, and the second information is information corresponding to the current coding block. In other words, the second information is indication information at the coding block (or coding unit) level. In some implementations, the second information can be represented by cu.lfnstIdx or lfnstIdx.
[0138] If the current transform block uses the transform skip mode, it means that the current transform block does not undergo a transform operation, but instead directly undergoes a quantization operation. For example, the residual information can be directly quantized to obtain the coefficients of the transform block. Therefore, if the current transform block uses the transform skip mode, it means that the current transform block does not use a non-separable transform.
[0139] If the luma component and the chroma component use different partitioning trees, the luma component and the chroma component have their own independent coding units; if the luma component and the chroma component do not use different partitioning trees, that is, the luma component and the chroma component use the same partitioning tree, the luma component and the chroma component share a coding unit. Whether the luma component and the chroma component use different partitioning trees can be determined by whether the luma component and the chroma component use DT partitioning. If the luma component and the chroma component use DT partitioning, the luma component and the chroma component use different partitioning trees; if the luma component and the chroma component do not use DT partitioning, the luma component and the chroma component do not use different partitioning trees.
[0140] The type of the current transform block may include a chroma transform block and a luma transform block.
[0141] The type of the current coding block may include one or more of the following: an intra-frame coding block, an inter-frame coding block, and an intra-block copy (IBC) coding block. In traditional schemes, the non-separable transform is only applied to intra-frame coding blocks, while the scheme of the embodiments of the present application can be extended to inter-frame coding blocks and / or IBC coding blocks.
[0142] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode. For example, if the current transform block uses a transform skip mode, the first information indicates that the current transform block does not use an inseparable transform.
[0143] In some implementations, the first information may be determined based on the second information. For example, if the second information indicates that the non-separable transform is not used, the first information indicates that the current transform block does not use the non-separable transform.
[0144] In some implementations, the first information may be determined based on the type of the current coding block. For example, if the current coding block does not belong to a preset type of coding block, the first information indicates that the current coding block does not use an inseparable transform.
[0145] In some embodiments, a preset type of coding block may refer to a coding block that can use an inseparable transform. The preset type of coding block may include one or more of the following: an intra-frame coding block, an inter-frame coding block, and an IBC coding block. Taking the example that the preset type of coding block includes an intra-frame coding block, but does not include an inter-frame coding block and an IBC coding block, if the current coding block does not belong to an intra-frame coding block, that is, the current coding block belongs to an inter-frame coding block or an IBC coding block, the first information indicates that the current coding block does not use an inseparable transform. Taking the example that the preset type of coding block includes an intra-frame coding block and an inter-frame coding block, if the current coding block belongs to an IBC coding block, the first information indicates that the current coding block does not use an inseparable transform. Taking the example that the preset type of coding block includes an intra-frame coding block and an IBC coding block, if the current coding block belongs to an inter-frame coding block, the first information indicates that the current coding block does not use an inseparable transform.
[0146] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode, the second information, and whether the luma component and the chroma component use different partitioning trees. For example, if the current transform block does not use a transform skip mode and the second information indicates that a non-separable transform mode is used, then if a first condition is satisfied, the first information indicates that a non-separable transform is used in the current transform block, where the first condition includes that the luma component and the chroma component use different partitioning trees. In other words, if the current transform block does not use a transform skip mode, the second information indicates that a non-separable transform is used, and the luma component and the chroma component use different partitioning trees, then the first information indicates that a non-separable transform is used in the current transform block.
[0147] In some implementations, the first information may be determined based on the type of the current coding block, whether the current transform block uses a transform skip mode, the second information, and whether different partitioning trees are used for the luminance component and the chrominance component. For example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, then if the first condition is satisfied, then the first information indicates that an inseparable transform is used for the current transform block, wherein the first condition includes that different partitioning trees are used for the luminance component and the chrominance component. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, and different partitioning trees are used for the luminance component and the chrominance component, then the first information indicates that an inseparable transform is used for the current transform block.
[0148] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode, the second information, whether the luma component and the chroma component use different partitioning trees, and the type of the current transform block. For example, if the current transform block does not use a transform skip mode and the second information indicates that a non-separable transform mode is used, then if a first condition is satisfied, the first information indicates that the current transform block uses a non-separable transform, wherein the first condition includes that the luma component and the chroma component do not use different partitioning trees and the current transform block is a luma block. In other words, if the current transform block does not use a transform skip mode, the second information indicates that a non-separable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block, then the first information indicates that the current transform block (i.e., the luma transform block) uses a non-separable transform.
[0149] For another example, if the current transform block does not use the transform skip mode and the second information indicates that the non-separable transform mode is used, if the second condition is satisfied, the first information indicates that the current transform block does not use the non-separable transform. The second condition includes that the luma component and the chroma component do not use different partitioning trees, and the current transform block is a chroma transform block. In other words, if the current transform block does not use the transform skip mode, the second information indicates that the non-separable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a chroma transform block, the first information indicates that the current transform block (i.e., the chroma transform block) does not use the non-separable transform.
[0150] In some implementations, the first information may be determined based on the type of the current coding block, whether the current transform block uses a transform skip mode, the second information, whether the luma component and the chroma component use different partitioning trees, and the type of the current transform block. For example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, then if the first condition is satisfied, then the first information indicates that the current transform block uses an inseparable transform, wherein the first condition includes that the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block, then the first information indicates that the current transform block (i.e., the luma transform block) uses an inseparable transform.
[0151] For another example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if the second condition is satisfied, the first information indicates that the current transform block does not use an inseparable transform. The second condition includes that the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance transform block. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance transform block, the first information indicates that the current transform block (i.e., the chrominance transform block) does not use an inseparable transform.
[0152] By determining the first information in the above manner, it is possible to accurately determine whether the current transform block uses an inseparable transform, thereby determining appropriate transform block coefficients, which is beneficial to improving transform performance.
[0153] In conventional solutions, the non-separable transform is only applied to intra-coded blocks, while the embodiments of the present application can be applied to one or more of intra-coded blocks, inter-coded blocks, and IBC-coded blocks.
[0154] Before parsing the second information, the type of the current transform block may be determined. If the third condition is met, the second information is parsed; if the third condition is not met, the second information may not be parsed.
[0155] The third condition may include that the current coding block belongs to a preset type of coding block. The preset type of coding block is a coding block that can use an inseparable transform. The preset type of coding block may include one or more of the following: intra-frame coding block, inter-frame coding block, and IBC coding block.
[0156] That is, if the current coding block belongs to a preset type of coding block, the second information may be parsed; if the current coding block does not belong to the preset type of coding block, the second information may not be parsed.
[0157] Taking the preset type of coding block including the intra-frame coding block as an example, if the preset type of coding block includes the intra-frame coding block, the decoder may parse the second information when the current coding block is the intra-frame coding block.
[0158] Taking the preset type of coding block including the inter-frame coding block as an example, if the preset type of coding block includes the inter-frame coding block, the decoder may parse the second information when the current coding block is the inter-frame coding block.
[0159] Taking the preset type of coding block including the IBC coding block as an example, if the preset type of coding block includes the IBC coding block, the decoder may parse the second information when the current coding block is the IBC coding block.
[0160] If the decoder does not parse the second information, it may be assumed that the current transform block does not use the inseparable transform.
[0161] After obtaining the coefficients of the transform block, residual information can be determined based on the coefficients of the current transform block. For example, the coefficients of the current transform block can be dequantized and inverse transformed to obtain residual information. The inverse transform method used depends on whether the current transform block uses a non-separable transform. If the current transform block uses a non-separable transform, the inverse transform method includes a non-separable inverse transform. If the current transform block does not use a non-separable transform, the inverse transform method includes a basic inverse transform and does not include a non-separable inverse transform.
[0162] After the residual information is obtained, reconstruction information can be determined based on the residual information.
[0163] In some implementations, residual information may be determined based on a transform kernel and coefficients of a current transform block. For example, the coefficients of the current transform block may be dequantized to obtain transformed coefficients; and then the transformed coefficients may be dequantized using the transform kernel to obtain residual information.
[0164] In some implementations, if the current transform block uses a non-separable transform, or if the first information indicates that the current transform block uses a non-separable transform, the transform kernel may be determined based on the angular information of the current transform block. If the current transform block is an intra-frame transform block, the transform kernel may be determined based on the intra-frame angular direction. If the current transform block is an inter-frame transform block or an IBC transform block, the angular information of the current transform block may be determined based on a gradient histogram. The specific determination method is described below.
[0165] Of course, in addition to using the gradient histogram to determine the angle information of the transform block, embodiments of the present application may also use other methods to determine the angle information of the transform block. For example, a list of candidate intra-frame prediction modes may be obtained from around the current transform block through template matching, and an angle may be determined by sorting the template costs.
[0166] The decoding method provided by the embodiment of the present application is described in detail above in conjunction with Figure 5. The encoding method provided by the embodiment of the present application is described in detail below in conjunction with Figure 6.
[0167] Figure 6 is a flow chart of an encoding method provided by an embodiment of the present application. The method of Figure 6 can be applied to an encoder.
[0168] Referring to FIG. 6 , in step S610 , first information is determined. The first information indicates whether a non-separable transform is used for the current transform block. The current transform block may refer to the current transform block to be decoded. In some implementations, the current transform block is a luma block. In other implementations, the current transform block may also be a chroma block. The current transform block may be an inter-frame prediction block or an intra-frame prediction block.
[0169] The non-separable transform can be LFNST or NSPT.
[0170] There are multiple ways to determine the first information, and the following will introduce the ways to determine the first information in detail.
[0171] 6 , in step S620 , the neighboring information of the current transform block is determined based on the first information. The neighboring information may be, for example, the neighboring information of the spatial position of the current transform block or the neighboring information of the current transform block in the scanning order.
[0172] In some implementations, if the first information indicates that the current transform block uses a non-separable transform, the neighboring information of the current transform block is neighboring information of the current transform block in a scanning order. In other implementations, if the first information indicates that the current transform block does not use a non-separable transform, the neighboring information of the current transform block is neighboring information of the spatial position of the current transform block.
[0173] Continuing to refer to FIG. 6 , in step S630 , the coefficients of the current transform block are encoded according to the neighboring information.
[0174] The embodiment of the present application encodes the coefficients of the current transform block according to appropriate neighboring information, so that the coefficients of the transform block can be encoded in an appropriate manner, which is beneficial to improving transform performance.
[0175] The embodiment of the present application does not specifically limit the encoding method of step S630. In some implementations, the coefficients of the current transform block can be directly encoded based on the adjacent information. In other implementations, the context model index and / or Rice parameter can be determined based on the adjacent information; then, the coefficients of the current transform block can be encoded based on the context model index and / or Rice parameter.
[0176] The coefficients of the current transform block can be represented by target parameters, which can be superimposed and / or combined to represent the coefficients of the current transform block. The target parameters may include one or more of the following parameters: a first parameter, a second parameter, a third parameter, a fourth parameter, and a fifth parameter. These five parameters can be superimposed and combined to represent the coefficients of the current transform block. These five parameters are explained below.
[0177] The first parameter may be used to indicate whether the coefficients of the current transform block are zero. The first parameter may also be referred to as a non-zero flag. The first parameter may be represented by sig_coeff_flag (of course, the first parameter may also be represented by any other letters and / or numbers). If the coefficients of the current transform block are zero, the value of the first parameter may be 1; if the coefficients of the current transform block are non-zero, the value of the first parameter may be 0.
[0178] The second parameter can be used to indicate whether the absolute value of the coefficients of the current transform block is greater than the first value. The second parameter can be represented by abs_level_gtx_flag (of course, the second parameter can also be represented by any other letters and / or numbers), and the first value can be x. The value of x can be 1, 2, 3, etc. If the absolute value of the coefficients of the current transform block is greater than the first value, the value of the second parameter can be 1; if the absolute value of the coefficients of the current transform block is not greater than the first value, the value of the second parameter can be 0.
[0179] The third parameter may be used to indicate the parity of the coefficients of the current transform block. The third parameter may be represented by par_level_flag (of course, the third parameter may also be represented by any other letters and / or numbers). If the coefficients of the current transform block are odd, the value of the third parameter may be 1; if the coefficients of the current transform block are even, the value of the third parameter may be 0.
[0180] The fourth parameter may be used to represent the remainder of the coefficients of the current transform block. For example, if the absolute value of the level of the transform block coefficient is greater than the second value, the fourth parameter may be used to represent the remaining transform block coefficients. The fourth parameter may be represented by abs_remainder (of course, the fourth parameter may also be represented by any other letters and / or numbers).
[0181] The fifth parameter may be used to represent the absolute value of the coefficient of the current transform block. When the context model-based identifier used for encoding and decoding the current block exceeds a threshold, the absolute value of the coefficient may be represented by directly encoding and decoding the fifth parameter. The fifth parameter may be represented by dec_abs_level (of course, the fifth parameter may also be represented by any other letters and / or numbers).
[0182] In some implementations, encoding the coefficients of the current transform block may refer to encoding the coefficients of the current transform block into target parameters.
[0183] In some implementations, target parameters can be written into the bitstream based on the context model index and / or Rice parameters, and the target parameters can be used to represent the coefficients of the current transform block.
[0184] The following describes a method for determining the first information.
[0185] The first information may be determined based on one or more of the following information: the second information; whether the current transform block uses a transform skip mode; whether different partitioning trees are used for the luma component and the chroma component; the type of the current transform block; and the type of the current coding block. The above information is explained below.
[0186] The second information is used to indicate whether a non-separable transform is used, and the second information is information corresponding to the current coding block. In other words, the second information is indication information at the coding block (or coding unit) level. In some implementations, the second information can be represented by cu.lfnstIdx or lfnstIdx.
[0187] If the current transform block uses the transform skip mode, it means that the current transform block does not undergo a transform operation, but instead directly undergoes a quantization operation. For example, the residual information can be directly quantized to obtain the coefficients of the transform block. Therefore, if the current transform block uses the transform skip mode, it means that the current transform block does not use a non-separable transform.
[0188] If the luma component and the chroma component use different partitioning trees, the luma component and the chroma component have their own independent coding units; if the luma component and the chroma component do not use different partitioning trees, that is, the luma component and the chroma component use the same partitioning tree, the luma component and the chroma component share a coding unit. Whether the luma component and the chroma component use different partitioning trees can be determined by whether the luma component and the chroma component use DT partitioning. If the luma component and the chroma component use DT partitioning, the luma component and the chroma component use different partitioning trees; if the luma component and the chroma component do not use DT partitioning, the luma component and the chroma component do not use different partitioning trees.
[0189] The type of the current transform block may include a chroma transform block and a luma transform block.
[0190] The type of the current coding block may include one or more of the following: an intra-frame coding block, an inter-frame coding block, and an intra-block copy (IBC) coding block. In traditional schemes, the non-separable transform is only applied to intra-frame coding blocks, while the scheme of the embodiments of the present application can be extended to inter-frame coding blocks and / or IBC coding blocks.
[0191] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode. For example, if the current transform block uses a transform skip mode, the first information indicates that the current transform block does not use an inseparable transform.
[0192] In some implementations, the first information may be determined based on the second information. For example, if the second information indicates that the non-separable transform is not used, the first information indicates that the current transform block does not use the non-separable transform.
[0193] In some implementations, the first information may be determined based on the type of the current coding block. For example, if the current coding block does not belong to a preset type of coding block, the first information indicates that the current coding block does not use an inseparable transform.
[0194] In some embodiments, a preset type of coding block may refer to a coding block that can use an inseparable transform. The preset type of coding block may include one or more of the following: an intra-frame coding block, an inter-frame coding block, and an IBC coding block. Taking the example that the preset type of coding block includes an intra-frame coding block, but does not include an inter-frame coding block and an IBC coding block, if the current coding block does not belong to an intra-frame coding block, that is, the current coding block belongs to an inter-frame coding block or an IBC coding block, the first information indicates that the current coding block does not use an inseparable transform. Taking the example that the preset type of coding block includes an intra-frame coding block and an inter-frame coding block, if the current coding block belongs to an IBC coding block, the first information indicates that the current coding block does not use an inseparable transform. Taking the example that the preset type of coding block includes an intra-frame coding block and an IBC coding block, if the current coding block belongs to an inter-frame coding block, the first information indicates that the current coding block does not use an inseparable transform.
[0195] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode, the second information, and whether the luma component and the chroma component use different partitioning trees. For example, if the current transform block does not use a transform skip mode and the second information indicates that a non-separable transform mode is used, then if a first condition is satisfied, the first information indicates that a non-separable transform is used in the current transform block, where the first condition includes that the luma component and the chroma component use different partitioning trees. In other words, if the current transform block does not use a transform skip mode, the second information indicates that a non-separable transform is used, and the luma component and the chroma component use different partitioning trees, then the first information indicates that a non-separable transform is used in the current transform block.
[0196] In some implementations, the first information may be determined based on the type of the current coding block, whether the current transform block uses a transform skip mode, the second information, and whether different partitioning trees are used for the luminance component and the chrominance component. For example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, then if the first condition is satisfied, then the first information indicates that an inseparable transform is used for the current transform block, wherein the first condition includes that different partitioning trees are used for the luminance component and the chrominance component. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, and different partitioning trees are used for the luminance component and the chrominance component, then the first information indicates that an inseparable transform is used for the current transform block.
[0197] In some implementations, the first information may be determined based on whether the current transform block uses a transform skip mode, the second information, whether the luma component and the chroma component use different partitioning trees, and the type of the current transform block. For example, if the current transform block does not use a transform skip mode and the second information indicates that a non-separable transform mode is used, then if a first condition is satisfied, the first information indicates that the current transform block uses a non-separable transform, wherein the first condition includes that the luma component and the chroma component do not use different partitioning trees and the current transform block is a luma block. In other words, if the current transform block does not use a transform skip mode, the second information indicates that a non-separable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block, then the first information indicates that the current transform block (i.e., the luma transform block) uses a non-separable transform.
[0198] For another example, if the current transform block does not use the transform skip mode and the second information indicates that the non-separable transform mode is used, if the second condition is satisfied, the first information indicates that the current transform block does not use the non-separable transform. The second condition includes that the luma component and the chroma component do not use different partitioning trees, and the current transform block is a chroma transform block. In other words, if the current transform block does not use the transform skip mode, the second information indicates that the non-separable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a chroma transform block, the first information indicates that the current transform block (i.e., the chroma transform block) does not use the non-separable transform.
[0199] In some implementations, the first information may be determined based on the type of the current coding block, whether the current transform block uses a transform skip mode, the second information, whether the luma component and the chroma component use different partitioning trees, and the type of the current transform block. For example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, then if the first condition is satisfied, then the first information indicates that the current transform block uses an inseparable transform, wherein the first condition includes that the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, the luma component and the chroma component do not use different partitioning trees, and the current transform block is a luma block, then the first information indicates that the current transform block (i.e., the luma transform block) uses an inseparable transform.
[0200] For another example, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if the second condition is satisfied, the first information indicates that the current transform block does not use an inseparable transform. The second condition includes that the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance transform block. In other words, if the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, the second information indicates that an inseparable transform is used, the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance transform block, the first information indicates that the current transform block (i.e., the chrominance transform block) does not use an inseparable transform.
[0201] By determining the first information in the above manner, it is possible to accurately determine whether the current transform block uses an inseparable transform, thereby determining appropriate transform block coefficients, which is beneficial to improving transform performance.
[0202] In conventional solutions, the non-separable transform is only applied to intra-coded blocks, while the embodiments of the present application can be applied to one or more of intra-coded blocks, inter-coded blocks, and IBC-coded blocks.
[0203] Before encoding the second information, the type of the current transform block may be determined. If the third condition is met, the second information is encoded; if the third condition is not met, the second information may not be encoded.
[0204] The third condition may include that the current coding block belongs to a preset type of coding block. The preset type of coding block is a coding block that can use an inseparable transform. The preset type of coding block may include one or more of the following: intra-frame coding block, inter-frame coding block, and IBC coding block.
[0205] That is, if the current coding block belongs to a preset type of coding block, the second information may be encoded; if the current coding block does not belong to the preset type of coding block, the second information may not be encoded.
[0206] Taking the preset type of coding block including the intra-frame coding block as an example, if the preset type of coding block includes the intra-frame coding block, the encoder may encode the second information when the current coding block is the intra-frame coding block.
[0207] Taking the preset type of coding block including the inter-frame coding block as an example, if the preset type of coding block includes the inter-frame coding block, the encoder may encode the second information when the current coding block is the inter-frame coding block.
[0208] Taking the preset type of coding block including the IBC coding block as an example, if the preset type of coding block includes the IBC coding block, the encoder may encode the second information when the current coding block is the IBC coding block.
[0209] If the encoder does not encode the second information, it may be assumed that the current transform block does not use the inseparable transform.
[0210] In some implementations, a transform may be performed on the current transform block to determine transform coefficients. The transform method may include a base transform or a non-separable transform. For example, the transform method may include a base transform and LFNST. Another example is that the transform method may include NSPT. After obtaining the transformed coefficients, the transformed coefficients may be quantized to determine the coefficients of the current transform block.
[0211] In some implementations, the current transform block may be transformed using a transform kernel to obtain transform coefficients.
[0212] In some implementations, if the current transform block uses a non-separable transform, or if the first information indicates that the current transform block uses a non-separable transform, the transform kernel may be determined based on the angular information of the current transform block. If the current transform block is an intra-frame transform block, the transform kernel may be determined based on the intra-frame angular direction. If the current transform block is an inter-frame transform block or an IBC transform block, the angular information of the current transform block may be determined based on a gradient histogram. The specific determination method is described below.
[0213] Of course, in addition to using the gradient histogram to determine the angle information of the transform block, embodiments of the present application may also use other methods to determine the angle information of the transform block. For example, a list of candidate intra-frame prediction modes may be obtained from around the current transform block through template matching, and an angle may be determined by sorting the template costs.
[0214] The following examples are used to describe the embodiments of the present application in more detail. It should be noted that the examples below are only intended to help those skilled in the art understand the embodiments of the present application, rather than to limit the embodiments of the present application to the specific numerical values or specific scenarios illustrated. It is apparent that those skilled in the art can make various equivalent modifications or changes based on the examples given below, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0215] In ECM, since the transform and inverse transform of LFNST and NSPT require the selection of the transform kernel based on the intra-frame angular direction, LFNST and NSPT are applied to blocks using intra-frame prediction mode. This means that the decoder can only further determine whether to decode lfnst_idx when a coding unit is an intra-frame type coding unit. The decoder parses the syntax element lfnst_idx at the coding unit level to determine whether there is a transform block in the current coding unit that requires LFNST or NSPT for inverse transform.
[0216] When the value of this syntax element is non-zero, the current coding unit contains transform blocks that need to be inverse transformed using LFNST or NSPT. When the value of this syntax element is zero, none of the transform blocks in the current coding unit use LFNST or NSPT for inverse transform. Since the transformation of LFNST and NSPT has angular dependence, in ECM, it is only used for transform blocks in coding units in intra-frame coding mode.
[0217] Regardless of whether luminance and chrominance are divided separately, the selection of the context model and the determination of the Rice parameters in the coefficient decoding process will be determined according to the following method. If lfnst_idx is 0, the information at the five positions on the left side of Figure 4 is used to derive the context model and Rice parameters. If lfnst_idx is 1, the information at the five positions on the right side of Figure 4 is used to derive the context model and Rice parameters.
[0218] Then, the quantized transform coefficients are constructed based on the parsed syntax elements. The quantized transform coefficients are inversely quantized to obtain transform coefficients. The transform coefficients are inversely transformed according to the selected transform mode (LFNST / NSPT, or other transform modes) to obtain residual coefficients. The obtained residual coefficients are added to the predicted values to obtain the reconstructed block.
[0219] In VCC, the steps for decoding the Coding unit syntax are as follows:
[0220] In transform_tree() in VCC, the syntax elements of each transform unit and the syntax elements related to the coefficient coding of the transform block are encoded and decoded. In VVC, lfnst_idx, which is used to represent the LFNST index, is encoded and decoded after transform_tree. In ECM (such as ECM-11.0), since the selection of the context model and the selection of Rice parameters in the encoding and decoding of the coefficients of the transform block need to be based on the value of lfnst_idx, the encoding and decoding content in transform_tree is divided into two parts. The first part is the syntax elements other than the coefficient values, and the second part is the syntax elements related to the coefficient values. Therefore, in ECM-11.0, the changes in the Coding unit syntax table are as follows.
[0221] When actually encoding and decoding coefficients, the lfnst_idx at the coding unit level can be used to determine the context model and Rice parameters.
[0222] When deriving Rice parameters, the process of calculating the sum of the absolute values of the coefficients at the five surrounding positions is as follows:
[0223] The process of exporting the necessary variables locNumSig and locSumAbsPass1 in sig_coeff_flag, abs_level_gtx_flag, and par_level_flag is as follows:
[0224] In ECM-11.0, lfnstidx was originally used to determine which neighboring information to use. However, as mentioned above, this approach can lead to the following situation: when the luma and chroma components do not apply different partitioning, even if lfnstidx is non-zero, the chroma transform block does not apply LFNST or NSPT. Therefore, in this embodiment, instead of using lfnstidx as a condition, the variable lfnstApplied is used to determine which neighboring information to use.
[0225] In ECM-11.0, the value of lfnstIdx is equal to the value of the coding unit syntax element lfnst_idx, and in this embodiment of the present application, the value of lfnstApplied can be obtained as follows:
[0226] lfnstApplied=(!transform_skip_flag||sh_ts_residual_coding_disabled_flag)&&lfnstIdx&&(treeType!=SINGLE_TREE?true:isLuma(compID))
[0227] lfnstApplied is used to determine whether the current transform block uses LFNST or NSPT as the transform mode, (! transform_skip_flag||sh_ts_residual_coding_disabled_flag) indicates that the current transform block does not use the transform skip mode, lfnstIdx indicates that the lfnst_idx of the current coding unit is non-zero,
[0228] (treeType!=SINGLE_TREE?true:isLuma(compID)) is used to obtain a Boolean variable (true or false). If treeType is a luma-chroma split, the Boolean variable true is obtained. If treeType is a luma-chroma split, it is necessary to further determine whether the color component of the current transform block is luma. If it is a luma component, the Boolean variable true is obtained; otherwise, the Boolean variable false is obtained.
[0229] Correspondingly, when deriving the Rice parameters, the process of calculating the sum of the absolute values of the coefficients of the five surrounding positions is as follows:
[0230] The process of exporting the necessary variables locNumSig and locSumAbsPass1 in sig_coeff_flag, abs_level_gtx_flag, and par_level_flag is as follows:
[0231] Figures 7 to 10 show several ways of determining neighbor information, which are described below. It should be noted that the solutions of Figures 7 to 10 are applicable to both the encoding end and the decoding end.
[0232] The solution shown in FIG. 7 is applicable to intra-frame coding blocks. That is, if the current coding block is an intra-frame coding block, the neighboring information can be determined in the manner shown in FIG. 7 .
[0233] Referring to Figure 7 , it is possible to determine whether the current coding block uses a luma and chroma splitting method. If the current coding block uses a luma and chroma splitting method, a conventional approach can be used. For example, neighbor information can be determined based on lfnst_idx. If the value of lfnst_idx is 0, the neighbor information is the neighbor information of the spatial position of the current transform block. If the value of lfnst_idx is not 0, the neighbor information is the neighbor information of the current transform block in the scanning order.
[0234] If the current coding block does not use the luma and chroma division method, the type of the current transform block can be further determined, such as whether the current transform block is a luma transform block. If the current transform block is a luma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order; if the current transform block is a chroma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order.
[0235] The solution shown in FIG8 is applicable to the case where the intra-frame coded blocks and the inter-frame coded blocks can use inseparable transforms.
[0236] Referring to Figure 8, it can be determined whether the current coding block is an intra-frame coding block. If the current coding block is an intra-frame coding block, it is further determined whether the current coding block adopts a method of dividing brightness and chrominance separately. If the current coding block adopts a method of dividing brightness and chrominance separately, the traditional approach can be used. For example, the adjacent information can be determined based on lfnst_idx. If the value of lfnst_idx is 0, the adjacent information is the adjacent information of the spatial position of the current transform block. If the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0237] If the current coding block does not use the luma and chroma division method, the type of the current transform block can be further determined, such as whether the current transform block is a luma transform block. If the current transform block is a luma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order; if the current transform block is a chroma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order.
[0238] If the current coding block is not an intra-frame coding block, it can be further determined whether the current coding block is an inter-frame coding block. If the current coding block is not an inter-frame coding block, the traditional approach can be used. If the current coding block is an inter-frame coding block, the type of the current transform block can be further determined, such as determining whether the current transform block is a luminance transform block. If the current transform block is a luminance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order; if the current transform block is a chrominance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0239] The solution shown in FIG9 is applicable to the case where the intra-coded blocks and the IBC-coded blocks can use inseparable transforms.
[0240] Referring to Figure 9, it can be determined whether the current coding block is an intra-frame coding block. If the current coding block is an intra-frame coding block, it is further determined whether the current coding block adopts a method of dividing brightness and chrominance separately. If the current coding block adopts a method of dividing brightness and chrominance separately, the traditional approach can be used. For example, the adjacent information can be determined based on lfnst_idx. If the value of lfnst_idx is 0, the adjacent information is the adjacent information of the spatial position of the current transform block. If the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0241] If the current coding block does not use the luma and chroma division method, the type of the current transform block can be further determined, such as whether the current transform block is a luma transform block. If the current transform block is a luma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order; if the current transform block is a chroma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order.
[0242] If the current coding block is not an intra-frame coding block, it can be further determined whether the current coding block is an IBC coding block. If the current coding block is not an IBC coding block, the traditional approach can be used. If the current coding block is an IBC coding block, the type of the current transform block can be further determined, such as determining whether the current transform block is a luminance transform block. If the current transform block is a luminance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order; if the current transform block is a chrominance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0243] The solution shown in FIG10 is applicable to the case where intra-frame coded blocks, inter-frame coded blocks and IBC coded blocks can use inseparable transforms.
[0244] Referring to Figure 10, it can be determined whether the current coding block is an intra-frame coding block or an IBC coding block. If the current coding block is an intra-frame coding block or an IBC coding block, it is further determined whether the current coding block adopts a method of dividing brightness and chrominance separately. If the current coding block adopts a method of dividing brightness and chrominance separately, the traditional approach can be used. For example, the adjacent information can be determined based on lfnst_idx. If the value of lfnst_idx is 0, the adjacent information is the adjacent information of the spatial position of the current transform block. If the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0245] If the current coding block does not use the luma and chroma division method, the type of the current transform block can be further determined, such as whether the current transform block is a luma transform block. If the current transform block is a luma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order; if the current transform block is a chroma transform block and the value of lfnst_idx is not 0, the neighboring information is the neighboring information of the current transform block in the scanning order.
[0246] If the current coding block is not an intra-frame coding block or an IBC coding block, it can be further determined whether the current coding block is an inter-frame coding block. If the current coding block is not an inter-frame coding block, the traditional approach can be used. If the current coding block is an inter-frame coding block, the type of the current transform block can be further determined, such as determining whether the current transform block is a luminance transform block. If the current transform block is a luminance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order; if the current transform block is a chrominance transform block and the value of lfnst_idx is not 0, the adjacent information is the adjacent information of the current transform block in the scanning order.
[0247] Determine the angle information of the current transform block
[0248] During inverse transformation, when the current transform block uses LFNST or NSPT, since LFNST and NSPT have the characteristics of needing to be derived together with the intra-frame prediction mode according to lfnst_idx, an angle can be derived using the gradient histogram. The gradient histogram angle calculation method is as follows.
[0249] As shown in Figure 11, the first step is to use a sliding 3x3 window to calculate the horizontal and vertical gradient values of each 3x3 window in the prediction block, g x and g y . g x and g y They are respectively composed of 3x3 horizontal gradient operators M x and the vertical gradient operator M y It is obtained by multiplying the predicted value within the window position.
[0250] Assuming that the prediction block is a block with a width and height of (w,h), the sliding 3x3 window can calculate the g of the (w-2)*(h-2) positions of the center of the interpolation filter prediction block. x and g y .
[0251] The second step is to calculate the g at each position. x and g y , calculate the traditional angle direction corresponding to each position according to the following formula
[0252] (angle), and calculate the amplitude value (Amp) of the gradient corresponding to the angle at each position.
[0253] Amp=|g x |+|g y | and
[0254] In some embodiments, the calculation process of atan can also be simplified, for example, by looking up a table or performing some transformations.
[0255] The third step is to accumulate the gradient amplitudes Amp at each position across the derived angle categories to create a gradient amplitude histogram. Finally, the angle with the largest cumulative amplitude is selected as the angle corresponding to the current predicted block. When the amplitudes derived from all angles are zero, the current block matches the category predicted by the traditional PLANAR model.
[0256] In some embodiments, instead of using the predicted value of the current transform block to construct the histogram-derived angle, the reconstructed values around the current transform block may be used to calculate the gradient histogram and derive the angle.
[0257] In some embodiments, a candidate intra prediction mode list may be obtained from around the current transform block by other template matching methods, and an angle for selecting a transform kernel of LFNST or NSPT may be determined by template cost sorting.
[0258] Non-separable transform applied to inter-coded blocks
[0259] In some embodiments, LFNST / NSPT can also be used in the residual transform process of the inter-frame coded transform block. Since the inter-frame coded block only exists in the inter-frame coded frame, and the inter-frame coded frame uses the same luminance and chrominance division by default, the context model selection in the coefficient encoding and decoding process on the inter-frame coded block is similar to the scheme described above for the Rice parameter. In these embodiments, as with the intra-frame prediction type coding unit, LFNST / NSPT is only used to transform the luminance transform block, not the chrominance transform block. The specific flowchart is shown in Figure 8.
[0260] The process for using LFNST / NSPT on inter-coded blocks is as follows. At the decoder, for inter-coded coding units, lfnst_idx is parsed to determine whether the coding unit contains a transform block using LFNST / NSPT. Unlike traditional methods, lfnst_idx parsing requires whether the current coding unit type is inter-frame prediction as a condition. An example of the lfnst_idx parsing process based on the VVC standard text is as follows.
[0261] CuPredMode[chType][x0][y0]==MODE_INTRA is used to determine whether the current coding block is an intra-frame coding block, and CuPredMode[chType][x0][y0]==MODE_INTER) is used to determine whether the current coding block is an inter-frame coding block.
[0262] The decoder can derive the context model and Rice parameters and parse the coefficients according to the method in Figure 8.
[0263] Inseparable transformation applied to IBC coded blocks
[0264] In some embodiments, LFNST / NSPT can also be extended to coding units of type IBC. When the IBC type coding unit is in a frame with the same luminance and chrominance divisions, the context model index and Rice parameters in the coefficient encoding and decoding process can also be determined in the same way as described above. The specific process is similar to the processing method of using inter-frame type coding units.
[0265] Similar to inter-frame type coding units, when parsing lfnst_idx, it is necessary to take whether the coding unit type is IBC as one of the conditions. An example of the parsing process of lfnst_idx based on the VVC standard text is as follows:
[0266] CuPredMode[chType][x0][y0]==MODE_INTRA is used to determine whether the current coding block is an intra-frame coding block, and CuPredMode[chType][x0][y0]==MODE_IBC is used to determine whether the current coding block is an IBC coding block.
[0267] When LFNST / NSPT is only applied to IBC coding units and not to inter-frame coding units, the execution flow is shown in Figure 9. When LFNST / NSPT can be applied to both IBC coding units and inter-frame coding units, the execution flow is shown in Figure 10.
[0268] Tables 2 and 3 show the test results of the embodiments of the present application. Table 2 shows the test results of the intra-frame coding unit after using the solution of the embodiment of the present application, and Table 3 shows the test results of the inter-frame coding unit after using the solution of the embodiment of the present application.
[0269] Table 2
[0270] Table 3
[0271] It can be seen from Table 2 and Table 3 that after the solution of the embodiment of the present application is applied to the inter-frame coding block, a certain improvement in the encoding and decoding performance of the chrominance components U and V can be achieved.
[0272] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 11. The device embodiment of the present application is described in detail below in conjunction with Figures 12 to 15. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0273] Figure 12 is a schematic diagram of the structure of a decoder provided by one embodiment of the present application. As shown in Figure 12, the decoder 1200 includes a first determination unit 1210, a second determination unit 1220, and a third determination unit 1230. The first determination unit 1210 is configured to determine first information indicating whether a non-separable transform is used for the current transform block; the second determination unit 1220 is configured to determine neighboring information of the current transform block based on the first information; and the third determination unit 1230 is configured to determine coefficients of the current transform block based on the neighboring information.
[0274] In some implementations, the first information indicates that the current transform block uses an inseparable transform, and the neighboring information is neighboring information of the current transform block in a scanning order.
[0275] In some implementations, the first information indicates that the current transform block does not use an inseparable transform, and the neighboring information is neighboring information of a spatial position of the current transform block.
[0276] In some implementations, the third determination unit 1230 is configured to determine the context model index and / or Rice parameter based on the neighboring information; and determine the coefficients of the current transform block based on the context model index and / or Rice parameter.
[0277] In some implementations, the third determination unit 1230 is configured to decode the target parameters in the code stream based on the context model index and / or Rice parameters, where the target parameters are used to represent the coefficients of the current block; and determine the coefficients of the current transform block based on the target parameters.
[0278] In some implementations, the first information is determined based on one or more of the following: second information, used to indicate whether an inseparable transform is used, and the second information is information corresponding to the current coding block; whether the current transform block uses a transform skip mode; whether the luminance component and the chrominance component use different partitioning trees; the type of the current transform block; and the type of the current coding block.
[0279] In some implementations, when the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a first condition is met, then the first information indicates that the current transform block uses an inseparable transform, and the first condition includes one of the following: the luminance component and the chrominance component use different partitioning trees; the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a luminance block.
[0280] In some implementations, when the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a second condition is met, then the first information indicates that the current transform block does not use an inseparable transform, and the second condition includes: the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance block.
[0281] In some implementations, when the current coding block does not belong to a preset type of coding block, the first information indicates that the current transform block does not use a non-separable transform.
[0282] In some implementations, the decoder 1200 further includes a parsing unit configured to parse the second information if a third condition is satisfied, where the third condition includes that the current coding block belongs to a preset type of coding block.
[0283] In some implementations, the preset type of coding block includes one or more of the following: an intra-frame coding block, an inter-frame coding block, and an intra-frame block copy (IBC) coding block.
[0284] In some implementations, the decoder 1200 further includes a fourth determining unit configured to determine residual information based on coefficients of the current transform block and a fifth determining unit configured to determine reconstruction information based on the residual information.
[0285] In some implementations, if the first information indicates that the current transform block uses an inseparable transform, the fourth determination unit is configured to determine a transform kernel based on the angle information of the current transform block; and determine the residual coefficient based on the transform kernel and the coefficient of the current transform block.
[0286] In some implementations, if the current coding block is an inter-frame coding block or an IBC coding block, the decoder further includes a sixth determining unit configured to determine the angle information of the current transform block using a gradient histogram.
[0287] In some implementations, the current transform block is a luma block or a chroma block.
[0288] In some implementations, the non-separable transform is LFNST or NSPT.
[0289] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0290] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0291] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the decoder 1200. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned decoding method is implemented.
[0292] Based on the composition of the above-mentioned decoder 1200 and the computer-readable storage medium, refer to Figure 13, which shows a specific hardware structure diagram of the decoder provided by an embodiment of the present application. As shown in Figure 13, the decoder 1300 may include: a communication interface 1310, a memory 1320 and a processor 1330; each component is coupled together through a bus system 1340. It can be understood that the bus system 1340 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 1340 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1340 in Figure 13. Among them,
[0293] The communication interface 1310 is used to receive and send signals when sending and receiving information with other external network elements.
[0294] The memory 1320 is used to store computer programs.
[0295] The processor 1330 is configured to, when running the computer program, execute:
[0296] determining first information, where the first information is used to indicate whether a non-separable transform is used for the current transform block;
[0297] determining, according to the first information, neighboring information of the current transform block;
[0298] Determine coefficients of the current transform block according to the neighboring information.
[0299] It is understood that the memory 1320 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1320 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0300] The processor 1330 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1330. The above-mentioned processor 1330 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1320 , and the processor 1330 reads the information in the memory 1320 and completes the steps of the above method in combination with its hardware.
[0301] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0302] Optionally, as another embodiment, the processor 1330 is further configured to execute the decoding method described in the above embodiment when running the computer program.
[0303] Figure 14 is a schematic diagram of the structure of an encoder provided by one embodiment of the present application. As shown in Figure 14, the encoder 1400 includes: a first determination unit 1410, a second determination unit 1420, and an encoding unit 1430. The first determination unit 1410 is configured to determine first information indicating whether a non-separable transform is used for the current transform block; the second determination unit 1420 is configured to determine neighboring information of the current transform block based on the first information; and the encoding unit 1430 is configured to encode the coefficients of the current transform block based on the neighboring information.
[0304] In some implementations, the first information indicates that the current transform block uses an inseparable transform, and the neighboring information is neighboring information of the current transform block in a scanning order.
[0305] In some implementations, the first information indicates that the current transform block does not use an inseparable transform, and the neighboring information is neighboring information of a spatial position of the current transform block.
[0306] In some implementations, the encoding unit 1430 is configured to determine a context model index and / or Rice parameter based on the neighboring information; and encode the coefficients of the current transform block based on the context model index and / or Rice parameter.
[0307] In some implementations, the encoding unit 1430 is configured to write target parameters into the bitstream based on the context model index and / or Rice parameters, where the target parameters are used to represent the coefficients of the current block.
[0308] In some implementations, the first information is determined based on one or more of the following: second information, used to indicate whether an inseparable transform is used, and the second information is information corresponding to the current coding block; whether the current transform block uses a transform skip mode; whether the luminance component and the chrominance component use different partitioning trees; the type of the current transform block; and the type of the current coding block.
[0309] In some implementations, when the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a first condition is met, then the first information indicates that the current transform block uses an inseparable transform, and the first condition includes one of the following: the luminance component and the chrominance component use different partitioning trees; the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a luminance block.
[0310] In some implementations, when the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a second condition is met, then the first information indicates that the current transform block does not use an inseparable transform, and the second condition includes: the luminance component and the chrominance component do not use different partitioning trees, and the current transform block is a chrominance block.
[0311] In some implementations, when the current coding block does not belong to a preset type of coding block, the first information indicates that the current transform block does not use a non-separable transform.
[0312] In some implementations, the encoding unit 1430 is further configured to encode the second information when a third condition is satisfied, where the third condition includes that the current coding block belongs to a preset type of coding block.
[0313] In some implementations, the preset type of coding block includes one or more of the following: an intra-frame coding block, an inter-frame coding block, and an intra-frame block copy (IBC) coding block.
[0314] In some implementations, the encoder 1400 further includes a third determining unit configured to transform the current transform block and determine transformed coefficients, and a fourth determining unit configured to quantize the transformed coefficients and determine coefficients of the current transform block.
[0315] In some implementations, if the first information indicates that the current transform block uses an inseparable transform, the third determination unit is further configured to determine a transform kernel based on the angle information of the current transform block; and transform the current transform block using the transform kernel to determine the transformed coefficients.
[0316] In some implementations, if the current transform block is an inter-frame transform block or an IBC transform block, the encoder further includes a fifth determining unit configured to determine the angle information of the current transform block using a gradient histogram.
[0317] In some implementations, the current transform block is a luma block or a chroma block.
[0318] In some implementations, the non-separable transform is LFNST or NSPT.
[0319] It is understandable that in the embodiments of the present application, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and of course it can also be a module, or it can be non-modular. Moreover, the various components in this embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional modules.
[0320] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0321] Therefore, an embodiment of the present application provides a computer-readable storage medium, which is applied to the encoder 1400. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the encoding method in the aforementioned embodiment is implemented.
[0322] Based on the composition of the above-mentioned encoder 1400 and the computer-readable storage medium, refer to Figure 15, which shows a specific hardware structure diagram of the encoder provided by an embodiment of the present application. As shown in Figure 15, the encoder 1500 may include: a communication interface 1510, a memory 1520 and a processor 1530; each component is coupled together through a bus system 1540. It can be understood that the bus system 1540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 1540 in Figure 15. Among them,
[0323] The communication interface 1510 is used to send and receive signals when sending and receiving information with other external network elements.
[0324] The memory 1520 is used to store computer programs.
[0325] The processor 1530 is configured to, when running the computer program, execute:
[0326] determining first information, where the first information is used to indicate whether a non-separable transform is used for the current transform block;
[0327] determining, according to the first information, neighboring information of the current transform block;
[0328] The coefficients of the current transform block are encoded according to the neighboring information.
[0329] It is understood that the memory 1520 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 1520 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0330] The processor 1530 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 1530. The above-mentioned processor 1530 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 1520 , and the processor 1530 reads the information in the memory 1520 and completes the steps of the above method in combination with its hardware.
[0331] It is to be understood that these embodiments described in the present application can be implemented with hardware, software, firmware, middleware, microcode or its combination.For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (Application Specific Integrated Circuits, ASIC), digital signal processor (Digital Signal Processing, DSP), digital signal processing equipment (DSP Device, DSPD), programmable logic device (Programmable Logic Device, PLD), field programmable gate array (Field-Programmable Gate Array, FPGA), general-purpose processor, controller, microcontroller, microprocessor, other electronic units for performing functions described in the present application or its combination.For software implementation, the technology described in the present application can be realized by the module (such as process, function etc.) that performs functions described in the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0332] Optionally, as another embodiment, the processor 1530 is further configured to execute the encoding method described in the above embodiment when running the computer program.
[0333] It should be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0334] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0335] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0336] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0337] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0338] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A decoding method, applied to a decoder, comprising: Determining first information, where the first information is used to indicate whether an inseparable transform is used for a current transform block; Determining adjacent information of the current transform block according to the first information; Determining coefficients of the current transform block according to the adjacent information.
2. The method according to claim 1, wherein The first information indicates that the inseparable transform is used for the current transform block, and the adjacent information is adjacent information of the current transform block in the scanning order.
3. The method according to claim 1, wherein, The first information indicates that the inseparable transform is not used for the current transform block, and the adjacent information is adjacent information of the current transform block in the spatial domain position.
4. The method according to claim 1, wherein, The determining the coefficients of the current transform block according to the adjacent information includes: Determining a context model index and / or a Rice parameter according to the adjacent information; Determining the coefficients of the current transform block according to the context model index and / or the Rice parameter.
5. The method according to claim 4, wherein The determining the coefficients of the current transform block according to the context model index and / or the Rice parameter includes: Decoding a target parameter in a bitstream according to the context model index and / or the Rice parameter, where the target parameter is used to represent coefficients of the current block; Determining the coefficients of the current transform block according to the target parameter.
6. The method according to any one of claims 1-5, wherein The first information is determined based on one or more of the following: Second information, which is used to indicate whether an inseparable transform is used, and the second information is information corresponding to a current coding block; Whether the current transform block uses a transform skip mode; Whether different partitioning trees are used for a luminance component and a chrominance component; The type of the current transform block; The type of the current coding block.
7. The method according to claim 6, wherein, When the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a first condition is satisfied, the first information indicates that the inseparable transform is used for the current transform block, and the first condition includes one of the following: Different partitioning trees are used for the luminance component and the chrominance component; Different partitioning trees are not used for the luminance component and the chrominance component, and the current transform block is a luminance block.
8. The method according to claim 6, wherein, When the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a second condition is satisfied, the first information indicates that the inseparable transform is not used for the current transform block, and the second condition includes: different partitioning trees are not used for the luminance component and the chrominance component, and the current transform block is a chrominance block.
9. The method according to claim 6, wherein When the current coding block does not belong to a coding block of a preset type, the first information indicates that the inseparable transform is not used for the current transform block.
10. The method according to any one of claims 6-9, wherein, The method further includes: Parsing the second information when a third condition is satisfied, and the third condition includes that the current coding block belongs to a coding block of a preset type.
11. The method according to any one of claims 7-10, wherein, The coding block of the preset type includes one or more of the following: an intra-coding block, an inter-coding block, and an intra block copy IBC coding block.
12. The method according to any one of claims 1-11, wherein, The method further includes: Determining residual information according to the coefficients of the current transform block; Determining reconstruction information according to the residual information.
13. The method according to claim 12, wherein If the first information indicates that the current transform block uses an inseparable transform, determining residual information according to the coefficients of the current transform block includes: Determining a transform kernel according to the angle information of the current transform block; Determining the residual coefficients according to the transform kernel and the coefficients of the current transform block.
14. The method according to claim 13, wherein If the current coding block is an inter-coded block or an IBC-coded block, the method further includes: Using a gradient histogram to determine the angle information of the current transform block.
15. The method according to any one of claims 1 - 14, wherein, The current transform block is a luminance block or a chrominance block.
16. The method according to any one of claims 1 to 15, wherein The inseparable transform is a low-frequency non-separable quadratic transform (LFNST) or a non-separable basic transform (NSPT).
17. An encoding method, applied to an encoder, includes: Determining first information for indicating whether an inseparable transform is used for a current transform block; Determining adjacent information of the current transform block according to the first information; Encoding the coefficients of the current transform block according to the adjacent information.
18. The method according to claim 17, wherein When the first information indicates that the current transform block uses an inseparable transform, the adjacent information is the adjacent information of the current transform block in the scanning order.
19. The method according to claim 17, wherein When the first information indicates that the current transform block does not use an inseparable transform, the adjacent information is the adjacent information of the current transform block in the spatial domain position.
20. The method according to claim 17, wherein, The encoding the coefficients of the current transform block according to the adjacent information includes: Determining a context model index and / or a Rice parameter according to the adjacent information; Encoding the coefficients of the current transform block according to the context model index and / or the Rice parameter.
21. The method according to claim 20, wherein, The encoding the coefficients of the current transform block according to the context model index and / or the Rice parameter includes: Writing a target parameter into a bitstream according to the context model index and / or the Rice parameter, where the target parameter is used to represent the coefficients of the current block.
22. The method according to any one of claims 17-20, wherein, The first information is determined based on one or more of the following: Second information for indicating whether an inseparable transform is used, and the second information is information corresponding to the current coding block; Whether the current transform block uses a transform skip mode; Whether different partitioning trees are used for the luminance component and the chrominance component; The type of the current transform block; The type of the current coding block.
23. The method according to claim 22, wherein, When the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a first condition is satisfied, the first information indicates that the current transform block uses an inseparable transform, and the first condition includes one of the following: Different partitioning trees are used for the luminance component and the chrominance component; Different partitioning trees are not used for the luminance component and the chrominance component, and the current transform block is a luminance block.
24. The method according to claim 22, wherein When the current coding block belongs to a coding block of a preset type, the current transform block does not use a transform skip mode, and the second information indicates that an inseparable transform mode is used, if a second condition is satisfied, the first information indicates that the current transform block does not use an inseparable transform, and the second condition includes: Different partitioning trees are not used for the luminance component and the chrominance component, and the current transform block is a chrominance block.
25. The method according to claim 22, wherein, In the case that the current coding block does not belong to a coding block of a preset type, the first information indicates that the current transform block does not use an inseparable transform.
26. The method according to any one of claims 22-25, wherein, The method further includes: Encoding the second information when a third condition is satisfied, where the third condition includes that the current coding block belongs to a coding block of a preset type.
27. The method according to any one of claims 23-26, wherein, The coding blocks of the preset type include one or more of the following: intra coding blocks, inter coding blocks, and intra block copy (IBC) coding blocks.
28. The method according to any one of claims 17-27, wherein, The method further includes: Performing a transform on the current transform block to determine the transformed coefficients; Quantizing the transformed coefficients to determine the coefficients of the current transform block.
29. The method according to claim 28, wherein, If the first information indicates that the current transform block uses an inseparable transform, the performing a transform on the current transform block to determine the transformed coefficients includes: Determining a transform kernel according to the angle information of the current transform block; Using the transform kernel to perform a transform on the current transform block to determine the transformed coefficients.
30. The method according to claim 29, wherein If the current transform block is an inter transform block or an IBC transform block, the method further includes: Using a gradient histogram to determine the angle information of the current transform block.
31. The method according to any one of claims 17 - 30, wherein The current transform block is a luminance block or a chrominance block.
32. The method according to any one of claims 17-31, wherein, The inseparable transform is a low-frequency non-separable quadratic transform (LFNST) or a non-separable basis transform (NSPT).
33. A decoder, comprising: A first determination unit configured to determine first information, where the first information is used to indicate whether an inseparable transform is used for a current transform block; A second determination unit configured to determine adjacent information of the current transform block according to the first information; A third determination unit configured to determine the coefficients of the current transform block according to the adjacent information.
34. A decoder, comprising: A memory for storing a computer program; A processor for, when running the computer program, executing the method according to any one of claims 1-16.
35. An encoder, comprising: A first determination unit configured to determine first information, where the first information is used to indicate whether an inseparable transform is used for a current transform block; A second determination unit configured to determine adjacent information of the current transform block according to the first information; An encoding unit for encoding the coefficients of the current transform block according to the adjacent information.
36. An encoder, comprising: A memory for storing a computer program; A processor for, when running the computer program, executing the method according to any one of claims 17-32.
37. A non-volatile computer-readable storage medium storing a bitstream, the bitstream being generated by an encoding method using an encoder, or the bitstream being decoded by a decoding method using a decoder, wherein, The decoding method is the method according to any one of claims 1-16, and the encoding method is the method according to any one of claims 17-32.
38. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1-16, or 17-32 is implemented.
39. A bitstream, where the bitstream includes a bitstream generated by the method according to any one of claims 17-32.
Citation Information
Patent Citations
Method and apparatus for video coding
CN113615182A
Context modeling for residual coding
CN113853785A
Contextual modeling of side information for reduced quadratic transforms in video
CN114223208A
Non-separable primary transform design method and apparatus
WO2023075353A1