Coding Method, Encoder, Decoder, and Storage Medium
By determining the mapping mode of the LFNST transform set based on the size parameter and limiting DIMD to smaller blocks, the method addresses the complexity issue in VVC, enhancing coding efficiency and reducing costs for high-resolution video applications.
Patent Information
- Application Number
- JP2024576768
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The high computational complexity of Decoder side Intra Mode Derivation (DIMD) technology in H.266/Versatile Video Coding (VVC) increases compression costs, making it inefficient for high-resolution and ultra-high-resolution video applications.
A method that reduces complexity by determining the mapping mode of the Low-Frequency Non-Separable Transform (LFNST) transform set based on the size parameter of the current block, limiting the use of DIMD for smaller blocks and using Matrix-Based Intra Prediction (MIP) for larger blocks, thereby simplifying the derivation process.
This approach reduces computational complexity and improves coding efficiency by minimizing the need for DIMD in larger image blocks, maintaining performance comparable to JVET-Z0048 while reducing software and hardware complexity.
Smart Images

Figure 2025520837000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of image processing, and in particular, to coding methods, encoders, decoders, and storage media.
Background Art
[0002] As people's demands for video display quality increase, new video application forms such as high-resolution videos and ultra-high-resolution videos have emerged. H.265 / High Efficiency Video Coding (HEVC) has become unable to meet the needs due to the rapid development of video applications. The Joint Video Exploration Team (JVET) has proposed H.266 / Versatile Video Coding (VVC), the next-generation video coding standard. The test mode corresponding to VVC is the VVC Test Model (VTM). The Enhanced Compression Model (ECM) begins to receive newer and more efficient compression algorithms based on VTM10.0.
[0003] Decoder side Intra Mode Derivation (DIMD) is an intra prediction technique of ECM. The main core of this technique is to derive the intra prediction mode using the same method as the encoding side on the decoding side, thereby achieving the purpose of saving bit overhead.
[0004] However, the application of DIMD technology brings relatively high complexity in both software and hardware, thereby increasing the compression cost.
Summary of the Invention
[0005] Embodiments of the present application provide a coding method, an encoder, a decoder, and a storage medium. Thereby, the complexity of calculation can be reduced, and thus the coding efficiency can be improved.
[0006] The technical solution of the embodiments of the present application can be realized as follows.
[0007] In a first aspect, embodiments of the present application provide a decoding method applied to a decoder. The method includes: decoding a bitstream to determine a prediction mode parameter; when the prediction mode parameter indicates that matrix-based intra prediction (MIP) is used to determine an intra prediction value, decoding the bitstream to determine the MIP parameter of the current block; decoding the bitstream to determine the transform coefficient and the low-frequency non-separable transform (LFNST) index of the current block; when the LFNST index indicates that LFNST is used for the current block, determining the mapping mode of the LFNST transform set based on the MIP parameter; selecting one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets based on the mapping mode of the LFNST transform set, and determining the LFNST transform kernel used for the current block from the selected LFNST transform kernel candidate set; and transforming the transform coefficient by using the LFNST transform kernel.
[0008] In a second aspect, embodiments of the present application provide an encoding method applied to an encoder. The method includes: determining a prediction mode parameter; when the prediction mode parameter indicates that MIP is used for the current block to determine an intra prediction value, determining the MIP parameter of the current block; Based on the MIP parameters, determine the intra prediction block of the current block, and calculate the residual block obtained by subtracting the intra prediction value from the current block. When LFNST is used for the current block, determine the mapping mode of the LFNST transform set based on the MIP parameters. Based on the mapping mode of the LFNST transform set, select one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets, and determine the LFNST transform kernel used for the current block from the selected LFNST transform kernel candidate set, set the LFNST index, and signal it to the video bitstream. Include transforming the residual block using the LFNST transform kernel.
[0009] In a third aspect, an embodiment of the present application provides an encoder. The encoder includes a first determination unit, an encoding unit, and a first conversion unit. The first determination unit determines prediction mode parameters. When the prediction mode parameter indicates that MIP is used for the current block to determine the intra prediction value, determine the MIP parameters of the current block. Based on the MIP parameters, determine the intra prediction block of the current block, calculate the residual block obtained by subtracting the intra prediction value from the current block. When LFNST is used for the current block, determine the mapping mode of the LFNST transform set based on the MIP parameters. Based on the mapping mode of the LFNST transform set, select one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets, and determine the LFNST transform kernel used for the current block from the selected LFNST transform kernel candidate set, and set the LFNST index. It is configured as follows. The encoding unit is configured to signal the LFNST index to the video bitstream. The first conversion unit is configured to convert the residual block using the LFNST conversion kernel.
[0010] In a fourth aspect, an embodiment of the present application provides an encoder. The encoder includes a first memory and a first processor. The first memory is configured to store a computer program executable by the first processor. When the first processor executes the computer program, it is configured to execute the method described in the second aspect.
[0011] In a fifth aspect, an embodiment of the present application provides a decoder. The decoder includes a second determination unit and a second conversion unit. The second determination unit decodes the bitstream to determine the prediction mode parameter. When the prediction mode parameter indicates that MIP is used to determine the intra prediction value, the bitstream is decoded to determine the MIP parameter of the current block. The bitstream is decoded to determine the transform coefficient and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. It is configured as follows. The second conversion unit is configured to convert the transform coefficient using the LFNST transform kernel.
[0012] In a sixth aspect, an embodiment of the present application provides a decoder. The decoder includes a second memory and a second processor, The second memory is configured to store a computer program executable by the second processor, When the second processor executes the computer program, it is configured to execute the method described in the first aspect.
[0013] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program is executed, the method described in the first aspect is realized, or the method described in the second aspect is realized.
[0014] Embodiments of the present application provide a coding method, an encoder, a decoder, and a storage medium. On the decoding side, a bitstream is decoded to determine a prediction mode parameter. When the prediction mode parameter indicates that MIP is used to determine an intra prediction value, the bitstream is decoded to determine the MIP parameter of the current block. The bitstream is decoded to determine the transform coefficient and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The transform coefficient is transformed using the LFNST transform kernel. On the encoding side, a prediction mode parameter is determined. When the prediction mode parameter indicates that MIP is used for the current block to determine an intra prediction value, the MIP parameter of the current block is determined. Based on the MIP parameter, the intra prediction block of the current block is determined, and a residual block obtained by subtracting the intra prediction value from the current block is calculated. When LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The LFNST index is set and signaled to the video bitstream. The residual block is transformed using the LFNST transform kernel. As can be seen from the above, in the embodiments of the present application, when deriving a prediction block, based on the size parameter in the MIP parameter of the current block, the mapping mode of the LFNST transform set is determined.For relatively large image blocks, DIMD may not be necessary to derive the mapping mode. Thereby, the computational complexity can be reduced, and thus, the coding efficiency can be improved.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0016] To more clearly understand the features and technical content of the embodiments of the present application, the realization of the embodiments of the present application will be described in detail below with reference to the drawings. The attached drawings are only used for explanation and do not limit the embodiments of the present application.
[0017] In the following description, although it relates to "some embodiments", which describe a subset of all possible embodiments, it should be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and they may be combined with each other as long as there is no contradiction. Also, note that the terms "first / second / third" in the embodiments of the present application are merely used to distinguish similar objects and do not mean a specific order of the objects. So that the embodiments of the present application described herein can be implemented in an order other than the order illustrated or described herein, "first / second / third" can exchange a specific order or priority when permitted.
[0018] In a video image, usually, a coding block (CB) is represented using a first color component, a second color component, and a third color component. These three color components are, respectively, one luma component, one blue chroma component, and one red chroma component. Specifically, the luma component is usually represented by the symbol Y, the blue chroma component is usually represented by the symbol Cb or U, and the red chroma component is usually represented by the symbol Cr or V. Thus, a video image can be represented in the YCbCr format or can also be represented in the YUV format.
[0019] In the embodiments of the present application, the first color component may be a luma component, the second color component may be a blue chroma component, and the third color component may be a red chroma component. It is not particularly limited in the embodiments of the present application.
[0020] The block-based hybrid coding framework is used in general video coding standards. Each image in a video is divided into the largest coding unit (LCU) or Coding Tree Unit (CTU) of a square with the same size (e.g., 128×128, 64×64, etc.). Each largest coding unit or coding tree unit can further be divided into rectangular coding units (CUs) based on rules. The coding unit may also be further divided into smaller prediction units (PUs), transform units (TUs), etc.
[0021] The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, inloop Filter, etc. The prediction module includes intra prediction and inter prediction. Inter prediction includes motion estimation and motion compensation. Since there is a strong correlation between adjacent samples in the video image, in video coding technology, the intra prediction method is used to eliminate the spatial redundancy between adjacent samples. However, since there is a strong similarity between adjacent images in a video, in video coding technology, the inter prediction method can be used to eliminate the temporal redundancy between adjacent images and improve the coding efficiency.
[0022] The basic flow of a video codec is executed as follows. On the encoding side, the video codec divides an image into blocks, generates a predicted block for the current block by performing intra prediction or inter prediction on the current block, subtracts the predicted block from the original block of the current block to obtain a residual block, transforms and quantizes the residual block to obtain a quantized coefficient matrix, and entropy-codes the quantized coefficient matrix and outputs it to a bitstream. On the decoding side, the video codec generates a predicted block for the current block by performing intra prediction or inter prediction on the current block. Also, it decodes the bitstream to obtain a quantized coefficient matrix, inverse quantizes and inverse-transforms the quantized coefficient matrix to obtain a residual block, and adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed block forms a reconstructed image, and the reconstructed image is loop-filtered based on the image or block to obtain a decoded image. On the encoding side as well, in order to obtain a decoded image, processing similar to that on the decoding side is required. The decoded image can be a reference image for inter prediction of subsequent images. Block partitioning information, mode information such as prediction, transformation, quantization, entropy coding, loop filtering, or parameter information, etc. determined on the encoding side is output to the bitstream as necessary. The decoding side decodes and analyzes the existing information to determine the same block partitioning information, mode information such as prediction, transformation, quantization, entropy coding, loop filtering, or parameter information as on the encoding side. Thereby, it is ensured that the decoded image obtained on the encoding side is the same as the decoded image obtained on the decoding side. The decoded image obtained on the encoding side is usually also called a reconstructed image. During prediction, the current block may be divided into prediction units, and during transformation, the current block may be divided into transformation units, and the partitioning of the prediction units and transformation units may be different. The above is the basic flow of video coding in a block-based hybrid coding framework.With the development of technology, some modules or some steps in some flows of the framework or flow may be optimized. Embodiments of the present application are applied to the basic flow of the video codec in the hybrid coding framework based on the block, but are not limited to the framework or flow.
[0023] The current block can be, for example, the current coding unit (CU) or the current prediction unit (PU).
[0024] The Joint Video Exploration Team (JVET) responsible for standardizing the next-generation video coding standard established a group to study a coding model beyond H.266 / VVC and named this model, i.e., the reference software, Enhanced Compression Model (ECM). Based on VTM10.0, ECM began to receive newer and more efficient compression algorithms and currently exceeds the coding performance of VVC by about 13%. ECM not only extends the size of the coding unit at a specific resolution but also integrates numerous intra prediction techniques and inter prediction techniques.
[0025] Hereinafter, technical solutions related to the matrix-based intra prediction (MIP) technique will be described.
[0026] The intra prediction technique based on matrices, i.e., the MIP technique, can be divided into three main steps: downsampling, matrix multiplication, and upsampling. In step 1, spatially adjacent reconstruction samples are downsampled, and the resulting downsampled sample sequence is used as the input vector for step 2. In step 2, the output vector of step 1 is used as the input for step 2, and the input for step 2 is multiplied by a preset matrix, and a bias vector is added to the result to output a calculated sample vector. In step 3, the output vector of step 2 is used as the input for step 3, and the input for step 3 is upsampled to obtain a final prediction block. Figure 1 is a schematic diagram showing the intra prediction technique based on matrices. As shown in Figure 1, in the MIP technique, in step 1, the reconstruction samples adjacent above the current coding unit are averaged to obtain a downsampled upper adjacent reconstruction sample vector, and the reconstruction samples adjacent to the left of the current coding unit are averaged to obtain a downsampled left adjacent reconstruction sample vector. The upper vector and the left vector are used as the input for the matrix multiplication in step 2. A k is a preset matrix, and b k is a preset bias vector, and k is the MIP mode index. In step 3, the result obtained in step 2 is upsampled by linear interpolation to obtain a prediction sample block having the number of samples that matches the actual number of samples of the coding unit. For coding units of different block sizes, the number of MIP modes is different. Taking H.266 / VVC as an example, for a coding unit of size 4×4, MIP has 16 prediction modes. For a coding unit of size 8×8, or a coding unit with a width or height equal to 4, MIP has 8 prediction modes. For coding units of other sizes, MIP has 6 prediction modes. Also, the MIP technology has a transpose function. For the prediction mode matching the current size, in MIP, the transpose calculation can be tried on the encoding side. If transposition is required, the sequence of inputting the upper input vector and the left input vector is swapped, and after performing matrix multiplication, the output is swapped.
[0027] Therefore, for MIP, a flag indicating whether the MIP technology is used for the current coding unit is required. Also, when the MIP technology is used for the current coding unit, it is necessary to further transmit the transpose flag and the MIP mode index to the decoding side.
[0028] In the VVC standard text, the transpose flag of MIP is binarized by the Fixed Length (FL) coding method and has a length of 1. The mode index of MIP is binarized by the Truncated Binary (TB) coding method.
[0029] Similarly, the low-frequency non-separable transform (LFNST) technology is also a technology adopted in the VVC text. Hereinafter, the technical solutions related to LFNST will be described.
[0030] LFNST is applied between the forward primary transform and quantization on the encoding side, and between the inverse primary transform and inverse quantization on the decoding side. After performing the primary transform on the residual of the current coding block, the coefficients in the frequency domain are obtained. Based on this, in LFNST, a frequency domain transform is performed on some of the coefficients, that is, after transforming some of the coefficients in the frequency domain to obtain the coefficients in other regions, operations such as quantization and entropy coding are performed. LFNST further removes statistical redundancy and has good performance in VTM, which is the reference software of VVC.
[0031] In LFNST, a secondary transform is mainly performed on a 4×4 or 8×8 area at the upper left corner of the transform block. Also, the transform kernels of LFNST are mainly classified into four transform sets in VVC, and each transform set has two candidate transform kernels. In ECM, the transform kernels of LFNST are extended from the original four transform sets to 35 transform sets. Since each of the original transform sets has two candidate transform kernels, it is extended to each transform set having three candidate transform kernels.
[0032] LFNST can act on both intra prediction and inter prediction. In LFNST, bit overhead can be saved by selecting a transform set corresponding to the intra prediction mode in intra prediction. Usually, intra prediction corresponds to an intra prediction mode, that is, it corresponds to the Direct Current (DC) mode, the PLANAR mode, or the angular prediction mode. These intra prediction modes are bound (or bound) to the LFNST transform set. For example, in VVC, the DC mode and the PLANAR mode correspond to the first transform set, as specifically shown in Table 1.
[0033]
Table 1
[0034] The predModeIntra may be an intra prediction mode indicator, and the SetIdx may be an LFNST index. The value of the LFNST index is set to indicate whether LFNST is used for the current block and the index of the LFNST transform kernel in the LFNST transform kernel candidate set. For example, when the LFNST transform set includes four transform kernel candidate sets (set0, set1, set2, set3), the corresponding values of SetIdx are 0, 1, 2, and 3, respectively.
[0035] Accordingly, after expanding the transform kernel of LFNST in the ECM, the LFNST transform sets corresponding to different intra prediction modes become finer. For example, FIG. 2 is a correspondence table between the intra prediction mode and the transform set. As shown in FIG. 2, after expansion, there are 35 transform sets.
[0036] Hereinafter, a technical solution related to the decoder-side intra mode derivation (DIMD) technique will be described.
[0037] DIMD is an intra prediction technology for ECM, which is not available in VVC. The main core of this technology is to derive the intra prediction mode on the decoding side using the same method as the encoding side. By doing so, it avoids transmitting the intra prediction mode index of the current coding unit in the bitstream and achieves the purpose of saving bit overhead. The specific process can be divided into two main steps. In Step 1, the prediction mode is derived, and on the decoding side and the encoding side, the same prediction mode intensity calculation method is used. On the encoding side, the Sobel operator is used to calculate the histogram of gradients for each prediction mode. The region of action is the reconstructed samples in the three adjacent rows above the current block, the three adjacent columns to the left of the current block, and the corresponding adjacent reconstructed samples in the upper left. By calculating the histogram of gradients within this L-shaped region, the first prediction mode corresponding to the largest amplitude (also referred to as magnitude) and the second prediction mode corresponding to the second largest amplitude in the histogram of gradients can be obtained. The decoding side derives the first and second prediction modes in the same step. In Step 2, the prediction block is derived, and on the decoding side and the encoding side, the same prediction block derivation method is used to obtain the prediction block of the current block. The encoding side judges the following two conditions. Condition 1: The gradient value of the second prediction mode is not 0. Condition 2: Neither the first prediction mode nor the second prediction mode is the PLANAR mode or the DC prediction mode. If the above two conditions do not hold simultaneously, the prediction sample values of the current block are calculated using only the first prediction mode, that is, the normal prediction process is applied to the first prediction mode. Otherwise, that is, when the above two conditions hold simultaneously, the current prediction block is derived using the weighted average method. The specific method is as follows. The PLANAR mode occupies a weight of 1 / 3.The weight of the first prediction mode is obtained by multiplying the ratio of the gradient strength of the first prediction mode to the sum of the gradient strengths of the first and second prediction modes by 2 / 3. The weight of the second prediction mode is obtained by multiplying the ratio of the gradient strength of the second prediction mode to the sum of the gradient strengths of the first and second prediction modes by 2 / 3. A weighted average is performed on the above three prediction modes, namely, the PLANAR mode, the first prediction mode, and the second prediction mode, to obtain the prediction block of the current coding unit. The decoder side also obtains the prediction block in the same steps. FIG. 3 is a schematic diagram showing the decoder-side intra mode derivation technique. The above specific process is shown as in FIG. 3.
[0038] In step 2 above, the specific weights are calculated as follows.
[0039]
Number
[0040]
Number
[0041]
Number
[0042] mode1 and mode2 represent the first prediction mode and the second prediction mode respectively, and amp1 and amp2 represent the gradient amplitude values of the first prediction mode and the second prediction mode respectively. In the DIMD technique, it is necessary to transmit a flag to the decoder side. The flag is used to indicate whether the DIMD technique is used for the current coding unit.
[0043] To improve LFNST and MIP, in VVC and ECM, the relationship between MIP and LFNST is simplified, and all MIP prediction modes are treated as the PLANAR mode by default before being mapped to the LFNST transform set. The reasons are as follows. In the initial stage of the design, for LFNST, the intra prediction mode is used as the training input, and the transform kernel coefficients of LFNST are obtained through deep learning training. However, the representation of the MIP prediction mode is different from that of the conventional intra prediction mode. The MIP prediction mode represents certain prediction matrix coefficients, while the conventional prediction mode represents directionality. Also, the prediction results of MIP are similar to those of the conventional PLANAR mode. Therefore, all MIP prediction modes are mapped to the LFNST transform set using PLANAR.
[0044] Optionally, for the MIP prediction block, the gradient amplitude values of each conventional intra prediction mode can be sorted using DIMD, and the optimal possible conventional intra prediction mode can be mapped to the LFNST transform set. Also, optionally, the original size range of the coding unit for which LFNST is permitted to be used in MIP can be extended. In VVC and ECM, LFNST is permitted to be used in MIP only when both the width and height of the current coding unit are 16 or more. After the extension, LFNST is permitted to be used in MIP when both the width and height of the current coding unit are 4 or more.
[0045] The above method successfully solves the problem of mapping the MIP prediction mode to LFNST and improves the encoding efficiency, but it also brings corresponding complexity. For example, it includes coding time in software or cache and timing problems in hardware implementation.
[0046] Regarding the extension of the range in which LFNST is used in the proposed MIP, the permitted condition is that if both the width and height are 4 or more, LFNST can be used. Thus, the range of the conventional intra prediction mode derived using DIMD for the MIP prediction block is extended, and the coding complexity continues to increase.
[0047] Compared with the encoding side, the cost of increasing the complexity on the decoding side is higher. Also, regarding LFNST, the usage range of the 4×4 coding unit is already supported in the original VVC and ECM, and it is considered that there are no additional concerns. However, the DIMD derivation process for the prediction block is a technology and operation not present in the current VVC and ECM. If the usage conditions of DIMD can be reduced or the steps simplified while maintaining the corresponding coding performance, the decoding efficiency on the decoding side can be significantly improved.
[0048] That is, a general coding technology solution based on DIMD technology brings relatively high complexity in both software and hardware, thereby increasing the compression cost and reducing the coding efficiency.
[0049] To solve the above problems, in the embodiments of the present application, on the decoding side, the bitstream is decoded to determine the prediction mode parameters. When the prediction mode parameters indicate that MIP is used to determine the intra prediction value, the bitstream is decoded to determine the MIP parameters of the current block. The bitstream is decoded to determine the transform coefficients and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, based on the MIP parameters, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The transform coefficients are transformed using the LFNST transform kernel. On the encoding side, the prediction mode parameters are determined. When the prediction mode parameters indicate that MIP is used for the current block to determine the intra prediction value, the MIP parameters of the current block are determined. Based on the MIP parameters, the intra prediction block of the current block is determined, and the residual block obtained by subtracting the intra prediction value from the current block is calculated. When LFNST is used for the current block, based on the MIP parameters, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The LFNST index is set and signaled to the video bitstream. The residual block is transformed using the LFNST transform kernel. As can be seen from the above, in the embodiments of the present application, when deriving the prediction block, based on the size parameter in the MIP parameters of the current block, the mapping mode of the LFNST transform set is determined.For relatively large image blocks, it is not necessary to use DIMD to derive the mapping mode. Thereby, the computational complexity can be reduced, and thus, the coding efficiency can be improved.
[0050] Referring to FIG. 4, FIG. 4 is a block diagram showing the configuration of a video encoding system according to an embodiment of the present application. As shown in FIG. 4, the video encoding system 10 can include a conversion / quantization unit 101, an intra prediction unit 102, an intra prediction unit 103, a motion compensation unit 104, a motion estimation unit 105, an inverse conversion / inverse quantization unit 106, a filter control analysis unit 107, a filtering unit 108, a coding unit 109, a decoding image buffer unit 110, and the like. The filtering unit 108 can implement deblocking filtering and sample adaptive offset (SAO) filtering. The coding unit 109 can implement header information coding and context-based adaptive binary arithmetic coding (CABAC). For the input original video signal, one video coding block can be obtained by dividing a coding tree unit (CTU). Next, for the residual sample information obtained by performing intra prediction or inter prediction, the conversion / quantization unit 101 performs a conversion on the video coding block, for example, including converting the residual information from the sample domain to the transform domain, and quantizing the obtained transform coefficients to further reduce the bit rate. The intra prediction unit 102 and the intra prediction unit 103 are used to perform intra prediction on the video coding block. Specifically, the intra prediction unit 102 and the intra prediction unit 103 are used to determine an intra prediction mode for coding the video coding block. The motion compensation unit 104 and the motion estimation unit 105 are used to perform inter prediction coding of the received video coding block with respect to one or more blocks in one or more reference images to provide temporal prediction information.The motion estimation performed by the motion estimation unit 105 is a process of generating a motion vector, and the motion of the video coding block can be estimated by the motion vector. Next, the motion compensation unit 104 performs motion compensation based on the motion vector specified by the motion estimation unit 105. After the intra prediction mode is specified, the intra prediction unit 103 is further used to provide the selected intra prediction data to the coding unit 109, and the motion estimation unit 105 transmits the motion vector data specified by calculation to the coding unit 109. Further, the inverse transform and inverse quantization unit 106 reconstructs the video coding block and is used to reconstruct the residual block within the sample region. The reconstructed residual block is filtered by the filter control analysis unit 107 and the filtering unit 108 to remove the blocking effect artifact, and then the reconstructed residual block is added to one prediction block in the image of the decoding image buffer unit 110 to generate a reconstructed video coding block. The coding unit 109 is used to code various coding parameters and the quantized transform coefficients. In the coding algorithm based on CABAC, the context content can be based on adjacent coding blocks and used to code the information indicating the specified intra prediction mode to output the bitstream of the video signal. The decoding image buffer unit 110 is used to store the reconstructed video coding block for prediction reference. As the video image coding is executed, new reconstructed video coding blocks are continuously generated, and all these reconstructed video coding blocks are stored in the decoding image buffer unit 110.
[0051] Referring to FIG. 5, FIG. 5 is a block diagram showing the configuration of a video decoding system according to an embodiment of the present application. As shown in FIG. 5, the video decoding system 20 includes a decoding unit 201, an inverse transform and inverse quantization unit 202, an intra prediction unit 203, a motion compensation unit 204, a filtering unit 205, a decoded image buffer unit 206, and the like. The decoding unit 201 can realize header information decoding and CABAC decoding. The filtering unit 205 can realize deblocking filtering and SAO filtering. After the input video signal is encoded as shown in FIG. 4, the bitstream of the video signal is output. The bitstream is input to the video decoding system 20, and first, the decoding unit 201 obtains the transform coefficients after decoding, and the inverse transform and inverse quantization unit 202 processes the transform coefficients to generate a residual block in the sample region. The intra prediction unit 203 can be used to generate prediction data for the current video decoding block based on the specified intra prediction mode and data from the blocks decoded before the current frame or image. The motion compensation unit 204 analyzes the motion vector and other related syntax elements to identify the prediction information for the video decoding block, and uses the prediction information to generate a prediction block for the decoded video decoding block. By adding the residual block from the inverse transform and inverse quantization unit 202 and the corresponding prediction block generated by the intra prediction unit 203 or the motion compensation unit 204, a decoded video block is formed. The decoded video signal can remove block effect artifacts by the filtering unit 205, thereby improving the video quality.Next, store the decoded video block in the decoding image buffer unit 206. The decoding image buffer unit 206 stores a reference image for subsequent intra prediction or motion compensation, and at the same time is used to output a video signal to obtain the restored original video signal.
[0052] The encoding method in the embodiments of the present application may be applied to the intra estimation unit 102 and the intra prediction unit 103 shown in FIG. 4. Also, the decoding method in the embodiments of the present application may be applied to the intra prediction unit 203 shown in FIG. 5. That is, the coding method in the embodiments of the present application may be applied to a video encoding system, may be applied to a video decoding system, and may further be applied to both a video encoding system and a video decoding system. It is not particularly limited in the embodiments of the present application. When the coding method is applied to a video encoding system, "the current block" specifically refers to the current encoding block in intra prediction. When the coding method is applied to a video decoding system, "the current block" specifically refers to the current decoding block in intra prediction.
[0053] With reference to the accompanying drawings of the embodiments of the present application, the technical solutions according to the embodiments of the present application will be clearly and completely described.
[0054] The embodiments of the present application provide a decoding method. FIG. 6 is a flowchart of the decoding method according to the embodiments of the present application. As shown in FIG. 6, the decoding method executed by the decoder may include the following steps. Step 101: Decode the bitstream to determine the prediction mode parameter.
[0055] In the embodiments of the present application, the decoder can first decode the bitstream to determine the prediction mode parameter.
[0056] In the embodiments of the present application, the prediction mode parameter indicates the current block coding mode and the parameters related to the coding mode. The prediction mode usually includes a conventional intra prediction mode and a non-conventional intra prediction mode. The conventional intra prediction mode can include a direct current (DC) mode, a planar (PLANAR) mode, an angular mode, etc., and the non-conventional intra prediction mode can include a MIP mode, a cross-component linear model prediction (CCLM) mode, an intra block copy (IBC) mode, a palette (PLT) mode, etc.
[0057] As can be understood, in the embodiments of the present application, on the encoding side, predictive encoding can be performed on the current block. During the predictive encoding, the prediction mode of the current block can be determined, and the corresponding prediction mode parameter can be signaled to the bitstream. Thereby, the prediction mode parameter can be transmitted from the encoder to the decoder.
[0058] Accordingly, on the decoding side, by decoding the bitstream, the intra prediction mode of the current block, or the luminance component or chrominance component of the coding block where the current block is located, can be obtained. In this case, the value of predModeIntra (intra prediction mode indicator) can be determined, and this value is calculated as follows.
[0059]
Equation
[0060] The color component indicator (which can be represented by cIdx) is used to represent the luminance component or the chroma component of the current block. When the luminance component of the current block is predicted, cIdx is equal to 0, and when the chroma component of the current block is predicted, cIdx is equal to 1. Also, (xTbY, yTbY) are the coordinates of the sample at the upper left corner of the current block, IntraPredModeY[xTbY][yTbY] is the intra prediction mode of the luminance component, and IntraPredModeC[xTbY][yTbY] is the intra prediction mode of the chroma component.
[0061] Furthermore, in the embodiments of the present application, by obtaining the prediction mode parameter, based on the prediction mode parameter, it is possible to determine whether MIP is used to determine the intra prediction value during intra prediction.
[0062] Step 102: When the prediction mode parameter indicates that MIP is used to determine the intra prediction value, decode the bitstream to determine the MIP parameter of the current block.
[0063] In the embodiments of the present application, after determining the prediction mode parameter, when the prediction mode parameter indicates that MIP is used to determine the intra prediction value, the bitstream can be continuously decoded, thereby determining the MIP parameter of the current block.
[0064] Note that in the embodiments of the present application, the MIP parameter can include a MIP transpose indication parameter (which can be represented by isTransposed), a MIP mode index (which can be represented by modelId), the size of the current block, and the type of the current block (which can be represented by mipSizeId). The values of these parameters can be obtained by decoding the bitstream.
[0065] That is, in the embodiments of the present application, the MIP parameters determined by decoding the bitstream can indicate at least one of information such as the MIP transposition instruction parameter, the MIP mode index, the size of the current block, and the type of the current block.
[0066] Furthermore, in the embodiments of the present application, by decoding the bitstream, the value of isTransposed can be determined. When the value of isTransposed is equal to 1, it can be determined that it is necessary to transpose the sample input vector used in the MIP mode. When the value of isTransposed is equal to 0, it can be determined that it is not necessary to transpose the sample input vector used in the MIP mode. That is, the MIP transposition instruction parameter isTransposed can be used to indicate whether to transpose the sample input vector used in the MIP mode.
[0067] Furthermore, in the embodiments of the present application, by decoding the bitstream, the MIP mode index modeId can be further determined. The MIP mode index may be used to indicate the MIP mode used for the current block, and the MIP mode may also be used to indicate the calculation and derivation method for determining the intra prediction block of the current block using MIP. That is, different MIP modes have different corresponding values of the MIP mode index. The value of the MIP mode index may be 0, 1, 2, 3, 4, or 5.
[0068] Furthermore, in the embodiments of the present application, by decoding the bitstream, parameter information such as the size of the current block, the ratio of the width to the height of the current block (i.e., the aspect ratio), and the type mipSizeId of the current block can be determined. In this way, after determining the MIP parameters, it becomes easy to continue to select the LFNST transform kernel (which can be represented by kernel) used for the current block based on the determined MIP parameters.
[0069] That is, in the embodiments of the present application, the MIP parameter may be used to determine the size parameter of the current block. The size parameter may represent the size of the current block, may be the height and width of the current block, or may be the aspect ratio of the current block.
[0070] Step 103: Decode the bitstream to determine the transform coefficients and the LFNST index of the current block.
[0071] In the embodiments of the present application, when the prediction mode parameter indicates that MIP is used to determine the intra prediction value, after determining the MIP parameter of the current block, the transform coefficients and the LFNST index of the current block can be determined by continuously decoding the bitstream.
[0072] Note that in the embodiments of the present application, the value of the LFNST index may be used to indicate whether LFNST is used for the current block, and may also be used to indicate the index of the LFNST transform kernel in the LFNST transform kernel candidate set.
[0073] That is, in the embodiments of the present application, after decoding the LFNST index, if the value of the LFNST index is equal to 0, it indicates that LFNST is not used for the current block. If the value of the LFNST index is greater than 0, it indicates that LFNST is used for the current block. In this case, the index of the transform kernel may be equal to the value of the LFNST index, or the index of the transform kernel may be equal to the value obtained by subtracting 1 from the value of the LFNST index.
[0074] Step 104: When the LFNST index indicates that LFNST is used for the current block, determine the mapping mode of the LFNST transform set based on the MIP parameter.
[0075] In an embodiment of the present application, after determining the transformation coefficient and the LFNST index of the current block, if the LFNST index indicates that LFNST is used for the current block, the mapping mode of the LFNST transformation set can be further determined based on the MIP parameter.
[0076] Note that in an embodiment of the present application, the MIP parameter may be the size parameter of the current block. The size parameter may represent the size of the current block, may be the height and width of the current block, or may be the aspect ratio of the current block.
[0077] That is, in an embodiment of the present application, when determining the mapping mode of the LFNST transformation set, the size parameter of the current block may be referred to. For example, based on the height and width of the current block, the mapping mode of the LFNST transformation set is determined, or based on the aspect ratio of the current block, the mapping mode of the LFNST transformation set is determined.
[0078] Furthermore, in an embodiment of the present application, when determining the mapping mode of the LFNST transformation set based on the size parameter of the current block, first, it can be determined whether the size parameter satisfies a first preset condition. If the size parameter satisfies the first preset condition, the first preset prediction mode can be determined as the mapping mode of the LFNST transformation set. If the size parameter does not satisfy the first preset condition, DIMD is used to determine the mapping mode of the LFNST transformation set.
[0079] In the embodiments of the present application, the first preset condition may be used to limit the size of the current block. The first preset condition corresponds to the size parameter of the current block. When the size parameter of the current block is the height and width of the current block, the first preset condition can be used to limit the height and width respectively. When the size parameter of the current block is the aspect ratio of the current block, the first preset condition can be used to limit the aspect ratio.
[0080] Exemplarily, in the embodiments of the present application, assuming that the size parameter of the current block is the height and width of the current block, the first preset condition can be set such that the width is greater than or equal to a preset width threshold, and / or the height is greater than or equal to a preset height threshold. For example, when the height of the current block is greater than or equal to the preset height threshold, or the width of the current block is greater than or equal to the preset width threshold, it can be determined that the size parameter satisfies the first preset condition. When the height of the current block is less than the preset height threshold and the width of the current block is less than the preset width threshold, it can be determined that the size parameter does not satisfy the first preset condition.
[0081] As can be understood, in the embodiments of the present application, both the preset width threshold and the preset height threshold may be any value greater than or equal to 0. For example, the preset width threshold is 32, and the preset height threshold is also 32. When the height or width of the current block is 32 or more, it can be determined that the current block satisfies the first preset condition. The preset width threshold is 32, and the preset height threshold is 16. When the height of the current block is 32 or more, or the width of the current block is 16 or more, it can be determined that the current block satisfies the first preset condition.
[0082] As can be understood, in the embodiments of the present application, the use of DIMD can be restricted under the first preset condition. That is, DIMD is only permitted to be used to determine the mapping mode of the LFNST conversion set when the size parameter of the current block does not meet the first preset condition.
[0083] Thus, according to the decoding method according to the embodiments of the present application, when determining the mapping mode of the LFNST conversion set, the size of the image block for which DIMD is used can be restricted using the first preset condition, and by only permitting the use of DIMD for a part of the image block, the computational complexity can be effectively reduced. For example, DIMD is only permitted to be used to determine the mapping mode of the LFNST conversion set for image blocks of a relatively small size.
[0084] Furthermore, in the embodiments of the present application, when the size parameter of the current block meets the first preset condition, the first preset prediction mode may be directly determined as the mapping mode of the LFNST conversion set. The first preset prediction mode may be the PLANAR mode or the DC mode.
[0085] That is, in the embodiments of the present application, when determining the mapping mode of the LFNST conversion set, in combination with the first preset condition, the mapping mode can be directly set for some image blocks. For example, for image blocks of a relatively large size, the PLANAR mode and the DC mode are directly determined as the mapping modes of the LFNST conversion set.
[0086] Furthermore, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set using DIMD, first, at least one intra prediction mode can be traversed, whereby at least one gradient information corresponding to the current block can be determined. Next, the mapping mode of the LFNST transform set may be determined based on the at least one gradient information. One intra prediction mode corresponds to one gradient information, and the gradient information may be a gradient histogram.
[0087] As can be understood, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set based on at least one gradient information, the gradient amplitude value corresponding to each intra prediction mode can be determined based on the at least one gradient information. Next, among the at least one intra prediction mode, the intra prediction mode with the maximum gradient amplitude value can be determined as the mapping mode of the LFNST transform set. One intra prediction mode corresponds to one gradient amplitude value.
[0088] That is, in the embodiments of the present application, based on the DIMD technology, the intra prediction mode can be derived on the decoding side using the same method as the encoding side, thereby saving bit overhead. This mainly includes two steps. In step 1, the prediction mode is derived, and the same method for calculating the prediction mode strength is used on the decoding side and the encoding side. For example, the Sobel operator is used to calculate the gradient histogram for each prediction mode. The region of interest consists of the reconstructed samples in the three adjacent rows above the current block, the three adjacent columns on the left of the current block, and the corresponding adjacent reconstructed samples in the upper left. By calculating the gradient histogram within the L-shaped region, the first prediction mode corresponding to the largest amplitude and the second prediction mode corresponding to the second largest amplitude in the gradient histogram can be obtained. In step 2, the prediction block is derived, and the same prediction block derivation method is used on the decoding side and the encoding side to obtain the prediction block of the current block. For example, the following two conditions are judged. Condition 1: The gradient value of the second prediction mode is not 0. Condition 2: Neither the first prediction mode nor the second prediction mode is the PLANAR mode or the DC prediction mode. If the above two conditions do not hold simultaneously, the prediction sample value of the current block is calculated using only the first prediction mode, that is, the normal prediction process is applied to the first prediction mode. Otherwise, that is, when the above two conditions hold simultaneously, the weighted average method is used to derive the current prediction block. The specific method is as follows. The PLANAR mode occupies a weight of 1 / 3. The weight of the first prediction mode is 2 / 3 multiplied by the ratio of the gradient strength of the first prediction mode to the sum of the gradient strengths of the first prediction mode and the second prediction mode, and the weight of the second prediction mode is 2 / 3 multiplied by the ratio of the gradient strength of the second prediction mode to the sum of the gradient strengths of the first prediction mode and the second prediction mode. Weighted averaging is performed on the above three prediction modes, that is, the PLANAR mode, the first prediction mode, and the second prediction mode, to obtain the prediction block of the current coding unit. The decoder also obtains the prediction block in the same steps.
[0089] Furthermore, in the embodiments of the present application, first, based on the MIP parameters, the downsampling vector of the current block can be determined. Next, based on the downsampling vector, a matrix multiplication operation is performed to obtain the MIP output vector. Furthermore, based on the MIP output vector, the MIP prediction block of the current block is determined. Finally, at least one intra prediction mode is traversed for the MIP prediction block to obtain at least one gradient information.
[0090] As can be understood, in the embodiments of the present application, based on the size parameter of the current block, half downsampling can be performed on the obtained surrounding reference reconstruction samples. The sampling step size is determined by the size parameter of the current block. The splicing order of the reference reconstruction samples after upsampling and the reference reconstruction samples after left downsampling is adjusted according to the MIP transpose instruction parameter obtained by decoding. When transposition is not required, after the reference reconstruction samples after left downsampling are spliced after the reference reconstruction samples after upsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upsampling are spliced after the reference reconstruction samples after left downsampling, the obtained vector is used as the input (downsampling vector).
[0091] Next, based on the MIP mode index obtained by decoding, the MIP matrix coefficients are acquired, and an output vector (MIP output vector) can be obtained by calculating the MIP matrix coefficients and the input (downsampling vector). Then, based on the number of samples of the output vector and the size parameter of the current block, the output vector is upsampled. When upsampling is not required, the vectors are arranged in order horizontally and output as the MIP prediction block of the current block. When upsampling is necessary, first, upsampling is performed horizontally, then downsampling is performed vertically, upsampling is continued until it reaches the same size as the template, and the MIP prediction block of the current block is output.
[0092] Next, when it is determined that DIMD is used for the current block, the optimal conventional intra prediction mode can be directly derived as the mapping mode of the LFNST transform set for the MIP prediction block of the current block using the DIMD method. That is, at least one intra prediction mode of the MIP prediction block of the current block is traversed to calculate the gradient information of at least one intra prediction mode in the MIP prediction block of the current block.
[0093] That is, in the embodiment of the present application, the DIMD calculation process may be executed after the MIP output vector is upsampled.
[0094] Furthermore, in the embodiment of the present application, first, based on the MIP parameter, the downsampling vector of the current block can be determined. Next, a matrix multiplication operation is executed based on the downsampling vector to obtain the MIP output vector. Finally, at least one intra prediction mode is traversed for the MIP output vector to obtain at least one gradient information.
[0095] For better understanding, in the embodiments of the present application, based on the size parameter of the current block, half-downsampling can be performed on the obtained surrounding reference reconstruction samples. The sampling step size is determined by the size parameter of the current block. The splicing order of the reference reconstruction samples after upper downsampling and the reference reconstruction samples after left downsampling is adjusted according to the MIP transposition instruction parameter obtained by decoding. When transposition is not required, after the reference reconstruction samples after left downsampling are spliced after the reference reconstruction samples after upper downsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upper downsampling are spliced after the reference reconstruction samples after left downsampling, the obtained vector is used as the input (downsampling vector).
[0096] Next, based on the MIP mode index obtained by decoding, the MIP matrix coefficient can be obtained, and the output vector (MIP output vector) can be obtained by calculating the MIP matrix coefficient and the input (downsampling vector).
[0097] Next, when it is determined that the current block DIMD is to be used, the optimal conventional intra prediction mode can be directly derived as the mapping mode of the LFNST transform set for the MIP output vector of the current block using the DIMD method. That is, at least one intra prediction mode is traversed for the MIP output vector to calculate the gradient information of at least one intra prediction mode in the MIP prediction block of the current block. Then, based on the number of samples of the output vector (MIP output vector) and the size parameter of the current block, the output vector is upsampled. If upsampling is not required, the vectors are arranged in order horizontally and output as the MIP prediction block of the current block. If upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, continue upsampling until it reaches the same size as the template, and output the MIP prediction block of the current block.
[0098] That is, in the embodiment of the present application, the DIMD calculation process may be executed before the MIP output vector is upsampled.
[0099] Note that in the embodiment of the present application, when determining the mapping mode of the LFNST transform set using DIMD, all 67 intra prediction modes can be traversed to obtain the corresponding 67 gradient information. It is also possible to traverse a part of the 67 intra prediction modes to obtain the corresponding gradient information.
[0100] That is, in the embodiment of the present application, in order to further reduce the complexity on the encoding side and the decoding side, for the use of DIMD, 67 intra prediction modes can be selectively skipped, thereby reducing the number of intra prediction modes to be traversed. For example, the selection can be executed by the step size being 1.
[0101] Step 105: Based on the mapping mode of the LFNST conversion set, select one LFNST conversion kernel candidate set from multiple LFNST conversion kernel candidate sets, and determine the LFNST conversion kernel used for the current block from the selected LFNST conversion kernel candidate set.
[0102] In the embodiment of the present application, when the LFNST index indicates that LFNST is used for the current block, after determining the mapping mode of the LFNST conversion set based on the MIP parameter, based on the mapping mode of the LFNST conversion set, one LFNST conversion kernel candidate set can be selected from multiple LFNST conversion kernel candidate sets, and the LFNST conversion kernel used for the current block can be determined from the selected LFNST conversion kernel candidate set.
[0103] Furthermore, in the embodiment of the present application, when determining the LFNST conversion kernel used for the current block, first, the index of the mapping mode of the LFNST conversion set can be determined, then, based on the value of the index, the value of the LFNST intra prediction mode index can be determined, further, based on the value of the LFNST intra prediction mode index, one LFNST conversion kernel candidate set can be selected from multiple LFNST conversion kernel candidate sets, and finally, the conversion kernel indicated by the LFNST index can be selected from the selected LFNST conversion kernel candidate set and set as the LFNST conversion kernel used for the current block.
[0104] That is, in the embodiments of the present application, after determining the mapping mode of the LFNST transform set, the index of the mapping mode of the LFNST transform set is further determined. Next, the value of the index of the mapping mode of the LFNST transform set can be converted into the value of the LFNST intra prediction mode index (which can be represented by predModeIntra). Then, based on the value of predModeIntra, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets to determine the transform kernel candidate set. And from the selected LFNST transform kernel candidate set, the transform kernel indicated by the LFNST index is selected and set as the LFNST transform kernel used for the current block.
[0105] Note that in the embodiments of the present application, when determining the LFNST transform kernel candidate set and the corresponding LFNST transform kernel, the plurality of LFNST transform kernel candidate sets can include four LFNST transform kernel candidate sets, and each LFNST transform kernel candidate set can include two LFNST transform kernels. Accordingly, the value of the LFNST intra prediction mode index corresponding to the value of the index can be determined using the first look-up table.
[0106] As can be understood, in the embodiments of the present application, based on the first look-up table, the DC mode, the PLANAR mode, or the angular prediction mode can be bound to the LFNST transform set. For example, the first look-up table shown in Table 1 can be cited.
[0107] In addition, in the embodiments of the present application, when determining the LFNST transformation kernel candidate set and the corresponding LFNST transformation kernel, the multiple LFNST transformation kernel candidate sets can include 35 LFNST transformation kernel candidate sets, and each LFNST transformation kernel candidate set includes 3 LFNST transformation kernels. Accordingly, using the second look-up table, the value of the LFNST intra prediction mode index corresponding to the value of the index can be determined.
[0108] As can be understood, in the embodiments of the present application, based on the second look-up table, the LFNST transformation sets corresponding to different intra prediction modes become finer. For example, the second look-up table shown in FIG. 2 can be cited.
[0109] Furthermore, in the embodiments of the present application, the MIP parameter may further include an MIP transposition instruction parameter, and the value of the MIP transposition instruction parameter is used to indicate whether to transpose the sample input vector used in the MIP mode. Accordingly, when the value of the MIP transposition instruction parameter indicates to transpose the sample input vector used in the MIP mode, a matrix transposition can be performed on the transformation kernel indicated by the LFNST index to obtain the LFNST transformation kernel used for the current block.
[0110] As can be understood, in the embodiments of the present application, when the value of the MIP transposition instruction parameter is equal to 1, it can be considered that the value of the MIP transposition instruction parameter indicates to transpose the sample input vector used in the MIP mode. In this case, it is necessary to perform the corresponding matrix transposition on the selected transformation kernel, whereby the LFNST transformation kernel used for the current block can be obtained.
[0111] Step 106: Use the LFNST transformation kernel to transform the transformation coefficient.
[0112] In the embodiments of the present application, based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and after determining the LFNST transform kernel to be used for the current block from the selected LFNST transform kernel candidate set, the LFNST transform kernel can be used to transform the transform coefficients.
[0113] Furthermore, in the embodiments of the present application, the LFNST transform kernel determined from the selected LFNST transform kernel candidate set is the LFNST transform kernel to be used for the current block, and the LFNST transform kernel may be a transformation matrix for transforming the transform coefficients. Then, with the secondary transform coefficient vector as the input, the primary transform coefficient vector can be obtained by multiplying the input by the transformation matrix (transform kernel). In this way, after the matrix multiplication operation, the transformation of the transform coefficients can be realized.
[0114] Exemplarily, in one possible embodiment, on the decoding side, the decoder can decode the CU-level type flag. If the CU-level type flag indicates the intra prediction mode, the decoder obtains the MIP usage permission flag (prediction mode parameter) by decoding. The MIP usage permission flag may be a sequence-level flag and is currently used to indicate whether the use of the MIP technology is permitted for the decoder. The sequence-level flag can be represented in the form of sps_mip_enable_flag.
[0115] Next, if the MIP usage permission flag is true, the MIP usage flag of the current coding unit (current block) is decoded. Otherwise, in the current decoding process, it is not necessary to decode the CU-level MIP usage flag, and the CU-level MIP usage flag defaults to false.
[0116] If the MIP usage flag of the current coding unit is true, obtain the MIP parameters of the current coding unit during decoding. The MIP parameters can include at least one of information such as the MIP transposition instruction parameter, the MIP mode index, the size of the current block, and the type of the current block. Otherwise, continue to decode information such as the usage flag or index of other intra prediction techniques, and based on the decoded information, obtain the final prediction block of the current coding unit.
[0117] After obtaining the MIP parameters during decoding, half-downsampling can be performed on the obtained surrounding reference reconstruction samples based on the size of the current coding unit (the size parameter of the current block). The sampling step size is determined by the size of the coding unit. Also, in combination with the MIP transposition instruction parameter obtained during decoding, adjust the splicing order of the reference reconstruction samples after upper-side downsampling and the reference reconstruction samples after left-side downsampling. When transposition is not required, after the reference reconstruction samples after left-side downsampling are spliced after the reference reconstruction samples after upper-side downsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upper-side downsampling are spliced after the reference reconstruction samples after left-side downsampling, the obtained vector is used as the input (downsampling vector).
[0118] Next, obtain MIP matrix coefficients based on the MIP prediction mode obtained by decoding, and obtain an output vector (MIP output vector) by calculating the MIP matrix coefficients and the input (downsampling vector). Then, upsample the output vector based on the number of samples of the output vector and the size parameter of the current coding unit. If upsampling is not required, arrange the vectors in order horizontally and output them as the prediction block of the current coding unit (MIP prediction block of the current block). If upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, perform upsampling until it reaches the same size as the template, and output the prediction block of the current coding unit (MIP prediction block of the current block).
[0119] Note that when the size parameter of the current block satisfies the first preset condition, directly determine the first preset prediction mode as the mapping mode of the LFNST transform set. For example, when both the width and height of the current coding unit are 32 or more, the PLANAR mode (the first preset prediction mode) can be used as the mapping mode of the LFNST transform set. When the size parameter of the current block satisfies the first preset condition, an optimal conventional intra prediction mode can be derived as the mapping mode of the LFNST transform set for the current MIP prediction block using the DIMD method.
[0120] Furthermore, when deriving a conventional intra prediction mode as a mapping mode of the LFNST transform set using the DIMD method, traverse 67 intra prediction modes (or some of the intra prediction modes) in the current VVC and ECM for the current MIP prediction block, calculate the gradient information of each intra prediction mode in the current MIP prediction block, and further, based on the gradient information, the corresponding gradient amplitude value can be determined. Next, sort the traversed intra prediction modes based on the gradient amplitude value. The intra prediction mode with the maximum amplitude value is the optimal mode, that is, it is used as the mapping mode of the LFNST transform set in the inverse transform process of the subsequent steps.
[0121] After completing the determination of the mapping mode of the LFNST transform set, continue to decode information such as the usage flag or index of other intra prediction techniques, and based on the decoded information, obtain the final prediction block of the current coding unit. Furthermore, decode the bitstream to obtain residual information, obtain the temporal domain residual information through inverse quantization and inverse transformation, and add the final prediction block and the temporal domain residual information to obtain a reconstructed sample block. Perform techniques such as in-loop filtering on all reconstructed sample blocks to obtain the final reconstructed image. The final reconstructed image may be used as a video output or as a reference for subsequent decoding.
[0122] Exemplarily, in other possible embodiments, on the decoding side, the decoder can decode the CU-level type flag. If the CU-level type flag indicates an intra prediction mode, the decoder obtains a MIP usage permission flag (prediction mode parameter) through decoding. The MIP usage permission flag may be a sequence-level flag and is currently used to indicate whether the decoder is permitted to use the MIP technique. The sequence-level flag can be represented in the form of sps_mip_enable_flag.
[0123] Next, if the MIP usage permission flag is true, decode the MIP usage flag of the current coding unit (current block). Otherwise, in the current decoding process, there is no need to decode the CU-level MIP usage flag, and the CU-level MIP usage flag defaults to false.
[0124] If the MIP usage flag of the current coding unit is true, obtain the MIP parameters of the current coding unit during decoding. The MIP parameters can include at least one of information such as the MIP transpose instruction parameter, the MIP mode index, the size of the current block, and the type of the current block. Otherwise, continue to decode information such as the usage flag or index of other intra prediction techniques, and based on the decoded information, obtain the final prediction block of the current coding unit.
[0125] After obtaining the MIP parameters during decoding, half-downsampling can be performed on the obtained surrounding reference reconstruction samples based on the size of the current coding unit (the size parameter of the current block). The sampling step size is determined by the size of the coding unit. Also, in combination with the MIP transpose instruction parameter obtained during decoding, adjust the splicing order of the reference reconstruction samples after upper downsampling and the reference reconstruction samples after left downsampling. When transposition is not required, after the reference reconstruction samples after left downsampling are spliced after the reference reconstruction samples after upper downsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upper downsampling are spliced after the reference reconstruction samples after left downsampling, the obtained vector is used as the input (downsampling vector).
[0126] Next, obtain MIP matrix coefficients based on the MIP prediction mode obtained by decoding, and obtain an output vector (MIP output vector) by calculating the MIP matrix coefficients and the input (downsampling vector).
[0127] If the size parameter of the current block satisfies the first preset condition, directly determine the first preset prediction mode as the mapping mode of the LFNST transform set. For example, when both the width and height of the current coding unit are 32 or more, the PLANAR mode (the first preset prediction mode) can be used as the mapping mode of the LFNST transform set. If the size parameter of the current block satisfies the first preset condition, an optimal conventional intra prediction mode can be derived as the mapping mode of the LFNST transform set for the current MIP output vector using the DIMD method.
[0128] Furthermore, when deriving a conventional intra prediction mode as the mapping mode of the LFNST transform set using the DIMD method, traverse 67 intra prediction modes (or some of the intra prediction modes) in the current VVC and ECM for the current MIP output vector (MIP output vector), calculate the gradient information of each intra prediction mode in the current MIP output vector, and further determine the corresponding gradient amplitude value based on the gradient information. Next, sort the intra prediction modes traversed based on the gradient amplitude value. The intra prediction mode with the maximum amplitude value is the optimal mode, that is, it is used as the mapping mode of the LFNST transform set in the inverse transform process of the subsequent steps.
[0129] Furthermore, based on the number of samples of the output vector and the size parameter of the current coding unit, the output vector can be upsampled. If upsampling is not required, the vectors are arranged in order horizontally and output as the prediction block of the current coding unit (the MIP prediction block of the current block). If upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, continue upsampling until it reaches the same size as the template, and output the prediction block of the current coding unit (the MIP prediction block of the current block).
[0130] After completing the determination of the mapping mode of the LFNST transform set, continue to decode information such as the usage flag or index of other intra prediction techniques, and obtain the final prediction block of the current coding unit based on the decoded information. Furthermore, decode the bitstream to obtain residual information, obtain the time-domain residual information through inverse quantization and inverse transformation, add the final prediction block and the time-domain residual information to obtain a reconstructed sample block. Perform techniques such as in-loop filtering on all reconstructed sample blocks to obtain the final reconstructed image. The final reconstructed image may be used as a video output or as a reference for subsequent decoding.
[0131] Note that the coding method according to the embodiments of the present application can be applied to intra prediction on the encoding side and decoding. After integrating the technical solution of the embodiments of the present application into JVET-Z008, the test results under normal test condition AI are shown as in Table 2 and Table 3.
[0132]
Table 2
[0133]
Table 3
[0134] By using the technical solution according to the embodiment of the present application and performing both anchor and test in class D, relatively accurate results can be provided. The results are shown as in Tables 4 and 5.
[0135]
Table 4
[0136]
Table 5
[0137] Using the technical solution according to the embodiment of the present application, the test results with ECM4.0 as the anchor (of the method provided by JVET-Z0048+) are shown as in Table 6.
[0138]
Table 6
[0139] As can be seen from the above test results, according to the decoding method according to the embodiment of the present application, while reducing the software and hardware complexity in JVET-Z0048, the performance similar to that of JVET-Z0048 is maintained. For example, there is no change in the performance of the luminance component. Compared with ECM4.0, the same performance as JVET-Z0048 is maintained.
[0140] Furthermore, in the embodiments of the present application, considering that the hardware decoder has different requirements for I-frames and B-frames, the decoding method according to the embodiments of the present application may be used only for B-frames or may be used for both I-frames and B-frames. Alternatively, the decoding method according to the embodiments of the present application may be used only for I-frames. Or, the conditions under which the use of the decoding method according to the embodiments of the present application is permitted are different for B-frames or I-frames. For example, in the case of I-frames, it is permitted to use the decoding method according to the embodiments of the present application for coding units of all sizes, and in the case of B-frames, it is permitted to use the decoding method according to the embodiments of the present application only for coding units of small sizes.
[0141] Furthermore, in the embodiments of the present application, when the MIP prediction mode is used for the luminance component of the current coding unit (current block) and the MIP prediction mode is not used for the chrominance component and the conventional intra prediction mode is not used, the LFNST transform set of the chrominance component may inherit the LFNST transform set of the luminance component.
[0142] Furthermore, in the embodiments of the present application, when the conventional intra prediction mode is not used for the current coding unit (current block), the LFNST transform sets of both the luminance component and the chrominance component can be obtained by the decoding method according to the embodiments of the present application.
[0143] As described above, the decoding method according to the embodiment of the present application relates to a method of deriving a MIP prediction block using DIMD and a method of mapping an LFNST transform set. On the other hand, it is proposed to limit the size of the coding unit in which DIMD is used. For an image block with a relatively large size, the upsampling of the MIP output vector is relatively complicated and the direction information is not clear. Therefore, the computational complexity is reduced by skipping the process of deriving the conventional intra prediction mode using DIMD. On the other hand, after proposing to limit the size of the coding unit in which DIMD is used, the computational complexity is further reduced, the MIP output vector before upsampling is used as the input of DIMD, and the optimal conventional intra prediction mode is derived.
[0144] Embodiments of the present application provide a decoding method. On the decoding side, a bitstream is decoded to determine a prediction mode parameter. When the prediction mode parameter indicates that MIP is used to determine an intra prediction value, the bitstream is decoded to determine the MIP parameter of the current block. The bitstream is decoded to determine the transform coefficient and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The transform coefficient is transformed using the LFNST transform kernel. As can be seen from the above, in the embodiments of the present application, when deriving a prediction block, based on the size parameter in the MIP parameter of the current block, the mapping mode of the LFNST transform set is determined. For an image block with a relatively large size, it may not be necessary to use DIMD to derive the mapping mode. Thereby, the computational complexity can be reduced, and thus the coding efficiency can be improved.
[0145] Based on the above embodiments, the present application provides an encoding method. FIG. 7 is a flowchart of the encoding method according to the embodiments of the present application. As shown in FIG. 7, the encoding method executed by the encoder may include the following steps. Step 201: Determine a prediction mode parameter.
[0146] In the embodiments of the present application, the encoder can first determine a prediction mode parameter.
[0147] In the embodiments of the present application, a video image can be divided into a plurality of image blocks. Each image block waiting for coding currently may be referred to as a coding block (CB). Each coding block may include a first color component, a second color component, and a third color component. The current block is a coding block waiting for prediction of the first color component, the second color component, or the third color component in the video image.
[0148] If a prediction of the first color component is performed on the current block and the first color component is a luminance component, that is, assuming that the color component waiting for prediction is a luminance component, the current block can be called a luminance block. Also, if a prediction of the second color component is performed on the current block and the second color component is a chroma component, that is, assuming that the color component waiting for prediction is a chroma component, the current block can be called a chroma block.
[0149] Note that the prediction mode parameter indicates the coding mode of the current block and the parameters related to that mode. Usually, the prediction mode parameter of the current block can be determined using rate distortion optimization (RDO).
[0150] Exemplarily, in the embodiments of the present application, when determining the prediction mode parameter of the current block, first, the color component waiting for prediction of the current block can be determined. Next, based on the parameters of the current block, predictive coding is performed on the color component waiting for prediction using a plurality of prediction modes, and the rate distortion cost results corresponding to each of the plurality of prediction modes are calculated. Finally, the minimum rate distortion cost result can be selected from the calculated plurality of rate distortion cost results, and the prediction mode corresponding to the minimum rate distortion cost result can be determined as the prediction mode parameter of the current block.
[0151] That is, in the embodiment of the present application, on the encoding side, for the current block, a plurality of prediction modes can be used to encode the color components waiting for prediction respectively. The plurality of prediction modes usually include a conventional intra prediction mode and a non-conventional intra prediction mode. The conventional intra prediction mode can include a direct current (DC) mode, a planar (PLANAR) mode, an angular mode, etc., and the non-conventional intra prediction mode can include a MIP mode, a cross-component linear model prediction (CCLM) mode, an intra block copy (IBC) mode, a palette (PLT) mode, etc.
[0152] In this way, after encoding the current block using a plurality of prediction modes respectively, rate-distortion cost results corresponding to each prediction mode can be obtained. Next, the minimum rate-distortion cost result is selected from the obtained plurality of rate-distortion cost results, and the prediction mode corresponding to the minimum rate-distortion cost result is determined as the prediction mode parameter of the current block. In this way, finally, the current block can be encoded using the determined prediction mode, and in this prediction mode, the residual can be reduced, thereby improving the encoding efficiency.
[0153] As can be understood, in the embodiment of the present application, on the encoding side, predictive encoding can be performed on the current block, and during the predictive encoding, the prediction mode of the current block can be determined, and the corresponding prediction mode parameter can be signaled to the bitstream. Thereby, the prediction mode parameter can be transmitted from the encoder to the decoder.
[0154] Accordingly, on the encoding side, by decoding the bitstream, the intra prediction mode of the luminance component or the chrominance component of the current block, or the coding block where the current block is located, can be obtained. In this case, the value of predModeIntra (intra prediction mode indicator) can be determined.
[0155] Furthermore, in the embodiments of the present application, by obtaining the prediction mode parameters, it is possible to determine whether MIP is used to determine the intra prediction value during intra prediction based on the prediction mode parameters.
[0156] Step 202: When the prediction mode parameter indicates that MIP is used for the current block to determine the intra prediction value, determine the MIP parameter of the current block.
[0157] In the embodiments of the present application, after determining the prediction mode parameters, when the prediction mode parameter indicates that MIP is used to determine the intra prediction value, the MIP parameter of the current block can continue to be determined.
[0158] Note that in the embodiments of the present application, the MIP parameter can include a MIP transpose indication parameter (which can be represented by isTransposed), a MIP mode index (which can be represented by modelId), the size of the current block, and the type of the current block (which can be represented by mipSizeId).
[0159] That is, in the embodiments of the present application, at least one of the information such as the MIP transpose indication parameter, the MIP mode index, the size of the current block, and the type of the current block can be indicated by the determined MIP parameter.
[0160] Furthermore, in the embodiments of the present application, the MIP parameter may include a MIP transpose indication parameter (which can be represented by isTransposed). The value of the MIP transpose indication parameter is used to indicate whether to transpose the sample input vector used in the MIP mode.
[0161] Specifically, in the MIP mode, an adjacent reference sample set can be obtained based on the reference sample value corresponding to the left adjacent reference sample of the current block and the reference sample value corresponding to the upper adjacent reference sample of the current block. In this way, after obtaining the adjacent reference sample set, an input reference sample set can be constructed, that is, a sample input vector used in the MIP mode can be constructed. However, the construction of the input reference sample set is different between the encoding side and the decoding side, which is mainly related to the value of the MIP transpose instruction parameter.
[0162] When applied to the encoding side, the value of the MIP transpose instruction parameter can be determined using rate distortion optimization. Specifically, the following content can be included. Calculate the first cost value when transposition is performed and the second cost value when transposition is not performed, respectively. If the first cost value is smaller than the second cost value, it can be determined that the value of the MIP transpose instruction parameter is equal to 1. If the first cost value is greater than or equal to the second cost value, it can be determined that the value of the MIP transpose instruction parameter is equal to 0.
[0163] Furthermore, when the value of the MIP transposition indication parameter is 0, in the buffer area, in the adjacent reference sample set, the upper adjacent reference sample value can be stored before storing the left adjacent reference sample value. In this case, there is no need to transpose, that is, there is no need to transpose the sample input vector used in the MIP mode, and the buffer area can be directly determined as the input reference sample set. When the value of the MIP transposition indication parameter is 1, in the buffer area, in the adjacent reference sample set, the upper adjacent reference sample value can be stored after storing the left adjacent reference sample value. In this case, the buffer area needs to be transposed, that is, the sample input vector used in the MIP mode needs to be transposed, and the transposed buffer area is determined as the input reference sample set. In this way, after obtaining the input reference sample set, it can be used in the process of determining the intra prediction value corresponding to the current block in the MIP mode.
[0164] Note that on the encoding side, after determining the value of the MIP transposition indication parameter, it is also necessary to signal the determined value of the MIP transposition indication parameter to the bitstream, thereby facilitating subsequent decoding on the decoding side.
[0165] Furthermore, in an embodiment of the present application, the MIP parameter may further include a MIP mode index (which can be represented by modeId). The MIP mode index is used to indicate the MIP mode used for the current block, and the MIP mode is used to indicate the calculation and derivation method for determining the intra prediction block of the current block using MIP. That is, in the MIP mode, since the MIP mode can include a plurality of MIP modes, these plurality of MIP modes can be distinguished by the MIP mode index, that is, different MIP modes have different MIP mode indices. In this way, based on the method of calculating and deriving to determine the intra prediction block of the current block using MIP, a specific MIP mode can be determined, and the corresponding MIP mode index can be obtained. In an embodiment of the present application, the value of the MIP mode index may be 0, 1, 2, 3, 4, or 5.
[0166] Furthermore, in an embodiment of the present application, the MIP parameter can include parameters such as the size and aspect ratio of the current block. Based on the size of the current block (i.e., the width and height of the current block), the type of the current block (which can be represented by mipSizeId) can also be determined.
[0167] For example, when both the width and height of the current block are equal to 4, the value of mipSizeId can be set to 0. Otherwise, when one of the width and height of the current block is equal to 4, or when both the width and height of the current block are equal to 8, the value of mipSizeId can be set to 1. Otherwise, when the current block is a block of other sizes, the value of mipSizeId can be set to 2.
[0168] For example, when both the width and height of the current block are equal to 4, the value of mipSizeId can be set to 0. If not, when one of the width and height of the current block is equal to 4, the value of mipSizeId can be set to 1. If not, when the current block is a block of other sizes, the value of mipSizeId can be set to 2.
[0169] In this way, in the process of determining the intra prediction value using MIP, the MIP parameters can also be determined. Thereby, based on the determined MIP parameters, it becomes easy to determine the LFNST transform kernel (which can be represented by kernel) used for the current block.
[0170] That is, in the embodiments of the present application, the MIP parameters may be used to determine the size parameters of the current block. The size parameters may represent the size of the current block, may be the height and width of the current block, or may be the aspect ratio of the current block.
[0171] Step 203: Based on the MIP parameters, determine the intra prediction block of the current block, and calculate the residual block obtained by subtracting the intra prediction value from the current block.
[0172] In the embodiments of the present application, when the prediction mode parameter indicates that MIP is used to determine the intra prediction value, after determining the MIP parameters of the current block, based on the MIP parameters, further determine the intra prediction block of the current block, and calculate the residual block obtained by subtracting the intra prediction value from the current block.
[0173] In the embodiments of the present application, for the MIP mode, the input data for MIP prediction includes the position (xTbCmp, yTbCmp) of the current block, the MIP prediction mode applied to the current block (which can be represented by modelId), the height of the current block (represented by nTbH), the width of the current block (represented by nTbW), and a transposition indication flag indicating whether transposition is required (which can be represented by isTransposed). The output data of MIP prediction includes the predicted block of the current block. The intra prediction value corresponding to the sample coordinates [x][y] in the predicted block is predSamples[x][y], where x = 0, 1, … nTbW - 1 and y = 0, 1, … nTbH - 1.
[0174] As can be understood, in the embodiments of the present application, the MIP prediction process can be divided into four steps: core parameter configuration, acquisition of reference samples, construction of input samples, and generation of predicted values. Regarding the core parameter configuration, based on the size of the current block in the image, the current block can be classified into three types, and the type of the current block is recorded by mipSizeId. When the types of the current blocks are different, the number of reference samples and the number of matrix multiplication output samples (also referred to as matrix multiplication output samples) are also different. Regarding the acquisition of reference samples, when predicting the current block, both the upper block and the left block of the current block are encoded blocks. The original reference samples of the MIP technology are the reconstructed values of the samples in the row above and the column to the left of the current block. The process of obtaining the upper adjacent reference sample (represented by refT) and the left adjacent reference sample (represented by refL) of the current block is the process of obtaining reference samples. Regarding the construction of input samples, this step is used for the input of matrix multiplication and mainly includes the process of obtaining reference samples, the process of constructing a reference sample buffer, and the process of deriving the input samples of matrix multiplication. The process of obtaining reference samples is a downsampling process, and the construction of the reference sample buffer can include the padding method when transposition is not required and the padding method when transposition is required. Regarding the generation of predicted values, this step is used to obtain the MIP predicted value of the current block and mainly includes the process of constructing a matrix multiplication output sample block, the process of clipping the matrix multiplication output samples, the process of transposing the matrix multiplication output samples, and the process of generating the MIP final predicted value. The process of constructing a matrix multiplication output sample block can include the process of obtaining a weighting matrix, the process of obtaining a shift factor and an offset factor, and the process of performing a matrix multiplication operation.The process of generating the MIP final prediction value includes a process of generating a prediction value that does not require upsampling and a process of generating a prediction value that requires upsampling. Thus, after the above four steps, the intra prediction block of the current block can be obtained.
[0175] As can be understood, in the embodiments of the present application, after determining the intra prediction block of the current block, a difference can be obtained by subtracting the intra prediction value from the actual value of the samples of the current block, and the obtained difference can be used as the residual block. Thereby, it becomes easier to transform the subsequent residual block.
[0176] Step 204: When LFNST is used for the current block, determine the mapping mode of the LFNST transform set based on the MIP parameter.
[0177] In the embodiments of the present application, when it is determined that LFNST is used for the current block, the mapping mode of the LFNST transform set can be further determined based on the MIP parameter.
[0178] Note that in the embodiments of the present application, it is not possible to perform LFNST on any current block. In one possible embodiment, LFNST can be performed on the current block only when the current block simultaneously satisfies the following conditions. These conditions include the following. (a) Both the width and height of the current block are 4 or more. (b) Both the width and height of the current block are less than or equal to the maximum size of the transform block. (c) The prediction mode of the current block or the coding block in which the current block is located is the intra prediction mode. (d) The primary transforms of the current block in both the horizontal and vertical directions are both two-dimensional forward primary transforms (for example, two-dimensional discrete cosine transform (DCT2)). (e) The intra prediction mode of the current block or the coding block in which the current block is located is the non-MIP mode, or the prediction mode of the transform block is the MIP mode and both the width and height of the transform block are 16 or more.
[0179] Furthermore, when it is determined that LFNST can be performed on the current block, it is also necessary to determine the LFNST transform kernel (which can be represented by kernel) used for the current block.
[0180] Note that in the embodiments of the present application, the MIP parameter may be the size parameter of the current block. The size parameter may represent the size of the current block, may be the height and width of the current block, or may be the aspect ratio of the current block.
[0181] That is, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set, the size parameter of the current block may be referred to. For example, based on the height and width of the current block, the mapping mode of the LFNST transform set is determined, or based on the aspect ratio of the current block, the mapping mode of the LFNST transform set is determined.
[0182] Furthermore, in the embodiments of the present application, when determining the mapping mode of the LFNST conversion set based on the size parameter of the current block, first, it can be determined whether the size parameter satisfies a first preset condition. If the size parameter satisfies the first preset condition, the first preset prediction mode can be determined as the mapping mode of the LFNST conversion set. If the size parameter does not satisfy the first preset condition, DIMD is used to determine the mapping mode of the LFNST conversion set.
[0183] Note that in the embodiments of the present application, the first preset condition may be used to limit the size of the current block. The first preset condition corresponds to the size parameter of the current block. When the size parameter of the current block is the height and width of the current block, the first preset condition can be used to limit the height and width respectively. When the size parameter of the current block is the aspect ratio of the current block, the first preset condition can be used to limit the aspect ratio.
[0184] Exemplarily, in the embodiments of the present application, assuming that the size parameter of the current block is the height and width of the current block, the first preset condition can be set such that the width is greater than or equal to a preset width threshold, and / or the height is greater than or equal to a preset height threshold. For example, when the height of the current block is greater than or equal to the preset height threshold, or the width of the current block is greater than or equal to the preset width threshold, it can be determined that the size parameter satisfies the first preset condition. When the height of the current block is less than the preset height threshold and the width of the current block is less than the preset width threshold, it can be determined that the size parameter does not satisfy the first preset condition.
[0185] As can be understood, in the embodiments of the present application, both the preset width threshold and the preset height threshold may be any value greater than or equal to 0. For example, the preset width threshold is 32, and the preset height threshold is also 32. When the height or width of the current block is 32 or more, it can be determined that the current block satisfies the first preset condition. The preset width threshold is 32, and the preset height threshold is 16. When the height of the current block is 32 or more, or the width of the current block is 16 or more, it can be determined that the current block satisfies the first preset condition.
[0186] As can be understood, in the embodiments of the present application, the use of DIMD can be restricted under the first preset condition. That is, DIMD is only permitted to be used to determine the mapping mode of the LFNST transform set when the size parameter of the current block does not satisfy the first preset condition.
[0187] Thus, according to the encoding method according to the embodiments of the present application, when determining the mapping mode of the LFNST transform set, the size of the image block for which DIMD is used can be restricted by using the first preset condition, and by permitting the use of DIMD only for a part of the image block, the computational complexity can be effectively reduced. For example, DIMD is permitted to be used to determine the mapping mode of the LFNST transform set only for image blocks of a relatively small size.
[0188] Furthermore, in the embodiments of the present application, when the size parameter of the current block satisfies the first preset condition, the first preset prediction mode may be directly determined as the mapping mode of the LFNST transform set. The first preset prediction mode may be the PLANAR mode or the DC mode.
[0189] That is, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set, in combination with the first preset condition, the mapping mode can be directly set for some image blocks. For example, for image blocks with a relatively large size, the PLANAR mode and the DC mode are directly determined as the mapping modes of the LFNST transform set.
[0190] Furthermore, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set using DIMD, first, at least one intra prediction mode can be traversed, whereby at least one gradient information corresponding to the current block can be determined. Next, the mapping mode of the LFNST transform set may be determined based on the at least one gradient information. One intra prediction mode corresponds to one gradient information, and the gradient information may be a gradient histogram.
[0191] As can be understood, in the embodiments of the present application, when determining the mapping mode of the LFNST transform set based on at least one gradient information, the gradient amplitude value corresponding to each intra prediction mode can be determined based on the at least one gradient information. Next, among the at least one intra prediction mode, the intra prediction mode with the maximum gradient amplitude value can be determined as the mapping mode of the LFNST transform set. One intra prediction mode corresponds to one gradient amplitude value.
[0192] That is, in the embodiments of the present application, based on the DIMD technology, the intra prediction mode can be derived on the decoding side using the same method as the encoding side, thereby saving bit overhead. This mainly includes two steps. In step 1, the prediction mode is derived, and the same calculation method for the prediction mode strength is used on the decoding side and the encoding side. For example, the Sobel operator is used to calculate the gradient histogram for each prediction mode. The region of interest consists of the reconstructed samples in the three adjacent rows above the current block, the reconstructed samples in the three adjacent columns to the left, and the corresponding adjacent reconstructed samples in the upper left, forming an L-shaped region. By calculating the gradient histogram within the L-shaped region, the first prediction mode corresponding to the largest amplitude and the second prediction mode corresponding to the second largest amplitude in the gradient histogram can be obtained. In step 2, the prediction block is derived, and the same prediction block derivation method is used on the decoding side and the encoding side to obtain the prediction block of the current block. For example, the following two conditions are judged. Condition 1: The gradient value of the second prediction mode is not 0. Condition 2: Neither the first prediction mode nor the second prediction mode is the PLANAR mode or the DC prediction mode. If the above two conditions do not hold simultaneously, the prediction sample value of the current block is calculated using only the first prediction mode, that is, the normal prediction process is applied to the first prediction mode. Otherwise, that is, when the above two conditions hold simultaneously, the weighted average method is used to derive the current prediction block. The specific method is as follows. The PLANAR mode occupies a weight of 1 / 3. The weight of the first prediction mode is 2 / 3 multiplied by the ratio of the gradient strength of the first prediction mode to the sum of the gradient strengths of the first prediction mode and the second prediction mode, and the weight of the second prediction mode is 2 / 3 multiplied by the ratio of the gradient strength of the second prediction mode to the sum of the gradient strengths of the first prediction mode and the second prediction mode. Weighted averaging is performed on the above three prediction modes, that is, the PLANAR mode, the first prediction mode, and the second prediction mode, to obtain the prediction block of the current coding unit. The decoder also obtains the prediction block in the same steps.
[0193] Furthermore, in the embodiments of the present application, first, based on the MIP parameters, the downsampling vector of the current block can be determined. Next, a matrix multiplication operation is performed based on the downsampling vector to obtain the MIP output vector. Furthermore, based on the MIP output vector, the MIP prediction block of the current block is determined. Finally, at least one intra prediction mode is traversed for the MIP prediction block to obtain at least one gradient information.
[0194] As can be understood, in the embodiments of the present application, based on the size parameter of the current block, half downsampling can be performed on the obtained surrounding reference reconstruction samples. The sampling step size is determined by the size parameter of the current block. The splicing order of the reference reconstruction samples after upsampling and the reference reconstruction samples after left downsampling is adjusted by the MIP transpose instruction parameter. When transposition is not required, after the reference reconstruction samples after left downsampling are spliced after the reference reconstruction samples after upsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upsampling are spliced after the reference reconstruction samples after left downsampling, the obtained vector is used as the input (downsampling vector).
[0195] Next, using the traversed prediction mode as an index, obtain the MIP matrix coefficients, and through the calculation of the MIP matrix coefficients and the input (downsampling vector), an output vector (MIP output vector) can be obtained. Then, based on the number of samples of the output vector and the size parameter of the current block, upsample the output vector. When upsampling is not required, arrange the vectors in order horizontally and output them as the MIP prediction block of the current block. When upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, continue upsampling until it reaches the same size as the template, and output the MIP prediction block of the current block.
[0196] Next, when it is determined that the current block DIMD is used, for the MIP prediction block of the current block, the optimal conventional intra prediction mode can be directly derived as the mapping mode of the LFNST transform set using the DIMD method. That is, traverse at least one intra prediction mode for the MIP prediction block of the current block and calculate the gradient information of at least one intra prediction mode in the MIP prediction block of the current block.
[0197] That is, in the embodiment of the present application, the DIMD calculation process may be executed after the MIP output vector is upsampled.
[0198] Furthermore, in the embodiment of the present application, first, based on the MIP parameters, the downsampling vector of the current block can be determined. Next, perform a matrix multiplication operation based on the downsampling vector to obtain the MIP output vector. Finally, traverse at least one intra prediction mode for the MIP output vector to obtain at least one gradient information.
[0199] As can be understood, in the embodiments of the present application, based on the size parameter of the current block, half-downsampling can be performed on the obtained surrounding reference reconstruction samples. The sampling step size is determined by the size parameter of the current block. The splicing order of the reference reconstruction samples after upsampling and the reference reconstruction samples after left-downsampling is adjusted by the MIP transposition instruction parameter. When transposition is not required, after the reference reconstruction samples after left-downsampling are spliced after the reference reconstruction samples after upsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upsampling are spliced after the reference reconstruction samples after left-downsampling, the obtained vector is used as the input (downsampling vector).
[0200] Next, the MIP matrix coefficients can be obtained using the traversed prediction mode as an index, and the output vector (MIP output vector) can be obtained by calculating the MIP matrix coefficients and the input (downsampling vector).
[0201] Next, when it is determined that the current block DIMD is to be used, the optimal conventional intra prediction mode can be directly derived as the mapping mode of the LFNST transform set for the MIP output vector of the current block using the DIMD method. That is, at least one intra prediction mode is traversed for the MIP output vector to calculate the gradient information of at least one intra prediction mode in the MIP prediction block of the current block. Then, based on the number of samples of the output vector (MIP output vector) and the size parameter of the current block, the output vector is upsampled. If upsampling is not required, the vectors are arranged in order horizontally and output as the MIP prediction block of the current block. If upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, continue upsampling until it reaches the same size as the template, and output the MIP prediction block of the current block.
[0202] That is, in the embodiments of the present application, the DIMD calculation process may be executed before the MIP output vector is upsampled.
[0203] Note that in the embodiments of the present application, when determining the mapping mode of the LFNST transform set using DIMD, all 67 intra prediction modes can be traversed to obtain the corresponding 67 gradient information. It is also possible to traverse a part of the 67 intra prediction modes to obtain the corresponding gradient information.
[0204] That is, in the embodiments of the present application, in order to further reduce the complexity on the encoding side and the decoding side, for the use of DIMD, 67 intra prediction modes can be selectively skipped, thereby reducing the number of intra prediction modes to be traversed. For example, the selection can be executed by the step size being 1.
[0205] Step 205: Based on the mapping mode of the LFNST transform set, select one set of LFNST transform kernel candidates from multiple sets of LFNST transform kernel candidates, determine the LFNST transform kernel used for the current block from the selected set of LFNST transform kernel candidates, set the LFNST index, and signal it to the video bitstream.
[0206] In the embodiments of the present application, when the LFNST is used for the current block, after determining the mapping mode of the LFNST transform set based on the MIP parameter, based on the mapping mode of the LFNST transform set, one set of LFNST transform kernel candidates can be selected from multiple sets of LFNST transform kernel candidates, and the LFNST transform kernel used for the current block can be determined from the selected set of LFNST transform kernel candidates. Thereby, the LFNST index can be set and the LFNST index can be signaled to the video stream.
[0207] Furthermore, in the embodiments of the present application, when determining the LFNST transform kernel used for the current block, first, the index of the mapping mode of the LFNST transform set can be determined, then, based on the value of the index, the value of the LFNST intra prediction mode index can be determined, and further, based on the value of the LFNST intra prediction mode index, one set of LFNST transform kernel candidates can be selected from multiple sets of LFNST transform kernel candidates, and finally, the LFNST transform kernel used for the current block can be selected from the selected set of LFNST transform kernel candidates. Thereby, the LFNST index can be set and the LFNST index can be signaled to the video stream.
[0208] That is, in the embodiments of the present application, after determining the mapping mode of the LFNST transform set, the index of the mapping mode of the LFNST transform set is further determined. Next, the value of the index of the mapping mode of the LFNST transform set can be converted into the value of the LFNST intra prediction mode index (which can be represented by predModeIntra). Then, based on the value of predModeIntra, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets to determine the transform kernel candidate set. And from the selected LFNST transform kernel candidate set, the LFNST transform kernel used for the current block is selected.
[0209] Note that in the embodiments of the present application, when determining the LFNST transform kernel candidate set and the corresponding LFNST transform kernel, the plurality of LFNST transform kernel candidate sets can include four LFNST transform kernel candidate sets, and each LFNST transform kernel candidate set can include two LFNST transform kernels. Accordingly, using the first look-up table, the value of the LFNST intra prediction mode index corresponding to the index value can be determined.
[0210] As can be understood, in the embodiments of the present application, based on the first look-up table, the DC mode, the PLANAR mode, or the angular prediction mode can be bound to the LFNST transform set. For example, the first look-up table shown in Table 1 can be cited.
[0211] Note that in the embodiments of the present application, when determining the LFNST transform kernel candidate set and the corresponding LFNST transform kernel, the plurality of LFNST transform kernel candidate sets can include thirty-five LFNST transform kernel candidate sets, and each LFNST transform kernel candidate set can include three LFNST transform kernels. Accordingly, using the second look-up table, the value of the LFNST intra prediction mode index corresponding to the index value can be determined.
[0212] For better understanding, in the embodiments of the present application, based on the second look-up table, the LFNST transform sets corresponding to different intra prediction modes become finer. For example, the second look-up table shown in FIG. 2 can be cited.
[0213] In addition, in the embodiments of the present application, the LFNST transform kernel can be understood as the transform matrix of LFNST, and is a matrix having a plurality of fixed coefficients obtained by training.
[0214] In addition, in the embodiments of the present application, since the LFNST transform kernel candidate set includes two or more preset transform kernels used for MIP, in this case, rate distortion optimization can be used to select the transform kernel used for the current block. Specifically, the rate-distortion cost (RDCost) of each transform kernel is calculated respectively using rate distortion optimization, and then the transform kernel with the minimum rate-distortion cost is selected as the transform kernel used for the current block.
[0215] That is, on the encoding side, based on the RDCost, a group of LFNST transform kernels can be selected, and the index corresponding to the LFNST transform kernel (which can be represented by lfnst_idx) is signaled to the video stream and transmitted to the decoding side. When selecting the first group of LFNST transform kernels (i.e., the first group of transform matrices) from the LFNST transform kernel candidate set, lfnst_idx is set to 1. When selecting the second group of LFNST transform kernels (i.e., the second group of transform matrices) from the LFNST transform kernel candidate set, lfnst_idx is set to 2.
[0216] Note that, in the embodiments of the present application, the value of the LFNST index may be used to indicate whether the LFNST is used for the current block, and may also be used to indicate the index of the LFNST conversion kernel in the LFNST conversion kernel candidate set.
[0217] That is, in the embodiments of the present application, regarding the value of the LFNST index (i.e., lfnst_idx), when the value of the LFNST index is equal to 0, the LFNST is not used. When the value of the LFNST index is greater than 0, the LFNST is used, and the index of the conversion kernel may be equal to the value of the LFNST index, or equal to the value obtained by subtracting 1 from the value of the LFNST index. In this way, based on the LFNST index, the LFNST conversion kernel used for the current block can be determined.
[0218] Furthermore, in the embodiments of the present application, the MIP parameter may further include a MIP transposition instruction parameter, and the value of the MIP transposition instruction parameter is used to indicate whether to transpose the sample input vector used in the MIP mode. Accordingly, when the value of the MIP transposition instruction parameter indicates to transpose the sample input vector used in the MIP mode, a matrix transposition can be performed on the selected conversion kernel to obtain the LFNST conversion kernel used for the current block.
[0219] As can be understood, in the embodiments of the present application, when the value of the MIP transposition instruction parameter is equal to 1, it can be considered that the value of the MIP transposition instruction parameter indicates to transpose the sample input vector used in the MIP mode. In this case, it is necessary to perform the corresponding matrix transposition on the selected conversion kernel, whereby the LFNST conversion kernel used for the current block can be obtained.
[0220] Step 206: Use the LFNST conversion kernel to convert the residual block.
[0221] In an embodiment of the present application, based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and after determining the LFNST transform kernel to be used for the current block from the selected LFNST transform kernel candidate set, the LFNST transform kernel is used, that is, the residual block is transformed using the transform matrix selected for the current block.
[0222] Exemplarily, in one possible embodiment, on the encoding side, the encoder traverses the prediction mode. When the current coding unit (current block) is in the intra prediction mode, the encoder obtains the usage permission flag in the coding method according to the embodiment of the present application, that is, obtains the MIP usage permission flag (prediction mode parameter). The MIP usage permission flag may be a sequence-level flag and is currently used to indicate whether the decoder is permitted to use the MIP technology. The sequence-level flag can be represented in the form of sps_mip_enable_flag.
[0223] Next, when the MIP usage permission flag is true, the encoder tries the MIP prediction method and calculates the corresponding rate-distortion cost, which is recorded as cost1. When the MIP usage permission flag is false, the encoder does not try the MIP prediction method, continues to traverse other intra prediction techniques, and calculates the corresponding rate-distortion costs, which are recorded as cost2...costN.
[0224] When the MIP usage permission flag is true, based on the size of the current coding unit (the size parameter of the current block), half-downsampling can be performed on the obtained surrounding reference reconstruction samples. The sampling step size is determined by the size of the coding unit. Also, in combination with the MIP transposition instruction parameter, the splicing order of the reference reconstruction samples after upsampling on the upper side and the reference reconstruction samples after downsampling on the left side can be adjusted. When transposition is not required, after the reference reconstruction samples after downsampling on the left side are spliced after the reference reconstruction samples after downsampling on the upper side, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after downsampling on the upper side are spliced after the reference reconstruction samples after downsampling on the left side, the obtained vector is used as the input (downsampling vector).
[0225] Next, the MIP matrix coefficients are obtained using the traversed prediction mode as an index, and an output vector (MIP output vector) is obtained by calculating the MIP matrix coefficients and the input (downsampling vector). Then, based on the number of samples in the output vector and the size parameter of the current coding unit, the output vector is upsampled. When upsampling is not required, the vectors are arranged in order horizontally and output as the prediction block of the current coding unit (the MIP prediction block of the current block). When upsampling is required, first, upsampling is performed horizontally, then downsampling is performed vertically, upsampling is continued until it reaches the same size as the template, and the prediction block of the current coding unit (the MIP prediction block of the current block) is output.
[0226] In addition, when the size parameter of the current block satisfies the first preset condition, the first preset prediction mode is directly determined as the mapping mode of the LFNST transform set. For example, when both the width and height of the current coding unit are 32 or more, the PLANAR mode (the first preset prediction mode) can be used as the mapping mode of the LFNST transform set. When the size parameter of the current block satisfies the first preset condition, for the current MIP prediction block, the optimal conventional intra prediction mode can be derived as the mapping mode of the LFNST transform set by using the DIMD method.
[0227] Furthermore, when deriving the conventional intra prediction mode as the mapping mode of the LFNST transform set by using the DIMD method, traverse the 67 intra prediction modes (or some of the intra prediction modes) in the current VVC and ECM for the current MIP prediction block, calculate the gradient information of each intra prediction mode in the current MIP prediction block, and further, based on the gradient information, the corresponding gradient amplitude value can be determined. Next, sort the intra prediction modes traversed based on the gradient amplitude value. The intra prediction mode with the maximum amplitude value is the optimal mode. That is, use the optimal mode to map the LFNST transform set of the current coding unit.
[0228] After completing the determination of the mapping mode of the LFNST transform set, subtract the prediction block (the MIP prediction block of the current block) from the original image block of the current coding unit to obtain the residual block of the current coding unit (the current block), perform a primary transform on the residual block to obtain the coefficient block in the frequency domain, and perform a secondary transform on the region of interest (ROI) of the coefficient block in the frequency domain by using LFNST. The mapping prediction mode of the LFNST transform set is determined by the above method. Then, through processes such as quantization, inverse quantization, and inverse transform, calculate the rate-distortion cost of the current coding unit and record it as cost1.
[0229] Furthermore, continue to traverse other intra prediction techniques and calculate the corresponding rate-distortion costs, which can be denoted as cost2...costN. If cost1 is the minimum among all the rate-distortion costs, the MIP technique is used for the current coding unit, the MIP usage flag and the corresponding MIP transposition flag (MIP transposition indication parameter) of the current coding unit are set to true, and signaled in the bitstream. If cost1 is not the minimum rate-distortion cost, other intra prediction techniques are used for the current coding unit, the MIP usage flag of the current coding unit is set to false, and signaled in the bitstream. Information such as flags or indexes of other intra prediction techniques is transmitted according to the definition.
[0230] Exemplarily, in other possible embodiments, the encoder traverses the prediction modes. When the current coding unit (current block) is in the intra prediction mode, the encoder obtains the usage permission flag in the coding method according to the embodiments of the present application, that is, obtains the MIP usage permission flag (prediction mode parameter). The MIP usage permission flag may be a sequence-level flag and is currently used to indicate whether the use of the MIP technique is permitted for the decoder. The sequence-level flag can be represented in the form of sps_mip_enable_flag.
[0231] Next, when the MIP usage permission flag is true, the encoder tries the MIP prediction method and calculates the corresponding rate-distortion cost, denoted as cost1. When the MIP usage permission flag is false, the encoder does not try the MIP prediction method, continues to traverse other intra prediction techniques, calculates the corresponding rate-distortion costs, and denotes them as cost2...costN.
[0232] When the MIP usage permission flag is true, based on the size of the current coding unit (the size parameter of the current block), half-downsampling can be performed on the obtained peripheral reference reconstruction samples. The sampling step size is determined by the size of the coding unit. Also, in combination with the MIP transposition instruction parameter, the splicing order of the reference reconstruction samples after upper downsampling and the reference reconstruction samples after left downsampling can be adjusted. When transposition is not required, after the reference reconstruction samples after left downsampling are spliced after the reference reconstruction samples after upper downsampling, the obtained vector is used as the input (downsampling vector). When transposition is required, after the reference reconstruction samples after upper downsampling are spliced after the reference reconstruction samples after left downsampling, the obtained vector is used as the input (downsampling vector).
[0233] Next, the MIP matrix coefficients are obtained using the traversed prediction mode as an index, and an output vector (MIP output vector) is obtained by calculating the MIP matrix coefficients and the input (downsampling vector).
[0234] Note that when the size parameter of the current block satisfies the first preset condition, the first preset prediction mode is directly determined as the mapping mode of the LFNST transform set. For example, when both the width and height of the current coding unit are 32 or more, the PLANAR mode (the first preset prediction mode) can be used as the mapping mode of the LFNST transform set. When the size parameter of the current block satisfies the first preset condition, an optimal conventional intra prediction mode can be derived as the mapping mode of the LFNST transform set for the current MIP output vector using the DIMD method.
[0235] Furthermore, when deriving a conventional intra prediction mode as a mapping mode of the LFNST transform set using the DIMD method, traverse 67 intra prediction modes (or some of the intra prediction modes) in the current VVC and ECM for the current MIP output vector (MIP output vector), calculate the gradient information of each intra prediction mode in the current MIP output vector, and further, based on the gradient information, the corresponding gradient amplitude value can be determined. Next, sort the traversed intra prediction modes based on the gradient amplitude value. The intra prediction mode with the maximum amplitude value is the optimal mode. That is, map the LFNST transform set of the current coding unit using the optimal mode.
[0236] Furthermore, based on the number of samples of the output vector and the size parameter of the current coding unit, the output vector can be upsampled. When upsampling is not required, arrange the vectors in order horizontally and output them as the prediction block of the current coding unit (the MIP prediction block of the current block). When upsampling is required, first perform upsampling horizontally, then perform downsampling vertically, perform upsampling until it reaches the same size as the template, and output the prediction block of the current coding unit (the MIP prediction block of the current block).
[0237] After completing the determination of the mapping mode of the LFNST transform set, subtract the prediction block (the MIP prediction block of the current block) from the original image block of the current coding unit to obtain the residual block of the current coding unit (the current block), perform a primary transform on the residual block to obtain a coefficient block in the frequency domain, and perform a secondary transform on the region of interest (ROI) of the coefficient block in the frequency domain using LFNST. The mapping prediction mode of the LFNST transform set is determined by the above method. Then, through processes such as quantization, inverse quantization, and inverse transform, calculate the rate-distortion cost of the current coding unit and denote it as cost1.
[0238] Furthermore, other intra prediction techniques can be continuously traversed, and the corresponding rate-distortion costs can be calculated and denoted as cost2...costN. If cost1 is the minimum among all the rate-distortion costs, the MIP technique is used for the current coding unit, the MIP usage flag and the corresponding MIP transposition flag (MIP transposition indication parameter) of the current coding unit are set to true, and signaled in the bitstream. If cost1 is not the minimum rate-distortion cost, other intra prediction techniques are used for the current coding unit, the MIP usage flag of the current coding unit is set to false, and signaled in the bitstream. Information such as the flag or index of other intra prediction techniques is transmitted according to the definition.
[0239] According to the encoding method according to the embodiment of the present application, while reducing the software and hardware complexity in JVET-Z0048, the performance similar to that of JVET-Z0048 is maintained. For example, there is no change in the performance of the luminance component. Compared with ECM4.0, the same performance as JVET-Z0048 is maintained.
[0240] Furthermore, in the embodiment of the present application, considering that the hardware decoder has different requirements for I frames and B frames, the encoding method according to the embodiment of the present application may be used only for B frames, or may be used for both I frames and B frames. Alternatively, the encoding method according to the embodiment of the present application may be used only for I frames. Or, the conditions for permitting the use of the encoding method according to the embodiment of the present application are different for B frames or I frames. For example, in I frames, it is permitted to use the encoding method according to the embodiment of the present application for coding units of all sizes, and in B frames, it is permitted to use the encoding method according to the embodiment of the present application only for coding units of small sizes.
[0241] Furthermore, in an embodiment of the present application, when the MIP prediction mode is used for the luminance component of the current coding unit (current block), and the MIP prediction mode is not used for the chrominance component and the conventional intra prediction mode is not used, the LFNST transform set of the chrominance component may inherit the LFNST transform set of the luminance component.
[0242] Furthermore, in an embodiment of the present application, when the conventional intra prediction mode is not used for the current coding unit (current block), the LFNST transform sets of both the luminance component and the chrominance component can be obtained by the encoding method according to the embodiment of the present application.
[0243] As described above, the encoding method according to the embodiment of the present application relates to a method for deriving a MIP prediction block using DIMD and a method for mapping an LFNST transform set. On the other hand, it is proposed to limit the size of the coding unit in which DIMD is used. For an image block with a relatively large size, the upsampling of the MIP output vector is relatively complex and the direction information is not clear. Therefore, the computational complexity is reduced by skipping the process of deriving the conventional intra prediction mode using DIMD. On the other hand, after proposing to limit the size of the coding unit in which DIMD is used, the computational complexity is further reduced, the MIP output vector before upsampling is used as the input of DIMD, and the optimal conventional intra prediction mode is derived.
[0244] Embodiments of the present application provide an encoding method. On the encoding side, a prediction mode parameter is determined. When the prediction mode parameter indicates that MIP is used for the current block to determine the intra prediction value, the MIP parameter of the current block is determined. Based on the MIP parameter, the intra prediction block of the current block is determined, and a residual block obtained by subtracting the intra prediction value from the current block is calculated. When LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and from the selected LFNST transform kernel candidate set, the LFNST transform kernel used for the current block is determined, the LFNST index is set and signaled to the video bitstream. The residual block is transformed using the LFNST transform kernel. As can be seen from the above, in the embodiments of the present application, when deriving the prediction block, based on the size parameter in the MIP parameter of the current block, the mapping mode of the LFNST transform set is determined. For an image block with a relatively large size, it is not necessary to use DIMD to derive the mapping mode, thereby reducing the computational complexity, and thus improving the coding efficiency.
[0245] Based on the above embodiments, in a further embodiment of the present application, based on the same inventive concept as the foregoing embodiments, FIG. 8 is a schematic diagram 1 showing the structure of an encoder. As shown in FIG. 8, the encoder 110 may include a first determination unit 111, an encoding unit 112, and a first conversion unit 113. The first determination unit 111 determines prediction mode parameters. When the prediction mode parameters indicate that the MIP is used for the current block to determine the intra prediction value, the MIP parameters of the current block are determined. Based on the MIP parameters, the intra prediction block of the current block is determined, and a residual block obtained by subtracting the intra prediction value from the current block is calculated. When the LFNST is used for the current block, based on the MIP parameters, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set, and the LFNST index is set. The encoding unit 112 is configured to signal the LFNST index to the video bitstream. The first transformation unit 113 is configured to transform the residual block using the LFNST transform kernel.
[0246] In this embodiment, it can be understood that the "unit" can be a part of a circuit, a part of a processor, a part of a program, or a part of software, etc. Of course, the "unit" can also be a module or a non-module. Further, each constituent unit according to this embodiment may be integrated into one processing unit, each unit may physically exist alone, or two or more units may be integrated into one unit. The above integrated unit can be realized in the form of a hardware or software function module.
[0247] The integrated unit may be stored in a computer-readable storage medium when it is implemented as a software functional module, rather than being sold or used as an independent product. According to this understanding, for the technical solution of this application, the essential part, or the part that can contribute to the prior art, or all or part of the technical solution, can be expressed as a software product. This computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various types of media capable of storing program codes, such as a universal serial bus (USB) flash drive, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0248] Therefore, the embodiments of this application provide a computer-readable storage medium applied to the encoder 110. A computer program is stored in this computer-readable storage medium. When the computer program is executed by a first processor, the method described in any one of the foregoing embodiments is executed.
[0249] Based on the configuration of the above encoder 110 and the computer-readable storage medium, FIG. 9 is a schematic diagram 2 showing the structure of the encoder. As shown in FIG. 9, the encoder 110 can include a first memory 114, a first processor 115, a first communication interface 116, and a first bus system 117. The first memory 114, the first processor 115, and the first communication interface 116 are coupled together via the first bus system 117. As can be understood, the first bus system 117 is used to realize the connection and communication between these components. The first bus system 117 further includes a power bus, a control bus, and a status signal bus in addition to the data bus. However, for the sake of clarity in the description, in FIG. 9, the various buses are marked as the first bus system 117. The first communication interface 116 is used to transmit and receive signals in the process of transmitting and receiving information between other external network elements. The first memory 114 is used to store a computer program executable by the first processor. When the first processor 115 executes a computer program, it determines prediction mode parameters. When the prediction mode parameters indicate that the MIP is used for the current block to determine the intra prediction value, it determines the MIP parameters of the current block. Based on the MIP parameters, it determines the intra prediction block of the current block and calculates the residual block obtained by subtracting the intra prediction value from the current block. When the LFNST is used for the current block, based on the MIP parameters, it determines the mapping mode of the LFNST transform set. Based on the mapping mode of the LFNST transform set, it selects one set of LFNST transform kernel candidates from a plurality of sets of LFNST transform kernel candidates, and determines the LFNST transform kernel used for the current block from the selected set of LFNST transform kernel candidates, sets the LFNST index, and signals it to the video bitstream. The residual block is transformed using the LFNST transform kernel.
[0250] For clarity, the first memory 114 of the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) that functions as an external cache. By way of example and not limitation, various RAMs are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synch-link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The first memory 114 of the systems and methods described in the present application can include these and any other suitable types of memory, but is not limited thereto.
[0251] The first processor 115 can be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the form of hardware of the first processor 115 or instructions in the form of software. The first processor 115 can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates, or transistor logic devices, discrete hardware components. The processor can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any ordinary processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the technical field, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The storage medium is located in the first memory 114. The first processor 115 reads the information in the first memory 114 and completes the steps of the above method in combination with the hardware of the processor.
[0252] It can be understood that these embodiments described in the present application can be realized by hardware, software, firmware, middleware, microcode, or a combination thereof. When realized by hardware, the processing unit can be realized by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units used to execute the functions described in the present application, or a combination thereof. When realized by software, the technology described in the present application can be realized by modules (for example, procedures, functions, etc.) for executing the functions described in the present application. The software code is stored in the memory and executed by the processor. The memory can be realized inside or outside the processor.
[0253] Optionally, as another embodiment, the first processor 115 is further configured to execute the method according to any one of the above embodiments when executing a computer program.
[0254] FIG. 10 is a schematic diagram 1 showing the structure of a decoder. As shown in FIG. 10, the decoder 120 can include a second determination unit 121 and a second conversion unit 122. The second determination unit 121 decodes the bitstream to determine a prediction mode parameter. When the prediction mode parameter indicates that MIP is used to determine the intra prediction value, it decodes the bitstream to determine the MIP parameters of the current block, decodes the bitstream to determine the transform coefficients and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, it determines the mapping mode of the LFNST transform set based on the MIP parameters, selects one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets based on the mapping mode of the LFNST transform set, and determines the LFNST transform kernel to be used for the current block from the selected LFNST transform kernel candidate set. The second transform unit 122 is configured to transform the transform coefficients using the LFNST transform kernel.
[0255] In this embodiment, it can be understood that the "unit" can be a part of a circuit, a part of a processor, a part of a program, or a part of software, etc. Of course, the "unit" can also be a module or a non-module. Further, each constituent unit according to this embodiment may be integrated into one processing unit, each unit may physically exist alone, or two or more units may be integrated into one unit. The above integrated unit can be realized in the form of a hardware or software function module.
[0256] The integrated unit may be stored in a computer-readable storage medium when it is implemented as a software function module, rather than being sold or used as an independent product. According to this understanding, for the technical solution of this application, the essential part, or the part that can contribute to the prior art, or all or part of this technical solution, can be expressed as a software product. This computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various types of media capable of storing program codes, such as a universal serial bus (USB) flash drive, a mobile hard disk, a read only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0257] Therefore, the embodiments of this application provide a computer-readable storage medium applied to the decoder 120. A computer program is stored in this computer-readable storage medium. When the computer program is executed by a first processor, the method described in any one of the aforementioned embodiments is executed.
[0258] Based on the configuration of the decoder 120 and the computer-readable storage medium, FIG. 11 is a schematic diagram 2 showing the structure of the decoder. As shown in FIG. 11, the decoder 120 can include a second memory 123, a second processor 124, a second communication interface 125, and a second bus system 126. The second memory 123, the second processor 124, and the second communication interface 125 are coupled together via the second bus system 126. As can be understood, the second bus system 126 is used to realize the connection and communication between these components. The second bus system 126 further includes a power bus, a control bus, and a status signal bus in addition to the data bus. However, for the sake of clarity in the description, in FIG. 11, various buses are marked as the second bus system 126. The second communication interface 125 is used to transmit and receive signals in the process of transmitting and receiving information with other external network elements. The second memory 123 is used to store computer programs executable by the second processor. When the second processor 124 executes a computer program, it decodes a bitstream to determine prediction mode parameters. When the prediction mode parameters indicate that MIP is used to determine the intra prediction value, it decodes the bitstream to determine the MIP parameters of the current block. It decodes the bitstream to determine the transform coefficients and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, it determines the mapping mode of the LFNST transform set based on the MIP parameters. Based on the mapping mode of the LFNST transform set, it selects one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets, and determines the LFNST transform kernel used for the current block from the selected LFNST transform kernel candidate set. It uses the LFNST transform kernel to transform the transform coefficients.
[0259] As can be understood, the second memory 123 of the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) that functions as an external cache. By way of example and not limitation, various RAMs are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synch-link DRAM (SLDRAM), direct rambus RAM (DRRAM). The second memory 123 of the systems and methods described in the present application can include these and any other suitable types of memory, but is not limited thereto.
[0260] The second processor 124 can be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit in the form of hardware of the second processor 124 or instructions in the form of software. The second processor 124 can be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates, or transistor logic devices, discrete hardware components. The processor can implement or execute various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any ordinary processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the technical field, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The storage medium is located in the second memory 123. The second processor 124 reads the information in the second memory 123 and completes the steps of the above method in combination with the hardware of the processor.
[0261] It can be understood that these embodiments described in the present application can be implemented by hardware, software, firmware, middleware, microcode, or a combination thereof. When implemented by hardware, the processing unit can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units used to execute the functions described in the present application, or a combination thereof. When implemented by software, the technology described in the present application can be implemented by modules (e.g., procedures, functions, etc.) for executing the functions described in the present application. The software code is stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.
[0262] Embodiments of the present application provide an encoder and a decoder. When deriving a prediction block, based on the size parameter in the MIP parameter of the current block, the mapping mode of the LFNST transform set is determined. For an image block with a relatively large size, it may not be necessary to use DIMD to derive the mapping mode. Thereby, the computational complexity can be reduced, and thus the coding efficiency can be improved.
[0263] In the embodiments of the present application, terms such as "including", "comprising", or other variants do not exclude including other components and are intended to cover. Therefore, a process, method, article, or device including a series of elements may include not only those elements but also other elements not explicitly listed or other elements specific to the process, method, article, or device. Unless there are more restrictions, the presence of another same element in a process, method, article, or device including elements limited by the phrase "comprising..." is not excluded.
[0264] The sequence numbers of the embodiments of the present application described above do not indicate the superiority or inferiority of the embodiments and are only used for explanation.
[0265] The methods disclosed in some method embodiments according to the present application can be arbitrarily combined to obtain new method embodiments as long as there is no contradiction.
[0266] The features disclosed in some product embodiments according to the present application can be arbitrarily combined to obtain new product embodiments as long as there is no contradiction.
[0267] The features disclosed in some method or device embodiments according to the present application can be arbitrarily combined to obtain new method embodiments or device embodiments as long as there is no contradiction.
[0268] The above are only specific embodiments of the present application, and the protection scope of the present application is not limited thereto. All changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be included within the protection scope of the present application. Therefore, the protection scope of the present application should be determined by the protection scope of the claims.
Industrial Applicability
[0269] Embodiments of the present application provide a coding method, an encoder, a decoder, and a storage medium. On the decoding side, a bitstream is decoded to determine a prediction mode parameter. When the prediction mode parameter indicates that MIP is used to determine an intra prediction value, the bitstream is decoded to determine the MIP parameter of the current block. The bitstream is decoded to determine the transform coefficient and the LFNST index of the current block. When the LFNST index indicates that LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The transform coefficient is transformed using the LFNST transform kernel. On the encoding side, a prediction mode parameter is determined. When the prediction mode parameter indicates that MIP is used for the current block to determine an intra prediction value, the MIP parameter of the current block is determined. Based on the MIP parameter, the intra prediction block of the current block is determined, and a residual block obtained by subtracting the intra prediction value from the current block is calculated. When LFNST is used for the current block, based on the MIP parameter, the mapping mode of the LFNST transform set is determined. Based on the mapping mode of the LFNST transform set, one LFNST transform kernel candidate set is selected from a plurality of LFNST transform kernel candidate sets, and the LFNST transform kernel used for the current block is determined from the selected LFNST transform kernel candidate set. The LFNST index is set and signaled to the video bitstream. The residual block is transformed using the LFNST transform kernel. As can be seen from the above, in the embodiments of the present application, when deriving a prediction block, based on the size parameter in the MIP parameter of the current block, the mapping mode of the LFNST transform set is determined.For relatively large image blocks, it is not necessary to use DIMD to derive the mapping mode. Thereby, the computational complexity can be reduced, and thus, the coding efficiency can be improved.
Claims
1. A decoding method applied to a decoder, the method comprising: decoding a bitstream to determine a prediction mode parameter; when the prediction mode parameter indicates that intra prediction (MIP) based on a matrix is used to determine an intra prediction value, decoding the bitstream to determine the MIP parameter of the current block; decoding the bitstream to determine the transform coefficient and the low-frequency non-separable transform (LFNST) index of the current block; when the LFNST index indicates that LFNST is used for the current block, determining a mapping mode of the LFNST transform set based on the MIP parameter; selecting one LFNST transform kernel candidate set from a plurality of LFNST transform kernel candidate sets based on the mapping mode of the LFNST transform set, and determining the LFNST transform kernel used for the current block from the selected LFNST transform kernel candidate set; transforming the transform coefficient using the LFNST transform kernel; including a decoding method characterized by the above.
2. The MIP parameter includes a size parameter of the current block, and the method comprises: when the size parameter meets a first preset condition, determining a first preset prediction mode as the mapping mode of the LFNST transform set; when the size parameter does not meet the first preset condition, using decoder-side intra mode derivation (DIMD) to determine the mapping mode of the LFNST transform set; further including the method according to claim 1, characterized by the above.
3. The method comprises: traversing at least one intra prediction mode to determine at least one gradient information corresponding to the current block; determining the mapping mode of the LFNST transform set based on the at least one gradient information; further including the method according to claim 2, characterized by the above.
4. The method comprises: determining a gradient amplitude value corresponding to each intra prediction mode based on the at least one gradient information; Determining, as the mapping mode of the LF-NST transform set, the intra prediction mode among the at least one intra prediction mode that has the maximum gradient amplitude value; further comprising; The method according to claim 3, characterized in that.
5. The method comprises: determining a downsampling vector of the current block based on the MIP parameter; performing a matrix multiplication operation based on the downsampling vector to obtain a MIP output vector; determining a MIP prediction block of the current block based on the MIP output vector; traversing the at least one intra prediction mode for the MIP prediction block to obtain the at least one gradient information; further comprising; The method according to claim 3 or 4, characterized in that.
6. The method comprises: determining a downsampling vector of the current block based on the MIP parameter; performing a matrix multiplication operation based on the downsampling vector to obtain a MIP output vector; traversing the at least one intra prediction mode for the MIP output vector to obtain the at least one gradient information; further comprising; The method according to claim 3 or 4, characterized in that.
7. The size parameter includes the height and width of the current block, and the method comprises: when the height of the current block is greater than or equal to a preset height threshold and / or the width of the current block is greater than or equal to a preset width threshold, determining that the size parameter satisfies the first preset condition; when the height of the current block is less than the preset height threshold and / or the width of the current block is less than the preset width threshold, determining that the size parameter does not satisfy the first preset condition; further comprising; The method according to claim 2, characterized in that.
8. The first preset prediction mode is a PLANAR mode or a direct current (DC) mode. The method according to claim 2 or 7, characterized in that.
9. The method comprises: determining an index of the mapping mode of the LF-NST transform set; Based on the value of the index, determining the value of the LFST intra prediction mode index; Based on the value of the LFST intra prediction mode index, selecting one set of LFST transform kernel candidates from a plurality of sets of LFST transform kernel candidates; Selecting the transform kernel indicated by the LFST index from the selected set of LFST transform kernel candidates and setting it as the LFST transform kernel to be used for the current block; further comprising; The method according to claim 2, characterized in that.
10. The plurality of sets of LFST transform kernel candidates includes four sets of LFST transform kernel candidates, and each set of LFST transform kernel candidates includes two LFST transform kernels, Accordingly, using a first look-up table, the value of the LFST intra prediction mode index corresponding to the value of the index is determined; The method according to claim 9, characterized in that.
11. The method is The plurality of sets of LFST transform kernel candidates includes 35 sets of LFST transform kernel candidates, and each set of LFST transform kernel candidates includes three LFST transform kernels, Accordingly, using a second look-up table, determining the value of the LFST intra prediction mode index corresponding to the value of the index; further comprising; The method according to claim 9, characterized in that.
12. The MIP parameter includes a MIP transpose instruction parameter, and the value of the MIP transpose instruction parameter is used to indicate whether to transpose the sample input vector used in the MIP mode. The method according to claim 10 or 11, characterized in that.
13. The method is When the value of the MIP transpose instruction parameter indicates that the sample input vector used in the MIP mode is to be transposed, performing a matrix transpose on the transform kernel indicated by the LFST index to obtain the LFST transform kernel to be used for the current block; further comprising; The method according to claim 12, characterized in that.
14. An encoding method applied to an encoder, the method comprising determining a prediction mode parameter; When the prediction mode parameter indicates that intra prediction (MIP) based on a matrix is used to determine the intra prediction value for the current block, determining the MIP parameters of the current block, and Based on the MIP parameters, determining the intra prediction block of the current block and calculating a residual block obtained by subtracting the intra prediction value from the current block; When low-frequency non-separable transform (LFNST) is used for the current block, determining the mapping mode of the LFNST transform set based on the MIP parameters; Based on the mapping mode of the LFNST transform set, selecting one set of LFNST transform kernel candidates from a plurality of sets of LFNST transform kernel candidates, and determining the LFNST transform kernel used for the current block from the selected set of LFNST transform kernel candidates, setting the LFNST index, and signaling it to the video bitstream; Transforming the residual block using the LFNST transform kernel; including An encoding method characterized by the above.
15. The MIP parameters include size parameters of the current block, and the method When the size parameters meet a first preset condition, determining a first preset prediction mode as the mapping mode of the LFNST transform set; When the size parameters do not meet the first preset condition, using decoder-side intra mode derivation (DIMD) to determine the mapping mode of the LFNST transform set; further including The method according to claim 14, characterized by the above.
16. The method Traversing at least one intra prediction mode to determine at least one gradient information corresponding to the current block; Based on the at least one gradient information, determining the mapping mode of the LFNST transform set; further including The method according to claim 15, characterized by the above.
17. The method Based on the at least one gradient information, determining a gradient amplitude value corresponding to each intra prediction mode; Determining, as the mapping mode of the LF-NST transform set, the intra prediction mode among the at least one intra prediction mode having the maximum gradient amplitude value; further comprising The method according to claim 16, characterized in that. **Claim 18** The method comprises determining a downsampling vector of the current block based on the MIP parameter; executing a matrix multiplication operation based on the downsampling vector to obtain a MIP output vector; determining a MIP prediction block of the current block based on the MIP output vector; traversing the at least one intra prediction mode for the MIP prediction block to obtain the at least one gradient information; further comprising The method according to claim 16 or 17, characterized in that. **Claim 19** The method comprises determining a downsampling vector of the current block based on the MIP parameter; executing a matrix multiplication operation based on the downsampling vector to obtain a MIP output vector; traversing the at least one intra prediction mode for the MIP output vector to obtain the at least one gradient information; further comprising The method according to claim 16 or 17, characterized in that. **Claim 20** The size parameter includes the height and width of the current block, and the method comprises when the height of the current block is greater than or equal to a preset height threshold and / or the width of the current block is greater than or equal to a preset width threshold, determining that the size parameter satisfies the first preset condition; when the height of the current block is less than the preset height threshold and / or the width of the current block is less than the preset width threshold, determining that the size parameter does not satisfy the first preset condition; further comprising The method according to claim 15, characterized in that. **Claim 21** The first preset prediction mode is a PLANAR mode or a direct current (DC) mode. The method according to claim 15 or 19, characterized in that. **Claim 22** The method comprises determining an index of the mapping mode of the LF-NST transform set; Based on the value of the index, determining the value of the LFST intra prediction mode index; Based on the value of the LFST intra prediction mode index, selecting one set of LFST transform kernel candidates from a plurality of sets of LFST transform kernel candidates; Selecting a transform kernel for the current block from the selected set of LFST transform kernel candidates; further comprising; The method according to claim 15, characterized in that.
23. The plurality of sets of LFST transform kernel candidates include four sets of LFST transform kernel candidates, and each set of LFST transform kernel candidates includes two LFST transform kernels. Accordingly, using a first look-up table, the value of the LFST intra prediction mode index corresponding to the value of the index is determined. The method according to claim 22, characterized in that.
24. The method includes the plurality of sets of LFST transform kernel candidates include 35 sets of LFST transform kernel candidates, and each set of LFST transform kernel candidates includes three LFST transform kernels; Accordingly, using a second look-up table, determining the value of the LFST intra prediction mode index corresponding to the value of the index; further comprising; The method according to claim 22, characterized in that.
25. The MIP parameter includes a MIP transpose instruction parameter, and the value of the MIP transpose instruction parameter is used to indicate whether to transpose the sample input vector used in the MIP mode. The method according to claim 23 or 24, characterized in that.
26. The method includes when the value of the MIP transpose instruction parameter indicates transposing the sample input vector used in the MIP mode, performing matrix transposition on the selected transform kernel to obtain the LFST transform kernel used for the current block; further comprising; The method according to claim 23, characterized in that.
27. An encoder, the encoder includes a first determination unit, an encoding unit, and a first conversion unit, The first determination unit determines a prediction mode parameter, When the prediction mode parameter indicates that matrix-based intra prediction (MIP) for determining an intra prediction value is used for the current block, determine the MIP parameters of the current block, Based on the MIP parameters, determine the intra prediction block of the current block, and calculate a residual block obtained by subtracting the intra prediction value from the current block, When low-frequency non-separable transform (LFNST) is used for the current block, determine the mapping mode of the LFNST transform set based on the MIP parameters, Based on the mapping mode of the LFNST transform set, select one set of LFNST transform kernel candidates from a plurality of sets of LFNST transform kernel candidates, and determine the LFNST transform kernel used for the current block from the selected set of LFNST transform kernel candidates, and set the LFNST index, It is configured as follows, The encoding unit is configured to signal the LFNST index to a video bitstream, The first conversion unit is configured to convert the residual block using the LFNST transform kernel, An encoder characterized by this.
28. An encoder, The encoder includes a first memory and a first processor, The first memory is configured to store a computer program executable by the first processor, When the first processor executes the computer program, it is configured to execute the method according to any one of claims 14 to 26, An encoder characterized by this.
29. A decoder, The decoder includes a second determination unit and a second conversion unit, The second determination unit, Decode the bitstream to determine the prediction mode parameter, When the prediction mode parameter indicates that matrix-based intra prediction (MIP) for determining an intra prediction value is used, decode the bitstream to determine the MIP parameters of the current block, Decode the bitstream to determine the transform coefficient and the low-frequency non-separable transform (LFNST) index of the current block, When the LFST index indicates that LFST is used for the current block, based on the MIP parameter, determine the mapping mode of the LFST conversion set, Based on the mapping mode of the LFST conversion set, select one LFST conversion kernel candidate set from a plurality of LFST conversion kernel candidate sets, and determine the LFST conversion kernel used for the current block from the selected LFST conversion kernel candidate set, is configured as follows, The second conversion unit is configured to convert the conversion coefficient using the LFST conversion kernel. A decoder characterized by this.
30. A decoder, The decoder includes a second memory and a second processor, The second memory is configured to store a computer program executable by the second processor, When the second processor executes the computer program, it is configured to execute the method according to any one of claims 1 to 13. A decoder characterized by this.
31. A computer-readable storage medium, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, the method according to any one of claims 1 to 13 is realized, or the method according to any one of claims 14 to 26 is realized. A computer-readable storage medium characterized by this.
Citation Information
Patent Citations
Transforms for matrix-based intra-prediction in video coding
JP2022529686A
Video coding using transform indexes
JP2022529688A
Image encoding / decoding method and device for performing MIP and lfnst, and method for transmitting bitstream
WO2020226424A1