Decoding method, encoding method, decoder and encoder
The decoding method addresses the challenge of low compression efficiency in video clarity by aligning transformation sets with video block textures, improving decompression performance and efficiency.
Patent Information
- Application Number
- JP2024575832
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2042-07-04
AI Technical Summary
Existing digital video compression technologies face challenges in achieving high compression efficiency for improving video clarity, necessitating better decompression methods.
A decoding method that involves analyzing a bitstream to determine a first intra prediction mode, performing transformations based on a conversion set corresponding to this mode, and reconstructing blocks to enhance decompression performance.
Improves decompression performance by aligning transformation sets with the texture direction of video blocks, reducing computational complexity, and enhancing compression efficiency without increasing decoder complexity.
Smart Images

Figure 2025520767000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of video coding, and more specifically, to a decoding method, an encoding method, a decoder, and an encoder.
Background Art
[0002] Digital video compression technology is mainly a technology for compressing huge digital image and video data for transmission, storage, etc. With the rapid increase in Internet videos and the growing demand of people for video clarity, although video decompression technology can be realized with existing digital video compression standards, in order to improve the compression efficiency, better digital video decompression technology is needed.
Summary of the Invention
[0003] Embodiments of the present application provide a decoding method, an encoding method, a decoder, and an encoder, whereby the compression efficiency can be improved.
[0004] In a first aspect, the present application provides a decoding method. The method includes the following. Analyze the bitstream of the current sequence to obtain the FirstObtain a conversion coefficient. Determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of a current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. Perform a first conversion on a first conversion coefficient based on a conversion set corresponding to the first intra prediction mode to obtain a second conversion coefficient of the current block. Perform a second conversion on the second conversion coefficient to obtain a residual block of the current block. Determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.
[0005] In a second aspect, the present application provides an encoding method. The method includes the following. Obtain a residual block of a current block in a current sequence. Perform a third transformation on the residual block of the current block to obtain third transformation coefficients of the current block. Determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. Perform a fourth transformation on the third transformation coefficients based on a transformation set corresponding to the first intra prediction mode to obtain fourth transformation coefficients of the current block. Encode the fourth transformation coefficients.
[0006] In a third aspect, the present application provides a decoder. The decoder includes an analysis unit, a transformation unit, and a reconstruction unit. The analysis unit analyzes a bitstream of a current sequence to obtain FirstIt is configured to obtain a conversion coefficient. The conversion unit determines a first intra prediction mode, and based on a set of conversions corresponding to the first intra prediction mode, performs a first conversion on the first conversion coefficient to obtain a second conversion coefficient of the current block, and performs a second conversion on the second conversion coefficient to obtain a residual block of the current block. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template area adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The reconstruction unit is configured to determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.
[0007] In a fourth aspect, the present application provides an encoder. The encoder includes a residual unit, a conversion unit, and an encoding unit. The residual unit is configured to obtain a residual block of a current block in a current sequence. The conversion unit is configured to perform a third conversion on the residual block of the current block to obtain a third conversion coefficient of the current block, determine a first intra prediction mode, and perform a fourth conversion on the third conversion coefficient based on a conversion set corresponding to the first intra prediction mode to obtain a fourth conversion coefficient of the current block. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The encoding unit is configured to encode the fourth conversion coefficient.
[0008] In a fifth aspect, the present application provides a decoder. The decoder includes a processor and a computer-readable storage medium. The processor is configured to execute computer instructions. The computer-readable storage medium stores computer instructions, and the computer instructions are loaded by the processor and the decoding method in the first aspect or each of its embodiments is executed.
[0009] In one embodiment, the processor is one or more, and the memory is one or more.
[0010] In one embodiment, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium may be installed separately from the processor.
[0011] In a sixth aspect, the present application provides an encoder. The encoder includes a processor and a computer-readable storage medium. The processor is configured to execute computer instructions. Computer instructions are stored in the computer-readable storage medium, and the computer instructions are loaded by the processor to execute the encoding method in the second aspect or each of its embodiments.
[0012] In one embodiment, there may be one or more processors and one or more memories.
[0013] In one embodiment, the computer-readable storage medium may be integrated with the processor, or the computer-readable storage medium may be installed separately from the processor.
[0014] In a seventh aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium is configured to store computer instructions, and when the computer instructions are read and executed by a processor of a computer device, the computer device is caused to execute the decoding method according to the first aspect or the encoding method according to the second aspect.
[0015] In an eighth aspect, the present application provides a bitstream. The bitstream is the bitstream according to the first aspect or the bitstream according to the second aspect.
[0016] Based on the above technical solution, a first intra prediction mode is introduced, and a first transformation is performed on the first transformation coefficient of the current block based on the transformation set corresponding to the first intra prediction mode, so as to improve the decompression performance of the current block. In particular, when the decoder predicts the current block using a non-traditional intra prediction mode, it is possible to avoid directly performing the first transformation using the transformation set corresponding to the planar mode. The transformation set corresponding to the first intra prediction mode can reflect the texture direction of the current block to a certain extent, and thus the decompression performance of the current block can be improved.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, with reference to the drawings, the technical solutions in the embodiments of the present application will be described.
[0019] The solution according to the embodiment of the present application can be applied to the technical field of digital video coding. For example, the technical field includes, but is not limited to, the fields of image coding, video coding, hardware video coding, dedicated circuit video coding, and real-time video coding. Also, the solution according to the embodiment of the present application can be combined with an audio video coding standard (AVS), a second-generation AVS standard (AVS2), or a third-generation AVS standard (AVS3). Examples include, but are not limited to, the H.264 / audio video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Further, with the solution according to the embodiment of the present application, lossy compression or lossless compression can be performed on an image. The lossless compression may be visually lossless compression or mathematically lossless compression.
[0020] A block-based hybrid coding framework is used in video coding standards. Specifically, each image (frame) in a video is divided into largest coding units (LCUs) or coding tree units (CTUs) that are squares of the same size (e.g., 128×128, 64×64, etc.). Each largest coding unit or coding tree unit can also be divided into rectangular coding units (CUs) based on rules. The coding unit may further be divided into a prediction unit (PU), a transform unit (TU), etc. The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and loop filter. The prediction module includes intra prediction and inter prediction. Inter prediction includes motion estimation and motion compensation. Since there is a strong correlation between adjacent samples in the video image, in video coding technology, the intra prediction method is used to eliminate the spatial redundancy between adjacent samples. In intra prediction, only the information of the image in the same frame is referred to, and the sample information within the current divided block is predicted. Since there is a strong similarity between adjacent images in a video, in video coding technology, the inter prediction method can be used to eliminate the temporal redundancy between adjacent images and improve the coding efficiency. In inter prediction, by referring to the image information of different frames, motion estimation can be used to search for the motion vector information that best matches the current divided block. Through transformation, the predicted image block is transformed into the frequency domain and the energy is redistributed.By combining transformation and quantization, information that is not sensitive to the human eye can be removed and it is used to eliminate visual redundancy. By entropy coding, character redundancy can be removed based on the current context model and the probability information of the binary bit stream.
[0021] In the process of digital video encoding, the encoder first reads a black-and-white image or a color image from the original video sequence and then can encode the black-and-white image or the color image. Here, the black-and-white image can include samples of the luma component, and the color image can include samples of the chroma component. Optionally, the color image can further include samples of the luma component. The color format of the original video sequence may be, for example, the luma-chroma (YCbCr, YUV) format or the red-green-blue (RGB) format. Specifically, after reading the black-and-white image or the color image, the encoder divides it into blocks, generates a predicted block of the current block by performing intra prediction or inter prediction on the current block, subtracts the predicted block from the original block of the current block to obtain a residual block, transforms and quantizes the residual block to obtain a quantized coefficient matrix, and entropy-codes the quantized coefficient matrix and outputs it to a bit stream. In the process of digital video decoding, the decoding side generates a predicted block of the current block by performing intra prediction or inter prediction on the current block. Also, the decoding side decodes the bit stream to obtain a quantized coefficient matrix, inverse quantizes and inverse-transforms the quantized coefficient matrix to obtain a residual block, and adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed blocks form a reconstructed image. The decoding side loop filters the reconstructed image based on the image or blocks to obtain a decoded image.
[0022] The current block can be, for example, the current coding unit (CU) or the current prediction unit (PU).
[0023] Note that on the encoding side as well, in order to obtain the decoded image, processing similar to that on the decoding side is necessary. The decoded image can be a reference image for inter prediction of subsequent images. Block splitting information, mode information or parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc. determined on the encoding side is output to the bitstream as necessary. The decoding side analyzes and analyzes the existing information to determine the same block splitting information, mode information or parameter information such as prediction, transformation, quantization, entropy coding, loop filtering, etc. as on the encoding side. Thereby, it is ensured that the decoded image obtained on the encoding side is the same as the decoded image obtained on the decoding side. The decoded image obtained on the encoding side is usually also called the reconstructed image. During prediction, the current block may be divided into prediction units, and during transformation, the current block may be divided into transformation units. The division of the prediction unit and the transformation unit may be the same or different. The above is the basic flow of video coding in a block-based hybrid coding framework. With the development of technology, some modules or some steps in the framework may be optimized. This application is applicable to the basic flow of a video codec in the block-based hybrid coding framework.
[0024] To facilitate understanding, first, the encoding framework according to this application will be briefly described.
[0025] FIG. 1 is a block diagram showing an encoding framework 100 according to an embodiment of this application.
[0026] As shown in FIG. 1, the encoding framework 100 can include an intra prediction unit 180, an inter prediction unit 170, a residual unit 110, a transform and quantization unit 120, an entropy encoding unit 130, an inverse transform and inverse quantization unit 140, and a loop filtering unit 150. Optionally, the encoding framework 100 may further include a decoding image buffer unit 160. The encoding framework 100 is also referred to as a hybrid framework encoding mode.
[0027] The intra prediction unit 180 or the inter prediction unit 170 can predict an image block waiting for coding and output a prediction block. The residual unit 110 can calculate a residual block, that is, the difference between the prediction block and the image block waiting for coding, based on the prediction block and the image block waiting for coding. The transform and quantization unit 120 is used to perform operations such as transform and quantization on the residual block, thereby removing information that is not sensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before being transformed and quantized by the transform and quantization unit 120 may be called a temporal residual block, and the residual block in the temporal domain after being transformed and quantized by the transform and quantization unit 120 may be called a frequency residual block or a frequency-domain residual block. The entropy encoding unit 130 can output a bitstream based on the quantized transform coefficient output by the transform and quantization unit 120 after receiving the quantized transform coefficient. For example, the entropy encoding unit 130 can remove character redundancy based on a target context model and probability information of a binary bitstream. For example, the entropy encoding unit 130 can be used for context-based adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also be called a header information encoding unit. Optionally, in the present application, the image block waiting for coding may also be called an original image block or a target image block. The prediction block may also be called a predicted image block or an image prediction block, and may also be called a prediction signal or prediction information.The reconstruction block may also be referred to as a reconstructed image block or an image reconstruction block, and may also be referred to as a reconstruction signal or reconstruction information. Further, on the encoding side, the image block waiting for encoding may also be referred to as an encoding block or an encoding image block. On the decoding side, the image block waiting for encoding may also be referred to as a decoding block or a decoding image block. The image block waiting for encoding may be a CTU or a CU.
[0028] In the encoding framework 100, the difference between the prediction block and the image block waiting for encoding is calculated to obtain a residual block, and processes such as transformation and quantization are performed on the residual block, and the residual block is transmitted to the decoding side. Accordingly, after receiving and decoding the bitstream on the decoding side, a residual block is obtained through steps such as inverse transformation and inverse quantization, and a reconstruction block is obtained by adding the residual block to the prediction block obtained by being predicted by the decoding side.
[0029] Note that the inverse transformation and inverse quantization unit 140, loop filtering unit 150, and decoded image buffer unit 160 within the encoding framework 100 can be used to form a decoder. The intra prediction unit 180 or inter prediction unit 170 can predict an image block waiting for coding based on an existing reconstructed block, so that it can be ensured that the encoding side uses the reference frame in the same way as the decoding side. In other words, the encoder can replicate the processing loop of the decoder, thereby generating the same prediction as the decoding side. Specifically, the quantized transform coefficients are inverse-transformed and inverse-quantized by the inverse transformation and inverse quantization unit 140, and the approximate residual block on the decoding side is replicated. After a prediction block is added to this approximate residual block, through the loop filtering unit 150, it is possible to perform smoothing filtering on the effects such as block-based processing and blocking artifacts due to quantization. The image block output from the loop filtering unit 150 can be stored in the decoded image buffer unit 160 for use in subsequent image prediction.
[0030] It should be understood that FIG. 1 is only an example of this application and should not be construed as a limitation of this application.
[0031] For example, the loop filtering unit 150 within the encoding framework 100 may include a deblocking filter (DBF) and a sample adaptive offset (SAO). The role of the DBF is to deal with deblocking artifacts, and the role of the SAO is to remove the ringing effect. In other embodiments of the present application, a neural-network-based loop filtering algorithm can be used in the encoding framework 100 to improve the compression efficiency of video. Or, the encoding framework 100 can be a deep learning neural-network-based video coding hybrid framework. In one embodiment, based on the DBF and SAO, a convolutional neural network-based model can be used to calculate the result after sample filtering. The network structure in the luminance component of the loop filtering unit 150 and the network structure in the chrominance component may be the same or different. Considering that more visual information is included in the luminance component, the luminance component can be used to guide the filtering of the chrominance component in order to improve the reconstruction quality of the chrominance component.
[0032] The related content of intra prediction and inter prediction will be described below.
[0033] In inter prediction, in order to eliminate temporal redundancy, image information of different frames can be referred to, and motion estimation can be utilized to search for the motion vector information that best matches the image block waiting for coding. The images used in inter prediction may be P frames and / or B frames. A P frame means a forward predicted picture, and a B frame means a bidirectional predicted picture.
[0034] In intra prediction, in order to eliminate spatial redundancy, only the information of the same image is referred to, and the sample information in the image block waiting to be coded is predicted. The image used for intra prediction may be an I frame. For example, according to the coding order from left to right and from top to bottom, the upper left image block, the upper image block, and the left image block can be used as reference information to predict the image block waiting to be coded. The image block waiting to be coded is also used as the reference information for the next image block. In this way, the entire image can be predicted. When the input digital video is in a color format such as the YUV4:2:0 format, each 4 pixels of each image frame of the digital video consists of 4 Y components and 2 UV components. In the encoding framework, the Y component (i.e., the luminance block) and the UV component (i.e., the chrominance block) can be encoded respectively. Similarly, the decoding side can decode according to the format.
[0035] For the intra prediction process, in order to obtain a prediction block, intra prediction can predict the image block waiting to be coded using an angular prediction mode and a non-angular prediction mode. Based on the rate-distortion information calculated from the prediction block and the image block waiting to be coded, the optimal prediction mode of the image block waiting to be coded is selected, and the prediction mode is transmitted to the decoding side via the bitstream. The decoding side analyzes the prediction mode, predicts and obtains the prediction block of the target decoding block, and can obtain the reconstructed block by adding the temporal residual block obtained by bitstream transmission.
[0036] Through the development of successive digital video coding standards, the non-angular prediction modes have been relatively stable, including average mode and planar mode. The number of angular prediction modes has increased with the evolution of digital video coding standards. Taking the international digital video coding standard H series as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angular prediction mode, while the H.265 / HEVC standard is extended to 33 angular prediction modes and 2 non-angular prediction modes. In H.266 / VVC, the intra-prediction modes are further extended, and for luminance blocks, there are 67 conventional prediction modes, non-conventional prediction modes, and matrix weighted intra-frame prediction (MIP) mode, and the 67 conventional prediction modes include planar mode, direct current (DC) mode, and 65 angular modes. Planar mode is typically used to process blocks where a gradient exists in the texture, DC mode, as the name suggests, is used to process flat areas, and angular prediction mode is typically used to process blocks where the angular texture is relatively sharp.
[0037] It should be noted that in this application, the current block used for intra prediction may be a square block or a rectangular block.
[0038] Furthermore, since all intra prediction blocks are square, the probabilities of using each angular prediction mode are equal. When the length and width of the current block are different, for a horizontal block (where the width is greater than the height), the probability of using the upper reference sample is greater than the probability of using the left reference sample. For a vertical block (where the height is greater than the width), the probability of using the upper reference sample is less than the probability of using the left reference sample. When predicting a rectangular block, if the conventional angular prediction mode is converted to a wide-angle prediction mode and the rectangular block can be predicted using the wide-angle prediction mode, the prediction angle range of the current block is larger than the prediction angle range when predicting the rectangular block using the conventional angular prediction mode. Optionally, when using the wide-angle prediction mode, the signal can still be transmitted using the index of the conventional angular prediction mode. Accordingly, after receiving the signal, the decoding side can convert the conventional angular prediction mode back to the wide-angle prediction mode again, so that neither the total number of intra prediction modes nor the encoding method of the intra mode is changed.
[0039] Furthermore, based on the size of the current block, it is possible to determine or select the intra prediction mode to be executed. For example, based on the size of the current block, it is possible to determine or select the wide-angle prediction mode to perform intra prediction on the current block. For example, when the current block is a rectangular block (having different dimensions of width and height), the current block can be intra predicted using the wide-angle prediction mode. The aspect ratio of the current block can be used to determine the angular prediction mode before and after the wide-angle prediction mode is replaced. For example, when predicting the current block, any intra prediction mode with an angle not exceeding the diagonal of the current block (from the lower left corner to the upper right corner of the current block) can be selected as the angular prediction mode after replacement.
[0040] Next, other intra prediction modes according to this application will be introduced.
[0041] (1) Matrix based Intra Prediction (MIP) mode
[0042] The MIP mode can also be called the Matrix weighted intra prediction mode. The process related to the MIP mode can be divided into three main steps: a downsampling process, a matrix multiplication process, and an upsampling process. Specifically, first, in the downsampling process, spatially adjacent reconstructed samples are downsampled. Next, the obtained downsampled sample sequence is used as the input vector of the matrix multiplication process, that is, the output vector of the downsampling process is used as the input vector of the matrix multiplication process, and the input vector of the matrix multiplication process is multiplied by a preset matrix, and a bias vector is added to the result to output a calculated sample vector. The output vector of the matrix multiplication process is used as the input vector of the upsampling process, and the final prediction block is obtained by upsampling.
[0043] FIG. 2 is a schematic diagram showing the MIP mode according to an embodiment of the present application.
[0044] JPEG2025520767000026.jpg80168
[0045] In other words, in order to predict a block with width W and height H, in MIP, it is necessary to input H reconstructed samples in the column to the left of the current block and W reconstructed samples in the row above the current block. In MIP, a predicted block is generated based on three steps: averaging of reference samples, matrix-vector multiplication, and interpolation. The core of MIP is matrix-vector multiplication, which can be considered as a process of generating a predicted block by using input samples (reference samples) in the matrix-vector multiplication method. Various matrices are provided in MIP. The difference in prediction methods can be reflected in the difference in matrices. Different results can be obtained by using different matrices for the same input samples. Also, the processes of averaging and interpolation of reference samples are designed considering the trade-off between performance and complexity. For blocks with a large size, the effect of approximating downsampling can be realized by averaging the reference samples, so the input can be adapted to a relatively small matrix. The effect of upsampling can be obtained by interpolation. In this way, it is not necessary to provide MIP matrices for each block size. Instead, only one or several matrices with specific sizes need to be provided. As the need for compression performance increases and the performance of hardware improves, more complex MIPs may appear in the next-generation standards.
[0046] In the MIP mode, the MIP mode can be simplified from a neural network. For example, the matrix used in the MIP mode can be obtained based on training. Therefore, the MIP mode has a relatively strong generalization ability and a prediction effect that cannot be achieved by conventional prediction modes. The MIP mode is a model obtained by simplifying the complexity of hardware and software multiple times for an intra-prediction model based on a neural network. Based on a large number of training samples, multiple prediction modes represent multiple models and parameters and can better cover the texture of natural sequences.
[0047] The MIP mode is somewhat similar to the Planar mode. Obviously, the MIP mode is more complex and flexible than the Planar mode.
[0048] Note that for coding units with different block sizes, the number of MIP modes can be different. For example, for a coding unit with a size of 4×4, the MIP mode has 16 prediction modes. For a coding unit with a size of 8×8 or a width or height equal to 4, the MIP mode has 8 prediction modes. For coding units of other sizes, the MIP mode has 6 prediction modes. In addition, the MIP mode has a transpose function. For the prediction mode that matches the current size, in the MIP mode, a transpose calculation can be attempted on the encoding side. Therefore, the MIP mode requires a usage flag indicating whether the MIP mode is used for the current coding unit. Also, when the MIP mode is used for the current coding unit, it is necessary to further transmit a transpose flag and an index flag to the decoder. The transpose flag is binarized by a Fixed Length (FL) coding method and has a length of 1. The index flag is binarized by a Truncated Binary (TB) coding method. For example, for a coding unit with a size of 4×4, the MIP mode has 16 prediction modes. The index flag can be a 5-bit or 6-bit truncated binary flag.
[0049] (2) Decoder side Intra Mode Derivation (DIMD) mode
[0050] The main core of the DIMD mode is to derive the intra prediction mode on the decoder side using the same method as the encoder. Thereby, it avoids transmitting the intra prediction mode index of the current coding unit in the bitstream and achieves the purpose of saving bit overhead.
[0051] The specific process of the DIMD mode is mainly divided into the following two steps.
[0052] Step 1: Derive the prediction mode.
[0053] FIG. 3 is a schematic diagram showing the derivation of a prediction mode based on DIMD according to an embodiment of the present application.
[0054] As shown in FIG. 3(a), in DIMD, the prediction mode is derived using the samples within the template in the reconstruction region (the reconstruction samples on the left and above the current block). For example, the template can include the reconstruction samples of the three adjacent rows above the current block, the reconstruction samples of the three adjacent columns on the left, and the corresponding adjacent reconstruction samples in the upper left. Based on this, a plurality of gradient values corresponding to a plurality of adjacent reconstruction samples within the template are determined according to a window (for example, the window shown in FIG. 3( b ) or the window shown in FIG. 3( c ). Each gradient value is used to obtain one Intra prediction mode (IPM) that conforms to its gradient direction. Based on this, the encoder can set the prediction mode corresponding to the largest gradient value among the plurality of gradient values and the prediction mode corresponding to the second largest gradient value as the derived prediction mode. For example, as shown in FIG. 3(b), for a block of size 4×4, all adjacent reconstruction samples for which gradient values need to be determined are analyzed to obtain the corresponding gradient histogram. For example, as shown in FIG. 3(c), for blocks of other sizes, all adjacent reconstruction samples for which gradient values need to be determined are analyzed to obtain the corresponding gradient histogram. Finally, the prediction mode corresponding to the largest gradient in the gradient histogram and the prediction mode corresponding to the second largest gradient are set as the derived prediction mode.
[0055] Of course, the gradient histogram in the present application is only an exemplification for determining the derived prediction mode, and when specifically implemented, it can be implemented in various simple forms and is not particularly limited in the present application. Furthermore, in the present application, the method for obtaining the gradient histogram is not limited. For example, the gradient histogram can be obtained using a Sobel operator or other methods. Additionally, in other alternative embodiments, the gradient value according to the present application may be replaced with a gradient amplitude value, and the present application is not particularly limited thereto.
[0056] Step 2: Derive a prediction block.
[0057] FIG. 4 is a schematic diagram showing the derivation of a prediction block based on DIMD according to an embodiment of the present application.
[0058] As shown in FIG. 4, the encoder can weight the predicted values corresponding to three intra prediction modes (the planar mode and two intra prediction modes derived based on DIMD). The codec obtains the prediction block of the current block using the same prediction block derivation method. Assuming that the prediction mode corresponding to the largest gradient value is prediction mode 1 and the prediction mode corresponding to the second largest gradient value is prediction mode 2, the encoder determines the following two conditions. 1. The gradient value of prediction mode 2 is not 0. 2. Neither prediction mode 1 nor prediction mode 2 is the planar mode or the direct current (DC) prediction mode.
[0059] If the above two conditions do not hold simultaneously, calculate the predicted sample value of the current block using only prediction mode 1, that is, apply the normal prediction process to prediction mode 1. Otherwise, that is, when the above two conditions hold simultaneously, derive the predicted block of the current block using the weighted average method. The specific method is as follows. The planar mode occupies a weight of 1 / 3, and the remaining 2 / 3 is the total weight of prediction mode 1 and prediction mode 2. For example, divide the gradient amplitude value of prediction mode 1 by the sum of the gradient amplitude values of prediction mode 1 and prediction mode 2, and use the result as the weight of prediction mode 1. Divide the gradient amplitude value of prediction mode 2 by the sum of the gradient amplitude values of prediction mode 1 and prediction mode 2, and use the result as the weight of prediction mode 2. Perform weighted averaging on the predicted blocks obtained based on the above three prediction modes, that is, the predicted block 1, predicted block 2, and predicted block 3 obtained based on the planar mode, prediction mode 1, and prediction mode 2 respectively, to obtain the predicted block of the current coding unit. The decoder also obtains the predicted block in the same steps.
[0060] In other words, the specific weights in step 2 above are calculated as follows. JPEG2025520767000027.jpg30150mode1 and mode2 represent prediction mode 1 and prediction mode 2 respectively, and amp1 and amp2 represent the gradient amplitude values of prediction mode 1 and prediction mode 2 respectively. In the DIMD mode, it is necessary to transmit a flag to the decoder. The flag is used to indicate whether the DIMD mode is used for the current coder unit.
[0061] Of course, the above weighted average method is only an example of this application and should not be understood as a limitation of this application.
[0062] In summary, in DIMD, the intra prediction mode is selected using the gradient analysis of the reconstructed samples, and it is possible to weight two intra prediction modes and the planar mode according to the analysis results. The advantage of DIMD is that when the DIMD mode is selected for the current block, it is not necessary to specifically indicate which intra prediction mode is used in the bitstream, because it is derived by the decoder itself through the above process, so a certain amount of overhead can be saved.
[0063] (3) Template based Intra Mode Derivation (TIMD) mode
[0064] The technical principle of the TIMD mode is similar to the technical principle of the DIMD mode described above. In both cases, the codec derives the prediction mode with the same operation to save the overhead of transmitting the mode index. The TIMD prediction mode can be mainly understood in two parts. First, the cost information of each prediction mode is calculated based on the template, and the prediction mode corresponding to the smallest cost and the prediction mode corresponding to the second smallest cost are selected. The prediction mode corresponding to the smallest cost is denoted as prediction mode 1, and the prediction mode corresponding to the second smallest cost is denoted as prediction mode 2. The ratio of the second smallest cost (costMode2) to the smallest cost (costMode1) is When it meets the preset conditions such as JPEG2025520767000028.jpg15150, the prediction block corresponding to prediction mode 1 and the prediction block corresponding to prediction mode 2 are weighted and fused according to the weight corresponding to prediction mode 1 and the weight corresponding to prediction mode 2 to obtain the final prediction block.
[0065] Exemplarily, the weight corresponding to prediction mode 1 and the weight corresponding to prediction mode 2 are determined based on the following method. JPEG2025520767000029.jpg24150weight1 is the weight of the prediction block corresponding to prediction mode 1, and weight2 is the weight of the prediction block corresponding to prediction mode 2. When the ratio of the second smallest cost costMode2 to the smallest cost costMode1 does not meet the preset condition, weighted fusion between prediction blocks is not performed, and the prediction block corresponding to prediction mode 1 becomes the prediction block of TIMD.
[0066] When performing intra prediction on the current block using the TIMD mode, if the adjacent reconstruction samples available in the reconstruction sample template of the current block are not included, in the TIMD mode, the planar mode is selected to perform intra prediction on the current block, that is, weighted fusion is not performed. Similar to the DIMD mode, in the TIMD mode, it is necessary to transmit a flag to the decoder. The flag is used to indicate whether the TIMD mode is used for the current coding unit.
[0067] The process by which the encoder or decoder calculates the cost information of each prediction mode is mainly as follows. Based on the reconstruction samples adjacent to the upper side of the template area and the reconstruction samples adjacent to the left side, intra-mode prediction is performed on the samples within the template area, and the prediction process is the same as the original intra prediction mode. For example, when performing intra-mode prediction on the samples within the template area using the DC mode, the average value of the entire coding unit is calculated. As another example, when performing intra-mode prediction on the samples within the template area using the angular prediction mode, a corresponding interpolation filter is selected according to the mode, and prediction samples are obtained by interpolation according to the rules. In this case, based on the prediction samples and reconstruction samples within the template area, the distortion between the prediction samples and reconstruction samples within the area, that is, the cost information of the current prediction mode, can be calculated.
[0068] FIG. 5 is a schematic diagram showing a template used for TIMD according to an embodiment of the present application.
[0069] As shown in FIG. 5, when the current block is a coding unit with a width equal to M and a height equal to N, the codec can select a reference template of the current block from a coding unit with a width equal to 2(M + L1)+1 and a height equal to 2(N + L2)+1 to calculate the template of the current block. If the template of the current block does not include available adjacent reconstruction samples, in the TIMD mode, the planar mode is selected to perform intra prediction on the current block. For example, the template of the current block may be samples adjacent to the left side and the upper side of the current CU in FIG. 5, that is, there are no available reconstruction samples in the hatched area. That is, when there are no available adjacent reconstruction samples in the hatched area, in the TIMD mode, the planar mode is selected to perform intra prediction on the current block.
[0070] Note that, except for the case of the boundary, when encoding the current block, theoretically, the reconstruction values can be obtained from the left side and the upper side of the current block, that is, the template of the current block contains available adjacent reconstruction samples. In a specific embodiment, the decoder can use a certain intra prediction mode to predict the template, compare the predicted value with the reconstruction value, and obtain the cost of the intra prediction mode on the template. For example, Sum of Absolute Differences (SAD), Sum of Absolute Transformed Difference (SATD), Sum of Squared for Error (SSE), etc. can be mentioned. Since the template and the current block are adjacent to each other, there is a correlation between the reconstruction samples in the template and the samples in the current block. Therefore, the behavior of this prediction mode on the current block can be estimated by using the behavior of the prediction mode on the template. In TIMD, several candidate intra prediction modes are used to predict the template, obtain the costs of the candidate intra prediction modes on the template, and use the predicted values in one or two intra prediction modes with the lowest costs as the intra prediction values of the current block. When the difference between the costs corresponding to the two intra prediction modes on the template is not large, a weighted average can be performed on the predicted values of the two intra prediction modes to improve the compression performance. Optionally, the weights of the predicted values of the two prediction modes are related to the above costs. For example, the weights are inversely proportional to the costs.
[0071] In summary, in TIMD, it is possible to select the intra prediction mode by using the prediction effect of the intra prediction mode on the template and weight the two intra prediction modes according to the cost on the template. The advantage of TIMD is that when the TIMD mode is selected for the current block, it is not necessary to specifically indicate which intra prediction mode is used in the bitstream, because it is derived by the decoder itself through the above process, so a certain amount of overhead can be saved.
[0072] Through the brief introduction of the above several intra prediction modes, the following can be understood. The technical principles of the DIMD mode and the TIMD mode are similar. Both utilize the fact that the decoder performs the same operations as the encoder to estimate the prediction mode of the current coding unit. In such prediction modes, when the complexity is acceptable, the transmission of the prediction mode index can be omitted, saving overhead and improving compression efficiency. However, due to the limitations of the available reference information and the fact that it does not improve the prediction quality much itself, the DIMD mode and the TIMD mode are more effective in a wide area where the texture characteristics are consistent. When the texture changes slightly or the template area cannot be covered, the prediction effect of such prediction modes is inferior.
[0073] Also, in both the DIMD mode and the TIMD mode, the prediction blocks obtained based on multiple conventional prediction modes are fused or weighted. By fusing the prediction blocks, effects that cannot be achieved with a single prediction mode can be produced. In the DIMD mode, the planar mode is introduced as an additional weighted prediction mode to enhance the spatial correlation between adjacent reconstructed samples and prediction samples, and the prediction effect of intra prediction can be improved. However, since the prediction principle of the planar mode is relatively simple, for prediction blocks with an obvious difference between the upper right corner and the lower left corner, using the planar mode as an additional weighted prediction mode may have a reverse effect.
[0074] The following describes the content related to the transformation of the residual block.
[0075] In the encoding process, first, the current block is predicted. During prediction, spatial or temporal dependencies are utilized to obtain an image that is the same as or similar to the current block. For one block, it is possible that the predicted block and the current block are exactly the same, but it is difficult to guarantee that all blocks in a video are the same. Especially in natural videos or videos shot by a camera, due to factors such as complex texture of the image and the presence of noise in the image, usually the predicted block and the current block are similar but actually different. Also, due to irregular movements, distortions and deformations, occlusions, luminance changes, etc. in the video, it is difficult for the current block to be completely predicted. Therefore, in the hybrid coding framework, a residual image is obtained by subtracting the predicted image from the original image of the current block, or a residual block is obtained by subtracting the predicted block from the current block. The residual block is usually much simpler than the original image, so prediction can significantly improve the compression efficiency. Instead of directly encoding the residual block, usually the residual block is first transformed. The transformation is to transform the residual image from the spatial domain to the frequency domain and remove the correlation of the residual image. After the residual image is transformed to the frequency domain, since the energy often concentrates in the low-frequency region, the non-zero coefficients after transformation often concentrate in the upper left corner, and then the residual block is further compressed by quantization. Optionally, since the human eye is not sensitive to high frequencies, a larger quantization step size can be used in the high-frequency region.
[0076] Image conversion technology is to convert the original image so that the original image can be represented by orthogonal functions or orthogonal matrices, and this conversion is two-dimensional linear and reversible. Generally, the original image is called a spatial domain image, and the converted image is called a converted domain image (also called a frequency domain image), and the converted domain image can be inversely converted to the spatial domain image. Through image conversion, the characteristics of the image itself can be more effectively reflected, while the energy can be concentrated in a small amount of data, which is beneficial for image storage, transmission, and processing.
[0077] The following describes the technology related to the conversion according to the present application.
[0078] In the technical field of video coding, after obtaining a residual block, an encoder can convert the residual block. The conversion methods include primary conversion and secondary conversion. The primary conversion methods include, but are not limited to, Discrete Cosine Transform (DCT) and Discrete Sine Transform (DST). The DCTs that can be used in video coding include, but are not limited to, DCT type 2 and DCT type 8. The DST that can be used in video coding includes, but is not limited to, DST type 7. Since DCT has strong energy concentration characteristics, after the original image is DCT-converted, non-zero coefficients exist only in a partial region (for example, the upper left corner region). Of course, in video coding, an image is divided into blocks for processing, so the conversion is also performed based on blocks.
[0079] It should be noted that since all images are two-dimensional, if a two-dimensional conversion is directly performed, the amount of calculation and memory overhead are unacceptable for hardware conditions. Therefore, the above-mentioned DCT type 2, DCT type 8, and DST type 7 conversions are usually divided into one-dimensional conversions in the horizontal and vertical directions, that is, performed in two steps. For example, the horizontal conversion is performed first and then the vertical conversion, or the vertical conversion is performed first and then the horizontal conversion. The above conversion method is effective for horizontal and vertical textures, but has a poor effect on diagonal textures. Since horizontal and vertical textures are the most common, the above conversion method is very useful for improving compression efficiency.
[0080] To further improve the compression efficiency, the encoder can perform a secondary conversion on top of the primary transform.
[0081] The primary transform is used to process textures in the horizontal and vertical directions. The primary transform is also called the base transform. For example, the primary transform includes, but is not limited to, the DCT2 type, DCT8 type, and DST7 type transforms described above. The secondary transform is used to process diagonal textures. For example, the secondary transform includes, but is not limited to, the low frequency non-separable transform (LFNST). On the encoding side, the secondary transform is executed after the primary transform and before quantization. On the decoder side, the secondary transform is executed after inverse quantization and before inverse primary transform.
[0082] FIG. 6 is an illustration of LFNST according to an embodiment of the present application.
[0083] As shown in FIG. 6, on the encoding side, the LFNST secondarily transforms the low frequency coefficients in the upper left corner after the base transform. The primary transform concentrates the energy in the upper left corner by removing the correlation from the image. The secondary transform removes the correlation again for the low frequency coefficients of the primary transform. On the encoding side, when 16 coefficients are input to a 4×4 LFNST, 8 coefficients are output, and when 64 coefficients are input to an 8×8 LFNST, 16 coefficients are output. On the decoder side, when 8 coefficients are input to a 4×4 inverse LFNST, 16 coefficients are output, and when 16 coefficients are input to an 8×8 inverse LFNST, 64 coefficients are output.
[0084] When the encoder secondarily transforms the current block in the current image, the residual block of the current block can be transformed using one of the transform cores in the selected transform set. Taking the case where the secondary transform is LFNST as an example, conversion setcan refer to a set of conversion cores used to convert a certain diagonal texture, or the conversion set can include a set of conversion cores used to convert several similar diagonal textures. Of course, in other alternative embodiments, the conversion core may be referred to by terms having similar or the same meaning such as a conversion matrix, a conversion core type, or a basis function, or may be replaced by these terms, and the present application is not particularly limited.
[0085] FIG. 7 is an illustration of a conversion set of LFNST according to an embodiment of the present application.
[0086] As shown in FIGS. 7(a) to 7(d), LFNST can have four conversion sets, and the conversion cores of the same conversion set have similar diagonal textures. For example, the conversion set shown in FIG. 7(a) can be a conversion set with an index of 0, the conversion set shown in FIG. 7(b) can be a conversion set with an index of 1, the conversion set shown in FIG. 7(c) can be a conversion set with an index of 2, and the conversion set shown in FIG. 7(d) can be a conversion set with an index of 3.
[0087] Hereinafter, a correlation scheme for applying LFNST to an intra-encoded block will be described.
[0088] Intra prediction predicts the current block by referring to the reconstructed samples around the current block. Since the current video is encoded from left to right and from top to bottom, the available reference samples for the current block are usually on the left and above. Angular prediction uses the reference samples tiled at a specified angle to the current block as the prediction value. This means that there is an obvious directional texture in the predicted block, and the residual of the current block after angular prediction also shows obvious angular characteristics statistically. Therefore, the selected transform set for LFNST can be bound to the intra prediction mode. That is, after determining the intra prediction mode, LFNST can use a transform set whose texture direction corresponds to the angular characteristics of the intra prediction mode in order to save bit overhead.
[0089] Exemplarily, assume that there are four transform sets in LFNST, and each transform set has two transform cores. Table 1 shows the correspondence between the intra prediction mode and the transform set.
[0090]
Table 1
[0091] As shown in Table 1, intra prediction modes 0 to 81 can be associated with the indexes of the four transform sets.
[0092] It should be noted that the cross-component prediction modes used for chroma intra prediction are 81 to 83, and these modes do not exist for luma intra prediction. The transform sets of LFNST can process more angles using one transform set by transposition. For example, both intra prediction modes 13 to 23 and intra prediction modes 45 to 55 correspond to transform set 2. However, intra prediction modes 13 to 23 are obviously closer to the horizontal mode, and intra prediction modes 45 to 55 are obviously closer to the vertical mode. The transform corresponding to intra prediction modes 45 to 55 is adapted by transposition.
[0093] In a specific implementation, since there are four conversion sets in the LFNST, on the encoding side, based on the intra prediction mode used for the current block, the conversion set to be used in the LFNST can be determined, and further, the conversion core to be used can be determined from one of the determined conversion sets. Since the correlation between the intra prediction mode and the conversion set of the LFNST can be utilized, the transmission of the selected conversion set of the LFNST in the bitstream can be reduced. Whether the current block uses the LFNST, and if the LFNST is used, whether to use the first conversion core or the second conversion core of one conversion set can be determined by the bitstream and some conditions.
[0094] Of course, considering that there are 67 types of normal intra prediction modes and only four conversion sets in the LFNST, and each conversion set needs to occupy storage space to store the coefficients of the conversion cores in the conversion set, from a trade-off consideration of performance and complexity, multiple approximate angular prediction modes can only correspond to one conversion set. As the requirement for compression efficiency increases and the hardware capability improves, the LFNST can also be designed more complexly. For example, use larger conversion sets, more conversion sets, and each conversion set uses more conversion cores.
[0095] Exemplarily, Table 2 shows another correspondence between the intra prediction mode and the conversion set.
[0096]
Table 2
[0097] As shown in Table 2, 35 conversion sets were used, and each conversion set used three conversion cores. The correspondence between the conversion sets and the intra prediction modes is as follows. Intra prediction modes 0 to 34 correspond in the forward direction to conversion sets 0 to 34, that is, the larger the number of the prediction mode, the larger the index of the conversion set. For intra prediction modes 35 to 67, due to transposition, they correspond in the reverse direction to conversion sets 2 to 33, that is, the larger the number of the prediction mode, the smaller the index of the conversion set. The remaining prediction modes can be uniformly corresponded to the conversion set with an index of 2. That is, without considering transposition, one intra prediction mode corresponds to one conversion set. According to such a design, the residuals corresponding to each intra prediction mode can obtain more appropriate conversion sets, and the compression performance is also improved.
[0098] Of course, theoretically, the wide-angle prediction mode and the conversion set can also correspond one-to-one, but the cost performance of such a design is lower, and the present application does not specifically describe this. Note that LFNST is only an example of the secondary conversion and should not be understood as a limitation of the secondary conversion. For example, LFNST is an inseparable secondary conversion, and in other alternative embodiments, a separable secondary conversion can also be used to improve the compression efficiency of the residual of the diagonal texture, and the present application is not specifically limited thereto.
[0099] FIG. 8 is a block diagram showing a decoding framework 200 according to an embodiment of the present application.
[0100] As shown in FIG. 8, the decoding framework 200 may include an entropy decoding unit 210, an inverse transform and inverse quantization unit 220, a residual unit 230, an intra prediction unit 240, an inter prediction unit 250, a loop filtering unit 260, and a decoded image buffer unit 270. The entropy decoding unit 210 receives and analyzes a bitstream to obtain a prediction block and a frequency-domain residual block. For the frequency-domain residual block, steps such as inverse transform and inverse quantization can be performed through the inverse transform and inverse quantization unit 220 to obtain a time-domain residual block. The residual unit 230 can obtain a reconstructed block by adding the prediction block obtained by being predicted by the intra prediction unit 240 or the inter prediction unit 250 to the time-domain residual block obtained after inverse transform and inverse quantization are performed through the inverse transform and inverse quantization unit 220.
[0101] FIG. 9 is a flowchart showing a decoding method 300 according to an embodiment of the present application. Note that the decoding method 300 can be executed by a decoder. For example, the decoding method 300 is applied to the decoding framework 200 shown in FIG. 8 To facilitate the explanation, the decoder will be described as an example hereinafter.
[0102] As shown in FIG. 9, the decoding method 300 may include the following steps. S310: Analyze the bitstream of the current sequence to obtain the First transformation coefficients of the current block. S320: Determine the first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of a current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. S330: Based on a conversion set corresponding to the first intra prediction mode, perform a first conversion on the first conversion coefficient to obtain a second conversion coefficient of the current block. S340: Perform a second conversion on the second conversion coefficient to obtain a residual block of the current block. S350: Determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.
[0103] Exemplarily, when the decoder uses an intra prediction mode derived from the DIMD mode for the reconstructed samples in the first template region, first, calculate the gradient values of the reconstructed samples in the first template region, and then, an intra prediction mode that matches the gradient direction of the reconstructed sample with the largest gradient value in the first template region can be determined as the intra prediction mode derived from the DIMD mode. Alternatively, the decoder can calculate the gradient values corresponding to each intra prediction mode by traversing the intra prediction modes based on the reconstructed samples in the first template region, and determine the intra prediction mode with the largest gradient value as the intra prediction mode derived from the DIMD mode.
[0104] Exemplarily, when the decoder uses an intra prediction mode derived from the DIMD mode for the prediction block of the current block (or the output vector of the optimal MIP mode), first, the gradient value of the prediction samples of the prediction block of the current block (or the output vector of the optimal MIP mode) is calculated, and then, the intra prediction mode that matches the gradient direction of the prediction sample with the largest gradient value among the prediction samples of the prediction block of the current block (or the output vector of the optimal MIP mode) can be determined as the intra prediction mode derived from the DIMD mode. Or, the decoder can calculate the gradient value corresponding to each intra prediction mode by traversing the intra prediction mode based on the prediction samples of the prediction block of the current block (or the output vector of the optimal MIP mode), and determine the intra prediction mode with the largest gradient value as the intra prediction mode derived from the DIMD mode.
[0105] Exemplarily, the first transformation is used to process the diagonal texture in the current block.
[0106] Exemplarily, the second transformation is used to process the horizontal and vertical textures in the current block.
[0107] It should be understood that the first transformation is the inverse transformation of the second transformation on the encoding side, and the second transformation is the inverse transformation of the basic transformation on the encoding side. For example, the first transformation can be an inverse LFNST, and the second transformation can also be an inverse DCT type 2, inverse DCT type 8, or inverse DCT type 7, etc.
[0108] Of course, the method of adapting the TMMIP technology and LFNST is also applicable to other secondary transformation methods. For example, LFNST is an inseparable secondary transformation, and in other alternative embodiments, the TMMIP technology is also applicable to separable secondary transformations, and the present application is not particularly limited thereto.
[0109] What needs to be noted is that when the encoder or decoder predicts the current block, it may perform LFNST using the conversion set corresponding to the PLANAR mode, and the reasons are as follows. Since the conversion core used for LFNST is obtained through deep learning training with the dataset of the conventional intra prediction mode, in the normal intra prediction process, the conversion core used for LFNST is usually also the conversion core selected from the conversion set of LFNST corresponding to the conventional intra prediction mode. However, the encoder or decoder may predict the current block using a non-traditional intra prediction mode. In this case, considering that the planar mode is usually used to process blocks where there is a gradient in the texture, and LFNST is used to process diagonal textures, the texture information of the predicted block usually output in the planar mode and the texture information of the planar mode in the conventional intra prediction mode can be processed as the same type of texture. That is, when the encoder or decoder predicts the current block using a non-traditional intra prediction mode, both perform LFNST using the conversion set corresponding to the planar mode. For example, when the encoder predicts the current block using the MIP mode, it performs LFNST using the conversion set corresponding to the planar mode. However, the meaning represented by the MIP mode is different from the meaning represented by the conventional intra prediction mode. That is, the conventional intra prediction mode has an obvious directionality, while the MIP mode is only an index of the matrix coefficient. Therefore, although the planar mode is used to process blocks where there is a gradient in the texture, it does not necessarily match the texture information of the current block. That is, the texture direction of the conversion set used by LFNST does not necessarily match the texture direction of the current block, which reduces the decompression performance of the current block.
[0110] In view of this, an embodiment of the present application introduces a first intra prediction mode, and performs a first transformation on the first transformation coefficient of the current block based on a transformation set corresponding to the first intra prediction mode, so as to further improve the decompression performance of the current block. In particular, when the decoder predicts the current block using a non-traditional intra prediction mode, it is possible to avoid directly performing the first transformation using the transformation set corresponding to the planar mode. The transformation set corresponding to the first intra prediction mode can reflect the texture direction of the current block to a certain extent, and thus can improve the decompression performance of the current block.
[0111] The following 3 and Table 4 Based on the test results, the beneficial effects of the technical solution provided by the present application will be described.
[0112] Table 3 is the result obtained by testing a test sequence when the current block is weighted predicted using the optimal MIP mode and the sub-optimal MIP mode, and the first intra prediction mode is designed as an intra prediction mode derived from the DIMD mode used for the predicted block of the current block. Table 4 is the result obtained by testing a test sequence when the current block is weighted predicted using the optimal MIP mode and an intra prediction mode derived from the decoder-side intra mode derivation (DIMD) mode used for the reconstruction samples within the first template region adjacent to the current block, and the first intra prediction mode is designed as an intra prediction mode used for the predicted block of the current block.
[0113] [Table 3]
[0114] [Table 4]
[0115] Table 3 and Table 4 As shown in, a negative incremental bit rate (BD-rate) represents an improvement in the performance of the solution provided by the present application with respect to the test results of ECM2.0. From the test results, under the general test conditions, the test results in Table 3 and Table 4 can all provide a luminance performance gain of 0.20% on average, and it was found that the 4K sequences are not vulgar. It should be noted that the TIMD prediction mode of ECM2.0 integration has a high complexity based on ECM1.0 and only a 0.4% performance gain. When it is difficult to improve the current intra-coding performance, the solution provided by the present application can bring good performance gains without increasing the complexity of the decoder. In particular, for 4K-type video sequences, the performance gains are obvious. Also, due to server load, even if the encoding time and decoding time vary slightly, theoretically the decoding time hardly increases.
[0116] In some embodiments, the output vector of the optimal MIP mode is the vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is the vector after upsampling the output vector of the optimal MIP mode.
[0117] In other words, the process of using the intra prediction mode derived from the DIMD mode for the output vector of the optimal MIP mode may be executed before upsampling the output vector of the optimal MIP mode, or may be executed after upsampling the output vector of the optimal MIP mode, and is not particularly limited in the present application.
[0118] After the decoder inputs the reference sample into the prediction matrix in the optimal MIP mode to obtain the output vector, the output vector in the optimal MIP mode has a maximum of 64 prediction samples. Compared with the maximum of thousands of prediction samples that the upsampled prediction block has, before upsampling the output vector in the optimal MIP mode, the decoder is derived from the DIMD mode ta i The intra prediction mode can be used to reduce the computational complexity and further improve the decompression performance of the current block. For example, the decoder can effectively reduce the computational complexity by calculating the gradient amplitude value of each conventional prediction mode using DIMD before upsampling.
[0119] In some embodiments, S350 may include that the decoder determines a first intra prediction mode based on the prediction mode for predicting the current block.
[0120] Exemplarily, the decoder determines the first intra prediction mode based on the mode type of the prediction mode for predicting the current block.
[0121] Exemplarily, the decoder determines the first intra prediction mode based on the derivation mode of the prediction mode for predicting the current block.
[0122] Exemplarily, the derivation mode of the prediction mode for predicting the current block includes, but is not limited to, the MIP mode, the DIMD mode, and the TIMD mode.
[0123] In some embodiments, when the prediction mode for predicting the current block includes an optimal MIP mode and a sub-optimal MIP mode for predicting the current block, the decoder determines, as a first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determines, as a first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode.
[0124] In other words, when the decoder performs weighted prediction of the current block using the optimal MIP mode and the sub-optimal MIP mode, the decoder determines, as a first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determines, as a first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode.
[0125] In this embodiment, when the prediction mode for predicting the current block includes an optimal MIP mode and a sub-optimal MIP mode, the decoder first sets, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block. Then, the texture direction of the transform set corresponding to the first intra prediction mode simultaneously conforms to the texture characteristics indicated by the predicted block of the current block by the optimal MIP mode and the texture characteristics indicated by the predicted block of the current block by the sub-optimal MIP mode. to doIt is possible to further improve the decompression performance of the current block as much as possible. When the decoder first sets the intra prediction mode derived from the DIMD mode, which is used for the output vector of the optimal MIP mode, as the first intra prediction mode, during the process of determining the optimal MIP mode, the output vector of the optimal MIP mode can be directly obtained, reducing the decompression complexity. Moreover, the texture direction of the conversion set corresponding to the first intra prediction mode can conform to the texture characteristics indicated by the predicted block of the current block in the optimal MIP mode, and thus the decompression performance of the current block can be improved as much as possible.
[0126] Of course, in other alternative embodiments, when the prediction modes for predicting the current block include the optimal MIP mode and the quasi-optimal MIP mode, the decoder can also determine the intra prediction mode derived from the DIMD mode or the intra prediction mode derived from the TIMD mode, which is used for the reconstructed samples in the first template area, as the first intra prediction mode, and it is not specifically limited in this application.
[0127] In some embodiments, when the prediction modes for predicting the current block include the optimal MIP mode and the intra prediction mode derived from the TIMD mode, the decoder determines the intra prediction mode derived from the DIMD mode, which is used for the predicted block of the current block, as the first intra prediction mode, or the decoder determines the intra prediction mode derived from the TIMD mode as the first intra prediction mode.
[0128] In other words, when the second intra prediction mode is the intra prediction mode derived from the TIMD mode, the decoder determines the intra prediction mode derived from the DIMD mode, which is used for the predicted block of the current block, as the first intra prediction mode, or the decoder determines the intra prediction mode derived from the TIMD mode as the first intra prediction mode.
[0129] In this embodiment, when the second intra prediction mode includes an intra prediction mode derived from the TIMD mode, the decoder first sets the intra prediction mode derived from the DIMD mode, which is used for the prediction block of the current block, as the first intra prediction mode. Then, the texture direction of the transform set corresponding to the first intra prediction mode can simultaneously conform to the texture characteristics indicated by the prediction block of the current block in the optimal MIP mode and the texture characteristics indicated by the prediction block of the current block in the intra prediction mode derived from the TIMD mode. to do This can further improve the decompression performance of the current block as much as possible. If the decoder first sets the intra prediction mode derived from the TIMD mode as the first intra prediction mode, it can directly determine the second intra prediction mode as the first intra prediction mode. After reducing the decompression complexity, the texture direction of the transform set corresponding to the first intra prediction mode can conform to the texture characteristics indicated by the prediction block of the current block in the optimal MIP mode, and thus the decompression performance of the current block can be improved as much as possible.
[0130] Of course, in other alternative embodiments, when the second intra prediction mode is the intra prediction mode derived from the TIMD mode, the decoder can also determine, as the first intra prediction mode, the intra prediction mode derived from the DIMD mode used for the output vector of the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstructed samples in the first template area. The present application is not particularly limited in this regard.
[0131] In some embodiments, if the prediction mode for predicting the current block includes the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, the decoder determines the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or determines the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region as the first intra prediction mode.
[0132] In other words, if the second intra prediction mode is the intra prediction mode derived from the DIMD mode, the decoder determines the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or the decoder determines the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region as the first intra prediction mode.
[0133] In this embodiment, when the second intra prediction mode includes the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, the decoder first sets the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode. Then, the texture direction of the transform set corresponding to the first intra prediction mode simultaneously conforms to the texture characteristics indicated by the predicted block of the current block in the optimal MIP mode and the texture characteristics indicated by the predicted block of the current block in the intra prediction mode derived from the DIMD mode. to doIt is possible to further improve the decompression performance of the current block as much as possible. If the decoder first sets the intra prediction mode derived from the DIMD mode used for the reconstruction samples in the first template area as the first intra prediction mode, it can directly determine the second intra prediction mode as the first intra prediction mode. After reducing the decompression complexity, the texture direction of the transform set corresponding to the first intra prediction mode can conform to the texture characteristics indicated by the prediction block of the current block by the optimal MIP mode, and thus the decompression performance of the current block can be improved as much as possible.
[0134] Of course, in other alternative embodiments, when the second intra prediction mode is the intra prediction mode derived from the DIMD mode used for the reconstruction samples in the first template area, the decoder can also determine the intra prediction mode derived from the DIMD mode and the intra prediction mode derived from the TIMD mode used for the output vector of the optimal MIP mode as the first intra prediction mode, and the present application is not particularly limited thereto.
[0135] In some embodiments, the decoding method 300 may further include the following. Determine a second intra prediction mode. The second intra prediction mode includes any one of a quasi-optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for the reconstruction samples in the first template area, and an intra prediction mode derived from the TIMD mode. Predict the current block based on the optimal MIP mode and the second intra prediction mode to obtain a prediction block of the current block.
[0136] Exemplarily, the process in which the decoder predicts the current block based on the optimal MIP mode and the second intra prediction mode can also be abbreviated as Template Matching MIP (TMMIP) technology, a TMMIP-based prediction mode derivation method, or a TMMIP fusion enhancement technology. That is, after obtaining the residual block of the current block, the decoder can enhance the performance of the prediction process of the current block based on the derived optimal MIP mode and the second intra prediction mode. In other words, in the TMMIP technology, at least one of the quasi-optimal MIP prediction mode, the intra prediction mode derived from the TIMD mode, the intra prediction mode derived from the DIMD mode used for the reconstructed samples in the first template region adjacent to the current block, and the optimal MIP prediction mode can be utilized to enhance the performance of the prediction process of the current block.
[0137] In this embodiment, the decoder predicts the current block based on the optimal MIP mode and the second intra prediction mode. The optimal MIP mode is determined based on the distortion costs corresponding to a plurality of MIP modes and is designed to be the optimal MIP mode for predicting the current block. The second intra prediction mode is determined based on the distortion costs corresponding to a plurality of MIP modes and is designed to include at least one of the quasi-optimal MIP mode for predicting the current block, the intra prediction mode derived from the DIMD mode used for the reconstructed samples in the first template region adjacent to the current block, and the intra prediction mode derived from the TIMD mode. This helps to avoid the decoder obtaining the MIP mode by analyzing the bitstream. Compared with the conventional MIP technology, in the present application, the bit overhead at the coding unit level can be effectively reduced, thereby improving the decompression efficiency of the current block.
[0138] Specifically, the bit overhead of the MIP mode is larger than that of other intra prediction modes. Not only is a flag indicating whether the MIP mode is used required, but also a flag indicating whether the MIP mode is transposed is needed. Finally, the part with the largest overhead is that truncated binary coding needs to be used to represent the index of the MIP mode. The MIP mode is a technology simplified based on neural network technology and is quite different from the conventional interpolation filtering prediction technology. For some special textures, the MIP mode is more effective than the conventional intra prediction mode, but its large flag overhead is a defect of the MIP mode. For example, in a 4×4-sized coding unit, there are a total of 16 prediction modes in the MIP mode, and its bit overhead includes a usage flag for one MIP mode, a transposition flag for one MIP mode, and a 5-bit or 6-bit truncated binary flag. In view of this, in this application, a method is utilized in which the decoder autonomously determines the optimal MIP mode for predicting the current block and determines the intra prediction mode of the current block based on the optimal MIP mode, so that an overhead of up to 5 bits or 6 bits can be saved, the bit overhead at the coding unit level can be effectively reduced, and thereby the decompression efficiency can be improved.
[0139] Also, saving an overhead of up to 5 bits or 6 bits for each coding unit is premised on the fact that the template matching-based prediction mode derivation algorithm needs to be very accurate. If the accuracy of the template matching-based prediction mode derivation algorithm is too low, decoding the MIP mode derived by the decoder side will be different from the MIP mode derived by the encoding side, and furthermore, the coding performance will deteriorate. Or, the coding performance depends on the accuracy of the template matching-based prediction mode derivation algorithm.
[0140] However, both the template-based derivation algorithm in the conventional intra prediction mode and the template matching-based derivation algorithm in the inter prediction mode do not meet the expected accuracy. Although they can save bit overhead and improve compression efficiency, as the number of prediction modes of the template matching-based prediction mode derivation algorithm increases, the additional bit overhead at the coding unit level brought by the template matching-based prediction mode derivation algorithm cannot enable subsequent technologies to improve the compression efficiency relying only on the template matching-based prediction mode derivation algorithm. Therefore, the template matching-based prediction mode derivation algorithm needs to improve the coding performance while improving the compression efficiency. As one possible embodiment, by saving the bit overhead at the coding unit level and creating different new prediction blocks, prediction diversity and selection diversification can be guaranteed, and the coding performance can be improved. In view of this, in this application, by fusing the optimal MIP mode and the second intra prediction mode, that is, by performing a fusion prediction on the current block based on the optimal MIP mode and the second intra prediction mode, it is possible to avoid completely replacing the optimal prediction mode calculated based on the rate distortion cost with the optimal MIP mode, and to achieve both prediction accuracy and prediction diversity. Furthermore, the decompression performance can be improved.
[0141] In particular, since the TMMIP technology predicts the current block by combining the optimal MIP mode and the second intra prediction mode, the predicted blocks obtained by predicting the current block using different prediction modes may have different texture characteristics. Therefore, when the current block selects the TMMIP technology, the predicted block of the current block shows one texture characteristic in the optimal MIP mode, and the predicted block of the current block shows another texture characteristic in the second intra prediction mode. In other words, after predicting the current block, statistically speaking, the residual block of the current block can also show two texture characteristics. That is, the residual block of the current block does not necessarily conform to the rules that a certain prediction mode can embody. At this time, for the TMMIP technology, when the first intra prediction mode is an intra prediction mode derived from the DIMD mode used for the predicted block of the current block, the texture direction of the conversion set corresponding to the first intra prediction mode can simultaneously conform to the texture characteristics shown by the predicted block of the current block in the optimal MIP mode and the texture characteristics shown by the predicted block of the current block in the second intra prediction mode to do This can improve the decompression performance of the current block. Furthermore, when the first intra prediction mode is an intra prediction mode derived from the DIMD mode or an intra prediction mode derived from the TIMD mode used for the reconstructed samples in the first template area, in the process of determining the second intra prediction mode, the first intra prediction mode can be directly determined. After reducing the decompression complexity, the texture direction of the conversion set corresponding to the first intra prediction mode can simultaneously conform to the texture characteristics shown by the predicted block of the current block in the optimal MIP mode and the texture characteristics shown by the predicted block of the current block in the second intra prediction mode to do This can improve the decompression efficiency.
[0142] In some embodiments, the decoder first predicts the current block based on the optimal MIP mode to obtain a first predicted block. Next, the decoder predicts the current block based on a second intra prediction mode to obtain a second predicted block. Then, based on the weight of the optimal MIP mode and the weight of the second intra prediction mode, a weighting process is performed on the first predicted block and the second predicted block to obtain a predicted block of the current block.
[0143] In some embodiments, before performing a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block, the decoding method 300 may further include the following steps. When the prediction mode for predicting the current block includes the optimal MIP mode, a sub-optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, determine the weight of the optimal MIP mode and the weight of the second intra prediction mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode. When the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode used for the reconstructed samples within the first template region, determine that both the weight of the optimal MIP mode and the weight of the second intra prediction mode are preset values.
[0144] In some embodiments, the decoder predicts the current block based on the optimal MIP mode to obtain a first predicted block. The decoder predicts the current block based on a second intra prediction mode to obtain a second predicted block. Then, the decoder performs a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block.
[0145] Exemplarily, the decoder can directly intra-predict the current block based on the optimal MIP mode to obtain a first predicted block. Further, the decoder can directly obtain an optimal prediction mode and a sub-optimal prediction mode based on the TIMD mode, predict the current block, and obtain a second predicted block. For example, neither the optimal prediction mode nor the sub-optimal prediction mode is a DC mode (also called an average value mode) or a plane mode (also called a flat mode), and when the distortion cost of the sub-optimal prediction mode is smaller than twice the distortion cost of the optimal prediction mode, a prediction block fusion operation is required. That is, the decoder can first intra-predict the current block based on the optimal prediction mode to obtain an optimal predicted block, then intra-predict the current block based on the sub-optimal prediction mode to obtain a sub-optimal predicted block, calculate the weight value of the optimal predicted block and the weight value of the sub-optimal predicted block using the ratio of the distortion cost of the optimal prediction mode to the distortion cost of the sub-optimal prediction mode, and finally, perform weighted fusion on the optimal predicted block and the sub-optimal predicted block to obtain a second predicted block. Further, when the optimal prediction mode or the sub-optimal prediction mode is a plane mode or a DC mode, or when the distortion cost of the sub-optimal prediction mode is greater than twice the distortion cost of the optimal prediction mode, a prediction block fusion operation is not required. That is, only the optimal predicted block obtained based on the optimal prediction mode can be directly used as the second predicted block. After obtaining the first predicted block and the second predicted block, the decoder performs a weighting process on the first predicted block and the second predicted block to obtain a predicted block of the current block.
[0146] In some embodiments, if the prediction mode for predicting the current block includes the optimal MIP mode, a sub-optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, the decoder determines the weight of the optimal MIP mode and the weight of the second intra prediction mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode. If the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, the decoder determines that both the weight of the optimal MIP mode and the weight of the second intra prediction mode are preset values.
[0147] In some embodiments, S320 may include the following. The decoder analyzes the bitstream of the current sequence to obtain a first flag. When the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the second intra prediction mode, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0148] Exemplarily, when the value of the first flag is a first value, the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the second intra prediction mode. When the value of the first flag is a second value, the first flag is used to indicate that it is not permitted to predict the image block in the current sequence using the optimal MIP mode and the second intra prediction mode. In one embodiment, the first value is 1 and the second value is 0. In another embodiment, the first value is 0 and the second value is 1. Of course, the first value and the second value may be other values, and the present application is not limited thereto.
[0149] Exemplarily, when the first flag is true, the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode. When the first flag is false, the first flag is used to indicate that it is not permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode.
[0150] Exemplarily, the decoder analyzes a block-level flag. When an intra prediction mode is used for the current block, the decoder analyzes or obtains the first flag. When the first flag is true, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0151] Exemplarily, the first flag is denoted as sps_timd_enable_flag, in which case the decoder analyzes or obtains the sps_timd_enable_flag. When the sps_timd_enable_flag is true, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0152] Exemplarily, the first flag is a sequence-level flag.
[0153] The description that the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode can be replaced with a description having a similar or identical meaning. For example, in other alternative embodiments, the description that the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode can also be replaced with any of the following. The first flag is used to indicate that it is permitted to determine the intra prediction mode of an image block in the current sequence using the TMMIP technique. The first flag is used to indicate that it is permitted to perform intra prediction on an image block in the current sequence using the TMMIP technique. The first flag is used to indicate that it is permitted to use the TMMIP technique for an image block in the current sequence. The first flag is used to indicate that it is permitted to predict an image block in the current sequence using the MIP mode determined based on a plurality of MIP modes.
[0154] Furthermore, in other alternative embodiments, when combining the TMMIP technique with other techniques, it is also possible to indirectly indicate whether the TMMIP technique is permitted to be used in the current sequence by an enable flag of the other technique. For example, taking the TIMD technique as an example, when the first flag is used to indicate that it is permitted to use the TIMD technique in the current sequence, it also indicates that it is permitted to use the TMMIP technique in the current sequence. In other words, when the first flag is used to indicate that it is permitted to use the TIMD technique in the current sequence, it indicates that it is permitted to use both the TIMD technique and the TMMIP technique in the current sequence. Thereby, further saving bit overhead.
[0155] In some embodiments, when the first flag is used to indicate that predicting an image block in the current sequence using the optimal MIP mode and the second intra prediction mode is permitted, the decoder analyzes the bitstream to obtain the second flag. When the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0156] Exemplarily, the decoder analyzes a block-level flag. When an intra prediction mode is used for the current block, the decoder analyzes or obtains the first flag. When the first flag is true, the decoder analyzes or obtains the second flag. When the second flag is true, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0157] Exemplarily, when the value of the second flag is a third value, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted. When the value of the second flag is a fourth value, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is not permitted. In one embodiment, the third value is 1 and the fourth value is 0. In another embodiment, the third value is 0 and the fourth value is 1. Of course, the third value and the fourth value may be other values, and the present application is not limited thereto.
[0158] Exemplarily, when the second flag is true, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted. When the second flag is false, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is not permitted.
[0159] Exemplarily, the first flag is denoted as sps_timd_enable_flag, and the second flag is denoted as cu_timd_enable_flag. In this case, the decoder analyzes or obtains the sps_timd_enable_flag. When the sps_timd_enable_flag is true, the decoder analyzes or obtains the cu_timd_enable_flag. When the cu_timd_enable_flag is true, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0160] Exemplarily, the second flag is a block-level flag or a coding unit-level flag.
[0161] The description that the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted can be replaced with a description having a similar or identical meaning. For example, in other alternative embodiments, the description that the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted can also be replaced with any of the following. The second flag is used to indicate that it is permitted to determine the intra prediction mode of the current block using the TMMIP technology. The second flag is used to indicate that it is permitted to perform intra prediction on the current block using the TMMIP technology. The second flag is used to indicate that it is permitted to use the TMMIP technology for the image block of the current block. The second flag is used to indicate that it is permitted to predict the current block using the MIP mode determined based on a plurality of MIP modes.
[0162] Furthermore, in other alternative embodiments, when combining the TMMIP technology with other technologies, it is also possible to indirectly indicate whether it is permitted to use the TMMIP technology for the current block by the permission flag of the other technology. For example, taking the TIMD technology as an example, when the second flag is used to indicate that it is permitted to use the TIMD technology for the current block, it also indicates that it is permitted to use the TMMIP technology for the current block. In other words, when the second flag is used to indicate that it is permitted to use the TIMD technology for the current block, it indicates that it is permitted to use both the TIMD technology and the TMMIP technology for the current block. Thereby, further saving bit overhead.
[0163] Also, when the decoding side analyzes the second flag, it can analyze the second flag before analyzing the residual block of the current block, or it can also analyze the second flag after analyzing the residual block of the current block. This application does not particularly limit it in this regard.
[0164] In some embodiments, the method 300 may further include the following. The decoder determines an optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes. The distortion costs corresponding to the plurality of MIP modes include distortion costs obtained by predicting samples within a second template region adjacent to the current block using the plurality of MIP modes.
[0165] Exemplarily, before determining an optimal MIP mode for predicting the current block based on distortion costs corresponding to a plurality of MIP modes, the decoder needs to calculate the distortion cost corresponding to each of the plurality of MIP modes and sort the plurality of MIP modes based on the distortion cost corresponding to each MIP mode. The MIP mode with the minimum cost is the optimal prediction result.
[0166] Also, the distortion cost related to the decoder in this application is different from the rate-distortion cost (RDcost) related to the encoder. The rate-distortion cost is the distortion cost used when the encoding side determines a specific intra prediction technique from a plurality of intra prediction techniques, and the rate-distortion cost can be the cost obtained by comparing the distorted image and the original image. Since the decoder cannot obtain the original image, the distortion cost related to the decoder can be the distortion cost between the reconstructed sample and the predicted sample. For example, it can be the SATD (Sum of Absolute Transformed Difference) cost between the reconstructed sample and the predicted sample, or the cost that can be used to calculate the difference between the reconstructed sample and the predicted sample.
[0167] Of course, in other alternative embodiments, the decoder first determines the array order of a plurality of MIP modes based on the distortion costs corresponding to the plurality of MIP modes, then determines the coding method used for the optimal MIP mode based on the array order of the plurality of MIP modes, and then decodes the bitstream of the current sequence based on the coding method used for the optimal MIP mode to obtain the index of the optimal MIP mode.
[0168] For example, the codeword length of the coding method used for the first n MIP modes in the array order is smaller than the codeword length of the coding method used for the MIP modes following the nth MIP mode in the array order, and / or a variable-length coding method is used for the first n MIP modes, and a truncated binary coding method is used for the MIP modes following the nth MIP mode. Exemplarily, n can be any value greater than or equal to 1. In the conventional MIP technology, the index of the MIP mode is usually binarized by selecting a truncated binary method similar to equal probability coding. That is, in the truncated binary method, all prediction modes are divided into two segments, one segment is represented by N codewords, and the other segment is represented by N + 1 codewords. In view of this, in the present application, before determining the optimal MIP mode for predicting the current block based on the distortion costs corresponding to the plurality of MIP modes, the decoder can calculate the distortion cost corresponding to each of the plurality of MIP modes and sort the plurality of MIP modes based on the distortion cost corresponding to each MIP mode. Finally, the decoder selects and uses a more flexible variable-length coding method according to the array order of the plurality of MIP modes. By flexibly setting the coding method of the MIP mode compared with equal probability coding, the bit overhead of the index of the MIP mode can be saved.
[0169] Note that the arrangement order is the order obtained by the decoder arranging a plurality of MIP modes in ascending order of distortion cost. The smaller the distortion cost corresponding to the MIP mode, the higher the probability that the encoder performs intra prediction on the current block using that MIP mode. Therefore, the codeword length of the coding method used for the first n MIP modes in the arrangement order is designed to be smaller than the codeword length of the coding method used for the MIP modes following the nth MIP mode in the arrangement order, and / or the coding method used for the first n MIP modes is designed as a variable-length coding method, and the coding method used for the MIP modes following the nth MIP mode is designed as a truncated binary coding method. In this way, the MIP modes that are highly likely to be used by the encoder use relatively short codeword lengths or variable-length coding methods. Thereby, the bit overhead of the index of the MIP mode can be saved, and the decompression performance can be improved.
[0170] In some embodiments, the decoding method 300 may further include the following. When the second intra prediction mode is a sub-optimal MIP mode, the decoder determines whether to adopt the sub-optimal MIP mode to predict the current block based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the sub-optimal MIP mode. If it is determined not to adopt the sub-optimal MIP mode, the decoder can directly predict the current block based on the optimal MIP mode. If it is determined to adopt the sub-optimal MIP mode, the decoder can predict the current block based on the optimal MIP mode and the sub-optimal MIP mode to obtain the predicted block of the current block.
[0171] Exemplarily, when the second intra prediction mode is the quasi-optimal MIP mode, if the ratio of the distortion cost corresponding to the optimal MIP mode to the distortion cost corresponding to the quasi-optimal MIP mode is less than or equal to a preset ratio, the decoder can directly predict the current block based on the optimal MIP mode to obtain the predicted block of the current block. Or, when the second intra prediction mode is the quasi-optimal MIP mode, if the ratio of the distortion cost corresponding to the quasi-optimal MIP mode to the distortion cost corresponding to the optimal MIP mode is greater than or equal to a preset ratio, the decoder can directly predict the current block based on the optimal MIP mode to obtain the predicted block of the current block. For example, when the distortion cost corresponding to the quasi-optimal MIP mode is a multiple (e.g., 2 times) or more of the distortion cost corresponding to the optimal MIP mode, the quasi-optimal MIP mode already has a large distortion and is not suitable for the current block, that is, it can be interpreted that it is possible to predict the current block using only the optimal MIP mode without the fusion enhancement technology.
[0172] In this embodiment, the decoder determines whether to adopt the quasi-optimal MIP mode to predict the current block based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the quasi-optimal MIP mode, which is equivalent to the decoder determining whether to adopt the quasi-optimal MIP mode to improve the performance of the optimal MIP mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the quasi-optimal MIP mode. Thereby, it is possible to avoid carrying a flag for determining whether to adopt the quasi-optimal MIP mode to improve the performance of the optimal MIP mode in the bitstream, save bit overhead, and as a result, improve the decompression performance.
[0173] In some embodiments, the second template area and the first template area may be the same or different.
[0174] Exemplarily, the size of the second template area can be predefined according to the size of the current block. For example, the width of the area adjacent to the upper side of the current block within the second template area is equal to the width of the current block, and its height is equal to at least the height of one row of samples. The height of the area adjacent to the left side of the current block within the second template area is equal to the height of the current block, and its width is equal to the width of 2 column samples. Of course, in other alternative embodiments, the second template area can be realized as a second template area of other sizes or magnitudes, and the present application is not specifically limited thereto.
[0175] In some embodiments, the method 300 may further include the following. The decoder predicts samples within the second template area based on a third flag and a plurality of MIP modes, and obtains distortion costs corresponding to the plurality of MIP modes under each state of the third flag. The third flag is used to indicate whether to transpose the input vector and the output vector corresponding to the MIP mode. The decoder determines an optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes under each state of the third flag.
[0176] Exemplarily, the decoder predicts samples within the second template area based on a third flag and a plurality of MIP modes, and obtains distortion costs corresponding to the plurality of MIP modes under each state of the third flag, before determining an optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes.
[0177] As described above, in the conventional MIP technology, its bit overhead is larger compared to other intra prediction tools. Not only is a flag indicating whether the MIP technology is used required, but also a flag indicating whether the MIP is transposed is needed. Finally, the part with the largest overhead is that truncated binary coding needs to be used to represent the MIP prediction mode. The MIP technology is a technology simplified based on neural network technology and is quite different from the conventional interpolation filtering prediction technology. For some special textures, the MIP prediction mode is more effective than the conventional intra prediction mode, but its large flag overhead is a defect of the MIP technology. For example, a 4×4 size coding unit has a total of 16 prediction samples, but its bit overhead includes one MIP usage flag, one MIP transpose flag, and a 5-bit or 6-bit truncated binary flag. In view of this, in the present application, when determining the optimal MIP mode, by traversing each state of the third flag, considering the transpose function of the MIP mode, the overhead of one MIP transpose flag can be saved, and the decompression efficiency can be further improved.
[0178] Exemplarily, the decoder traverses each state of the third flag and a plurality of MIP modes, determines the distortion cost corresponding to the plurality of MIP modes under each state of the third flag, and determines the optimal MIP mode based on the distortion cost corresponding to the plurality of MIP modes under each state of the third flag. Alternatively, the decoder traverses each state of the third flag and a plurality of MIP modes, determines the distortion cost under each state of the third flag corresponding to the plurality of MIP modes, and determines the optimal MIP mode based on the distortion cost under each state of the third flag corresponding to the plurality of MIP modes. That is, on the decoding side, it may first traverse a plurality of MIP modes, or first traverse the state of the third flag.
[0179] Exemplarily, when the value of the third flag is the fifth value, the third flag is used to indicate that the input vector and output vector corresponding to the MIP mode are to be transposed. When the value of the third flag is the sixth value, the third flag is used to indicate that the input vector and output vector corresponding to the MIP mode are not to be transposed. In this case, each state of the third flag can be replaced by each value of the third flag. In one embodiment, the fifth value is 1 and the sixth value is 0. In another embodiment, the fifth value is 0 and the sixth value is 1. Of course, the fifth value and the sixth value may be other values, and the present application is not limited thereto.
[0180] Exemplarily, when the third flag is true, the third flag is used to indicate that the input vector and output vector corresponding to the MIP mode are to be transposed. When the third flag is false, the third flag is used to indicate that the input vector and output vector corresponding to the MIP mode are not to be transposed. In this case, whether the third flag is true or false is a state of the third flag.
[0181] Exemplarily, the third flag is a sequence-level flag, a block-level flag, or a coding unit-level flag.
[0182] Exemplarily, the third flag may also be referred to as a transpose message, a transpose flag, or a MIP transpose flag.
[0183] Note that the description that the third flag is used to indicate whether to transpose the input vector and output vector corresponding to the MIP mode can be replaced by descriptions having a similar or identical meaning. For example, in other alternative embodiments, the third flag is used to indicate whether it is necessary to transpose the input and output corresponding to the MIP mode. The third flag is used to indicate whether the input vector and output vector corresponding to the MIP mode are the transposed vectors. The third flag is used to indicate whether to transpose.
[0184] In some embodiments, the decoding method 300 may further include the following. When the size of the current block is a preset size, the decoder determines an optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes.
[0185] Exemplarily, the preset size may include a size with a preset width and a preset height. In other words, when the width of the current block is a preset width and the height is a preset height, the decoder determines an optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes.
[0186] Exemplarily, the preset size can be realized by pre-storing other ways that can be used to indicate corresponding codes, tables, or related information in a device (for example, including a decoder and an encoder). In this application, the specific embodiments thereof are not limited. For example, the preset size can refer to a size defined in a protocol. Optionally, the "protocol" can refer to a standard protocol in the field of coding technology, and can include related protocols such as the VCC protocol or the ECM protocol, etc.
[0187] Of course, in other alternative embodiments, the decoder can also determine whether to determine an optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes based on a preset size through other means. In this application, it is not specifically limited thereto.
[0188] For example, the decoder can determine whether to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes only based on the width or height of the current block. In one embodiment, when the width of the current block is a preset width or the height is a preset height, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. In another example, the decoder can determine whether to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes by comparing the size of the current block with a preset size. In one embodiment, when the size of the current block is larger or smaller than the preset size, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. In another embodiment, when the width of the current block is larger or smaller than the preset width, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. In another embodiment, when the height of the current block is larger or smaller than the preset height, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0189] In some embodiments, the method 300 may include the following. When the frame in which the current block is located is an I-frame and the size of the current block is a preset size, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0190] Exemplarily, when the frame in which the current block is located is an I-frame, the width of the current block is a preset width, and the height of the current block is a preset height, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. That is, only when the frame in which the current block is located is an I-frame, the decoder determines whether to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes based on the size of the current block.
[0191] In some embodiments, the method 300 may further include the following. When the frame in which the current block is located is a B-frame, the decoder determines an optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes.
[0192] Exemplarily, when the frame in which the current block is located is a B-frame, the decoder can directly determine an optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes. That is, when the frame in which the current block is located is a B-frame, regardless of the size of the current block, the decoder can directly determine an optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes.
[0193] In some embodiments, before S320, the method 300 may further include the following. The decoder obtains the MIP mode used for an adjacent block adjacent to the current block. The decoder determines the MIP mode used for the adjacent block as a plurality of MIP modes.
[0194] Exemplarily, the adjacent block may be an image block adjacent to at least one of the upper side, left side, lower left, upper right, and upper left of the current block. For example, the decoder can determine the image blocks obtained in the order of the upper side, left side, lower left, upper right, and upper left of the current block as adjacent blocks. Optionally, the plurality of MIP modes can be used by the decoder to determine available MIP modes for predicting the current block or to construct a list of available MIP modes. Thereby, the decoder determines an optimal MIP mode by predicting samples in the second template region from the available MIP modes or the list of available MIP modes.
[0195] In some embodiments, the method 300 may further include the following. The decoder performs reconstruction sample padding on a reference region adjacent to the outside of the second template region to obtain reference rows and reference columns of the second template region. The decoder uses each of a plurality of MIP modes with the reference rows and reference columns as inputs to predict samples within the second template region, obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes. The decoder determines distortion costs corresponding to the plurality of MIP modes based on the plurality of prediction blocks and reconstruction blocks within the second template region.
[0196] Exemplarily, the decoder performs reconstruction sample padding on a reference region adjacent to the outside of the second template region before determining an optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes.
[0197] Exemplarily, the width of the region adjacent to the upper side of the second template region within the reference region is equal to the width of the second template region. The height of the region adjacent to the left side of the second template region within the reference region is equal to the height second template region. When the width of the region adjacent to the upper side of the second template region within the reference region is greater than the width of the second template region, the decoder can perform downsampling or dimensionality reduction processing on the region adjacent to the upper side of the second template region within the reference region to obtain reference rows. When the height of the region adjacent to the left side of the second template region within the reference region is greater than the height second template region, the decoder can perform downsampling or dimensionality reduction processing on the region adjacent to the left side of the second template region within the reference region to obtain reference columns.
[0198] Exemplarily, the second template region may be the template region used for the TIMD mode described above, and the reference region may be the reference template used for the TIMD mode. For example, referring to FIG. 5 for description, if the current block is a coding unit with a width equal to M and a height equal to N, the decoder pads the reference region composed of coding units with a width equal to 2(M + L1)+1 and a height equal to 2(N + L2)+1 with reconstruction samples, performs downsampling or dimensionality reduction processing on the padded reference region to obtain reference rows and reference columns, and further constructs an input vector corresponding to the MIP mode based on the reference rows and reference columns.
[0199] Exemplarily, after obtaining the reference row and the reference column, the decoder uses each of a plurality of MIP modes with the reference row and the reference column as inputs to predict samples in the second template region, thereby obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes. That is, the decoder predicts samples in the second template region of the current block by traversing a plurality of MIP modes based on the reconstructed samples in the reference template of the current block. Taking the currently traversed MIP mode as an example, the decoder takes the reference row, the reference column, the index of the currently traversed MIP mode, and the third flag as inputs to obtain a prediction block corresponding to the currently traversed MIP mode. The reference row and the reference column are used to construct an input vector corresponding to the currently traversed MIP mode. The index of the currently traversed MIP mode is used to determine a matrix and / or a bias vector corresponding to the currently traversed MIP mode. The third flag is used to indicate whether to transpose the input vector and the output vector corresponding to the MIP mode. For example, when the third flag is used to indicate that the input vector and the output vector corresponding to the MIP mode are not transposed, the reference column is spliced after the reference row to form an input vector corresponding to the currently traversed MIP mode. When the third flag is used to indicate that the input vector and the output vector corresponding to the MIP mode are transposed, the reference row is spliced after the reference column to form an input vector corresponding to the currently traversed MIP mode. Accordingly, when the third flag is used to indicate that the input vector and the output vector corresponding to the MIP mode are transposed, the decoder transposes the output of the currently traversed MIP mode to obtain a prediction block in the second template region.By traversing a plurality of MIP modes, the decoder obtains a plurality of prediction blocks corresponding to the plurality of MIP modes, and then, based on the distortion cost between the plurality of prediction blocks and the reconstructed samples in the second template region, selects the MIP mode with the minimum cost according to the principle of the minimum distortion cost, and can determine it as the optimal MIP mode of the current block under the template matching-based MIP mode.
[0200] In some embodiments, when the decoder predicts samples in the second template region using each of the plurality of MIP modes, first, it downsamples the reference row and the reference column to obtain an input vector, and then, using the input vector as an input, predicts the samples in the second template region by traversing the plurality of MIP modes to obtain output vectors corresponding to the plurality of MIP modes, and finally, upsamples the output vectors corresponding to the plurality of MIP modes to obtain prediction blocks corresponding to the plurality of MIP modes.
[0201] Exemplarily, the reference row and the reference column satisfy the input conditions of the plurality of MIP modes. If the reference row and the reference column do not satisfy the input conditions of the plurality of MIP modes, first, the reference row and / or the reference column are processed to become input samples that satisfy the input conditions of the plurality of MIP modes, and then, based on the input samples that satisfy the input conditions of the plurality of MIP modes, the input vectors corresponding to the plurality of MIP modes can be determined. For example, taking the input condition as a specified number of input samples as an example, if the reference row and the reference column do not satisfy the number of input samples of the MIP mode, the decoder performs Haar-downsampling or the like on the reference row and / or the reference column to reduce the dimension of the reference row and / or the reference column to the specified number of input samples, and determines the input vectors corresponding to the plurality of MIP modes based on the specified number of input samples with reduced dimension.
[0202] In some embodiments, S320 may include the following. The decoder determines the optimal MIP mode based on the sum of absolute transformed differences (SATD) corresponding to a plurality of MIP modes within the second template region.
[0203] In this embodiment, when the decoder determines the optimal MIP mode based on the distortion cost corresponding to a plurality of MIP modes within the second template region, it designs the distortion cost corresponding to the plurality of MIP modes into the SATD corresponding to the plurality of MIP modes. Compared with directly calculating the rate-distortion cost corresponding to a plurality of MIP modes, not only can it be realized to determine the optimal MIP mode based on the distortion cost corresponding to a plurality of MIP modes within the second template region, but also the complexity of calculating the distortion cost corresponding to a plurality of MIP modes can be simplified. As a result, the decompression performance of the decoder can be improved.
[0204] In summary, in the scheme according to the present application, an idea of fusion enhancement is proposed based on the optimal MIP mode. That is, the decoder not only needs to determine the optimal MIP mode for predicting the current block, but also needs to fuse another prediction block to achieve different prediction effects. Thereby, not only can bit overhead be saved, but also new prediction techniques can be generated. Regarding the fusion process, actually, since the optimal MIP mode cannot completely replace the optimal prediction mode calculated based on the rate-distortion cost by the encoding side, it is a method that makes both prediction accuracy and prediction diversity compatible using the fusion method.
[0205] Exemplarily, the main idea of the template matching-based MIP mode derivation method by the decoder can be divided into the following several parts.
[0206] First, pad the reference region (e.g., the reference template shown in FIG. 5) with the reconstruction samples, i.e., the reference reconstruction samples required when predicting the samples within the second template region (e.g., the template shown in FIG. 5). Optionally, the width and height of the reference region do not need to exceed the width and height of the second template region. If the width and height of the reference region padded with samples exceed the width and height of the second template region, it is necessary to perform downsampling or other dimensionality reduction methods until the requirements for the input dimensions of the MIP are met.
[0207] Next, the decoder takes as input the reference reconstruction samples within the reference region, the indices of multiple MIP modes, and the MIP transpose flag, predicts the samples within the second template region, and obtains prediction blocks corresponding to multiple MIP modes. Optionally, the reference reconstruction samples within the reference region need to meet the input conditions of the MIP mode. For example, dimensionality reduction is performed by half-downsampling until the specified number of input samples. The indices of multiple MIP modes determine the matrix indices of the MIP technology and are further used to obtain the MIP prediction matrix coefficients. The MIP transpose flag is used to indicate whether it is necessary to transpose the input and output.
[0208] Next, for the prediction blocks corresponding to multiple MIP modes, traverse all combinations of all MIP modes and whether to transpose the MIP to obtain the predicted samples within the second template region under each state of each MIP mode and the MIP transpose flag, calculate the distortion between the predicted samples and the reconstruction samples within the second template region, and record the cost information. Finally, according to the principle of minimizing distortion, select the MIP mode with the minimum cost and its corresponding MIP transpose information, and set the MIP mode with the minimum cost as the optimal MIP mode of the current block under the template matching-based MIP prediction derivation mode.
[0209] Finally, the decoder predicts the current block by using each of the optimal MIP prediction mode and the second intra prediction mode, to obtain a first predicted block and a second predicted block respectively, and performs a weighting calculation on the first predicted block and the second predicted block based on the weight of the optimal MIP prediction mode and the weight of the second intra prediction mode, to obtain a predicted block of the current block.
[0210] Note that some calculations according to this application can be replaced by a LookUp Table or a shift method. The LookUp Table method may result in some errors compared with directly performing division, but it is beneficial for hardware implementation and control of coding cost. Examples of the above-mentioned some calculations include calculations related to distortion cost or calculations related to determining the optimal MIP mode.
[0211] The decoding method according to the embodiment of this application has been described in detail from the perspective of the decoder. Next, with reference to FIG. 10 The encoding method according to the embodiment of this application will be described from the perspective of the encoder.
[0212] FIG. 10 is a flowchart showing an encoding method 400 according to the embodiment of this application. Note that the encoding method 400 can be executed by an encoder. For example, the encoding method 400 is applied to the encoding framework 100 shown in FIG. 1. For ease of explanation, the encoder will be described as an example hereinafter.
[0213] As shown in FIG. 10, the encoding method 400 may include the following steps. S410: Obtain a residual block of the current block in the current sequence. S420: Perform a third transformation on the residual block of the current block to obtain third transformation coefficients of the current block. S430: Determine the first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode that is used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode that is used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode that is used for reconstructed samples within a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. S440: Based on a conversion set corresponding to the first intra prediction mode, perform a fourth conversion on the third conversion coefficient to obtain the fourth conversion coefficient of the current block. S450: Encode the fourth conversion coefficient.
[0214] It should be understood that the first conversion on the decoding side is the inverse conversion of the fourth conversion on the encoding side, and the second conversion on the decoding side is the inverse conversion of the third conversion on the encoding side. For example, the third conversion is the basic conversion or the primary conversion described above, the fourth conversion is the secondary conversion described above, and correspondingly, the first conversion is the inverse conversion of the secondary conversion, and the second conversion is the inverse conversion of the basic conversion or the primary conversion. For example, the first conversion can be an inverse LFNST, the second conversion can be an inverse DCT type 2, an inverse DCT type 8, an inverse DST type 7, etc., and correspondingly, the third conversion can be a DCT type 2, a DCT type 8, a DST type 7, etc., and the fourth conversion can be an LFNST.
[0215] In some embodiments, the output vector of the optimal MIP mode is the vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is the vector after upsampling the output vector of the optimal MIP mode.
[0216] In some embodiments, S430 may include determining a first intra prediction mode based on a prediction mode for predicting a current block.
[0217] In some embodiments, when the prediction mode for predicting a current block includes an optimal MIP mode and a sub-optimal MIP mode for predicting the current block, determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode.
[0218] In some embodiments, when the prediction mode for predicting a current block includes an optimal MIP mode and an intra prediction mode derived from the TIMD mode, determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determine, as the first intra prediction mode, an intra prediction mode derived from the TIMD mode.
[0219] In some embodiments, when the prediction mode for predicting a current block includes an optimal MIP mode and an intra prediction mode derived from the DIMD mode that is used for the reconstructed samples within the first template region, determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the reconstructed samples within the first template region.
[0220] In some embodiments, the second template region and the first template region may be the same or different.
[0221] In some embodiments, S410 may further include the following. Determine a second intra prediction mode. The second intra prediction mode includes any one of an almost optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstruction samples within the first template region, and an intra prediction mode derived from the TIMD mode. Predict the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block. Obtain a residual block of the current block based on the predicted block of the current block.
[0222] In some embodiments, predict the current block based on the optimal MIP mode to obtain a first predicted block. Predict the current block based on the second intra prediction mode to obtain a second predicted block. Perform a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block.
[0223] In some embodiments, when the prediction mode for predicting the current block includes the optimal MIP mode, an almost optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, determine the weight of the optimal MIP mode and the weight of the second intra prediction mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode. When the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode used for reconstruction samples within the first template region, determine that the weights of the optimal MIP mode and the second intra prediction mode are both preset values.
[0224] In some embodiments, the encoder obtains a first flag, and when the first flag is used to indicate that it is permitted to predict an image block in the current sequence using an optimal MIP mode and a second intra prediction mode, the encoder determines the second intra prediction mode. S450 may include encoding a fourth transform coefficient and the first flag.
[0225] In some embodiments, when the first flag is used to indicate that it is permitted to predict an image block in the current sequence using an optimal MIP mode and a second intra prediction mode, the current block is predicted based on the optimal MIP mode and the second intra prediction mode to obtain a first rate-distortion cost. The current block is predicted based on at least one intra prediction mode to obtain at least one rate-distortion cost. When the first rate-distortion cost is less than or equal to the minimum value of the at least one rate-distortion cost, the predicted block obtained by predicting the current block based on the optimal MIP mode and the second intra prediction mode is determined as the predicted block of the current block. S450 may include the following. Encode a fourth transform coefficient, the first flag, and a second flag. When the first rate-distortion cost is less than or equal to the minimum value of the at least one rate-distortion cost, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the second intra prediction mode, and when the first rate-distortion cost is greater than the minimum value of the at least one rate-distortion cost, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the second intra prediction mode.
[0226] In some embodiments, the method 400 may further include the following. Based on the distortion costs corresponding to a plurality of MIP modes, determine an optimal MIP mode. The distortion costs corresponding to the plurality of MIP modes include the distortion costs obtained by predicting samples in a second template region adjacent to the current block using the plurality of MIP modes.
[0227] In some embodiments, the second template region and the first template region may be the same or different.
[0228] In some embodiments, predict samples in the second template region based on a third flag and a plurality of MIP modes to obtain distortion costs corresponding to the plurality of MIP modes under each state of the third flag. The third flag is used to indicate whether to transpose the input vector and the output vector corresponding to the MIP mode. Based on the distortion costs corresponding to the plurality of MIP modes under each state of the third flag, determine an optimal MIP mode.
[0229] In some embodiments, before determining an optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes, the method 400 may further include the following. Obtain the MIP mode used for an adjacent block adjacent to the current block. Determine the MIP mode used for the adjacent block as the plurality of MIP modes.
[0230] In some embodiments, before determining the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes, the method 400 may further include the following. Reconfiguration sample padding is performed on a reference region adjacent to the outside of the second template region to obtain the reference rows and reference columns of the second template region. Using each of the multiple MIP modes with the reference rows and reference columns as inputs, samples within the second template region are predicted to obtain multiple prediction blocks corresponding to the multiple MIP modes. Based on the multiple prediction blocks and the reconstruction blocks within the second template region, the distortion costs corresponding to the multiple MIP modes are determined.
[0231] In some embodiments, the reference rows and reference columns are downsampled to obtain an input vector. Using the input vector as an input, samples within the second template region are predicted by traversing multiple MIP modes to obtain output vectors corresponding to the multiple MIP modes. The output vectors corresponding to the multiple MIP modes are upsampled to obtain prediction blocks corresponding to the multiple MIP modes.
[0232] In some embodiments, the optimal MIP mode is determined based on the sum of absolute transform differences (SATD) corresponding to multiple MIP modes within the second template region.
[0233] It should be noted that the encoding method can be understood as the reverse process of the decoding method. Therefore, for the specific scheme of the encoding method 400, reference can be made to the relevant content of the decoding method 300, and for the sake of convenience of description, it will not be elaborated in detail in this application.
[0234] Hereinafter, the solution of this application will be described in connection with specific embodiments.
[0235] <Example 1>
[0236] In this embodiment, the second intra prediction mode is a quasi-optimal prediction mode. That is, the encoder or decoder can perform intra prediction on the current block based on the optimal MIP mode and the quasi-optimal MIP mode to obtain the predicted block of the current block.
[0237] The encoder traverses the prediction modes. When the intra mode is used for the current block, the encoder obtains a sequence-level enable flag such as sps_tmmip_enable_flag. The sequence-level enable flag is used to indicate whether the template matching-based MIP mode derivation technique is permitted to be used in the current sequence. When all the tmmip enable flags are true, it indicates that the TMMIP technique is currently permitted to be used by the encoder.
[0238] Exemplarily, the process of the encoder can be realized as the following process.
[0239] Step 1: If sps_tmmip_enable_flag is true, the encoder tries the TMMIP technique, that is, executes Step 2. If sps_tmmip_enable_flag is false, the encoder cannot try the TMMIP technique, that is, skips Step 2 and directly executes Step 3.
[0240] Step 2: First, the encoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template area. The padding process is the same as the padding method of the original intra prediction process. For example, the encoder can traverse and pad from the lower left corner to the upper right corner. If all reconstruction samples are available, pad them in order with all available reconstruction samples. If all reconstruction samples are unavailable, pad them all with the average value. If some reconstruction samples are available, first pad with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the encoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to pad the previous unavailable positions. Next, the encoder uses the permitted MIP mode with the reconstruction samples outside the padded second template area as input to predict the samples within the second template area.
[0241] Exemplarily, for a 4×4 size block, there are 16 permitted MIP modes. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 permitted MIP modes. For blocks of other sizes, there are 6 permitted MIP modes. Also, the MIP transposition function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0242] Exemplarily, the specific prediction calculation process includes the following. First, the encoder performs Haar downsampling on the reconstruction samples. For example, the encoder determines the down-sampling step size based on the block size. Next, according to the information of whether to transpose or not, the encoder adjusts the splicing order of the reconstruction samples after upper downsampling and the reconstruction samples after left downsampling. When transposition is not required, the reconstruction samples after left downsampling are spliced after the reconstruction samples after upper downsampling, and the obtained vector is used as the input. When transposition is required, the reconstruction samples after upper downsampling are spliced after the reconstruction samples after left downsampling, and the obtained vector is used as the input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as the index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the encoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a prediction block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template, and output the prediction block within the second template area.
[0243] Next, the encoder calculates a distortion cost based on the predicted block within the second template region obtained by traversing each MIP mode and the reconstructed samples within the second template region, and records the distortion cost under each prediction mode and the transposition information. After traversing all the permitted prediction modes and transposition information, according to the principle of the minimum cost, the optimal MIP mode and its corresponding transposition information, and the sub-optimal MIP mode and its corresponding transposition information are selected. The encoder determines whether fusion enhancement is necessary based on the relationship between the cost of the optimal MIP mode and the cost of the sub-optimal MIP mode. If the cost of the sub-optimal MIP mode is less than twice the cost of the optimal MIP mode, it is necessary to perform fusion enhancement on the optimal MIP prediction block and the sub-optimal MIP prediction block. If the cost of the sub-optimal MIP prediction mode is twice or more the cost of the optimal MIP mode, there is no need for fusion enhancement.
[0244] Finally, if fusion enhancement is required, the encoder obtains a prediction block corresponding to the optimal MIP mode and a prediction block corresponding to the sub-optimal MIP mode based on the optimal MIP mode, the sub-optimal MIP mode, the transposed information of the optimal MIP mode, and the transposed information of the sub-optimal MIP mode. Specifically, first, the encoder downsamples the reconstruction samples adjacent to the upper side and the left side of the current block as appropriate, splices them according to the transposed information to obtain an input vector, reads out the matrix coefficients under the current mode using the MIP mode as an index, and then obtains an output vector by calculating the input vector and the matrix coefficients. The encoder transposes the output based on the transposed information, upsamples the output vector based on the size of the current block and the number of samples of the output vector to obtain an optimal MIP prediction block and a sub-optimal MIP prediction block of the same size as the current block, and also performs a weighted average on the optimal MIP prediction block and the sub-optimal MIP prediction block based on the calculated weight value of the optimal MIP mode and the weight value of the sub-optimal MIP mode to obtain a new prediction block as the final prediction block of the current block. If fusion enhancement is not required, the encoder can calculate an optimal MIP prediction block based on the optimal MIP mode and its transposed information, and the calculation process is the same as the above. Finally, the encoder uses the optimal MIP prediction block as the prediction block of the current block.
[0245] Furthermore, the encoder obtains the rate-distortion cost of the current block and denotes it as cost1.
[0246] The encoder determines the intra prediction mode derived from the DIMD mode used for the prediction block of the current block as the first intra prediction mode, or the encoder determines the intra prediction mode derived from the DIMD mode used for the output vector of the optimal MIP mode as the first intra prediction mode.
[0247] Step 3: The encoder continues to traverse other intra prediction techniques, calculates the corresponding rate-distortion cost, and records it as cost2...costN.
[0248] Step 4: If cost1 is the minimum among all rate-distortion costs, the TMMIP technique is used for the current block. The encoder sets the TMMIP usage flag of the current block to true and writes it to the bitstream. If cost1 is not the minimum rate-distortion cost, other intra prediction techniques are used for the current block. The encoder sets the TMMIP usage flag of the current block to false and writes it to the bitstream. Note that information such as the flags or indices of other intra prediction techniques is transmitted based on the definition and is not detailed here.
[0249] Step 5: The encoder determines the residual block of the current block based on the predicted block of the current block and the original block of the current block, performs a base transformation on the residual block of the current block, and performs a secondary transformation on the transformed coefficients after the base transformation based on the first intra prediction mode. Then, operations such as quantization, entropy coding, and loop filtering are performed on the transformed coefficients after the secondary transformation. Note that for the specific process of that quantization, the above related content can be referred to, and to avoid duplication, it is not detailed here.
[0250] The related scheme of the decoder in this embodiment will be described below.
[0251] The decoder analyzes block-level flags. If the intra mode is used for the current block, the decoder analyzes or obtains a sequence-level permission flag such as sps_tmmip_enable_flag. The sequence-level permission flag is used to indicate whether the template matching-based MIP mode derivation technique is permitted to be used for the current sequence. If any of the tmmip permission flags are true, it indicates that the TMMIP technique is currently permitted to be used by the decoder.
[0252] Exemplarily, the process of the decoder can be realized as the following process.
[0253] Step 1: If sps_tmmip_enable_flag is true, the decoder analyzes the TMMIP usage flag of the current block. Otherwise, in the current decoding process, the block-level TMMIP usage flag does not need to be decoded, and the block-level TMMIP usage flag is default set to false. If the TMMIP usage flag of the current block is true, Step 2 is executed. Otherwise, Step 3 is executed.
[0254] Step 2: First, the decoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template area. The padding process is the same as the padding method in the original intra prediction process. For example, the decoder can traverse and pad from the lower left corner to the upper right corner. If all reconstruction samples are available, padding is performed in order with all available reconstruction samples. If all reconstruction samples are unavailable, all are padded with the average value. If some reconstruction samples are available, first padding is performed with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the decoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to perform padding on the previous unavailable position. Next, the decoder uses the permitted MIP mode with the reconstruction samples outside the padded second template area as input to predict the samples within the second template area.
[0255] Exemplarily, for a 4×4 size block, there are 16 permitted MIP modes. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 permitted MIP modes. For blocks of other sizes, there are 6 permitted MIP modes. Also, the MIP transposition function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0256] Exemplarily, the specific prediction calculation process includes the following. First, the decoder performs half-downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, the decoder adjusts the splicing order of the reconstructed samples after upper downsampling and the reconstructed samples after left downsampling according to the information of whether to transpose or not. When transposition is not required, the reconstructed samples after left downsampling are spliced after the reconstructed samples after upper downsampling, and the obtained vector is used as the input. When transposition is required, the reconstructed samples after upper downsampling are spliced after the reconstructed samples after left downsampling, and the obtained vector is used as the input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the decoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a predicted block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template, and output the predicted block within the second template area.
[0257] Next, the decoder calculates the distortion cost based on the predicted block within the second template region obtained by traversing each MIP mode and the reconstructed samples within the second template region, and records the distortion cost under each prediction mode and the transposition information. After traversing all the permitted prediction modes and transposition information, according to the principle of the minimum cost, the optimal MIP mode and its corresponding transposition information, and the sub-optimal MIP mode and its corresponding transposition information are selected. The decoder determines whether fusion enhancement is necessary based on the relationship between the cost of the optimal MIP mode and the cost of the sub-optimal MIP mode. If the cost of the sub-optimal MIP mode is less than twice the cost of the optimal MIP mode, it is necessary to perform fusion enhancement on the optimal MIP prediction block and the sub-optimal MIP prediction block. If the cost of the sub-optimal prediction mode is twice or more the cost of the optimal MIP mode, there is no need for fusion enhancement.
[0258] Finally, if fusion enhancement is required, the decoder obtains a prediction block corresponding to the optimal MIP mode and a prediction block corresponding to the sub-optimal MIP mode based on the optimal MIP mode, the sub-optimal MIP mode, the transposed information of the optimal MIP mode, and the transposed information of the sub-optimal MIP mode. Specifically, first, the decoder downsamples the reconstructed samples adjacent to the upper and left sides of the current block as appropriate, splices them according to the transposed information to obtain an input vector, reads out the matrix coefficients under the current mode using the MIP mode as an index, and then obtains an output vector by calculating the input vector and the matrix coefficients. The decoder transposes the output based on the transposed information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain an optimal MIP prediction block and a sub-optimal MIP prediction block of the same size as the current block, and also performs a weighted average on the optimal MIP prediction block and the sub-optimal MIP prediction block based on the calculated weight value of the optimal MIP mode and the weight value of the sub-optimal MIP mode to obtain a new prediction block as the final prediction block of the current block. If fusion enhancement is not required, the decoder can calculate an optimal MIP prediction block based on the optimal MIP mode and its transposed information, and the calculation process is the same as the above. Finally, the decoder uses the optimal MIP prediction block as the prediction block of the current block.
[0259] Note that the decoder determines the intra prediction mode derived from the DIMD mode used for the prediction block of the current block as the first intra prediction mode, or determines the intra prediction mode derived from the DIMD mode used for the output vector of the optimal MIP mode as the first intra prediction mode.
[0260] Step 3: The decoder continues to analyze information such as the usage flag or index of other intra prediction techniques, and obtains the final prediction block of the current block based on the analyzed information.
[0261] Step 4: The decoder analyzes the bitstream to obtain the frequency-domain residual block of the current block (also referred to as frequency-domain residual information), and performs inverse quantization and inverse transformation on the frequency-domain residual block of the current block (first, perform inverse transformation on the secondary transformation based on the first intra prediction mode, and then perform inverse transformation on the base transformation or primary transformation), and obtain the residual block of the current block (also referred to as the time-domain residual block or time-domain residual information). Next, the decoder adds the prediction block of the current block to the residual block of the current block to obtain a reconstructed sample block. conversion (perform), and obtain the residual block of the current block (also referred to as the time-domain residual block or time-domain residual information).
[0262] Step 5: After techniques such as loop filtering are performed on all the reconstructed sample blocks in the current image, a final reconstructed image is obtained.
[0263] Optionally, the reconstructed image may be used as a video output, or may be used as a reference for subsequent decoding.
[0264] In this embodiment, the size of the second template region used by the encoder or decoder with the TMMIP technology can be predefined according to the size of the current block. For example, the width of the region adjacent to the upper side of the current block in the second template region is equal to the width of the current block, and its height is equal to the height of two rows of samples. The height of the region adjacent to the left side of the current block in the second template region is equal to the height of the current block, and its width is 2 column sample widths. Of course, in other alternative embodiments, it can be realized as a second template region of other sizes, and the present application does not specifically limit it in this regard.
[0265] In this embodiment, the TMMIP technology is adopted.
[0266] <Example 2>
[0267] In this embodiment, the second intra prediction mode is an intra prediction mode derived from the TIMD mode. That is, the encoder or decoder can perform intra prediction on the current block based on the optimal MIP mode and the intra prediction mode derived from the TIMD mode to obtain the predicted block of the current block.
[0268] That is, in the MIP mode derivation fusion enhancement technology based on template matching, not only can two derived MIP prediction blocks be fused, but also prediction blocks generated by other template matching-based derivation technologies can be fused. In this application, a method of fusing the TMMIP technology and the TIMD technology to fuse the derived conventional prediction block and the prediction block based on the matrix is obtained. In the TIMD technology, the idea of template matching is used on the encoding side and the decoding side to derive the optimal conventional intra prediction mode, and also, by the TIMD technology, offset expansion can be performed on this prediction mode to obtain an updated intra prediction mode. Also, in the TMMIP technology, the idea of template matching is used on the encoding side and the decoding side to derive the optimal MIP mode. By fusing these two optimal prediction modes, the directivity of the conventional prediction block and the unique texture characteristics of the MIP prediction can be made compatible, generating a completely new prediction block and improving the coding efficiency.
[0269] The encoder traverses the prediction mode. When the intra mode is used for the current block, the encoder obtains a sequence-level permission flag such as sps_tmmip_enable_flag. The sequence-level permission flag is used to indicate whether the template matching-based MIP mode derivation technology is permitted to be used in the current sequence. When all the permission flags of tmmip are true, it indicates that the TMMIP technology is currently permitted to be used by the encoder.
[0270] Exemplarily, the encoder process can be realized as the following process.
[0271] Step 1: If sps_tmmip_enable_flag is true, the encoder tries the TMMIP technology, that is, it executes Step 2. If sps_tmmip_enable_flag is false, the encoder cannot try the TMMIP technology, that is, it skips Step 2 and directly executes Step 3.
[0272] Step 2: First, the encoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template area. The padding process is the same as the padding method of the original intra prediction process. For example, the encoder can traverse and pad from the lower left corner to the upper right corner. If all reconstruction samples are available, all available reconstruction samples are padded in order. If all reconstruction samples are unavailable, all are padded with the average value. If some reconstruction samples are available, first pad with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the encoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to pad the previous unavailable positions. Next, the encoder uses the permitted MIP mode with the reconstruction samples outside the padded second template area as input to predict the samples within the second template area.
[0273] Exemplarily, for a 4×4 size block, there are 16 MIP modes permitted for use. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 MIP modes permitted for use. For blocks of other sizes, there are 6 MIP modes permitted for use. Also, the MIP transposition function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0274] Exemplarily, the specific prediction calculation process includes the following content. First, the encoder performs half-downsampling on the reconstructed samples. For example, the encoder determines the downsampling step size based on the block size. Next, according to the information on whether to transpose or not, the encoder adjusts the splicing order of the reconstructed samples after upper downsampling and the reconstructed samples after left downsampling. When transposition is not required, the reconstructed samples after left downsampling are spliced after the reconstructed samples after upper downsampling, and the obtained vector is used as the input. When transposition is required, the reconstructed samples after upper downsampling are spliced after the reconstructed samples after left downsampling, and the obtained vector is used as the input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the encoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a predicted block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template, and output the predicted block within the second template area.
[0275] Furthermore, the encoder needs to try the template matching calculation process of TIMD, obtain different interpolation filters based on different prediction mode indexes, and obtain the predicted samples in the template by interpolating the reference samples.
[0276] Next, the encoder calculates the distortion cost based on the predicted samples in the second template area and the reconstructed samples in the second template area obtained by traversing each MIP mode, and records the distortion cost under each prediction mode and transposition information. Based on the distortion cost under each prediction mode and transposition information, according to the principle of the minimum cost, the optimal MIP mode and the corresponding transposition information are selected. Furthermore, the encoder traverses all the permitted intra prediction modes in TIMD, calculates the predicted samples in the template, calculates the distortion cost using the predicted samples in the template and the reconstructed samples in the template, and according to the principle of the minimum cost, records the optimal prediction mode, sub-optimal prediction mode derived from the TIMD technology, the distortion cost corresponding to the optimal prediction mode, and the distortion cost corresponding to the sub-optimal prediction mode.
[0277] Finally, based on the obtained optimal MIP mode and transposition information, the encoder may downsample the reconstructed samples adjacent to the upper side and the left side of the current block, splice them based on the transposition information to obtain an input vector, read out the matrix coefficients under the current mode with the MIP mode as the index, and then obtain an output vector by calculating the input vector and the matrix coefficients. The encoder transposes the output based on the transposition information, upsamples the output vector based on the size of the current block and the number of samples of the output vector to obtain an output of the same size as the current block, and can use it as the optimal MIP prediction block of the current block.
[0278] Regarding the optimal prediction mode and the sub-optimal prediction mode derived from the TIMD technology, when neither the optimal prediction mode nor the sub-optimal prediction mode is the average value (DC) mode or the planar (PLANAR) mode, and the distortion cost corresponding to the sub-optimal prediction mode is less than twice the distortion cost corresponding to the optimal prediction mode, the encoder needs to fuse the prediction blocks. First, the encoder obtains the interpolation filtering coefficients based on the optimal prediction mode, performs interpolation filtering on the reconstructed samples adjacent to the upper and left sides, obtains the prediction samples at all positions within the current block, and records it as the optimal prediction block. Next, the encoder obtains the interpolation filtering coefficients based on the sub-optimal prediction mode, performs interpolation filtering on the reconstructed samples adjacent to the upper and left sides, obtains the prediction samples at all positions within the current block, and records it as the sub-optimal prediction block. Furthermore, the encoder calculates the weight value of the optimal prediction block and the weight value of the sub-optimal prediction block by using the ratio of the cost corresponding to the optimal prediction mode to the cost corresponding to the sub-optimal prediction mode. Finally, the encoder performs weighted fusion on the optimal prediction block and the sub-optimal prediction block to obtain the prediction block of the current block and outputs it. Also, when the optimal prediction mode or the sub-optimal prediction mode is the average value mode (DC) or the planar mode (PLANAR), or when the cost corresponding to the sub-optimal prediction mode is greater than twice the cost corresponding to the optimal prediction mode, the encoder does not need to fuse the prediction blocks, and the optimal prediction block obtained by performing interpolation filtering on the reconstructed samples adjacent to the upper and left sides using only the optimal prediction mode is taken as the optimal TIMD prediction block of the current block.
[0279] Finally, the encoder performs a weighted average on the optimal MIP prediction block and the optimal TIMD prediction block based on the calculated weight value of the optimal MIP mode and the weight value of the prediction mode derived from the TIMD technology to obtain a new prediction block. That new prediction block is the prediction block of the current block.
[0280] Furthermore, the encoder obtains the rate-distortion cost of the current block and records it as cost1.
[0281] Note that the encoder determines the intra prediction mode derived from the DIMD mode used for the prediction block of the current block as the first intra prediction mode, or the encoder determines the intra prediction mode derived from the TIMD mode as the first intra prediction mode.
[0282] Note that the template region in the TIMD technology and the second template region (i.e., the template region in the TMMIP technology) can be set to be the same. That is, since the template regions for calculating the distortion cost are the same, the cost information of the template region in the TIMD technology and the cost information of the template region in the TMMIP technology can be equivalent or at the same comparison level. In this case, it can be determined whether fusion enhancement is performed based on the cost information, which is not specifically limited in this application.
[0283] Step 3: The encoder continues to traverse other intra prediction techniques, calculates the corresponding rate-distortion costs, and records them as cost2...costN.
[0284] Step 4: If cost1 is the smallest among all the rate-distortion costs, the TMMIP technology is used for the current block. The encoder sets the TMMIP usage flag of the current block to true and writes it to the bitstream. If cost1 is not the smallest rate-distortion cost, other intra prediction techniques are used for the current block. The encoder sets the TMMIP usage flag of the current block to false and writes it to the bitstream. Note that information such as the flags or indexes of other intra prediction techniques is transmitted based on the definition and is not elaborated here.
[0285] Step 5: The encoder determines the residual block of the current block based on the predicted block of the current block and the original block of the current block, performs a base transform on the residual block of the current block, and performs a secondary transform on the transform coefficients after the base transform based on the first intra prediction mode. Then, operations such as quantization, entropy coding, and loop filtering are performed on the transform coefficients after the secondary transform. Note that for the specific process of quantization, the above related content can be referred to, and to avoid duplication, it will not be elaborated here.
[0286] The related scheme of the decoder in this embodiment will be described below.
[0287] The decoder analyzes the block-level flag. When the intra mode is used for the current block, the decoder analyzes or obtains the sequence-level permission flag such as sps_tmmip_enable_flag. The sequence-level permission flag is used to indicate whether the template matching-based MIP mode derivation technology is permitted to be used in the current sequence. When all the permission flags of tmmip are true, it indicates that currently, the decoder is permitted to use the TMMIP technology.
[0288] Exemplarily, the process of the decoder can be realized as follows.
[0289] Step 1: When sps_tmmip_enable_flag is true, the decoder analyzes the TMMIP usage flag of the current block. Otherwise, in the current decoding process, it is not necessary to decode the block-level TMMIP usage flag, and the block-level TMMIP usage flag is default set to false. When the TMMIP usage flag of the current block is true, Step 2 is executed. Otherwise, Step 3 is executed.
[0290] Step 2: First, the decoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template area. The padding process is the same as the padding method in the original intra prediction process. For example, the decoder can traverse and pad from the lower left corner to the upper right corner. If all reconstruction samples are available, pad them in order with all available reconstruction samples. If all reconstruction samples are unavailable, pad them all with the average value. If some reconstruction samples are available, first pad with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the decoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to pad the previous unavailable positions. Next, the decoder uses the reconstruction samples outside the padded second template area as input and utilizes the permitted MIP mode to predict the samples within the second template area.
[0291] Exemplarily, for a 4×4 size block, there are 16 permitted MIP modes. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 permitted MIP modes. For blocks of other sizes, there are 6 permitted MIP modes. Also, the MIP transpose function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0292] Exemplarily, the specific prediction calculation process includes the following. First, the decoder performs half downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, the decoder adjusts the splicing order of the reconstructed samples after upper downsampling and the reconstructed samples after left downsampling according to the information of whether to transpose or not. When transposition is not required, the reconstructed samples after left downsampling are spliced after the reconstructed samples after upper downsampling, and the obtained vector is used as the input. When transposition is required, the reconstructed samples after upper downsampling are spliced after the reconstructed samples after left downsampling, and the obtained vector is used as the input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as the index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the decoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a predicted block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template, and output the predicted block within the second template area.
[0293] Furthermore, the decoder needs to try the template matching calculation process of TIMD, obtain different interpolation filters based on different prediction mode indexes, and obtain the predicted samples within the template by interpolating the reference samples.
[0294] Next, the decoder calculates the distortion cost based on the predicted samples within the second template region obtained by traversing each MIP mode and the reconstructed samples within the second template region, and records the distortion cost under each prediction mode and transposition information. Based on the distortion cost under each prediction mode and transposition information, following the principle of the minimum cost, the optimal MIP mode and the corresponding transposition information are selected. Further, the decoder traverses all the permitted intra prediction modes in the TIMD, calculates the predicted samples within the template, calculates the distortion cost using the predicted samples within the template and the reconstructed samples within the template, and needs to record the optimal prediction mode derived from the TIMD technology, the sub-optimal prediction mode, the distortion cost corresponding to the optimal prediction mode, and the distortion cost corresponding to the sub-optimal prediction mode following the principle of the minimum cost.
[0295] Finally, based on the obtained optimal MIP mode and transposition information, the decoder optionally downsamples the reconstructed samples adjacent to the upper and left sides of the current block, splices them based on the transposition information to obtain an input vector, reads out the matrix coefficients under the current mode using the MIP mode as an index, and then obtains an output vector by calculating the input vector and the matrix coefficients. The decoder can transpose the output based on the transposition information, upsample the output vector based on the size of the current block and the number of samples of the output vector to obtain an output of the same size as the current block and use it as the optimal MIP prediction block of the current block.
[0296] For the optimal prediction mode and the sub-optimal prediction mode derived from the TIMD technology, when neither the optimal prediction mode nor the sub-optimal prediction mode is the average value (DC) mode or the planar (PLANAR) mode, and the distortion cost corresponding to the sub-optimal prediction mode is less than twice the distortion cost corresponding to the optimal prediction mode, the decoder needs to fuse the prediction blocks. First, the decoder obtains the interpolation filtering coefficients based on the optimal prediction mode, performs interpolation filtering on the reconstructed samples adjacent to the upper side and the left side, obtains the prediction samples at all positions within the current block, and records it as the optimal prediction block. Next, the decoder obtains the interpolation filtering coefficients based on the sub-optimal prediction mode, performs interpolation filtering on the reconstructed samples adjacent to the upper side and the left side, obtains the prediction samples at all positions within the current block, and records it as the sub-optimal prediction block. Further, the decoder calculates the weight value of the optimal prediction block and the weight value of the sub-optimal prediction block by using the ratio of the cost corresponding to the optimal prediction mode to the cost corresponding to the sub-optimal prediction mode. Finally, the decoder performs weighted fusion on the optimal prediction block and the sub-optimal prediction block to obtain the prediction block of the current block and outputs it. Also, when the optimal prediction mode or the sub-optimal prediction mode is the average value mode (DC) or the planar mode (PLANAR), or the cost corresponding to the sub-optimal prediction mode is greater than twice the cost corresponding to the optimal prediction mode, the decoder does not need to fuse the prediction blocks, and the optimal prediction block obtained by performing interpolation filtering on the reconstructed samples adjacent to the upper side and the left side using only the optimal prediction mode is used as the optimal TIMD prediction block of the current block.
[0297] Finally, the decoder performs a weighted average on the optimal MIP prediction block and the optimal TIMD prediction block based on the calculated weight value of the optimal MIP mode and the weight value of the prediction mode derived from the TIMD technology to obtain a new prediction block. That new prediction block is the prediction block of the current block.
[0298] Note that the decoder determines the intra prediction mode derived from the DIMD mode used for the prediction block of the current block as the first intra prediction mode, or the decoder determines the intra prediction mode derived from the TIMD mode as the first intra prediction mode.
[0299] Step 3: The decoder continues to analyze information such as the usage flag or index of other intra prediction techniques, and obtains the final prediction block of the current block based on the analyzed information.
[0300] Step 4: The decoder analyzes the bitstream to obtain the frequency domain residual block of the current block (also referred to as frequency domain residual information), and performs inverse quantization and inverse transformation on the frequency domain residual block of the current block (first, perform inverse transformation on the secondary transformation based on the first intra prediction mode, and then perform inverse transformation on the base transformation or primary transformation), to obtain the residual block of the current block (also referred to as the time domain residual block or time domain residual information). Next, the decoder adds the prediction block of the current block to the residual block of the current block to obtain the reconstructed sample block. conversion Step 5: After techniques such as loop filtering are performed on all the reconstructed sample blocks in the current image, the final reconstructed image is obtained.
[0301] Optionally, the reconstructed image may be used as a video output, or may be used as a reference for subsequent decoding.
[0302] Optionally, the reconstructed image may be used as a video output, or may be used as a reference for subsequent decoding.
[0303] In this embodiment, for the process of calculating the weight value for weighted fusion of the TIMD prediction block, reference can be made to the description content of the TIMD technology described above. To avoid duplication, it will not be elaborated here. Furthermore, the encoder or decoder can determine whether fusion enhancement is to be performed based on the optimal prediction mode derived from TIMD. For example, when the optimal prediction mode derived from TIMD is the DC mode or the PLANAR mode, the encoder or decoder may not need to use fusion enhancement. That is, the encoder or decoder uses only the prediction block generated by the optimal MIP mode derived from the TMMIP technology as the prediction block of the current block. Furthermore, the size of the second template region used by the encoder or decoder in the TMMIP technology can be predefined according to the size of the current block. For example, the definition regarding the second template region in the TMMIP technology may be the same as or different from the definition regarding the template region in the TIMD technology. For example, when the width of the current block is 8 or less, the height of the region adjacent to the upper side of the current block within the second template region is equal to the height of 2 rows of samples. Otherwise, the height is equal to the height of 4 rows of samples. Similarly, when the height of the current block is 8 or less, the width of the region adjacent to the left side of the current block within the second template region is equal to the width width of 2 columns of samples. Otherwise, the width is equal to the width width of 4 columns of samples.
[0304] <Example 3>
[0305] In this embodiment, the second intra prediction mode described above is an intra prediction mode derived from the DIMD mode. That is, the encoder or decoder can perform intra prediction on the current block based on the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstructed samples in the first template region adjacent to the current block, to obtain the prediction block of the current block.
[0306] Similar to Example 2, the TMMIP technology can also be fused and enhanced together with the DIMD technology.
[0307] Note that both the prediction mode derived from the DIMD technology and the prediction mode derived from the TIMD technology are conventional intra prediction modes. However, since the derivation methods are different, the prediction modes obtained by the two are not necessarily the same. Also, the method of fusing and enhancing the TMMIP technology and the DIMD technology is different from the method of fusing and enhancing the TMMIP technology and the TIMD technology. For example, in the TMMIP technology and the TIMD technology, the size of the second template region is generally the same, and the calculated cost information is basically SATD (Sum of Absolute Transformed Difference), so it is also called the distortion cost based on the Hadamard transform. In the TMMIP technology and the TIMD technology, the fusion weight can be directly calculated based on this cost information. However, the size of the second template region in the DIMD technology is generally different from the size of the second template region in the TMMIP technology (or TIMD technology), and the rule of the DIMD-derived prediction mode is based on the gradient amplitude value. Since the gradient amplitude value is not equivalent to the SATD cost, the weight cannot be calculated simply by referring to the scheme for fusing the TMMIP technology and the TIMD technology.
[0308] The encoder traverses the prediction mode. When the intra mode is used for the current block, the encoder obtains a sequence-level permission flag such as sps_tmmip_enable_flag. The sequence-level permission flag is used to indicate whether the template matching-based MIP mode derivation technology is permitted to be used for the current sequence. When all the permission flags for tmmip are true, it indicates that the TMMIP technology is currently permitted to be used by the encoder.
[0309] Exemplarily, the encoder process can be realized as the following process.
[0310] Step 1: If sps_tmmip_enable_flag is true, the encoder tries the TMMIP technology, that is, it executes Step 2. If sps_tmmip_enable_flag is false, the encoder cannot try the TMMIP technology, that is, it skips Step 2 and directly executes Step 3.
[0311] Step 2: First, the encoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template area. The padding process is the same as the padding method of the original intra prediction process. For example, the encoder can traverse and pad from the lower left corner to the upper right corner. If all the reconstruction samples are available, padding is performed in order with all the available reconstruction samples. If all the reconstruction samples are unavailable, padding is performed with the average value for all. If some of the reconstruction samples are available, first padding is performed with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the encoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to perform padding on the previous unavailable position. Next, the encoder uses the reconstruction samples outside the padded second template area as input and utilizes the permitted MIP modes to predict the samples within the second template area.
[0312] Exemplarily, for a 4×4 size block, there are 16 permitted MIP modes. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 permitted MIP modes. For blocks of other sizes, there are 6 permitted MIP modes. Also, the MIP transpose function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0313] Exemplarily, the specific prediction calculation process includes the following. First, the encoder performs half-downsampling on the reconstruction samples. For example, the encoder determines the downsampling step size based on the block size. Next, the encoder adjusts the splicing order of the reconstruction samples after upper-side downsampling and the reconstruction samples after left-side downsampling according to the information of whether to transpose or not. When transposition is not required, the reconstruction samples after left-side downsampling are spliced after the reconstruction samples after upper-side downsampling, and the obtained vector is used as the input. When transposition is required, the reconstruction samples after upper-side downsampling are spliced after the reconstruction samples after left-side downsampling, and the obtained vector is used as the input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the encoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a prediction block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template, and output the prediction block within the second template area.
[0314] Furthermore, the encoder uses the DIMD technology to derive the optimal intra prediction mode, that is, to derive the optimal DIMD mode. In the DIMD technology, based on the Sobel operator, the gradient values of the reconstruction samples within the first template area are calculated, and the gradient values are converted based on the angle values corresponding to different prediction modes to obtain the amplitude values under the corresponding prediction modes.
[0315] Next, the encoder calculates the distortion cost based on the predicted blocks of the templates obtained by traversing each MIP mode and the reconstructed samples within the templates, and records the optimal MIP mode and transposition information according to the principle of the minimum cost. Further, the encoder traverses all the intra prediction modes allowed for use, calculates the amplitude values under each intra prediction mode, and records the optimal DIMD prediction mode according to the principle of the maximum amplitude.
[0316] Finally, based on the obtained optimal MIP mode and transposition information, the encoder may downsample the reconstructed samples adjacent to the upper side and the left side of the current block, splice them based on the transposition information to obtain an input vector, read out the matrix coefficients under the current mode with the MIP mode as the index, and then obtain an output vector by calculating the input vector and the matrix coefficients. The encoder transposes the output based on the transposition information, upsamples the output vector based on the size of the current block and the number of samples of the output vector to obtain an output of the same size as the current block and can use it as the optimal MIP prediction block of the current block. Further, for the optimal DIMD prediction mode, the encoder obtains the corresponding interpolation filtering coefficients, performs interpolation filtering on the reconstructed samples adjacent to the upper side and the left side to obtain the predicted samples at all positions within the current block, and records it as the optimal DIMD prediction block. The encoder obtains a new predicted block by weighted averaging of each predicted sample in the optimal MIP prediction block and the optimal DIMD prediction block according to the preset weights. The new predicted block is the predicted block of the current block.
[0317] Furthermore, the encoder obtains the rate-distortion cost of the current block and records it as cost1.
[0318] Note that the encoder determines the intra prediction mode derived from the DIMD mode used for the prediction block of the current block as the first intra prediction mode, or determines the intra prediction mode derived from the DIMD mode used for the reconstructed samples within the first template area as the first intra prediction mode.
[0319] Step 3: The encoder continues to traverse other intra prediction techniques, calculates the corresponding rate-distortion costs, and records them as cost2...costN.
[0320] Step 4: If cost1 is the minimum among all the rate-distortion costs, the TMMIP technique is used for the current block. The encoder sets the TMMIP usage flag of the current block to true and writes it into the bitstream. If cost1 is not the minimum rate-distortion cost, other intra prediction techniques are used for the current block. The encoder sets the TMMIP usage flag of the current block to false and writes it into the bitstream. Note that information such as the flags or indices of other intra prediction techniques is transmitted based on the definition and is not detailed here.
[0321] Step 5: The encoder determines the residual block of the current block based on the prediction block of the current block and the original block of the current block, performs a base transformation on the residual block of the current block, and performs a secondary transformation on the transformed coefficients after the base transformation based on the first intra prediction mode. Then, operations such as quantization, entropy coding, and loop filtering are performed on the transformed coefficients after the secondary transformation. Note that for the specific process of that quantization, the above related content can be referred to, and to avoid duplication, it is not detailed here.
[0322] Next, the related solutions of the decoder in this embodiment will be described.
[0323] The decoder analyzes block-level flags. If the intra mode is used for the current block, the decoder analyzes or obtains sequence-level permission flags such as the sps_tmmip_enable_flag. The sequence-level permission flag is used to indicate whether the template matching-based MIP mode derivation technique is permitted to be used in the current sequence. If any of the tmmip permission flags are true, it indicates that the TMMIP technique is currently permitted to be used by the decoder.
[0324] Exemplarily, the process of the decoder can be realized as follows.
[0325] Step 1: If the sps_tmmip_enable_flag is true, the decoder analyzes the TMMIP usage flag of the current block. Otherwise, in the current decoding process, the block-level TMMIP usage flag does not need to be decoded, and the block-level TMMIP usage flag is default set to false. If the TMMIP usage flag of the current block is true, Step 2 is executed. Otherwise, Step 3 is executed.
[0326] Step 2: First, the decoder performs reconstruction sample padding on the rows and columns adjacent to the outside of the second template region. The padding process is the same as the padding method of the original intra prediction process. For example, the decoder can traverse and pad from the lower left corner to the upper right corner. If all reconstruction samples are available, padding is performed in order with all available reconstruction samples. If all reconstruction samples are unavailable, all are padded with the average value. If some reconstruction samples are available, first pad with the available reconstruction samples, and for the remaining unavailable reconstruction samples, the decoder traverses in the order from the lower left corner to the upper right corner until the first available reconstruction sample appears, and uses the first available reconstruction sample to pad the previous unavailable positions. Next, the decoder uses the reconstruction samples outside the padded second template region as input and utilizes the permitted MIP mode to predict the samples within the second template region.
[0327] Exemplarily, for a 4×4 size block, there are 16 permitted MIP modes. For a block with a width or height equal to 4, or an 8×8 size block, there are 8 permitted MIP modes. For blocks of other sizes, there are 6 permitted MIP modes. Also, the MIP transposition function can be used for blocks of any size, and the above TMMIP prediction mode is the same as the MIP technology.
[0328] Exemplarily, the specific prediction calculation process includes the following. First, the decoder performs half-downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, according to the information of whether to transpose or not, the decoder adjusts the splicing order of the reconstructed samples after the upper downsampling and the reconstructed samples after the left downsampling. When transposition is not required, the reconstructed samples after the left downsampling are spliced after the reconstructed samples after the upper downsampling, and the obtained vector is used as the input. When transposition is required, the reconstructed samples after the upper downsampling are spliced after the reconstructed samples after the left downsampling, and the obtained vector is used as the input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input vector. Finally, the decoder upsamples the output vector according to the number of samples of the output vector and the size of the current template. When upsampling is not required, the output vector is arranged in order horizontally and output as a predicted block within the template area. When upsampling is required, first perform upsampling horizontally, and then upsampling perform upsampling until it reaches the same size as the template size, and output the predicted block within the second template area.
[0329] Furthermore, the decoder uses the DIMD technology to derive the optimal intra prediction mode, that is, to derive the optimal DIMD mode. In the DIMD technology, based on the Sobel operator, the gradient values of the reconstructed samples within the first template area are calculated, and the gradient values are converted based on the angle values corresponding to different prediction modes to obtain the amplitude values under the corresponding prediction modes.
[0330] Next, the decoder calculates the distortion cost based on the predicted block of the template obtained by traversing each MIP mode and the reconstructed samples within the template, and records the optimal MIP mode and transposition information according to the principle of the minimum cost. Further, the decoder traverses all the intra prediction modes allowed to be used, calculates the amplitude value under each intra prediction mode, and records the optimal DIMD prediction mode according to the principle of the maximum amplitude.
[0331] Finally, based on the obtained optimal MIP mode and transposition information, the decoder may downsample the reconstructed samples adjacent to the upper side and the left side of the current block, splice them based on the transposition information to obtain an input vector, read out the matrix coefficients under the current mode with the MIP mode as the index, and then obtain an output vector through the calculation of the input vector and the matrix coefficients. The decoder transposes the output based on the transposition information, upsamples the output vector based on the size of the current block and the number of samples of the output vector to obtain an output of the same size as the current block and can use it as the optimal MIP prediction block of the current block. Further, for the optimal DIMD prediction mode, the decoder obtains the corresponding interpolation filtering coefficients, performs interpolation filtering on the reconstructed samples adjacent to the upper side and the left side to obtain the predicted samples at all positions within the current block, and records it as the optimal DIMD prediction block. The decoder weights and averages each predicted sample in the optimal MIP prediction block and the optimal DIMD prediction block according to the preset weights to obtain a new predicted block. The new predicted block is the predicted block of the current block.
[0332] Note that the decoder determines the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or the decoder determines the intra prediction mode derived from the DIMD mode used for the reconstructed samples within the first template area as the first intra prediction mode.
[0333] Step 3: The decoder continues to analyze information such as the usage flag or index of other intra prediction techniques, and determines the final prediction block of the current block based on the analyzed information.
[0334] Step 4: The decoder analyzes the bitstream to obtain the frequency domain residual block of the current block (also referred to as frequency domain residual information), and performs inverse quantization and inverse transformation on the frequency domain residual block of the current block (first, perform inverse transformation on the secondary transformation based on the first intra prediction mode, and then perform inverse conversion transformation on the base transformation or primary transformation) to obtain the residual block of the current block (also referred to as the time domain residual block or time domain residual information). Next, the decoder adds the prediction block of the current block to the residual block of the current block to obtain the reconstructed sample block.
[0335] Step 5: After performing techniques such as loop filtering on all the reconstructed sample blocks in the current image, the final reconstructed image is obtained.
[0336] Optionally, the reconstructed image may be used as a video output, or may be used as a reference for subsequent decoding.
[0337] In this embodiment, for the calculation process of the optimal DIMD prediction block, reference may be made to the description content of the above-mentioned DIMD technology. To avoid duplication, it will not be described in detail here. Furthermore, the fusion weight of the optimal MIP prediction block and the optimal DIMD prediction block can be a preset value. For example, the optimal MIP prediction block occupies 5 / 9, and the optimal DIMD prediction block occupies 4 / 9. Of course, in other alternative embodiments, the fusion weight of the optimal MIP prediction block and the optimal DIMD prediction block may be other values, which are not specifically limited in this application. Note that the second template area and the first template area may be the same or different, which is not specifically limited in this application.
[0338] As described above, the preferred embodiments of the present application have been described in detail with reference to the accompanying drawings. However, the present application is not limited to the detailed content of the above embodiments. Within the scope of the technical idea of the present application, various simple modifications can be made to the technical solutions of the present application, and all of these simple modifications belong to the protection scope of the present application. For example, each specific technical feature described in the above specific embodiments may be combined by any appropriate means without contradiction. To avoid unnecessary duplication, various possible combinations are not described again in the present application. Also, for example, among various different embodiments of the present application, any combination should be regarded as disclosed in the present application as long as it does not conflict with the idea of the present application. It should be understood that in various method embodiments of the present application, the magnitude of the sequence number of each process above does not mean the execution order. The execution order of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0339] As described above, the method embodiments of the present application have been described in detail. Hereinafter, with reference to FIGS. 11 to 13, the apparatus embodiments of the present application will be described in detail.
[0340] FIG. 11 is a block diagram showing a decoder 500 according to an embodiment of the present application.
[0341] As shown in FIG. 11, the decoder 500 can include an analysis unit 510, a conversion unit 520, and a reconstruction unit 530. The analysis unit 510 is configured to analyze the bitstream of the current sequence to obtain the First transformation coefficients of the current block. The conversion unit 520 is configured to determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode that is used for a prediction block of a current block, an intra prediction mode derived from the DIMD mode that is used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode that is used for reconstructed samples within a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The transform unit 520 is configured to perform a first transform on the first transform coefficients based on a transform set corresponding to the first intra prediction mode to obtain second transform coefficients of the current block. The transform unit 520 is configured to perform a second transform on the second transform coefficients to obtain a residual block of the current block. The reconstruction unit 530 is configured to determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block.
[0342] In some embodiments, the output vector of the optimal MIP mode is the vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is the vector after upsampling the output vector of the optimal MIP mode.
[0343] In some embodiments, the transform unit 520 is specifically configured to determine the first intra prediction mode based on a prediction mode for predicting the current block.
[0344] In some embodiments, when the prediction modes for predicting the current block specifically include an optimal MIP mode and a sub-optimal MIP mode for predicting the current block, the conversion unit 520 is configured to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode.
[0345] In some embodiments, when the prediction modes for predicting the current block specifically include an optimal MIP mode and an intra prediction mode derived from the TIMD mode, the conversion unit 520 is configured to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or to determine, as the first intra prediction mode, an intra prediction mode derived from the TIMD mode.
[0346] In some embodiments, when the prediction modes for predicting the current block specifically include an optimal MIP mode and an intra prediction mode derived from the DIMD mode that is used for the reconstructed samples within the first template region, the conversion unit 520 is configured to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the reconstructed samples within the first template region.
[0347] In some embodiments, the reconstruction unit 530 is further configured to determine a second intra prediction mode. The second intra prediction mode includes any one of an almost-optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples within the first template region, and an intra prediction mode derived from the TIMD mode. The reconstruction unit 530 is further configured to predict the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block.
[0348] In some embodiments, specifically, the reconstruction unit 530 predicts the current block based on the optimal MIP mode to obtain a first predicted block, predicts the current block based on the second intra prediction mode to obtain a second predicted block, and performs a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block.
[0349] In some embodiments, before the reconstruction unit 530 performs a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block, if the prediction mode for predicting the current block includes the optimal MIP mode, an almost-optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, the weight of the optimal MIP mode and the weight of the second intra prediction mode are determined based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode. if the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode used for reconstructed samples within the first template region, it is configured to determine that both the weight of the optimal MIP mode and the weight of the second intra prediction mode are preset values.
[0350] In some embodiments, the conversion unit 520 is specifically configured to analyze the bitstream of the current sequence to obtain a first flag, and when the first flag is used to indicate that it is permitted to predict an image block in the current sequence using an optimal MIP mode and a second intra prediction mode, determine the second intra prediction mode.
[0351] In some embodiments, the conversion unit 520 is specifically configured to when the first flag is used to indicate that it is permitted to predict an image block in the current sequence using an optimal MIP mode and a second intra prediction mode, analyze the bitstream to obtain a second flag, and when the second flag is used to indicate that it is permitted to predict the current block using an optimal MIP mode and a second intra prediction mode, determine the second intra prediction mode.
[0352] In some embodiments, the reconstruction unit 530 is further configured to determine an optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes. The distortion costs corresponding to the plurality of MIP modes include distortion costs obtained by predicting samples in a second template region adjacent to the current block using the plurality of MIP modes.
[0353] In some embodiments, the second template region and the first template region may be the same or different.
[0354] In some embodiments, the reconstruction unit 530 is specifically configured to predict samples in the second template region based on a third flag and a plurality of MIP modes, and obtain distortion costs corresponding to the plurality of MIP modes under each state of the third flag. and determining an optimal MIP mode based on distortion costs corresponding to the plurality of MIP modes under each state of the third flag. The third flag is used to indicate whether to transpose the input and output vectors corresponding to the MIP mode.
[0355] In some embodiments, before determining the optimal MIP mode based on the distortion costs corresponding to the multiple MIP modes, the reconstruction unit 530 further Get the MIP mode used for the neighboring blocks adjacent to the current block, The MIP mode used for the neighboring blocks is determined as a plurality of MIP modes.
[0356] In some embodiments, before determining the optimal MIP mode based on the distortion costs corresponding to the multiple MIP modes, the reconstruction unit 530 further performing reconstruction sample padding on a reference region adjacent to an outer side of the second template region to obtain a reference row and a reference column of the second template region; Using the reference row and the reference column as input, predicting samples in the second template region using each of a plurality of MIP modes to obtain a plurality of predicted blocks corresponding to the plurality of MIP modes; and determining distortion costs corresponding to a plurality of MIP modes based on the plurality of predicted blocks and the reconstructed blocks in the second template region.
[0357] In some embodiments, the reconstruction unit 530 specifically: Downsample the reference row and the reference column to obtain the input vector, Taking the input vector as input, predicting samples in a second template region by traversing a plurality of MIP modes to obtain an output vector corresponding to the plurality of MIP modes; It is configured to upsample output vectors corresponding to a plurality of MIP modes to obtain prediction blocks corresponding to the plurality of MIP modes.
[0358] In some embodiments, the reconstruction unit 530 is specifically configured to determine an optimal MIP mode based on the sum of absolute differences of transform (SATD) corresponding to a plurality of MIP modes within the second template region.
[0359] FIG. 12 is a block diagram showing an encoder 600 according to an embodiment of the present application.
[0360] As shown in FIG. 12, the encoder 600 can include a residual unit 610, a conversion unit 620, and an encoding unit 630. The residual unit 610 is configured to obtain a residual block of the current block in the current sequence. The conversion unit 620 is configured to perform a third conversion on the residual block of the current block to obtain third conversion coefficients of the current block. The conversion unit 620 is configured to determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The conversion unit 620 is configured to perform a fourth conversion on the third conversion coefficients based on a conversion set corresponding to the first intra prediction mode to obtain fourth conversion coefficients of the current block. The encoding unit 630 is configured to encode a fourth conversion coefficient.
[0361] In some embodiments, the output vector of the optimal MIP mode is the vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is the vector after upsampling the output vector of the optimal MIP mode.
[0362] In some embodiments, the conversion unit 620 is specifically configured to determine a first intra prediction mode based on a prediction mode for predicting a current block.
[0363] In some embodiments, when the prediction mode for predicting the current block includes an optimal MIP mode and a sub-optimal MIP mode for predicting the current block, the conversion unit 620 is specifically configured to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode.
[0364] In some embodiments, when the prediction mode for predicting the current block includes an optimal MIP mode and an intra prediction mode derived from the TIMD mode, the conversion unit 620 is specifically configured to determine, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determine, as the first intra prediction mode, an intra prediction mode derived from the TIMD mode.
[0365] In some embodiments, specifically when the prediction modes for predicting the current block include the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, the conversion unit 620 is configured to determine the intra prediction mode derived from the DIMD mode as the first intra prediction mode to be used for the prediction block of the current block, or to determine the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region as the first intra prediction mode.
[0366] In some embodiments, specifically, the residual unit 610 is configured to determine the second intra prediction mode, predict the current block based on the optimal MIP mode and the second intra prediction mode to obtain the prediction block of the current block, and obtain the residual block of the current block based on the prediction block of the current block. The second intra prediction mode includes any one of a quasi-optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, and an intra prediction mode derived from the TIMD mode.
[0367] In some embodiments, specifically, the residual unit 610 predicts the current block based on the optimal MIP mode to obtain the first prediction block, predicts the current block based on the second intra prediction mode to obtain the second prediction block, and is configured to perform a weighting process on the first prediction block and the second prediction block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain the prediction block of the current block.
[0368] In some embodiments, the residual unit 610 performs a weighting process on the first prediction block and the second prediction block based on the weights of the optimal MIP mode and the weights of the second intra prediction mode, and further, before obtaining the prediction block of the current block, when the prediction mode for predicting the current block includes the optimal MIP mode, the sub-optimal MIP mode for predicting the current block, or the intra prediction mode derived from the TIMD mode, the weights of the optimal MIP mode and the weights of the second intra prediction mode are determined based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode, when the prediction mode for predicting the current block includes the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, it is configured to determine that the weights of the optimal MIP mode and the weights of the second intra prediction mode are both preset values.
[0369] In some embodiments, specifically, the residual unit 610 obtains a first flag, and when the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the second intra prediction mode, it is configured to determine the second intra prediction mode. Specifically, the encoding unit 630 is configured to encode the fourth transform coefficient and the first flag.
[0370] In some embodiments, specifically, the residual unit 610 when the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the second intra prediction mode, predicts the current block based on the optimal MIP mode and the second intra prediction mode to obtain a first rate-distortion cost, Predict the current block based on at least one intra prediction mode to obtain at least one rate-distortion cost. When the first rate-distortion cost is less than or equal to the minimum value among the at least one rate-distortion cost, the predicted block obtained by predicting the current block based on the optimal MIP mode and the second intra prediction mode is determined as the predicted block of the current block. Specifically, the encoding unit 630 is configured to encode the fourth conversion coefficient, the first flag, and the second flag. When the first rate-distortion cost is less than or equal to the minimum value among the at least one rate-distortion cost, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the second intra prediction mode. When the first rate-distortion cost is greater than the minimum value among the at least one rate-distortion cost, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the second intra prediction mode.
[0371] In some embodiments, the residual unit 610 is further configured to determine the optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes. The distortion costs corresponding to the plurality of MIP modes include the distortion costs obtained by predicting samples in the second template region adjacent to the current block using the plurality of MIP modes.
[0372] In some embodiments, the second template region and the first template region may be the same or different.
[0373] In some embodiments, the residual unit 610 is specifically configured to predict samples within the second template region based on a third flag and a plurality of MIP modes, obtain distortion costs corresponding to the plurality of MIP modes under each state of the third flag, and determine an optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes under each state of the third flag. The third flag is used to indicate whether to transpose the input vector and the output vector corresponding to the MIP mode.
[0374] In some embodiments, before determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes, the residual unit 610 further obtains the MIP mode used for an adjacent block adjacent to the current block, and is configured to determine the MIP mode used for the adjacent block as one of the plurality of MIP modes.
[0375] In some embodiments, before determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes, the residual unit 610 further performs reconstruction sample padding on a reference region adjacent to the outside of the second template region to obtain a reference row and a reference column of the second template region, uses each of the plurality of MIP modes with the reference row and the reference column as inputs to predict samples within the second template region to obtain a plurality of prediction blocks corresponding to the plurality of MIP modes, and is configured to determine the distortion costs corresponding to the plurality of MIP modes based on the plurality of prediction blocks and the reconstruction blocks within the second template region.
[0376] In some embodiments, the residual unit 610 specifically downsamples the reference row and the reference column to obtain an input vector, Using the input vector as input, samples in the second template region are predicted by traversing a plurality of MIP modes to obtain output vectors corresponding to the plurality of MIP modes. It is configured to upsample the output vectors corresponding to the plurality of MIP modes to obtain prediction blocks corresponding to the plurality of MIP modes.
[0377] In some embodiments, the residual unit 610 is specifically configured to determine the optimal MIP mode based on the sum of absolute differences of transform (SATD) corresponding to a plurality of MIP modes in the second template region.
[0378] Note that the apparatus embodiments can correspond to the method embodiments. For similar descriptions, reference can be made to the method embodiments. To avoid repetition, the description is omitted here. Specifically, the decoder 500 shown in FIG. 11 may correspond to the entity that executes the method 300 in the embodiments of the present application. Also, the foregoing and other operations and / or functions of each unit in the decoder 500 are respectively used to implement the corresponding processes in each method such as the method 300. Similarly, the encoder 600 shown in FIG. 12 may correspond to the entity that executes the method 400 in the embodiments of the present application. That is, the foregoing and other operations and / or functions of each unit in the encoder 600 are respectively used to implement the corresponding processes in each method such as the method 400.
[0379] Furthermore, each unit in the decoder 500 or the encoder 600 according to the embodiments of the present application may be integrated, either individually or in whole, with one or several other units, or some of them may be further divided into a plurality of functionally smaller units. Thereby, the same operations can be realized without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logic functions. In actual applications, the function of one unit can be realized by a plurality of units, or the functions of a plurality of units can be realized by one unit. In other embodiments of the present application, the decoder 500 or the encoder 600 may include other units. In actual applications, these functions may be realized by the cooperation of other units, or may also be realized by the cooperation of a plurality of units. According to another embodiment of the present application, for example, by a general-purpose computer device including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read only memory (ROM), and executing a computer program (including program code) capable of executing each step according to the corresponding method, the decoder 500 or the encoder 600 according to the embodiments of the present application is constructed, and the encoding method or the decoding method according to the embodiments of the present application is realized. The computer program is recorded on a computer-readable storage medium, for example, and by being mounted on an electronic device via the computer-readable storage medium and operating therein, the corresponding method in the embodiments of the present application can be realized.
[0380] In other words, the unit related to the above may be implemented in the form of hardware, may be implemented by instructions in the form of software, or may be implemented by a combination of hardware and software. Specifically, each step of the method embodiment in the embodiments of the present application can be completed by an integrated logic circuit of hardware in a processor or instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software in the decoding processor. Optionally, the software can be located in a mature storage medium in the technical field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory. The processor reads the information in the memory and completes the steps of the above method embodiment together with the hardware of the processor.
[0381] FIG. 13 is a block diagram showing an electronic device 700 according to an embodiment of the present application.
[0382] As shown in FIG. 13, the electronic device 700 includes at least a processor 710 and a computer-readable storage medium 720. , transceiver 730 and Note that the processor 710 、 The computer-readable storage medium 720 , and transceiver 730 are connected in a bus or other manner. mutuallyIt may be connected. The computer-readable storage medium 720 is used to store a computer program 721 including computer instructions. The processor 710 is used to execute the computer instructions stored in the computer-readable storage medium 720. The processor 710 is the computing core and control core of the electronic device 700, and is suitable for implementing one or more computer instructions. Specifically, by loading and executing one or more computer instructions, it is suitable for implementing the process of the corresponding method or the corresponding function.
[0383] By way of example, the processor 710 may also be referred to as a CPU. The processor 710 may include, but is not limited to, a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic devices, discrete hardware components, etc.
[0384] By way of example, computer-readable storage medium 720 may be high-speed RAM memory, or may be non-volatile memory, such as at least one magnetic disk storage device. Optionally, computer-readable storage medium 720 may be at least one computer-readable storage medium remote from processor 710. Specifically, computer-readable storage medium 720 includes, but is not limited to, volatile memory and / or non-volatile memory. Non-volatile memory can be ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), or flash memory. Volatile memory can be RAM functioning as an external high-speed cache. By way of non-limiting example, various RAMs are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), direct rambus RAM (DRRAM).
[0385] In one embodiment, the electronic device 700 may be an encoder or an encoding framework according to the embodiments of the present application. The computer-readable storage medium 720 stores a first computer instruction. By loading and executing the first computer instruction stored in the computer-readable storage medium 720, the processor 710 realizes the corresponding steps in the encoding method according to the embodiments of the present application. In other words, the first computer instruction in the computer-readable storage medium 720 is loaded and executed by the processor 710 to execute the corresponding steps. To avoid repetition, the description is omitted here.
[0386] In one embodiment, the electronic device 700 may be a decoder or a decoding framework according to the embodiments of the present application. The computer-readable storage medium 720 stores a second computer instruction. By loading and executing the second computer instruction stored in the computer-readable storage medium 720, the processor 710 realizes the corresponding steps in the decoding method according to the embodiments of the present application. In other words, the second computer instruction in the computer-readable storage medium 720 is loaded and executed by the processor 710 to execute the corresponding steps. To avoid repetition, the description is omitted here.
[0387] According to another aspect of the present application, in the embodiments of the present application, a coding system is provided. The coding system includes the encoder and decoder as described above.
[0388] According to another aspect of the present application, in an embodiment of the present application, a computer-readable storage medium (Memory) is provided. The computer-readable storage medium is a storage device in the electronic device 700 and is used to store programs and data. For example, the computer-readable storage medium may be the computer-readable storage medium 720. As can be understood, the computer-readable storage medium 720 here may include a built-in storage medium in the electronic device 700, and of course, may also include an extended storage medium supported by the electronic device 700. The computer-readable storage medium provides a storage space in which the operating system of the electronic device 700 is stored. Further, one or more computer instructions suitable for being loaded and executed by the processor 710 are stored in the storage space, and these computer instructions may be one or more computer programs 721 (including program codes).
[0389] According to another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. For example, the computer instructions may be the computer program 721. In this case, electronic device 700 can be a computer, the processor 710 reads the computer instructions from the computer-readable storage medium 720, and the processor 710 executes the computer instructions to cause the computer to execute the encoding method or the decoding method according to the various selectable methods described above.
[0390] In other words, when implemented by software, all or part of the above embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes described in the embodiments of the present application are executed, or the functions described in the embodiments of the present application are realized. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL), etc.) or wirelessly (such as infrared, radio, microwave, etc.).
[0391] It is obvious to those skilled in the art that the present application can be realized by electronic hardware, or a combination of computer software and electronic hardware, in combination with each exemplary unit and process step described in the embodiments disclosed in this specification. Whether these functions are executed by hardware or software is determined by specific application cases of the technical solution and design limitations, etc. Those skilled in the art can use different methods for each specific application to realize the described functions, but these realizations should not be regarded as exceeding the scope of the present application.
[0392] Finally, the above content is only a specific embodiment of the present application, and the protection scope of the present application is not limited thereto. All changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be determined by the protection scope of the claims.
Claims
Claim 1 A decoding method, comprising: analyzing a bitstream of a current sequence to obtain first transform coefficients of a current block; determining a first intra prediction mode, wherein the first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode; performing a first transform on the first transform coefficients based on a transform set corresponding to the first intra prediction mode to obtain second transform coefficients of the current block; performing a second transform on the second transform coefficients to obtain a residual block of the current block; determining a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block; and characterized in that it is a decoding method. Claim 2 The output vector of the optimal MIP mode is a vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is a vector after upsampling the output vector of the optimal MIP mode. The method according to claim 1, characterized in that. Claim 3 Determining the first intra prediction mode includes: determining the first intra prediction mode based on a prediction mode for predicting the current block. The method according to claim 1 or 2, characterized in that. Claim 4 Determining the first intra prediction mode based on a prediction mode for predicting the current block includes: When the prediction mode for predicting the current block includes the optimal MIP mode and a sub-optimal MIP mode for predicting the current block, determining the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or determining the intra prediction mode derived from the DIMD mode used for the output vector of the optimal MIP mode as the first intra prediction mode, The method according to claim 3, characterized in that.
5. Determining the first intra prediction mode based on the prediction mode for predicting the current block, When the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the TIMD mode, determining the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or determining the intra prediction mode derived from the TIMD mode as the first intra prediction mode, The method according to claim 3, characterized in that.
6. Determining the first intra prediction mode based on the prediction mode for predicting the current block, When the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode used for the reconstructed samples within the first template area, determining the intra prediction mode derived from the DIMD mode used for the predicted block of the current block as the first intra prediction mode, or determining the intra prediction mode derived from the DIMD mode used for the reconstructed samples within the first template area as the first intra prediction mode, The method according to claim 3, characterized in that.
7. The method is, Determining a second intra prediction mode, wherein the second intra prediction mode includes any one of a quasi-optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode and used for reconstructed samples within the first template region, and an intra prediction mode derived from the TIMD mode, and performing the determination; Predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block; further comprising; The method according to any one of claims 1 to 6, characterized in that.
8. Predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block includes: Predicting the current block based on the optimal MIP mode to obtain a first predicted block; Predicting the current block based on the second intra prediction mode to obtain a second predicted block; Performing a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block; including; The method according to claim 7, characterized in that.
9. Before performing a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block, the method includes: When the prediction mode for predicting the current block includes the optimal MIP mode, a quasi-optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, determining the weight of the optimal MIP mode and the weight of the second intra prediction mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode; When the prediction mode for predicting the current block includes the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, it is determined that both the weight of the optimal MIP mode and the weight of the second intra prediction mode are preset values. Further comprising The method according to claim 8, characterized in that.
10. Determining the second intra prediction mode Analyzing the bitstream of the current sequence to obtain a first flag When the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode, determining the second intra prediction mode Including The method according to any one of claims 7 to 9, characterized in that.
11. Determining the second intra prediction mode When the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode, analyzing the bitstream to obtain a second flag When the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the second intra prediction mode, determining the second intra prediction mode Including The method according to claim 10, characterized in that.
12. The method Further includes determining the optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes The distortion costs corresponding to the plurality of MIP modes include the distortion costs obtained by predicting samples within a second template region adjacent to the current block using the plurality of MIP modes. The method according to any one of claims 1 to 11, characterized in that.
13. The second template region and the first template region may be the same or different. The method according to claim 12, characterized in that.
14. Determining the optimal MIP mode based on the distortion costs corresponding to a plurality of MIP modes Predicting samples within the second template region based on the third flag and the plurality of MIP modes, and obtaining distortion costs corresponding to the plurality of MIP modes under each state of the third flag, wherein the third flag is used to indicate whether to transpose the input vector and the output vector corresponding to the MIP mode, obtaining; Determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes under each state of the third flag; including; The method according to claim 12 or 13, characterized in that.
15. Before determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes, the method includes: Obtaining the MIP mode used for an adjacent block adjacent to the current block; Determining the MIP mode used for the adjacent block as the plurality of MIP modes; further including; The method according to any one of claims 12 to 14, characterized in that.
16. Before determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes, the method includes: Performing reconstruction sample padding on a reference region adjacent to the outside of the second template region to obtain a reference row and a reference column of the second template region; Using each of the plurality of MIP modes with the reference row and the reference column as inputs to predict samples within the second template region, and obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes; Determining the distortion costs corresponding to the plurality of MIP modes based on the plurality of prediction blocks and the reconstruction blocks within the second template region; further including; The method according to any one of claims 12 to 15, characterized in that.
17. Using each of the plurality of MIP modes with the reference row and the reference column as inputs to predict samples within the second template region, and obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes, includes: Downsampling the reference row and the reference column to obtain an input vector; Using the input vector as an input, traversing the plurality of MIP modes to predict samples within the second template region, and obtaining output vectors corresponding to the plurality of MIP modes; Upsampling the output vectors corresponding to the plurality of MIP modes to obtain prediction blocks corresponding to the plurality of MIP modes; including The method according to claim 16, characterized in that. **Claim 18** Determining the optimal MIP mode based on distortion costs corresponding to a plurality of MIP modes includes determining the optimal MIP mode based on the sum of absolute differences of differential transforms (SATD) corresponding to the plurality of MIP modes in the second template region. The method according to any one of claims 12 to 17, characterized in that. **Claim 19** An encoding method, wherein the method is applied to an encoder, and the method comprises: obtaining a residual block of a current block in a current sequence; performing a third transform on the residual block of the current block to obtain third transform coefficients of the current block; determining a first intra prediction mode, wherein the first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode; performing a fourth transform on the third transform coefficients based on a transform set corresponding to the first intra prediction mode to obtain fourth transform coefficients of the current block; encoding the fourth transform coefficients; including An encoding method, characterized in that. **Claim 20** The output vector of the optimal MIP mode is the vector before upsampling the output vector of the optimal MIP mode, or the output vector of the optimal MIP mode is the vector after upsampling the output vector of the optimal MIP mode. The method according to claim 19, characterized in that. **Claim 21** Determining the first intra prediction mode includes determining the first intra prediction mode based on a prediction mode for predicting the current block. The method according to claim 19 or 20, characterized in that. **Claim 22** Determining the first intra prediction mode based on a prediction mode for predicting the current block includes when the prediction mode for predicting the current block includes the optimal MIP mode and a sub-optimal MIP mode for predicting the current block, determining, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determining, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the output vector of the optimal MIP mode. The method according to claim 21, characterized in that. **Claim 23** Determining the first intra prediction mode based on a prediction mode for predicting the current block includes when the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the TIMD mode, determining, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determining, as the first intra prediction mode, an intra prediction mode derived from the TIMD mode. The method according to claim 21, characterized in that. **Claim 24** Determining the first intra prediction mode based on a prediction mode for predicting the current block includes when the prediction mode for predicting the current block includes the optimal MIP mode and an intra prediction mode derived from the DIMD mode that is used for reconstruction samples in the first template area, determining, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for the predicted block of the current block, or determining, as the first intra prediction mode, an intra prediction mode derived from the DIMD mode that is used for reconstruction samples in the first template area. The method according to claim 21, characterized in that...
25. Obtaining the residual block of the current block in the current sequence is... Determining a second intra prediction mode, where the second intra prediction mode includes any one of an almost optimal MIP mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples within the first template region, and an intra prediction mode derived from the TIMD mode; and determining... Predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block... Obtaining the residual block of the current block based on the predicted block of the current block... including... The method according to any one of claims 19 to 22, characterized in that...
26. Predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block is... Predicting the current block based on the optimal MIP mode to obtain a first predicted block... Predicting the current block based on the second intra prediction mode to obtain a second predicted block... Performing a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block... including... The method according to claim 25, characterized in that...
27. Before performing a weighting process on the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the second intra prediction mode to obtain a predicted block of the current block, the method is... If the prediction mode for predicting the current block includes the optimal MIP mode, an almost optimal MIP mode for predicting the current block, or an intra prediction mode derived from the TIMD mode, determining the weight of the optimal MIP mode and the weight of the second intra prediction mode based on the distortion cost corresponding to the optimal MIP mode and the distortion cost corresponding to the second intra prediction mode... When the prediction mode for predicting the current block includes the optimal MIP mode and the intra prediction mode derived from the DIMD mode used for the reconstruction samples within the first template region, it is determined that both the weight of the optimal MIP mode and the weight of the second intra prediction mode are preset values. Further comprising The method according to claim 26, characterized in that.
28. Determining the second intra prediction mode Obtaining a first flag When the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode, determining the second intra prediction mode. Encoding the fourth transform coefficient Including encoding the fourth transform coefficient and the first flag The method according to any one of claims 25 to 27, characterized in that.
29. Predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a predicted block of the current block When the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the second intra prediction mode, predicting the current block based on the optimal MIP mode and the second intra prediction mode to obtain a first rate-distortion cost. Predicting the current block based on at least one intra prediction mode to obtain at least one rate-distortion cost. When the first rate-distortion cost is less than or equal to the minimum value of the at least one rate-distortion cost, determining the predicted block obtained by predicting the current block based on the optimal MIP mode and the second intra prediction mode as the predicted block of the current block. Encoding the fourth transform coefficient and the first flag Including encoding the fourth transform coefficient, the first flag and the second flag When the first rate distortion cost is less than or equal to the minimum value among the at least one rate distortion cost, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is permitted; when the first rate distortion cost is greater than the minimum value among the at least one rate distortion cost, the second flag is used to indicate that predicting the current block using the optimal MIP mode and the second intra prediction mode is not permitted. The method according to claim 28, characterized in that.
30. The method includes further determining the optimal MIP mode based on rate distortion costs corresponding to a plurality of MIP modes, wherein the rate distortion costs corresponding to the plurality of MIP modes include rate distortion costs obtained by predicting samples in a second template region adjacent to the current block using the plurality of MIP modes. The method according to any one of claims 19 to 29, characterized in that.
31. The second template region and the first template region may be the same or different. The method according to claim 30, characterized in that.
32. Determining the optimal MIP mode based on rate distortion costs corresponding to a plurality of MIP modes includes predicting samples in the second template region based on a third flag and the plurality of MIP modes to obtain rate distortion costs corresponding to the plurality of MIP modes under each state of the third flag, where the third flag is used to indicate whether to transpose an input vector and an output vector corresponding to the MIP mode, and obtaining; determining the optimal MIP mode based on the rate distortion costs corresponding to the plurality of MIP modes under each state of the third flag; including The method according to claim 30 or 31, characterized in that.
33. Before determining the optimal MIP mode based on rate distortion costs corresponding to a plurality of MIP modes, the method includes obtaining the MIP mode used for an adjacent block adjacent to the current block; determining the MIP mode used for the adjacent block as the plurality of MIP modes; further including The method according to any one of claims 30 to 32, characterized in that.
34. Based on the distortion costs corresponding to a plurality of MIP modes, before determining the optimal MIP mode, the method comprises: Performing reconstruction sample padding on a reference region adjacent to the outside of the second template region to obtain a reference row and a reference column of the second template region; Using each of the plurality of MIP modes with the reference row and the reference column as inputs to predict samples within the second template region, obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes; Determining distortion costs corresponding to the plurality of MIP modes based on the plurality of prediction blocks and reconstruction blocks within the second template region; Further comprising: The method according to any one of claims 30 to 33, characterized in that.
35. Using each of the plurality of MIP modes with the reference row and the reference column as inputs to predict samples within the second template region, obtaining a plurality of prediction blocks corresponding to the plurality of MIP modes, comprises: Downsampling the reference row and the reference column to obtain an input vector; Using the input vector as an input to traverse the plurality of MIP modes to predict samples within the second template region, obtaining output vectors corresponding to the plurality of MIP modes; Upsampling the output vectors corresponding to the plurality of MIP modes to obtain prediction blocks corresponding to the plurality of MIP modes; Comprising: The method according to claim 34, characterized in that.
36. Based on the distortion costs corresponding to a plurality of MIP modes, determining the optimal MIP mode, comprises: Determining the optimal MIP mode based on the sum of absolute differences of differential transforms (SATD) corresponding to the plurality of MIP modes within the second template region. The method according to any one of claims 30 to 35, characterized in that.
37. A decoder, comprising: An analysis unit, a conversion unit, and a reconstruction unit; The analysis unit is configured to analyze a bitstream of a current sequence to obtain first conversion coefficients of a current block; The conversion unit is configured to determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The conversion unit is configured to perform a first conversion on the first conversion coefficient based on a conversion set corresponding to the first intra prediction mode to obtain a second conversion coefficient of the current block. The conversion unit is configured to perform a second conversion on the second conversion coefficient to obtain a residual block of the current block. The reconstruction unit is configured to determine a reconstructed block of the current block based on the prediction block of the current block and the residual block of the current block. A decoder characterized by the above.
38. An encoder, comprising a residual unit, a conversion unit, and an encoding unit, the residual unit is configured to obtain a residual block of a current block in a current sequence, the conversion unit is configured to perform a third conversion on the residual block of the current block to obtain a third conversion coefficient of the current block, the conversion unit is configured to determine a first intra prediction mode. The first intra prediction mode includes any one of an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode used for a prediction block of the current block, an intra prediction mode derived from the DIMD mode used for an output vector of an optimal matrix-based intra prediction (MIP) mode for predicting the current block, an intra prediction mode derived from the DIMD mode used for reconstructed samples in a first template region adjacent to the current block, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The conversion unit is configured to perform a fourth conversion on the third conversion coefficient based on a conversion set corresponding to the first intra prediction mode to obtain a fourth conversion coefficient of the current block. The encoding unit is configured to encode the fourth conversion coefficient. An encoder characterized by the above.
39. An electronic device comprising: a processor and a computer-readable storage medium; the processor is configured to execute a computer program; the computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 18 or the method according to any one of claims 19 to 36 is executed. An electronic device characterized by the above.
40. A computer-readable storage medium, the computer-readable storage medium is configured to store a computer program, and the computer program causes a computer to execute the method according to any one of claims 1 to 18 or the method according to any one of claims 19 to 36. A computer-readable storage medium characterized by the above.
41. A computer program product, the computer program product includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 18 or the method according to any one of claims 19 to 36 is executed. A computer program product characterized by the above.
42. A bit stream, wherein the bit stream is the bit stream in the method according to any one of claims 1 to 18, or the bit stream generated based on the method according to any one of claims 19 to 36, characterized bit stream.
Citation Information
Patent Citations
Method, apparatus and computer program for video coding
JP2022515849A
Video processing method, device, storage medium, and recording medium
JP2022526990A
Encoder, decoder, and corresponding method for reconciling matrix-based intra prediction and quadratic transform core selection
JP2022529030A
Decoder-side intra mode derivation
JP2023527662A
Method and apparatus for video coding
US20200389666A1