Decoding method, encoding method, decoder, and encoder
The proposed decoding and encoding methods using MIP, DIMD, and TIMD modes address the inefficiencies in existing video compression by reducing bit overhead and improving decompression efficiency through optimal prediction mode determination.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-04-08
- Publication Date
- 2026-06-01
AI Technical Summary
Existing digital video compression technologies face challenges in improving compression efficiency to handle the increasing demand for high-quality video content.
A decoding method and encoder that utilize matrix-based intra-prediction (MIP) modes, combined with decoder-side intra-mode derivation (DIMD) and template-based intra-mode derivation (TIMD) to determine optimal prediction modes, reducing bit overhead and enhancing prediction accuracy and diversity.
This approach effectively reduces bit overhead and improves decompression efficiency by autonomously determining optimal prediction modes, thereby enhancing video compression performance.
Smart Images

Figure 0007868170000005 
Figure 0007868170000006 
Figure 0007868170000007
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of image video coding, and more specifically, to a decoding method, an encoding method, a decoder, and an encoder.
Background Art
[0002] Digital video compression technology is mainly a technology for compressing huge digital image video data for transmission, storage, etc. With the rapid increase of Internet videos and the ever-growing demands of people for video sharpness, although video decompression technology can be realized with existing digital video compression standards, in order to improve the compression efficiency, better digital video decompression technology is needed.
Summary of the Invention
[0003] Embodiments of the present application provide a decoding method, an encoding method, a decoder, and an encoder, whereby the compression efficiency can be improved.
[0004] In a first embodiment, the present application provides a decoding method, the method comprising: analyzing a bitstream to obtain a residual block of the current block in the current sequence; determining the optimal MIP mode for predicting the current block based on strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, the strain costs corresponding to multiple MIP modes including strain costs obtained by predicting samples in a template region adjacent to the current block using multiple MIP modes; determining a first intra-prediction mode, the first intra-prediction mode being determined based on strain costs corresponding to multiple MIP modes and comprising at least one of a suboptimal MIP mode for predicting the current block, an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode, or an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode; predicting the current block based on the optimal MIP mode and the first intra-prediction mode to obtain a predicted block of the current block; and obtaining a reconstructed block of the current block based on the residual block of the current block and the predicted block of the current block.
[0005] In a second embodiment, the present application provides an encoding method comprising: determining the optimal MIP mode for predicting the current block in the current sequence based on the strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes; the strain costs corresponding to multiple MIP modes include the strain costs obtained by using multiple MIP modes to predict a sample in a template region adjacent to the current block; determining a first intra-prediction mode, which includes at least one of a suboptimal MIP mode determined based on the strain costs corresponding to multiple MIP modes and for predicting the current block, an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode, or an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode; predicting the current block based on the optimal MIP mode and the first intra-prediction mode to obtain a predicted block of the current block; obtaining a residual block of the current block based on the predicted block of the current block and the original block of the current block; encoding the residual block of the current block to obtain a bitstream of the current sequence.
[0006] In a third embodiment, the application provides a decoder comprising an analysis unit, a prediction unit, and a reconstruction unit. The analysis unit is configured to analyze a bitstream to obtain a residual block of the current block in the current sequence. The prediction unit is configured to determine an optimal MIP mode for predicting the current block based on strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, the strain costs corresponding to multiple MIP modes include strain costs obtained by using multiple MIP modes to predict samples in a template region adjacent to the current block. The prediction unit is configured to determine a first intra-prediction mode, the first intra-prediction mode being determined based on strain costs corresponding to multiple MIP modes and including at least one of an intra-optimal MIP mode for predicting the current block, an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode, and an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode. The prediction unit is configured to predict the current block based on the optimal MIP mode and the first intra-prediction mode to obtain a predicted block of the current block. The reconstruction unit is configured to obtain the reconstructed block of the current block based on the residual block of the current block and the predicted block of the current block.
[0007] In a fourth aspect, the application provides an encoder comprising a prediction unit, a residual unit, and an encoding unit. The prediction unit is configured to determine an optimal MIP mode for predicting the current block in the current sequence based on strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, the strain costs corresponding to multiple MIP modes include strain costs obtained by using multiple MIP modes to predict a sample in a template region adjacent to the current block. The prediction unit is configured to determine a first intra-prediction mode, the first intra-prediction mode being determined based on strain costs corresponding to multiple MIP modes and including at least one of an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode and an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode for predicting the current block. The prediction unit is configured to predict the current block based on the optimal MIP mode and the first intra-prediction mode to obtain a predicted block of the current block. The residual unit is configured to obtain the residual block of the current block based on the predicted block of the current block and the original block of the current block. The encoding unit is configured to encode the residual block of the current block to obtain the bitstream of the current sequence.
[0008] In a fifth embodiment, the application provides a decoder comprising a processor and a computer-readable storage medium. The processor is configured to execute computer instructions. Computer instructions are stored in the computer-readable storage medium, which is loaded by the processor and the decoding method of the first embodiment or each embodiment thereof is executed. In one embodiment, there is one or more processors and one or more memories. In one embodiment, the computer-readable storage medium may be integrated with the processor, or it may be installed separately from the processor.
[0009] In a sixth embodiment, the application provides an encoder comprising a processor and a computer-readable storage medium. The processor is configured to execute computer instructions. Computer instructions are stored in the computer-readable storage medium, which is loaded by the processor and the encoding method of the second embodiment or each embodiment thereof is executed. In one embodiment, there is one or more processors and one or more memories. In one embodiment, the computer-readable storage medium may be integrated with the processor, or it may be installed separately from the processor.
[0010] In a seventh embodiment, the present application provides a computer-readable storage medium. The computer-readable storage medium is configured to store computer instructions, which, when read and executed by the processor of a computer device, cause the computer device to execute the decoding method according to the first embodiment or the encoding method according to the second embodiment.
[0011] In the eighth embodiment, the present application provides a bitstream, which is either the bitstream according to the first embodiment or the bitstream according to the second embodiment.
[0012] Based on the above proposed technology, the decoder predicts the current block based on an optimal MIP mode and a first intra-prediction mode. The optimal MIP mode is designed to be the optimal MIP mode for predicting the current block, determined based on the strain costs corresponding to multiple MIP modes. The first intra-prediction mode is designed to include at least one of the following: a suboptimal MIP mode, an intra-prediction mode derived from a DIMD mode, or an intra-prediction mode derived from a TIMD mode, determined based on the strain costs corresponding to multiple MIP modes, for predicting the current block. This helps to avoid the decoder obtaining the MIP mode by analyzing the bitstream. Compared to conventional MIP technology, this application can effectively reduce bit overhead at the coding unit level, thereby improving decompression efficiency.
[0013] Furthermore, by merging the optimal MIP mode with the first intra-prediction mode, it is possible to avoid completely replacing the optimal prediction mode calculated based on rate distortion cost with the optimal MIP mode, thereby achieving both prediction accuracy and prediction diversity, and further improving decompression performance. [Brief explanation of the drawing]
[0014] [Figure 1] Figure 1 is a block diagram showing an encoding framework according to an embodiment of this application. [Figure 2] Figure 2 is a schematic diagram showing the MIP mode according to an embodiment of this application. [Figure 3] Figure 3 is a schematic diagram showing how to derive a prediction mode based on DIMD according to the embodiment of this application. [Figure 4] Figure 4 is a schematic diagram showing the derivation of a prediction block based on DIMD according to the embodiment of this application. [Figure 5] Figure 5 is a schematic diagram showing a template used in TIMD according to the embodiment of this application. [Figure 6] Figure 6 is a block diagram showing a decoding framework according to an embodiment of this application. [Figure 7] Figure 7 is a flowchart showing the decoding method according to an embodiment of this application. [Figure 8] Figure 8 is a flowchart showing the encoding method according to the present invention. [Figure 9] Figure 9 is a block diagram showing a decoder according to an embodiment of this application. [Figure 10] Figure 10 is a block diagram showing an encoder according to an embodiment of this application. [Figure 11] Figure 11 is a block diagram showing an electronic device according to an embodiment of this application. [Modes for carrying out the invention]
[0015] The technical proposal in the embodiments of this application will be described below with reference to the drawings.
[0016] The solutions according to the embodiments of this application may be applied to the technical field of digital video coding, which includes, but is not limited to, the image coding field, the video coding field, the hardware video coding field, the dedicated circuit video coding field, and the real-time video coding field. Furthermore, the solutions according to the embodiments of this application can be combined with audio video coding standards (AVS), second-generation AVS standards (AVS2), or third-generation AVS standards (AVS3). Examples include, but are not limited to, the H.264 / audio video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. In addition, the solutions according to the embodiments of this application can be used to perform lossy compression or lossless compression on images. The lossless compression in question may be visually lossless compression or mathematically lossless compression.
[0017] A mixed coding framework based on blocks is used in video coding standards. Each image (frame) in video is divided into the largest coding unit (LCU) or coding tree unit (CTU) of the same size square (e.g., 128x128, 64x64, etc.). Each largest coding unit or coding tree unit can also be divided into rectangular coding units (CU) based on rules. Coding units can be further divided into prediction units (PU), transform units (TU), etc. The mixed coding framework includes modules such as prediction, transform, quantization, entropy coding, and loop filtering. The prediction module includes intra-prediction and inter-prediction. Inter-prediction includes motion estimation and motion compensation. Because there is a strong correlation between adjacent samples in video images, video coding techniques utilize intra-prediction methods to eliminate spatial redundancy between adjacent samples. In intra-prediction, sample information within the current segment block is predicted by referring only to image information from the same frame. Because there is a strong similarity between adjacent images in video, video coding techniques can improve coding efficiency by using inter-prediction methods to eliminate temporal redundancy between adjacent images. Inter-prediction can refer to image information from different frames and use motion estimation to find the motion vector information that best matches the current segment block. Transformation transforms the predicted image block into the frequency domain and redistributes the energy. Combining transformation and quantization can remove information that is not sensitive to the human eye and is used to eliminate visual redundancy.Entropy coding can be used to remove character redundancy based on the current context model and the probability information of the binary bit stream.
[0018] In the process of digital video encoding, the encoder first reads a black-and-white image or a color image from the original video sequence, and then can encode the black-and-white image or the color image. Here, the black-and-white image can include samples of the luma component, and the color image can include samples of the chroma component. Optionally, the color image can further include samples of the luma component. The color format of the original video sequence can be, for example, the luma-chroma (YCbCr, YUV) format or the red-green-blue (RGB) format. Specifically, after reading the black-and-white image or the color image, the encoder divides it into blocks, generates a predicted block for the current block by performing intra prediction or inter prediction on the current block, subtracts the predicted block from the original block of the current block to obtain a residual block, transforms and quantizes the residual block to obtain a quantized coefficient matrix, and entropy codes the quantized coefficient matrix and outputs it to the bit stream. In the process of digital video decoding, the decoding side generates a predicted block for the current block by performing intra prediction or inter prediction on the current block. Also, the decoding side analyzes the bit stream to obtain a quantized coefficient matrix, inverse quantizes and inverse transforms the quantized coefficient matrix to obtain a residual block, adds the predicted block and the residual block to obtain a reconstructed block. The reconstructed block forms a reconstructed image. The decoding side loop filters the reconstructed image based on the image or block to obtain a decoded image.
[0019] The current block can be, for example, the current coding unit (CU) or the current prediction unit (PU).
[0020] On the encoding side as well, in order to obtain the decoded image, processing similar to that on the decoding side is necessary. The decoded image can be a reference image for inter-prediction of subsequent images. Block partitioning information, prediction, transformation, quantization, entropy coding, loop filtering, and other mode information or parameter information determined on the encoding side are output to the bitstream as needed. The decoding side analyzes the existing information to determine the same block partitioning information, prediction, transformation, quantization, entropy coding, loop filtering, and other mode information or parameter information as on the encoding side. Thereby, it is ensured that the decoded image obtained on the encoding side is the same as the decoded image obtained on the decoding side. The decoded image obtained on the encoding side is usually also called the reconstructed image. During prediction, the current block may be divided into prediction units, and during transformation, the current block may be divided into transformation units. The divisions of the prediction units and the transformation units may be the same or different. The above is the basic flow of video coding in the block-based hybrid coding framework. With the development of technology, some steps in some modules or flows in the framework may be optimized. This application is applicable to the basic flow of the video codec in the block-based hybrid coding framework.
[0021] To facilitate understanding, first, the encoding framework according to this application will be briefly described.
[0022] FIG. 1 is a block diagram showing an encoding framework 100 according to an embodiment of this application.
[0023] As shown in Figure 1, the encoding framework 100 may include an intra-prediction unit 180, an inter-prediction unit 170, a residual unit 110, a transform and quantization unit 120, an entropy encoding unit 130, an inverse transform and inverse quantization unit 140, and a loop filtering unit 150. Optionally, the encoding framework 100 may further include a decoded image buffer unit 160. The encoding framework 100 is also called a mixed framework encoding mode.
[0024] The intra-prediction unit 180 or inter-prediction unit 170 can predict image blocks awaiting coding and output predicted blocks. The residual unit 110 can calculate residual blocks, i.e., the difference between predicted blocks and image blocks awaiting coding, based on the predicted blocks and image blocks awaiting coding. The transformation and quantization unit 120 is used to perform operations such as transformation and quantization on the residual blocks, thereby removing information that is not sensitive to the human eye and eliminating visual redundancy. Selectively, residual blocks before transformation and quantization by the transformation and quantization unit 120 may be called temporal residual blocks, and temporal residual blocks after transformation and quantization by the transformation and quantization unit 120 may be called frequency residual blocks or frequency-domain residual blocks. The entropy encoding unit 130 can receive the quantized transform coefficient output by the transformation and quantization unit 120 and then output a bitstream based on the quantized transform coefficient. For example, the entropy encoding unit 130 can remove character redundancy based on a target context model and probability information of the binary bitstream. For example, the entropy encoding unit 130 can be used for context-based adaptive binary arithmetic coding (CABAC). The entropy encoding unit 130 may also be called a header information encoding unit. Optionally, in this application, the image block awaiting coding may also be called an original image block or a target image block. A prediction block may also be called a prediction image block or an image prediction block, and may also be called a prediction signal or prediction information.A reconstructed block may also be called a reconstructed image block or image reconstruction block, and may also be called a reconstructed signal or reconstructed information. Furthermore, on the encoding side, the image block awaiting coding may also be called an encoding block or encoded image block. On the decoding side, the image block awaiting coding may also be called a decoding block or decoded image block. The image block awaiting coding may be a CTU or a CU.
[0025] In the encoding framework 100, the difference between the predicted block and the image block awaiting coding is calculated to obtain a residual block. Processes such as transformation and quantization are performed on the residual block, and the residual block is transmitted to the decoding side. Accordingly, the decoding side receives and analyzes the bitstream, obtains a residual block through steps such as inverse transformation and inverse quantization, and obtains a reconstructed block by adding the residual block to the predicted block obtained by the decoding side.
[0026] Furthermore, the inverse transform and inverse quantization unit 140, the loop filtering unit 150, and the decoding image buffer unit 160 within the encoding framework 100 can be used to form a decoder. The intra-prediction unit 180 or inter-prediction unit 170 can predict image blocks awaiting coding based on existing reconstruction blocks, thereby ensuring that the encoding side can use the reference frame in the same manner as the decoding side. In other words, the encoder can replicate the decoder's processing loop, thereby generating the same predictions as the decoding side. Specifically, the quantized transformation coefficients are inversely transformed and inversely quantized by the inverse transform and inverse quantization unit 140 to replicate the approximate residual block on the decoding side. After the prediction block is added to this approximate residual block, the loop filtering unit 150 can be used to smooth out the effects of block-based processing and blocking artifacts due to quantization. The image blocks output from the loop filtering unit 150 can be stored in the decoding image buffer unit 160 for use in subsequent image prediction.
[0027] Figure 1 is merely an example of this application and should not be interpreted as a limitation of this application.
[0028] For example, the loop filtering unit 150 within the encoding framework 100 may include a deblocking filter (DBF) and a sample adaptive offset (SAO). The role of the DBF is to remove deblocking artifacts, and the role of the SAO is to remove ringing effects. In other embodiments of this application, a neural-network-based loop filtering algorithm can be used in the encoding framework 100 to improve the video compression efficiency. Alternatively, the encoding framework 100 may be a deep learning neural network-based video coding hybrid framework. In one embodiment, the results after sample filtering can be calculated using a convolutional neural network-based model based on the deblocking filter and the SAO. The network structure in the luminance component and the network structure in the saturation component of the loop filtering unit 150 may be the same or different. Given that the luminance component contains more visual information, the luminance component can be used to guide the filtering of the saturation component in order to improve the reconstruction quality of the saturation component.
[0029] The following explains the details related to intranet prediction.
[0030] Intra prediction predicts sample information within an image block awaiting coding by referencing only the information of the same image to eliminate spatial redundancy. The image used for intra prediction may be an I-frame. For example, following the coding order from left to right and top to bottom, the top-left image block, the image block above, and the left image block can be used as reference information to predict the image block awaiting coding. The image block awaiting coding is then used as reference information for the next image block. In this way, the entire image can be predicted. If the input digital video is in a color format such as YUV4:2:0, each of the four pixels in each image frame of the digital video consists of four Y components and two UV components. The encoding framework can encode the Y component (i.e., luminance block) and the UV component (i.e., chromaticity block), respectively. Similarly, the decoding side can decode according to the format.
[0031] The following describes other intra-prediction modes related to this application.
[0032] (1) Matrix-based Intra Prediction (MIP) mode
[0033] The MIP model process can be divided into three main steps: the downsampling process, the matrix multiplication process, and the upsampling process. Specifically, first, in the downsampling process, spatially adjacent reconstructed samples are downsampled. Next, the obtained downsampled sample sequence is used as the input vector for the matrix multiplication process, that is, the output vector of the downsampling process is used as the input vector for the matrix multiplication process, the input vector of the matrix multiplication process is multiplied by a pre-set matrix, the result is added to a bias vector, and the calculated sample vector is output. Finally, the output vector of the matrix multiplication process is used as the input vector for the upsampling process, and the final prediction block is obtained by upsampling.
[0034] Figure 2 is a schematic diagram showing the MIP mode according to an embodiment of this application.
[0035] JPEG0007868170000001.jpg69151
[0036] In other words, to predict a block with width W and height H, MIP requires H reconstructed samples from the left column of the current block and W reconstructed samples from the row above the current block as input. MIP generates the predicted block based on three steps: averaging of reference samples, matrix-vector multiplication, and interpolation. The core of MIP is matrix-vector multiplication, which can be considered the process of generating a predicted block using input samples (reference samples) in a matrix-vector multiplication manner. Various matrices are provided for MIP, and differences in prediction methods can be reflected in differences in matrices, resulting in different results when different matrices are used for the same input samples. Furthermore, the processes of averaging and interpolation of reference samples are designed to consider the trade-off between performance and complexity. For blocks with large sizes, averaging of reference samples can achieve an effect that approximates downsampling, allowing the input to be adapted to a relatively small matrix. Interpolation provides an upsampling effect. Thus, it is no longer necessary to provide a MIP matrix for each block size; instead, only one or a few matrices of specific sizes need to be provided. As the need for compression performance increases and hardware performance improves, more complex MIPs may appear in next-generation standards.
[0037] In MIP mode, the MIP mode can be simplified from a neural network; for example, the matrix used in the MIP mode can be obtained based on training. Therefore, the MIP mode has relatively strong generalization ability and predictive effects that cannot be achieved by conventional predictive modes. The MIP mode is a model obtained by performing multiple simplifications of hardware and software complexity on an intra-predictive model based on a neural network. Based on a large number of training samples, multiple predictive modes can represent multiple models and parameters, better covering the texture of natural sequences.
[0038] MIP mode is somewhat similar to planar mode, but clearly, MIP mode is more complex and flexible than planar mode.
[0039] The number of MIP modes can vary depending on the coding unit's block size. For example, a 4x4 coding unit has 16 prediction modes. An 8x8 coding unit, or a coding unit with a width equal to 4 or a height equal to 4, has 8 prediction modes. For other sizes of coding units, it has 6 prediction modes. MIP modes also have a transpose function. For prediction modes that match the current size, MIP modes allow the encoder to attempt a transpose calculation. Therefore, MIP modes require a flag indicating whether or not MIP modes are used for the current coding unit. If MIP modes are used for the current coding unit, a transpose flag must also be transmitted to the decoder.
[0040] The MIP inverted flag is binarized using a fixed-length (FL) coding scheme, with a length of 1. The MIP mode index is binarized using a truncated binary (TB) coding scheme.
[0041] (2) Decoder-side Intra-Mode Derivation (DIMD) Prediction Mode
[0042] The core of the DIMD prediction mode is that the decoder derives the intra-prediction mode using the same method as the encoder. This avoids transmitting the intra-prediction mode index of the current coding unit in the bitstream, thereby saving bit overhead.
[0043] DIMD The specific process of the prediction mode can be divided into the following two main steps.
[0044] Step 1: Derive the prediction mode.
[0045] The encoding side uses the Sobel operator to calculate the histogram of gradients for each prediction mode. The domain of operation is the three adjacent rows of reconstructed samples above the current block, the three adjacent columns of reconstructed samples to the left, and the corresponding adjacent reconstructed sample in the upper left. By calculating the gradient histogram within the aforementioned L-shaped domain, the prediction mode corresponding to the largest amplitude (also called magnitude) in the gradient histogram can be obtained, as well as the prediction mode corresponding to the second largest amplitude.
[0046] Figure 3 shows an embodiment of the present application. DIMD This is a schematic diagram showing how to derive the prediction mode based on [the given formula].
[0047] As shown in Figure 3(a), DIMD derives a prediction mode using samples in a template within the reconstruction region (reconstruction samples to the left and above the current block). For example, the template can include three adjacent rows of reconstruction samples above the current block, three adjacent columns of reconstruction samples to the left, and the corresponding adjacent reconstruction sample to the upper left. Based on this, the window (for example, in Figure 3) (b) Or Figure 3 (c) Multiple gradient values are determined within the template according to the window shown. Each gradient value is used to obtain one intra prediction mode (ipm) that fits the direction of the gradient. Based on this, the encoder can derive the prediction modes corresponding to the largest gradient value and the second largest gradient value among the multiple gradient values. For example, as shown in Figure 3(b), for a 4x4 block, all samples for which gradient values need to be determined are analyzed to obtain the corresponding gradient histogram. For example, as shown in Figure 3(c), for a block of other sizes, all samples for which gradient values need to be determined are analyzed to obtain the corresponding gradient histogram. Finally, the prediction modes corresponding to the largest gradient and the second largest gradient in the gradient histogram are derived as the prediction modes.
[0048] Of course, the gradient histogram in this application is merely an example for determining the derived prediction mode, and when actually implemented, it can be implemented in various simple forms, and this application is not particularly limited thereto. Furthermore, this application is not limited to the method of obtaining the gradient histogram, and for example, it can be obtained using the Sobel operator or other methods.
[0049] Step 2: Derive the predicted block.
[0050] Figure 4 is a schematic diagram showing the derivation of a prediction block based on DIMD according to the embodiment of this application.
[0051] As shown in Figure 4, the encoder can weight the predicted values corresponding to three intra-prediction modes (planar mode and two intra-prediction modes derived based on DIMD). The codec obtains the predicted block for the current block using the same prediction block derivation scheme. Assuming that the prediction mode corresponding to the largest gradient value is prediction mode 1 and the prediction mode corresponding to the second largest gradient value is prediction mode 2, the encoder determines the following two conditions: 1. The gradient value for prediction mode 2 is not 0. 2. Neither Prediction Mode 1 nor Prediction Mode 2 are planar mode or direct current (DC) prediction modes.
[0052] If the above two conditions are not met simultaneously, the predicted sample value for the current block is calculated using only prediction mode 1, i.e., the normal prediction process is applied to prediction mode 1. If not, i.e., if the above two conditions are met simultaneously, the predicted block for the current block is derived using a weighted averaging method. The specific method is as follows: Planar mode accounts for 1 / 3 of the weight, and the remaining 2 / 3 is the combined weight of prediction mode 1 and prediction mode 2. For example, the gradient amplitude value of prediction mode 1 is divided by the sum of the gradient amplitude values of prediction mode 1 and prediction mode 2, and the result is used as the weight for prediction mode 1. The gradient amplitude value of prediction mode 2 is divided by the sum of the gradient amplitude values of prediction mode 1 and prediction mode 2, and the result is used as the weight for prediction mode 2. Finally, a weighted averaging is performed on the predicted blocks obtained based on the above three prediction modes, i.e., predicted block 1, predicted block 2, and predicted block 3 obtained based on planar mode, prediction mode 1, and prediction mode 2, respectively, to obtain the predicted block for the current coding unit. The decoder also obtains its predicted block in the same steps.
[0053] In other words, the specific weights for step 2 above are calculated as follows: Weight(PLANAR) = 1 / 3, Weight(mode1) = 2 / 3 * (amp1 / (amp1+amp2)), Weight(mode2) = 1 - Weight(PLANAR) - Weight(mode1). mode1 and mode2 represent prediction mode 1 and prediction mode 2, respectively, and amp1 and amp2 represent the gradient amplitude values for prediction mode 1 and prediction mode 2, respectively. In DIMD prediction mode, a flag needs to be transmitted to the decoder. This flag is transmitted to the current coder unit. DIMD This is used to indicate whether or not the prediction mode is being used.
[0054] Of course, the weighted average method described above is merely one example of this application and should not be understood as a limitation of this application.
[0055] In summary, DIMD uses gradient analysis of the reconstructed sample to select an intra-prediction mode, and depending on the analysis results, it is possible to weight two intra-prediction modes and a planar mode. The advantage of DIMD is that, once a DIMD mode is selected for the current block, there is no need to specifically indicate which intra-prediction mode is being used in the bitstream, as this is derived by the decoder itself through the process described above, thus saving some overhead.
[0056] (3) Template-based Intra Mode Derivation (TIMD) Prediction Mode
[0057] In TIMD prediction mode, the codec saves the overhead of transmitting the mode index by using the same operation to derive the prediction mode. TIMD prediction mode can be understood in two main parts. First, cost information for each prediction mode is calculated based on a template, and the prediction mode corresponding to the smallest cost and the prediction mode corresponding to the second smallest cost are selected. The prediction mode corresponding to the smallest cost is denoted as prediction mode 1, and the prediction mode corresponding to the second smallest cost is denoted as prediction mode 2. If the ratio of the second smallest cost (costMode2) to the smallest cost (costMode1) satisfies a pre-set condition such as costMode2 < 2*costMode1, then the prediction block corresponding to prediction mode 1 and the prediction block corresponding to prediction mode 2 are weighted fused according to the weights corresponding to prediction mode 1 and prediction mode 2 to obtain the final prediction block.
[0058] For example, the weights corresponding to prediction mode 1 and prediction mode 2 are determined based on the following method. weight1 = costMode2 / (costMode1+ costMode2), weight2 = 1 - weight1. weight1 is the weight of the prediction block corresponding to prediction mode 1, and weight2 is the weight of the prediction block corresponding to prediction mode 2. If the ratio of the second smallest cost costMode2 to the smallest cost costMode1 does not satisfy a predetermined condition, weight fusion between prediction blocks does not occur, and the prediction block corresponding to prediction mode 1 becomes a TIMD prediction block.
[0059] Furthermore, when performing intraprediction on the current block using TIMD prediction mode, if the reconstruction sample template for the current block does not contain any available adjacent reconstruction samples, TIMD prediction mode will select planar mode and perform intraprediction on the current block, i.e., will not perform weighted fusion. Similar to DIMD prediction mode, TIMD prediction mode requires the transmission of a flag to the decoder. This flag is used to indicate whether or not TIMD prediction mode is being used for the current coder unit.
[0060] The encoder or decoder primarily calculates cost information for each prediction mode as follows: Intra-mode prediction is performed on samples within the template region based on reconstructed samples adjacent to the upper and left sides of the template region, and the prediction process is the same as for the original intra-prediction mode. For example, when performing intra-mode prediction on samples within the template region using DC mode, the average value of the entire coding unit is calculated. As another example, when performing intra-mode prediction on samples within the template region using angle prediction mode, a corresponding interpolation filter is selected according to the mode, and predicted samples are obtained by interpolation according to the rules. In this case, based on the predicted samples and reconstructed samples within the template region, the distortion between the predicted samples and reconstructed samples within that region, i.e., the cost information of the current prediction mode, can be calculated.
[0061] Figure 5 is a schematic diagram showing a template used in TIMD according to the embodiment of this application.
[0062] As shown in Figure 5, if the current block is a coding unit with width equal to M and height equal to N, the codec can predict the samples in the template region of the current block by selecting reconstructed samples from the current block's reference template from coding units with width equal to 2(M+L1)+1 and height equal to 2(N+L2)+1. Of course, if the template region of the current block does not contain any available adjacent reconstructed samples, the TIMD prediction mode selects planar mode and performs intra-prediction for the current block. For example, available adjacent reconstructed samples may be samples adjacent to the left and above the current CU in Figure 5, i.e., there are no available reconstructed samples in the shaded area. In other words, if there are no available reconstructed samples in the shaded area, the TIMD prediction mode selects planar mode and performs intra-prediction for the current block.
[0063] Except in the case of boundaries, when encoding the current block, theoretically, reconstructed values can be obtained from the left and above the current block; that is, the template of the current block contains available adjacent reconstructed samples. In a specific embodiment, the decoder predicts the template using a certain intra-prediction mode and compares the predicted value with the reconstructed value to obtain the cost of that intra-prediction mode on the template. Examples include the Sum of Absolute Differences (SAD), the Sum of Absolute Transformed Difference (SATD), and the Sum of Squared Error (SSE). Since the template and the current block are adjacent to each other, the reconstructed samples in the template and the samples in the current block are correlated. Therefore, the behavior of the prediction mode on the template can be used to estimate the behavior of this prediction mode on the current block. TIMD predicts the template using several candidate intra-prediction modes to obtain the costs of the candidate intra-prediction modes on the template, and the predicted value of the one or two intra-prediction modes with the lowest cost is taken as the intra-prediction value of the current block. If the difference between the two costs corresponding to the two intra-prediction modes on the template is not large, the compression performance can be improved by performing a weighted average on the predicted values of the two intra-prediction modes. Selectively, the weights of the predicted values of the two prediction modes are related to the above costs, for example, the weights are inversely proportional to the costs.
[0064] In summary, TIMD allows for the selection of an intra-prediction mode by leveraging the predictive effect of the intra-prediction mode on the template, and weighting two intra-prediction modes according to their cost on the template. The advantage of TIMD is that, once a TIMD mode is selected for the current block, there is no need to specifically indicate which intra-prediction mode is being used in the bitstream, as this is derived by the decoder itself through the process described above, thus saving some overhead.
[0065] Through a brief introduction to the above intra-prediction modes, the following can be observed: The technical principles of DIMD and TIMD prediction modes are similar. Both utilize the fact that the decoder performs the same operation as the encoder to estimate the prediction mode of the current coding unit. In such prediction modes, if complexity is acceptable, the transmission of the prediction mode index can be omitted, saving overhead and improving compression efficiency. However, due to the limitations of the available information and the fact that they do not significantly improve prediction quality, DIMD and TIMD prediction modes are more effective in large areas where texture characteristics are consistent. When the texture changes slightly or the template area cannot be covered, the prediction effect of such prediction modes is inferior.
[0066] Furthermore, both DIMD and TIMD prediction modes merge or weight prediction blocks obtained based on multiple conventional prediction modes. Merging prediction blocks can produce effects that cannot be achieved with a single prediction mode. In DIMD prediction mode, a planar mode can be introduced as an additional weighted prediction mode to enhance the spatial relevance between adjacent reconstructed samples and predicted samples, thereby improving the prediction effect of intra-prediction. However, because the prediction principle of the planar mode is relatively simple, using the planar mode as an additional weighted prediction mode for prediction blocks where there is a clear difference between the upper right corner and the lower left corner may have a counterproductive effect.
[0067] Figure 6 is a block diagram showing a decoding framework 200 according to an embodiment of this application.
[0068] As shown in Figure 6, the decoding framework 200 may include an entropy decoding unit 210, an inverse transform / inverse quantization unit 220, a residual unit 230, an intra prediction unit 240, an inter prediction unit 250, a loop filtering unit 260, and a decoding image buffer unit 270.
[0069] The entropy decoding unit 210 receives and analyzes the bitstream to obtain a prediction block and a frequency-domain residual block. The frequency-domain residual block can be subjected to steps such as inverse transformation and inverse quantization through the inverse transformation / inverse quantization unit 220 to obtain a time-domain residual block. The residual unit 230 can obtain a reconstructed block by adding the prediction block obtained by the intra-prediction unit 240 or the inter-prediction unit 250 to the time-domain residual block obtained after inverse transformation and inverse quantization through the inverse transformation / inverse quantization unit 220. For example, the intra-prediction unit 240 or the inter-prediction unit 250 can obtain a prediction block by decoding the header information of the bitstream.
[0070] Figure 7 is a flowchart illustrating a decoding method 300 according to an embodiment of this application. This decoding method 300 can be executed by a decoder. For example, this decoding method 300 is applied to the decoding framework 200 shown in Figure 6. For the sake of clarity, a decoder will be described below as an example.
[0071] As shown in Figure 7, the decoding method 300 may include some or all of the following: S310: Analyze the bitstream to obtain the residual block of the current block in the current sequence. S320: Determine the optimal MIP mode for predicting the current block based on the strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes. The strain costs corresponding to multiple MIP modes include the strain costs obtained by predicting samples in template regions adjacent to the current block using multiple MIP modes. S330: Determine the first intra-prediction mode. The first intra-prediction mode is determined based on the strain costs corresponding to multiple MIP modes and includes at least one of the following: a suboptimal MIP mode for predicting the current block; an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode; or an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode. S340: Predict the current block based on the optimal MIP mode and the first intra prediction mode to obtain the predicted block for the current block. S350: Based on the residual block of the current block and the predicted block of the current block, the reconstructed block of the current block is obtained.
[0072] Exemplary, in this application, S320 to S340 (i.e., the process by which the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes, and predicts the current block based on the optimal MIP mode and the first intra-prediction mode) may also be abbreviated as Template Matching MIP (TMMIP) technology, TMMIP-based prediction mode derivation method, or TMMIP fusion enhancement technology. That is, after obtaining the residual block of the current block, the decoder can enhance the performance of the prediction process for the current block based on the derived optimal MIP mode and the first intra-prediction mode. In other words, TMMIP technology can enhance the performance of the prediction process for the current block by utilizing at least one of a suboptimal MIP prediction mode, an intra-prediction mode derived from a TIMD mode, or an intra-prediction mode derived from a DIMD mode, along with the optimal MIP prediction mode.
[0073] In this embodiment, the decoder predicts the current block based on an optimal MIP mode and a first intra-prediction mode. The optimal MIP mode is designed to be the optimal MIP mode for predicting the current block, determined based on the strain costs corresponding to multiple MIP modes. The first intra-prediction mode is designed to include at least one of a suboptimal MIP mode, an intra-prediction mode derived from a DIMD mode, or an intra-prediction mode derived from a TIMD mode, determined based on the strain costs corresponding to multiple MIP modes, for predicting the current block. This helps to avoid the decoder obtaining the MIP mode by analyzing the bitstream. Compared to conventional MIP techniques, this application can effectively reduce bit overhead at the coding unit level, thereby improving decompression efficiency.
[0074] In other words, conventional MIP technology has more bit overhead than other intra-prediction tools, requires a flag to indicate whether MIP technology is used, as well as a flag to indicate whether MIP is transposed, and finally, the biggest overhead is that truncated binary coding must be used to represent the MIP prediction mode. MIP technology is a simplified technique based on neural network technology and differs significantly from conventional interpolation filtering prediction techniques. For some specific textures, the MIP prediction mode is more effective than conventional intra-prediction modes, but its large flag overhead is a drawback of MIP technology. For example, a 4x4 coding unit has a total of 16 prediction samples, but its bit overhead includes one MIP use flag, one MIP transpose flag, and a 5-bit or 6-bit truncated binary flag. In light of this, this application utilizes a method in which the decoder autonomously determines the optimal MIP mode for predicting the current block and determines the intra-prediction mode of the current block based on the optimal MIP mode, thereby saving up to 5 or 6 bits of overhead, effectively reducing the bit overhead at the coding unit level, and thereby improving decompression efficiency.
[0075] Furthermore, saving up to 5 or 6 bits of overhead per block presupposes that the template matching-based prediction mode derivation algorithm is extremely accurate. If the MIP prediction mode calculated by the fast template matching-based algorithm differs from the MIP prediction mode calculated by the encoding side using rate-distortion optimized and original sequences, it may not be the optimal choice. Therefore, the selection of the template matching-based MIP prediction mode depends on the degree of matching accuracy, i.e., the precision. However, both the template-based derivation algorithms in conventional intra-prediction modes and the template matching-based derivation algorithms in inter-prediction modes fall short of expectations in terms of accuracy, and although performance has improved considerably, the large additional bit overhead at the block level introduced by the template matching-based algorithm prevents subsequent technologies from achieving higher performance by relying solely on the template matching-based algorithm. In light of this, this application combines the optimal MIP mode with a first intra-prediction mode, that is, performs fusion prediction on the current block based on the optimal MIP mode and the first intra-prediction mode. This avoids completely replacing the optimal prediction mode calculated based on rate distortion cost with the optimal MIP mode, thereby achieving both prediction accuracy and prediction diversity, and further improving decompression performance.
[0076] Furthermore, the distortion cost related to the decoder in this application is different from the rate-distortion cost (RDcost) related to the encoder. The rate-distortion cost is the distortion cost used when the encoding side determines a specific intra-prediction technique from among multiple intra-prediction techniques, and the rate-distortion cost can be the cost obtained by comparing the distorted image with the original image. Since the decoder cannot obtain the original image, the distortion cost related to the decoder can be the distortion cost between the reconstructed sample and the predicted sample, and examples include the SATD cost between the reconstructed sample and the predicted sample, or a cost that can be used to calculate the difference between the reconstructed sample and the predicted sample.
[0077] Referring to the test results in Tables 1 and 2, the beneficial effects of the scheme relating to this application are described below. Table 1 shows the results obtained by testing the scheme of this application when the first intra-prediction mode is designed to be a suboptimal MIP mode for predicting the current block, which is determined based on the strain costs corresponding to multiple MIP modes. Table 2 shows the results obtained by testing the scheme of this application when the first intra-prediction mode is designed to be an intra-prediction mode derived from the TIMD mode.
[0078] [Table 1]
[0079] [Table 2]
[0080] JPEG0007868170000004.jpg82152
[0081] In some embodiments, S320 may include: analyzing the bitstream of the current sequence and obtaining a first flag; determining the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes, if the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and a first intra-prediction mode; and if the optimal MIP mode is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and a first intra-prediction mode.
[0082] For example, if the value of the first flag is a first value, the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode. If the value of the first flag is a second value, the first flag is used to indicate that it is not permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode. In one embodiment, the first value is 1 and the second value is 0. In another embodiment, the first value is 0 and the second value is 1. Of course, the first and second values may be other values, and are not limited thereto in this application.
[0083] For example, if the first flag is true, it is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode. If the first flag is false, it is used to indicate that it is not permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode.
[0084] For example, the decoder analyzes block-level flags. If the intra-predictive mode is used for the current block, the decoder analyzes or retrieves a first flag. If the first flag is true, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0085] For example, if the first flag is denoted as sps_timd_enable_flag, the decoder parses or retrieves sps_timd_enable_flag. If sps_timd_enable_flag is true, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0086] For example, the first flag is a sequence-level flag.
[0087] The statement that the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode can be replaced with a statement having a similar or identical meaning. For example, in other alternative embodiments, the statement that the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode can be replaced with any of the following: The first flag is used to indicate that it is permitted to determine the intra-prediction mode of image blocks in the current sequence using TMMIP technology. The first flag is used to indicate that it is permitted to perform intra-prediction for image blocks in the current sequence using TMMIP technology. The first flag is used to indicate that it is permitted to use TMMIP technology for image blocks in the current sequence. The first flag is used to indicate that it is permitted to predict image blocks in the current sequence using an MIP mode determined based on multiple MIP modes.
[0088] Furthermore, in other alternative embodiments, when TMMIP technology is combined with other technologies, an enable flag for the other technology can indirectly indicate whether or not the use of TMMIP technology is permitted in the current sequence. For example, taking TIMD technology as an example, if the first flag is used to indicate that the use of TIMD technology is permitted in the current sequence, it also indicates that the use of TMMIP technology is permitted in the current sequence. In other words, if the first flag is used to indicate that the use of TIMD technology is permitted in the current sequence, it indicates that both TIMD technology and TMMIP technology are permitted in the current sequence. This further saves bit overhead.
[0089] In some embodiments, if the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode, the bitstream is parsed to obtain a second flag. If the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode, the optimal MIP mode is determined based on the strain costs corresponding to multiple MIP modes.
[0090] For example, the decoder analyzes block-level flags. If the current block uses an intra-predictive mode, the decoder analyzes or retrieves a first flag. If the first flag is true, the decoder analyzes or retrieves a second flag. If the second flag is true, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0091] For example, if the value of the second flag is the third value, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode. If the value of the second flag is the fourth value, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode. In one embodiment, the third value is 1 and the fourth value is 0. In another embodiment, the third value is 0 and the fourth value is 1. Of course, the third and fourth values may be other values, and are not limited thereto in this application.
[0092] For example, if the second flag is true, it is used to indicate that predicting the current block using the optimal MIP mode and the first intra-prediction mode is permitted. If the second flag is false, it is used to indicate that predicting the current block using the optimal MIP mode and the first intra-prediction mode is not permitted.
[0093] For example, the first flag is denoted as sps_timd_enable_flag and the second flag as cu_timd_enable_flag, in which case the decoder parses or retrieves sps_timd_enable_flag. If sps_timd_enable_flag is true, the decoder parses or retrieves cu_timd_enable_flag. If cu_timd_enable_flag is true, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0094] For example, the second flag can be a block-level flag or a coding unit-level flag.
[0095] The statement that the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode can be replaced with a statement having a similar or identical meaning. For example, in other alternative embodiments, the statement that the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode can be replaced with any of the following: The second flag is used to indicate that it is permitted to determine the intra-prediction mode of the current block using TMMIP technology. The second flag is used to indicate that it is permitted to perform intra-prediction for the current block using TMMIP technology. The second flag is used to indicate that it is permitted to use TMMIP technology on the image block of the current block. The second flag is used to indicate that it is permitted to predict the current block using an MIP mode determined based on multiple MIP modes.
[0096] Furthermore, in other alternative embodiments, when TMMIP technology is combined with other technologies, the permission flag for the other technology can indirectly indicate whether or not the use of TMMIP technology is permitted in the current block. For example, taking TIMD technology as an example, if the second flag is used to indicate that the use of TIMD technology is permitted in the current block, it also indicates that the use of TMMIP technology is permitted in the current block. In other words, if the second flag is used to indicate that the use of TIMD technology is permitted in the current block, it indicates that both TIMD technology and TMMIP technology are permitted in the current block. This further saves bit overhead.
[0097] Furthermore, when the decoding side analyzes the second flag, it can analyze the second flag before analyzing the residual block of the current block, or it can analyze the second flag after analyzing the residual block of the current block. This application does not particularly limit this.
[0098] In some embodiments, method 300 may further include: determining the sequence order of multiple MIP modes based on the distortion costs corresponding to the multiple MIP modes; determining the coding scheme to be used for the optimal MIP mode based on the sequence order of the multiple MIP modes; and decoding the bitstream of the current sequence to obtain the index of the optimal MIP mode based on the coding scheme to be used for the optimal MIP mode.
[0099] For example, before determining the optimal MIP mode for predicting the current block based on the strain costs of multiple MIP modes, the decoder needs to calculate the strain cost for each of the multiple MIP modes and sort the multiple MIP modes based on the strain cost for each MIP mode. The MIP mode with the lowest cost is the optimal prediction.
[0100] In conventional MIP technology, the index of the MIP mode is typically binarized using a truncated binary method similar to equal probability coding. That is, in the truncated binary method, all prediction modes are divided into two segments, one segment represented by N codewords and the other segment represented by N+1 codewords. In view of this, in this application, the decoder can calculate the strain cost corresponding to each of the multiple MIP modes and sort the multiple MIP modes based on the strain cost corresponding to each MIP mode before determining the optimal MIP mode for predicting the current block based on the strain costs corresponding to the multiple MIP modes. Finally, the decoder selects and uses a more flexible variable-length coding method according to the arrangement order of the multiple MIP modes. Compared to equal probability coding, the bit overhead of the MIP mode index can be saved by flexibly setting the coding method for the MIP modes.
[0101] In some embodiments, the codeword length of the coding scheme used for the first n MIP modes in the sequence order is smaller than the codeword length of the coding scheme used for the MIP mode following the nth MIP mode in the sequence order, and / or a variable-length coding scheme is used for the first n MIP modes and a truncated binary coding scheme is used for the MIP mode following the nth MIP mode.
[0102] For example, N can be any value greater than or equal to 1.
[0103] Exemplary, the sequence order is the order obtained by the decoder arranging multiple MIP modes in ascending order of distortion cost. The codeword length of the coding scheme used for the first n MIP modes in the sequence order is smaller than the codeword length of the coding scheme used for the MIP mode following the nth MIP mode in the sequence order, and / or a variable-length coding scheme is used for the first n MIP modes and a truncated binary coding scheme is used for the MIP mode following the nth MIP mode.
[0104] In this embodiment, the smaller the distortion cost corresponding to a MIP mode, the higher the probability that the encoder will use that MIP mode to perform an intra prediction for the current block. Therefore, the coding scheme used for the first n MIP modes in the sequence order is designed to have a smaller codeword length than the coding scheme used for the MIP mode following the nth MIP mode in the sequence order, and / or the coding scheme used for the first n MIP modes is designed to be a variable-length coding scheme, and the coding scheme used for the MIP mode following the nth MIP mode is designed to be a truncated binary coding scheme. In this way, MIP modes that are used with a high probability by the encoder use relatively short codeword lengths or variable-length coding schemes. This saves bit overhead in the MIP mode index and improves decompression performance.
[0105] In some embodiments, before S330, method 300 may further include the following: If the first intra-prediction mode is a suboptimal MIP mode, the decoder determines whether to adopt the suboptimal MIP mode to predict the current block based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the suboptimal MIP mode. If it is determined not to adopt the suboptimal MIP mode, the decoder can directly predict the current block based on the optimal MIP mode. If it is determined to adopt the suboptimal MIP mode, the decoder can predict the current block based on the optimal MIP mode and the suboptimal MIP mode to obtain a predicted block for the current block.
[0106] For example, when the first intra-prediction mode is a suboptimal MIP mode, if the ratio of the strain cost corresponding to the optimal MIP mode to the strain cost corresponding to the suboptimal MIP mode is less than or equal to a preset ratio, the decoder can directly predict the current block based on the optimal MIP mode and obtain a predicted block for the current block. Alternatively, when the first intra-prediction mode is a suboptimal MIP mode, if the ratio of the strain cost corresponding to the suboptimal MIP mode to the strain cost corresponding to the optimal MIP mode is greater than or equal to a preset ratio, the decoder can directly predict the current block based on the optimal MIP mode and obtain a predicted block for the current block. For example, if the strain cost corresponding to the suboptimal MIP mode is a multiple (e.g., 2 times) or more of the strain cost corresponding to the optimal MIP mode, it can be interpreted that the suboptimal MIP mode already has a large strain and is unsuitable for the current block, i.e., it is possible to predict the current block using only the optimal MIP mode without fusion enhancement techniques.
[0107] In this embodiment, the decoder determining whether to adopt a suboptimal MIP mode to predict the current block based on the distortion cost of the optimal MIP mode and the distortion cost of the suboptimal MIP mode is equivalent to the decoder determining whether to adopt a suboptimal MIP mode to improve the performance of the optimal MIP mode based on the distortion cost of the optimal MIP mode and the distortion cost of the suboptimal MIP mode. This avoids carrying a flag in the bitstream to determine whether to adopt a suboptimal MIP mode to improve the performance of the optimal MIP mode, thereby saving bit overhead and improving decompression performance.
[0108] In some embodiments, S340 may include the following: The decoder predicts the current block based on the optimal MIP mode to obtain a first predicted block. The decoder also predicts the current block based on the first intra-prediction mode to obtain a second predicted block. Next, the decoder weights the first and second predicted blocks based on the weights of the optimal MIP mode and the first intra-prediction mode to obtain a predicted block for the current block.
[0109] For example, the decoder can directly perform an intra-prediction on the current block based on the optimal MIP mode to obtain a first predicted block. Furthermore, the decoder can directly obtain an optimal and a suboptimal predicted mode based on the TIMD mode, predict the current block, and obtain a second predicted block. For example, if neither the optimal nor the suboptimal predicted mode is a DC mode (also known as the mean mode) or a planar mode (also known as the planar mode), and the strain cost corresponding to the suboptimal predicted mode is less than twice the strain cost corresponding to the optimal predicted mode, then the predicted blocks need to be merged. That is, the decoder first performs an intra-prediction on the current block based on the optimal predicted mode to obtain an optimal predicted block. Next, the decoder performs an intra-prediction on the current block based on the suboptimal predicted mode to obtain a suboptimal predicted block. Furthermore, the decoder uses the ratio of the strain cost corresponding to the optimal predicted mode to the strain cost corresponding to the suboptimal predicted mode to calculate the weight values of the optimal predicted block and the suboptimal predicted block. Finally, the decoder weight-merges the optimal and suboptimal predicted blocks to obtain a second predicted block. In another example, if the optimal or suboptimal prediction mode is a planar mode or a DC mode, or if the strain cost corresponding to the suboptimal prediction mode is greater than twice the strain cost corresponding to the optimal prediction mode, it is not necessary to merge the prediction blocks. That is, only the optimal prediction block obtained based on the optimal prediction mode can be directly used as the second prediction block. After obtaining the first and second prediction blocks, the decoder performs a weighting process on the first and second prediction blocks to obtain the prediction block for the current block.
[0110] In some embodiments, S340 may include the following: If the first intra-prediction mode includes an intra-prediction mode derived from a suboptimal MIP mode or TIMD mode, the decoder determines the weights of the optimal MIP mode and the first intra-prediction mode based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra-prediction mode includes an intra-prediction mode derived from a DIMD mode, the decoder determines both the weights of the optimal MIP mode and the weights of the first intra-prediction mode to preset values.
[0111] In some embodiments, S320 may include the following: The decoder predicts samples in the template region based on a third flag and multiple MIP modes and obtains strain costs corresponding to multiple MIP modes under each state of the third flag. The third flag is used to indicate whether or not to transpose the input and output vectors corresponding to the MIP mode. The decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes under each state of the third flag.
[0112] As mentioned above, conventional MIP technology has a higher bit overhead compared to other intra-prediction tools, requiring not only a flag to indicate whether MIP technology is used, but also a flag to indicate whether MIP is transposed, and finally, the largest overhead is that truncated binary coding must be used to represent the MIP prediction mode. MIP technology is a simplified technology based on neural network technology and differs significantly from conventional interpolation filtering prediction technology. For some special textures, the MIP prediction mode is more effective than conventional intra-prediction modes, but its large flag overhead is a drawback of MIP technology. For example, a 4x4 coding unit has a total of 16 prediction samples, and its bit overhead includes one MIP use flag, one MIP transpose flag, and a 5-bit or 6-bit truncated binary flag. In light of this, this application proposes that when determining the optimal MIP mode, traversing each state of a third flag can take into account the transpose function of the MIP mode, saving the overhead of one MIP transpose flag and further improving decompression efficiency.
[0113] For example, the decoder traverses each state of the third flag and multiple MIP modes, determines the distortion cost corresponding to each of the multiple MIP modes under each state of the third flag, and determines the optimal MIP mode based on the distortion costs corresponding to each of the multiple MIP modes under each state of the third flag. Alternatively, the decoder traverses each state of the third flag and multiple MIP modes, determines the distortion cost under each state of the third flag corresponding to multiple MIP modes, and determines the optimal MIP mode based on the distortion costs under each state of the third flag corresponding to multiple MIP modes. In other words, the decoding side may traverse the multiple MIP modes first, or it may traverse the states of the third flag first.
[0114] For example, when the value of the third flag is the fifth value, the third flag is used to indicate that the input and output vectors corresponding to the MIP mode are transposed. When the value of the third flag is the sixth value, the third flag is used to indicate that the input and output vectors corresponding to the MIP mode are not transposed. In this case, each state of the third flag can be replaced by each value of the third flag. In one embodiment, the fifth value is 1 and the sixth value is 0. In another embodiment, the fifth value is 0 and the sixth value is 1. Of course, the fifth and sixth values may be other values, and this application is not limited thereto.
[0115] For example, when the third flag is true, it is used to indicate that the input and output vectors corresponding to the MIP mode are transposed. When the third flag is false, it is used to indicate that the input and output vectors corresponding to the MIP mode are not transposed. In this case, both true and false of the third flag are states of the third flag.
[0116] For example, the third flag may be a sequence-level flag, a block-level flag, or a coding unit-level flag.
[0117] For example, the third flag may also be called the inverted message, inverted flag, or MIP inverted flag.
[0118] The statement that the third flag is used to indicate whether or not to transpose the input and output vectors corresponding to the MIP mode can be replaced with a similar or identical statement. For example, in other alternative embodiments, the third flag is used to indicate whether or not it is necessary to transpose the input and output corresponding to the MIP mode. The third flag is used to indicate whether or not the input and output vectors corresponding to the MIP mode are transposed vectors. The third flag is used to indicate whether or not to transpose.
[0119] In some embodiments, S320 may include the following: If the current block size is a preset size, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0120] For example, a predefined size may include a size where the width is a predefined width and the height is a predefined height. In other words, if the current block's width is a predefined width and the height is a predefined height, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0121] Exemplary, a pre-set size can be realized by pre-storing a corresponding code, table, or other method that can be used in a device (e.g., including a decoder and encoder) to indicate the relevant information. This application is not limited to specific embodiments. For example, a pre-set size may refer to a size defined in a protocol. Optionally, “protocol” may refer to a standard protocol in the technical field of coding, and may include relevant protocols such as the VCC protocol or the ECM protocol.
[0122] Of course, in other alternative embodiments, the decoder may also determine, by other means, whether or not to determine the optimal MIP mode based on the strain cost corresponding to multiple MIP modes, based on a preset size. This application is not specifically limited thereto.
[0123] For example, a decoder can determine whether or not to determine the optimal MIP mode based on the strain costs corresponding to multiple MIP modes, based solely on the width or height of the current block. In one embodiment, if the width of the current block is a preset width, or if the height is a preset height, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes. In another embodiment, the decoder can determine whether or not to determine the optimal MIP mode based on the strain costs corresponding to multiple MIP modes by comparing the size of the current block with a preset size. In one embodiment, if the size of the current block is greater than or less than the preset size, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes. In another embodiment, if the width of the current block is greater than or less than the preset width, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes. In yet another embodiment, if the height of the current block is greater than or less than the preset height, the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes.
[0124] In some embodiments, S320 may include the following: If the image frame in which the current block is located is an I-frame and the size of the current block is a preset size, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0125] For example, if the image frame in which the current block is located is an I-frame, the width of the current block is a preset width, and the height of the current block is a preset height, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. That is, only when the image frame in which the current block is located is an I-frame, the decoder determines whether or not to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes, based on the size of the current block.
[0126] In some embodiments, S320 may include the following: If the image frame in which the current block is located is a B-frame, the decoder determines the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0127] For example, if the image frame in which the current block is located is a B-frame, the decoder can directly determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. That is, if the image frame in which the current block is located is a B-frame, regardless of the size of the current block, the decoder can directly determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0128] In some embodiments, prior to S320, method 300 may further include the following: The decoder obtains the MIP mode to be used for the adjacent block adjacent to the current block. The decoder determines the MIP mode to be used for the adjacent block to be one of several MIP modes.
[0129] Exemplary, an adjacent block may be an image block adjacent to at least one of the following: above, to the left, below left, above right, and above left of the current block. For example, the decoder may determine the image blocks acquired in the order of above, to the left, below left, above right, and above left of the current block as adjacent blocks. Selectively, multiple MIP modes may be used by the decoder to construct an available MIP mode or a list of available MIP modes to predict the current block and to determine the available MIP mode. Thereafter, the decoder determines the optimal MIP mode by predicting the sample in the template region from the available MIP mode or list of available MIP modes.
[0130] In some embodiments, prior to S320, method 300 may further include the following: The decoder performs reconstructed sample padding on a reference region adjacent to the outside of the template region to obtain reference rows and reference columns of the template region. The decoder takes the reference rows and reference columns as input and predicts samples in the template region using each of a plurality of MIP modes to obtain a plurality of prediction blocks corresponding to the plurality of MIP modes. Based on the plurality of prediction blocks and the reconstructed blocks in the template region, the decoder determines the strain costs corresponding to the plurality of MIP modes.
[0131] For example, the decoder pads with reference reconstruction samples necessary for predicting the template.
[0132] For example, the width of the area adjacent to the top of the template area within the reference area is equal to the width of the template area. The height of the area adjacent to the left of the template area within the reference area is equal to the width of the template area. height It is equal to. If the width of the region adjacent to the top of the template region within the reference region is greater than the width of the template region, the decoder can obtain the reference row by downsampling or reducing the dimensionality of the region adjacent to the top of the template region within the reference region. The height of the region adjacent to the left of the template region within the reference region is equal to the width of the template region height If it is larger, the decoder can obtain the reference column by downsampling or reducing the dimensionality of the region adjacent to the left of the template region within the reference region.
[0133] For example, the template region may be the template region used in the TIMD mode described above, and the reference region may be the reference of template used in the TIMD mode. For example, referring to Figure 5, if the current block is a coding unit with a width equal to M and a height equal to N, the decoder pads a reference region consisting of coding units with a width equal to 2(M+L1)+1 and a height equal to 2(N+L2)+1 with reconstructed samples, performs downsampling or dimensionality reduction on the padded reference region to obtain reference rows and reference columns, and then constructs an input vector corresponding to the MIP mode based on the reference rows and reference columns.
[0134] Exemplary, after obtaining the reference row and reference column, the decoder uses the reference row and reference column as input to predict samples in the template region using each of the multiple MIP modes, thereby obtaining multiple prediction blocks corresponding to the multiple MIP modes. That is, the decoder predicts samples in the template region of the current block by traversing the multiple MIP modes based on the reconstructed samples in the reference template of the current block. Using the currently traversed MIP mode as an example, the decoder uses the reference row, reference column, the index of the currently traversed MIP mode, and the third flag mentioned above as input to obtain a prediction block corresponding to the currently traversed MIP mode. The reference row and reference column are used to construct the input vector corresponding to the currently traversed MIP mode. The index of the currently traversed MIP mode is used to determine the matrix and / or bias vector corresponding to the currently traversed MIP mode. The third flag is used to indicate whether or not to transpose the input and output vectors corresponding to the MIP mode. For example, if the third flag is used to indicate that the input and output vectors corresponding to the MIP mode are not transposed, the reference column is spliced after the reference row to form the input vector corresponding to the currently traversed MIP mode. If the third flag is used to indicate that the input and output vectors corresponding to the MIP mode are transposed, the reference row is spliced after the reference column to form the input vector corresponding to the currently traversed MIP mode. Accordingly, if the third flag is used to indicate that the input and output vectors corresponding to the MIP mode are transposed, the decoder transposes the output of the currently traversed MIP mode to obtain a prediction block in the template region. After obtaining multiple prediction blocks corresponding to multiple MIP modes by traversing multiple MIP modes, the decoder can select the MIP mode with the lowest cost based on the strain cost between the multiple prediction blocks and the reconstructed samples in the template region, following the principle of minimum strain cost, and determine it as the optimal MIP mode under template matching-based MIP modes for the current block.
[0135] In some embodiments, when the decoder predicts a sample in a template region using each of multiple MIP modes, it first downsamples the reference row and reference column to obtain an input vector, then uses the input vector as input to predict a sample in the template region by traversing the multiple MIP modes to obtain an output vector corresponding to the multiple MIP modes, and finally upsamples the output vector corresponding to the multiple MIP modes to obtain a prediction block corresponding to the multiple MIP modes.
[0136] For example, a reference row and / or reference column satisfy the input conditions for multiple MIP modes. If a reference row and / or reference column do not satisfy the input conditions for multiple MIP modes, first, the reference row and / or reference column are processed to become input samples that satisfy the input conditions for multiple MIP modes. Then, based on the input samples that satisfy the input conditions for multiple MIP modes, the input vectors corresponding to multiple MIP modes can be determined. For example, if the input condition is a specified number of input samples, and the reference row and / or reference column do not meet the number of input samples for the MIP mode, the decoder reduces the dimensionality of the reference row and / or reference column to the specified number of input samples by performing Haar-downsampling or similar methods, and then determines the input vectors corresponding to multiple MIP modes based on the dimensionality-reduced specified number of input samples.
[0137] In some embodiments, S320 may include the following: The decoder determines the optimal MIP mode based on the differential absolute sum (SATD) corresponding to multiple MIP modes in the template region.
[0138] In this embodiment, when the decoder determines the optimal MIP mode based on the strain costs corresponding to multiple MIP modes within the template region, it designs the strain costs corresponding to multiple MIP modes into SATDs corresponding to multiple MIP modes. Compared to directly calculating the rate strain costs corresponding to multiple MIP modes, this not only enables determining the optimal MIP mode based on the strain costs corresponding to multiple MIP modes within the template region, but also simplifies the complexity of calculating the strain costs corresponding to multiple MIP modes, thereby improving the decompression performance of the decoder.
[0139] In summary, the scheme of this application proposes the idea of fusion enhancement based on the optimal MIP mode. That is, the decoder not only needs to determine the optimal MIP mode for predicting the current block, but also needs to fuse other predicted blocks to achieve different prediction effects. This not only saves bit overhead but also generates new prediction techniques. In fact, since the optimal MIP mode cannot completely replace the optimal prediction mode calculated by the encoding side based on rate distortion cost, the fusion method is used to achieve both prediction accuracy and prediction diversity.
[0140] For example, the main idea behind a decoder-based template matching method for deriving MIP modes can be divided into the following parts:
[0141] First, the reference region (for example, the reference template shown in Figure 5) is padded with reconstructed samples, i.e., with the reference reconstructed samples necessary to predict the samples within the template region (for example, the template shown in Figure 5). Selectively, the width and height of the reference region do not need to exceed the width and height of the template region. If the width and height of the sample-padded reference region exceed the width and height of the template region, downsampling or other dimensionality reduction methods must be performed until the input dimension requirements of the MIP are met.
[0142] Next, the decoder takes the reference reconstruction sample in the reference region, the indices of multiple MIP modes, and the MIP transpose flag as input to predict the sample in the template region and obtain prediction blocks corresponding to the multiple MIP modes. Selectively, the reference reconstruction sample in the reference region must satisfy the input conditions of the MIP mode, for example, by performing dimensionality reduction using half-down sampling down to a specified number of input samples. The indices of multiple MIP modes are used to determine the matrix index of the MIP technique and to obtain the MIP prediction matrix coefficients. The MIP transpose flag is used to indicate whether or not the input and output need to be transposed.
[0143] Next, for prediction blocks corresponding to multiple MIP modes, all combinations of MIP modes and whether or not to transpose the MIP can be traversed to obtain predicted samples within the template region under each state of the MIP mode and MIP transposition flag. The strain between the predicted samples and reconstructed samples within the template region is calculated, and its cost information is recorded. Finally, following the principle of least strain, the MIP mode with the lowest cost and its corresponding MIP transposition information are selected, and the MIP mode with the lowest cost is set as the optimal MIP mode under the template matching-based MIP prediction derivation mode for the current block.
[0144] Finally, the decoder uses the optimal MIP prediction mode and the first intra prediction mode to predict the current block, obtaining the first and second predicted blocks, respectively. The decoder then weights the first and second predicted blocks according to the weights of the optimal MIP prediction mode and the first intra prediction mode to obtain the predicted block for the current block.
[0145] Furthermore, some of the calculations in this application can be replaced with lookup table or shift methods. While the lookup table method may result in some errors compared to performing division directly, it is beneficial for controlling hardware implementation and coding costs. Examples of the above-mentioned calculations include calculations related to strain cost or calculations related to determining the optimal MIP mode.
[0146] The decoding method according to the embodiment of this application has been described in detail above from the perspective of the decoder. 8 Referring to the encoder angle, the encoding method according to the embodiment of this application will be described.
[0147] Figure 8 is a flowchart illustrating the encoding method 400 according to the present invention. The encoding method 400 can be executed by an encoder. For example, the encoding method 400 is applied to the encoding framework 100 shown in Figure 1. For the sake of clarity, an encoder will be described below as an example.
[0148] As shown in Figure 8, the encoding method 400 may include the following: S410: Determine the optimal MIP mode for predicting the current block in the current sequence based on the strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes. The strain costs corresponding to multiple MIP modes include the strain costs obtained by using multiple MIP modes to predict samples in template regions adjacent to the current block. S420: Determine the first intra-prediction mode. The first intra-prediction mode is determined based on the strain costs corresponding to multiple MIP modes and includes at least one of the following: a suboptimal MIP mode for predicting the current block; an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode; or an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode. S430: Predict the current block based on the optimal MIP mode and the first intra prediction mode to obtain the predicted block for the current block. S440: Based on the predicted block of the current block and the original block of the current block, the residual block of the current block is obtained. S450: Encode the residual block of the current block to obtain the bitstream of the current sequence.
[0149] In some embodiments, S410 may include: obtaining a first flag; determining the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes, if the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and a first intra-prediction mode. S450 may include encoding the residual block of the current block and the first flag to obtain a bitstream.
[0150] In some embodiments, S430 may include the following: If the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the first intra-prediction mode, the current block is predicted based on the optimal MIP mode and the first intra-prediction mode to obtain the first rate distortion cost. Predict the current block based on at least one intra-prediction mode and obtain at least one rate distortion cost. If the first rate distortion cost is less than or equal to the minimum of at least one rate distortion cost, the predicted block obtained by predicting the current block based on the optimal MIP mode and the first intra prediction mode is confirmed as the predicted block of the current block. S450 may include the following: Encode the residual block of the current block, the first flag, and the second flag to obtain a bitstream. If the first rate distortion cost is less than or equal to the minimum of at least one rate distortion cost, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode; if the first rate distortion cost is greater than the minimum of at least one rate distortion cost, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode.
[0151] In some embodiments, S450 may include the following: The arrangement order of multiple MIP modes is determined based on the distortion cost corresponding to each MIP mode. Based on the arrangement order of multiple MIP modes, the coding scheme to be used for the optimal MIP mode is determined. The residual blocks of the current block are encoded, and the index of the optimal MIP mode is encoded based on the coding scheme used for the optimal MIP mode to obtain the bitstream.
[0152] In some embodiments, the codeword length of the coding scheme used for the first n MIP modes in the sequence order is smaller than the codeword length of the coding scheme used for the MIP mode following the nth MIP mode in the sequence order, and / or The first n MIP modes use a variable-length coding scheme, and the MIP modes following the nth MIP mode use a truncated binary (TB) coding scheme.
[0153] In some embodiments, S430 may include the following: Based on the optimal MIP mode, the current block is predicted to obtain the first predicted block. Based on the first intra prediction mode, the current block is predicted to obtain the second predicted block. Based on the optimal MIP mode weight and the first intra-prediction mode weight, the first and second prediction blocks are weighted to obtain the prediction block for the current block.
[0154] In some embodiments, S430 may include the following: If the first intra-prediction mode includes an intra-prediction mode derived from a suboptimal MIP mode or TIMD mode, the weights of the optimal MIP mode and the first intra-prediction mode are determined based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra prediction mode includes an intra prediction mode derived from the DIMD mode, the weights of the optimal MIP mode and the weights of the first intra prediction mode are both set to pre-configured values.
[0155] In some embodiments, S410 may include the following: Based on the third flag and multiple MIP modes, samples within the template region are predicted, and the strain costs corresponding to the multiple MIP modes under each state of the third flag are obtained. The third flag is used to indicate whether or not to transpose the input and output vectors corresponding to the MIP modes. The optimal MIP mode is determined based on the distortion costs corresponding to multiple MIP modes under each state of the third flag.
[0156] In some embodiments, S410 may include the following: If the current block size is a preset size, the optimal MIP mode is determined based on the strain costs corresponding to multiple MIP modes.
[0157] In some embodiments, S410 may include the following: If the image frame in which the current block is located is an I-frame and the size of the current block is a preset size, the optimal MIP mode is determined based on the distortion costs corresponding to multiple MIP modes.
[0158] In some embodiments, S410 may include the following: If the image frame in which the current block is located is a B-frame, the optimal MIP mode is determined based on the distortion costs corresponding to multiple MIP modes.
[0159] In some embodiments, prior to S410, method 400 may further include the following: Get the MIP mode used for the adjacent blocks next to the current block. The MIP mode used for adjacent blocks is determined to be one of several MIP modes.
[0160] In some embodiments, prior to S410, method 400 may further include the following: Reconstruction sample padding is performed on the reference region adjacent to the outside of the template region to obtain the reference row and reference column of the template region. Using reference rows and columns as input, the system predicts samples within the template region using each of multiple MIP modes to obtain multiple prediction blocks corresponding to the multiple MIP modes. Based on multiple prediction blocks and reconstruction blocks within the template region, the strain costs corresponding to multiple MIP modes are determined.
[0161] In some embodiments, when predicting samples in a template region using each of multiple MIP modes, first, the reference row and reference column are downsampled to obtain an input vector; then, using the input vector as input, the samples in the template region are predicted by traversing the multiple MIP modes to obtain output vectors corresponding to the multiple MIP modes; and finally, the output vectors corresponding to the multiple MIP modes are upsampled to obtain prediction blocks corresponding to the multiple MIP modes.
[0162] In some embodiments, S410 may include the following: determining the optimal MIP mode based on the differential absolute sum (SATD) corresponding to multiple MIP modes within the template region.
[0163] Furthermore, the encoding method can be understood as the reverse process of the decoding method. Therefore, the specific scheme of the encoding method 400 can be found in the relevant section of the decoding method 300, and for the sake of clarity, it will not be described in detail in this application.
[0164] The scheme of this application will be described below in relation to specific embodiments.
[0165] <Example 1>
[0166] In this embodiment, the first intra prediction mode is a suboptimal prediction mode. That is, the encoder or decoder can perform intra prediction on the current block based on the optimal MIP mode and the suboptimal MIP mode to obtain a predicted block for the current block.
[0167] The encoder traverses the prediction mode. If intra-mode is used for the current block, the encoder obtains a sequence-level enable flag, such as sps_tmmip_enable_flag. Sequence-level enable flags are used to indicate whether template matching-based MIP mode derivation techniques are permitted for the current sequence. If all tmmip enable flags are true, it indicates that the encoder is currently permitted to use TMMIP techniques.
[0168] For example, the encoder process can be implemented as follows:
[0169] Step 1: If sps_tmmip_enable_flag is true, the encoder attempts the TMMIP technique, i.e., performs Step 2. If sps_tmmip_enable_flag is false, the encoder does not attempt the TMMIP technique, i.e., skips Step 2 and performs Step 3 directly.
[0170] Step 2: First, the encoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the encoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value for all of them. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the encoder traverses from the bottom left corner to the top right corner in order until the first available reconstructed sample appears, and then uses the first available reconstructed sample to pad for the previous unavailable position. Next, the encoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0171] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0172] As an example, the specific prediction calculation process includes the following: First, the encoder performs half-downsampling on the reconstructed sample. For example, the encoder determines the down-sampling step size based on the block size. Next, the encoder adjusts the splicing order of the upper and left down-sampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left down-sampled reconstructed sample is spliced after the upper down-sampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper down-sampled reconstructed sample is spliced after the left down-sampled reconstructed sample, and the resulting vector is used as input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the encoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks within the template region. If upsampling is required, first upsampling is performed horizontally, and then vertically. UpsamplingThis process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0173] Next, the encoder calculates the strain cost based on the predicted blocks and reconstructed samples within the template region obtained by traversing each MIP mode, and records the strain cost under each predicted mode and transpose information. After traversing all permitted predicted modes and transpose information, the encoder selects the optimal MIP mode and its corresponding transpose information, and the suboptimal MIP mode and its corresponding transpose information, according to the principle of least cost. The encoder determines whether fusion enhancement is necessary based on the relationship between the cost of the optimal MIP mode and the cost of the suboptimal MIP mode. If the cost of the suboptimal MIP mode is less than twice the cost of the optimal MIP mode, the optimal MIP predicted block and the suboptimal MIP predicted block need to be fused and enhanced. If the cost of the suboptimal predicted mode is twice or more the cost of the optimal MIP mode, fusion enhancement is not necessary.
[0174] Finally, if fusion enhancement is required, the encoder obtains prediction blocks corresponding to the optimal MIP mode and prediction blocks corresponding to the suboptimal MIP mode based on the optimal MIP mode, the suboptimal MIP mode, the transpose information of the optimal MIP mode, and the transpose information of the suboptimal MIP mode. Specifically, first, the encoder may downsample the reconstructed samples adjacent to the top and left of the current block, splice them according to the transpose information to obtain the input vector, read the matrix coefficients under the current mode using the MIP mode as the index, and then obtain the output vector by calculating the input vector and matrix coefficients. The encoder transposes the output based on the transpose information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain the optimal MIP prediction block and the suboptimal MIP prediction block of the same size as the current block, and also performs a weighted average on the optimal MIP prediction block and the suboptimal MIP prediction block based on the calculated weight values of the optimal MIP mode and the weight values of the suboptimal MIP mode to obtain a new prediction block which becomes the final prediction block for the current block. If fusion enhancement is not required, the encoder can calculate the optimal MIP prediction block based on the optimal MIP mode and its transpose information, and the calculation process is the same as described above. Finally, the encoder sets the optimal MIP prediction block as the prediction block for the current block. Furthermore, the encoder obtains the rate strain cost of the current block and denots it as cost1.
[0175] Step 3: The encoder continues to traverse other intra-prediction techniques and calculates the corresponding rate distortion costs, denoted as cost2...costN.
[0176] Step 4: If cost1 is the smallest of all rate distortion costs, use the TMMIP technique for the current block, set the TMMIP usage flag for the current block to true, and write it to the bitstream. If cost1 is not the smallest rate distortion cost, use another intra-prediction technique for the current block, set the TMMIP usage flag for the current block to false, and write it to the bitstream. Note that information such as flags or indices for other intra-prediction techniques is transmitted based on their definitions and is not described in detail here.
[0177] Step 5: The encoder determines the residual block of the current block based on the predicted block of the current block and the original block of the current block, and performs operations such as transformation and quantization, entropy coding, and loop filtering on the residual block of the current block. The specific process can be found in the related content above, and will not be described in detail here to avoid repetition.
[0178] The decoder-related scheme in this embodiment is described below.
[0179] The decoder parses block-level flags. If intra-mode is used for the current block, the decoder parses or retrieves sequence-level permission flags such as sps_tmmip_enable_flag. Sequence-level permission flags are used to indicate whether the current sequence is permitted to use template matching-based MIP mode derivation techniques. If all tmmip permission flags are true, it indicates that the decoder is currently permitted to use TMMIP techniques.
[0180] For example, the decoder process can be implemented as follows:
[0181] Step 1: If sps_tmmip_enable_flag is true, the decoder parses the TMMIP usage flag for the current block. Otherwise, the current decoding process does not need to decode the block-level TMMIP usage flag, and the block-level TMMIP usage flag is set to false by default. If the TMMIP usage flag for the current block is true, Step 2 is performed. Otherwise, Step 3 is performed.
[0182] Step 2: First, the decoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the decoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the decoder traverses from the bottom left corner to the top right corner until the first available reconstructed sample appears, and uses the first available reconstructed sample to pad for the previous unavailable position. Next, the decoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0183] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0184] As an example, the specific prediction calculation process includes the following: First, the decoder performs half-downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, the decoder adjusts the splicing order of the upper and left downsampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left downsampled reconstructed sample is spliced after the upper downsampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper downsampled reconstructed sample is spliced after the left downsampled reconstructed sample, and the resulting vector is used as input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the decoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks in the template region. If upsampling is required, upsampling is performed first according to the horizontal direction, and then according to the vertical direction. Upsampling This process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0185] Next, the decoder calculates the strain cost based on the predicted blocks and reconstructed samples within the template region obtained by traversing each MIP mode, and records the strain cost under each predicted mode and transpose information. After traversing all permitted predicted modes and transpose information, the decoder selects the optimal MIP mode and its corresponding transpose information, and the suboptimal MIP mode and its corresponding transpose information, according to the principle of least cost. The decoder determines whether fusion enhancement is necessary based on the relationship between the cost of the optimal MIP mode and the cost of the suboptimal MIP mode. If the cost of the suboptimal MIP mode is less than twice the cost of the optimal MIP mode, the optimal MIP predicted block and the suboptimal MIP predicted block need to be fused and enhanced. If the cost of the suboptimal predicted mode is twice or more the cost of the optimal MIP mode, fusion enhancement is not necessary.
[0186] Finally, if fusion enhancement is required, the decoder obtains prediction blocks corresponding to the optimal MIP mode and prediction blocks corresponding to the suboptimal MIP mode based on the optimal MIP mode, the suboptimal MIP mode, the transpose information of the optimal MIP mode, and the transpose information of the suboptimal MIP mode. Specifically, first, the decoder may downsample the reconstructed samples adjacent to the top and left of the current block, splice them according to the transpose information to obtain the input vector, read out the matrix coefficients under the current mode using the MIP mode as the index, and then obtain the output vector by calculating the input vector and matrix coefficients. The decoder transposes the output based on the transpose information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain the optimal MIP prediction block and the suboptimal MIP prediction block of the same size as the current block, and also performs a weighted average on the optimal MIP prediction block and the suboptimal MIP prediction block based on the calculated weight values of the optimal MIP mode and the weight values of the suboptimal MIP mode to obtain a new prediction block which becomes the final prediction block for the current block. If fusion enhancement is not required, the decoder can calculate the optimal MIP prediction block based on the optimal MIP mode and its transposition information, and the calculation process is the same as described above. Finally, the decoder sets the optimal MIP prediction block as the prediction block for the current block.
[0187] Step 3: The decoder continues to analyze information such as the use flags or indices of other intra-prediction techniques and determines the final predicted block for the current block based on the analyzed information.
[0188] Step 4: The decoder analyzes the bitstream to obtain the frequency-domain residual block (also called frequency-domain residual information) of the current block, and then performs inverse quantization and inverse transform on the frequency-domain residual block of the current block to obtain the residual block (also called time-domain residual block or time-domain residual information) of the current block. Next, the decoder adds the predicted block of the current block to the residual block of the current block to obtain the reconstructed sample block.
[0189] Step 5: After techniques such as loop filtering are performed on all reconstruction sample blocks in the current image, the final reconstructed image is obtained.
[0190] Selectively, the reconstructed image may be used as a video output or as a reference for subsequent decoding.
[0191] In this embodiment, the size of the template region used by the encoder or decoder in TMMIP technology can be predefined according to the size of the current block. For example, the width of the region adjacent to the top of the current block within the template region is equal to the width of the current block, and its height is equal to the height of two rows of samples. The height of the region adjacent to the left of the current block within the template region is equal to the height of the current block, and its width is equal to two rows of samples. column This is equal to the width of the sample. Of course, in other alternative embodiments, it can be implemented as a template region of other sizes, and this application is not specifically limited to that.
[0192] <Example 2>
[0193] In this embodiment, the first intra-prediction mode is an intra-prediction mode derived from the TIMD mode. That is, the encoder or decoder can perform an intra-prediction on the current block based on the optimal MIP mode and the intra-prediction mode derived from the TIMD mode, and obtain a predicted block for the current block.
[0194] In other words, the template matching-based MIP mode derivation fusion enhancement technology can not only fuse two derived MIP prediction blocks, but also fuse them with prediction blocks generated by other template matching-based derivation technologies. This application provides a method for fusing derived conventional prediction blocks with matrix-based prediction blocks by fusing TMMIP technology and TIMD technology. In TIMD technology, the idea of template matching is used on the coding side to derive the optimal conventional intra-prediction mode, and TIMD technology can also be used to perform offset expansion on this prediction mode to obtain an updated intra-prediction mode. In addition, TMMIP technology derives the optimal MIP mode using the idea of template matching on the coding side. By fusing these two optimal prediction modes, it is possible to reconcile the directivity of conventional prediction blocks with the unique texture characteristics of MIP prediction, thereby generating entirely new prediction blocks and improving coding efficiency.
[0195] The encoder traverses the prediction mode. If intra-mode is used for the current block, the encoder obtains a sequence-level permission flag such as sps_tmmip_enable_flag. Sequence-level permission flags are used to indicate whether the current sequence is permitted to use template matching-based MIP mode derivation techniques. If all tmmip permission flags are true, it indicates that the encoder is currently permitted to use TMMIP techniques.
[0196] For example, the encoder process can be implemented as follows:
[0197] Step 1: If sps_tmmip_enable_flag is true, the encoder attempts the TMMIP technique, i.e., performs Step 2. If sps_tmmip_enable_flag is false, the encoder does not attempt the TMMIP technique, i.e., skips Step 2 and performs Step 3 directly.
[0198] Step 2: First, the encoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the encoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value for all of them. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the encoder traverses from the bottom left corner to the top right corner in order until the first available reconstructed sample appears, and then uses the first available reconstructed sample to pad for the previous unavailable position. Next, the encoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0199] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0200] As an example, the specific prediction calculation process includes the following: First, the encoder performs half-downsampling on the reconstructed sample. For example, the encoder determines the downsampling step size based on the block size. Next, the encoder adjusts the splicing order of the upper and left downsampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left downsampled reconstructed sample is spliced after the upper downsampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper downsampled reconstructed sample is spliced after the left downsampled reconstructed sample, and the resulting vector is used as input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the encoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks in the template area. If upsampling is required, upsampling is performed first according to the horizontal direction, and then according to the vertical direction. Upsampling This process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0201] Furthermore, the encoder needs to attempt TIMD's template matching calculation process, obtaining different interpolation filters based on different prediction mode indices and retrieving predicted samples within the template by interpolating the reference samples.
[0202] Next, the encoder calculates the strain cost based on the predicted and reconstructed samples within the template region obtained by traversing each MIP mode, and records the strain cost under each predicted mode and transpose information. Based on the strain cost under each predicted mode and transpose information, the encoder selects the optimal MIP mode and its corresponding transpose information according to the principle of least cost. Furthermore, the encoder must traverse all permitted intra-prediction modes in TIMD, calculate the predicted samples within the template, calculate the strain cost using the predicted and reconstructed samples within the template, and record the optimal, suboptimal, strain cost corresponding to the optimal, and strain cost corresponding to the suboptimal prediction mode derived from the TIMD technique, according to the principle of least cost.
[0203] Finally, based on the obtained optimal MIP mode and transpose information, the encoder optionally downsamples the reconstructed samples adjacent to the top and left of the current block, splices them based on the transpose information to obtain the input vector, reads out the matrix coefficients under the current mode using the MIP mode as the index, and then obtains the output vector by calculating the input vector and matrix coefficients. The encoder transposes the output based on the transpose information, and upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain an output of the same size as the current block, which can be used as the optimal MIP prediction block for the current block.
[0204] For the optimal and suboptimal prediction modes derived from TIMD technology, if neither the optimal nor the suboptimal prediction mode is the mean value (DC) mode or the planar mode, and the strain cost corresponding to the suboptimal prediction mode is less than twice the strain cost corresponding to the optimal prediction mode, the encoder needs to merge the prediction blocks. First, the encoder obtains interpolation filtering coefficients based on the optimal prediction mode and performs interpolation filtering on the reconstructed samples adjacent to the upper and left to obtain prediction samples at all positions in the current block, which are recorded as the optimal prediction block. Next, the encoder obtains interpolation filtering coefficients based on the suboptimal prediction mode and performs interpolation filtering on the reconstructed samples adjacent to the upper and left to obtain prediction samples at all positions in the current block, which are recorded as the suboptimal prediction block. Furthermore, the encoder uses the ratio of the cost corresponding to the optimal prediction mode to the cost corresponding to the suboptimal prediction mode to calculate the weight values of the optimal prediction block and the weight values of the suboptimal prediction block. Finally, the encoder weight-merges the optimal and suboptimal prediction blocks to obtain the prediction block for the current block, which is output. Furthermore, if the optimal or suboptimal prediction mode is the mean mode (DC) or planar mode (PLANAR), or if the cost corresponding to the suboptimal prediction mode is greater than twice the cost corresponding to the optimal prediction mode, the encoder does not need to merge prediction blocks and uses only the optimal prediction mode to perform interpolation filtering on the reconstructed samples adjacent to the upper and left, and the resulting optimal prediction block is considered the optimal TIMD prediction block for the current block.
[0205] Finally, the encoder obtains a new prediction block by performing a weighted average on the optimal MIP prediction block and the optimal TIMD prediction block, based on the calculated optimal MIP mode weights and the prediction mode weights derived from the TIMD technique. This new prediction block is the prediction block for the current block. Furthermore, the encoder obtains the rate distortion cost of the current block and denots it as cost1.
[0206] Furthermore, the template region in TIMD technology and the template region in TMMIP technology can be set to be the same; that is, the region for calculating the strain cost corresponding to the template region is the same. Therefore, the cost information of the template regions in the two technologies can be equivalent and at the same comparison level. In this case, it can be determined whether or not fusion strengthening is performed based on the cost information, and this application is not specifically limited to this.
[0207] Step 3: The encoder continues to traverse other intra-prediction techniques and calculates the corresponding rate distortion costs, denoted as cost2...costN.
[0208] Step 4: If cost1 is the smallest of all rate distortion costs, use the TMMIP technique for the current block, set the TMMIP usage flag for the current block to true, and write it to the bitstream. If cost1 is not the smallest rate distortion cost, use another intra-prediction technique for the current block, set the TMMIP usage flag for the current block to false, and write it to the bitstream. Note that information such as flags or indices for other intra-prediction techniques is transmitted based on their definitions and is not described in detail here.
[0209] Step 5: The encoder determines the residual block of the current block based on the predicted block of the current block and the original block of the current block, and performs operations such as transformation and quantization, entropy coding, and loop filtering on the residual block of the current block. The specific process can be found in the related content above, and will not be described in detail here to avoid repetition.
[0210] The decoder-related scheme in this embodiment is described below.
[0211] The decoder parses block-level flags. If intra-mode is used for the current block, the decoder parses or retrieves sequence-level permission flags such as sps_tmmip_enable_flag. Sequence-level permission flags are used to indicate whether the current sequence is permitted to use template matching-based MIP mode derivation techniques. If all tmmip permission flags are true, it indicates that the decoder is currently permitted to use TMMIP techniques.
[0212] For example, the decoder process can be implemented as follows:
[0213] Step 1: If sps_tmmip_enable_flag is true, the decoder parses the TMMIP usage flag for the current block. Otherwise, the current decoding process does not need to decode the block-level TMMIP usage flag, and the block-level TMMIP usage flag is set to false by default. If the TMMIP usage flag for the current block is true, Step 2 is performed. Otherwise, Step 3 is performed.
[0214] Step 2: First, the decoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the decoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the decoder traverses from the bottom left corner to the top right corner until the first available reconstructed sample appears, and uses the first available reconstructed sample to pad for the previous unavailable position. Next, the decoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0215] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0216] As an example, the specific prediction calculation process includes the following: First, the decoder performs half-downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, the decoder adjusts the splicing order of the upper and left downsampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left downsampled reconstructed sample is spliced after the upper downsampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper downsampled reconstructed sample is spliced after the left downsampled reconstructed sample, and the resulting vector is used as input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the decoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks in the template region. If upsampling is required, upsampling is performed first according to the horizontal direction, and then according to the vertical direction. Upsampling This process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0217] Furthermore, the decoder needs to attempt TIMD's template matching calculation process, obtaining different interpolation filters based on different prediction mode indices and retrieving predicted samples within the template by interpolating the reference samples.
[0218] Next, the decoder calculates the strain cost based on the predicted samples and reconstructed samples within the template region obtained by traversing each MIP mode, and records the strain cost under each predicted mode and transpose information. Based on the strain cost under each predicted mode and transpose information, the decoder selects the optimal MIP mode and its corresponding transpose information according to the principle of least cost. Furthermore, the decoder must traverse all permitted intra-prediction modes in TIMD, calculate the predicted samples within the template, calculate the strain cost using the predicted samples and reconstructed samples within the template, and record the optimal prediction mode, suboptimal prediction mode, strain cost corresponding to the optimal prediction mode, and strain cost corresponding to the suboptimal prediction mode derived from the TIMD technique, according to the principle of least cost.
[0219] Finally, based on the obtained optimal MIP mode and transpose information, the decoder optionally downsamples the reconstructed samples adjacent to the top and left of the current block, splices them based on the transpose information to obtain the input vector, reads out the matrix coefficients under the current mode using the MIP mode as the index, and then obtains the output vector by calculating the input vector and matrix coefficients. The decoder transposes the output based on the transpose information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain an output of the same size as the current block, which can then be used as the optimal MIP prediction block for the current block.
[0220] For the optimal and suboptimal prediction modes derived from TIMD technology, if neither the optimal nor the suboptimal prediction mode is the mean value (DC) mode or the planar mode, and the strain cost corresponding to the suboptimal prediction mode is less than twice the strain cost corresponding to the optimal prediction mode, the decoder needs to merge the prediction blocks. First, the decoder obtains interpolation filtering coefficients based on the optimal prediction mode and performs interpolation filtering on the reconstructed samples adjacent to the upper and left to obtain prediction samples at all positions in the current block, which are recorded as the optimal prediction block. Next, the decoder obtains interpolation filtering coefficients based on the suboptimal prediction mode and performs interpolation filtering on the reconstructed samples adjacent to the upper and left to obtain prediction samples at all positions in the current block, which are recorded as the suboptimal prediction block. Furthermore, the decoder uses the ratio of the cost corresponding to the optimal prediction mode to the cost corresponding to the suboptimal prediction mode to calculate the weight values of the optimal prediction block and the weight values of the suboptimal prediction block. Finally, the decoder weight-merges the optimal and suboptimal prediction blocks to obtain the prediction block for the current block, which is output. Furthermore, if the optimal or suboptimal prediction mode is the mean mode (DC) or planar mode (PLANAR), or if the cost corresponding to the suboptimal prediction mode is greater than twice the cost corresponding to the optimal prediction mode, the decoder does not need to merge prediction blocks and uses only the optimal prediction mode to perform interpolation filtering on the reconstructed samples adjacent to the upper and left, and the resulting optimal prediction block is considered the optimal TIMD prediction block for the current block.
[0221] Finally, the decoder obtains a new prediction block by performing a weighted average on the optimal MIP prediction block and the optimal TIMD prediction block, based on the calculated optimal MIP mode weights and the prediction mode weights derived from the TIMD technique. This new prediction block is the prediction block for the current block.
[0222] Step 3: The decoder continues to analyze information such as the use flags or indices of other intra-prediction techniques and determines the final predicted block for the current block based on the analyzed information.
[0223] Step 4: The decoder analyzes the bitstream to obtain the frequency-domain residual block (also called frequency-domain residual information) of the current block, and then performs inverse quantization and inverse transform on the frequency-domain residual block of the current block to obtain the residual block (also called time-domain residual block or time-domain residual information) of the current block. Next, the decoder adds the predicted block of the current block to the residual block of the current block to obtain the reconstructed sample block.
[0224] Step 5: After techniques such as loop filtering are performed on all reconstruction sample blocks in the current image, the final reconstructed image is obtained.
[0225] Selectively, the reconstructed image may be used as a video output or as a reference for subsequent decoding.
[0226] In this embodiment, the process for calculating weight values for weighted fusion of TIMD prediction blocks can be found in the above description of the TIMD technology, and will not be detailed here to avoid duplication. Furthermore, the encoder or decoder can determine whether or not fusion enhancement is performed based on the optimal prediction mode derived from TIMD. For example, if the optimal prediction mode derived from TIMD is DC mode or PLANAR mode, the encoder or decoder does not need to use fusion enhancement. That is, the encoder or decoder uses only prediction blocks generated by the optimal MIP mode derived from the TIMD technology as the prediction blocks for the current block. Furthermore, the size of the template region used by the encoder or decoder in the TMMIP technology can be predefined according to the size of the current block. For example, the definition of the template region in the TMMIP technology may be the same as or different from the definition of the template region in the TIMD technology. For example, if the width of the current block is 8 or less, the height of the region adjacent to the top of the current block within the template region is equal to the height of two rows of samples. Otherwise, the height is equal to the height of four rows of samples. Similarly, if the height of the current block is 8 or less, the width of the area adjacent to the left of the current block within the template area is 2 columns of samples. width It is equal to . Otherwise, the width is 4 columns of samples width It is equal to.
[0227] <Example 3>
[0228] In this embodiment, the first intra prediction mode is an intra prediction mode derived from the DIMD mode. That is, the encoder or decoder can perform an intra prediction on the current block based on the optimal MIP mode and the intra prediction mode derived from the DIMD mode to obtain a predicted block for the current block.
[0229] Similar to Example 2, TMMIP technology can also be integrated and enhanced with DIMD technology.
[0230] Note that while both the prediction modes derived from DIMD technology and TIMD technology are conventional intra-prediction modes, the derivation methods differ, so the prediction modes obtained by both are not necessarily the same. Furthermore, the method of fusion enhancement between TMMIP technology and DIMD technology differs from the method of fusion enhancement between TMMIP technology and TIMD technology. For example, in TMMIP technology and TIMD technology, the size of the template region is generally the same, and the cost information calculated is basically SATD (also called distortion cost based on the Hadamard transform), so fusion weights can be directly calculated based on this cost information in both TMMIP technology and TIMD technology. However, the size of the template region in DIMD technology is generally different from that of TMMIP technology (or TIMD The template region size differs from that in the technology, and the rules for the DIMD derivation prediction mode are based on gradient amplitude values. Because gradient amplitude values are not equivalent to SATD costs, the weights cannot be easily calculated by referring to the scheme when fusing TMMIP and TIMD technologies.
[0231] The encoder traverses the prediction mode. If intra-mode is used for the current block, the encoder obtains a sequence-level permission flag such as sps_tmmip_enable_flag. Sequence-level permission flags are used to indicate whether the current sequence is permitted to use template matching-based MIP mode derivation techniques. If all tmmip permission flags are true, it indicates that the encoder is currently permitted to use TMMIP techniques.
[0232] For example, the encoder process can be implemented as follows:
[0233] Step 1: If sps_tmmip_enable_flag is true, the encoder attempts the TMMIP technique, i.e., performs Step 2. If sps_tmmip_enable_flag is false, the encoder does not attempt the TMMIP technique, i.e., skips Step 2 and performs Step 3 directly.
[0234] Step 2: First, the encoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the encoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value for all of them. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the encoder traverses from the bottom left corner to the top right corner in order until the first available reconstructed sample appears, and then uses the first available reconstructed sample to pad for the previous unavailable position. Next, the encoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0235] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0236] As an example, the specific prediction calculation process includes the following: First, the encoder performs half-downsampling on the reconstructed sample. For example, the encoder determines the downsampling step size based on the block size. Next, the encoder adjusts the splicing order of the upper and left downsampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left downsampled reconstructed sample is spliced after the upper downsampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper downsampled reconstructed sample is spliced after the left downsampled reconstructed sample, and the resulting vector is used as input. Next, the encoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the encoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks in the template area. If upsampling is required, upsampling is performed first according to the horizontal direction, and then according to the vertical direction. Upsampling This process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0237] Furthermore, the encoder utilizes DIMD technology to derive the optimal intra-prediction mode, i.e., the optimal DIMD mode. In DIMD technology, the gradient value of the reconstructed sample within the template region is calculated based on the Sobel operator, and the gradient value is transformed based on the angle value corresponding to different prediction modes to obtain the amplitude value under the corresponding prediction mode.
[0238] Next, the encoder calculates the strain cost using the predicted blocks of the template obtained by traversing each MIP mode and the reconstructed samples within the template, and records the optimal MIP mode and transpose information according to the principle of least cost. Furthermore, the encoder traverses all intra-prediction modes that are permitted to be used, calculates the amplitude value under each intra-prediction mode, and records the optimal DIMD prediction mode according to the principle of maximum amplitude.
[0239] Finally, based on the obtained optimal MIP mode and transpose information, the encoder optionally downsamples the reconstructed samples adjacent to the upper and left sides of the current block, splices them based on the transpose information to obtain the input vector, reads the matrix coefficients under the current mode using the MIP mode as the index, and then obtains the output vector by calculating the input vector and matrix coefficients. The encoder transposes the output based on the transpose information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain an output of the same size as the current block, which can be the optimal MIP prediction block for the current block. Furthermore, for the optimal DIMD prediction mode, the encoder obtains the corresponding interpolation filtering coefficients, performs interpolation filtering on the reconstructed samples adjacent to the upper and left sides to obtain prediction samples at all positions within the current block, and records them as the optimal DIMD prediction block. For each prediction sample, the encoder weights the optimal MIP prediction block and the optimal DIMD prediction block according to preset weights to obtain a new prediction block. The new prediction block is the prediction block for the current block. Furthermore, the encoder obtains the rate distortion cost of the current block and records it as cost1.
[0240] Step 3: The encoder continues to traverse other intra-prediction techniques and calculates the corresponding rate distortion costs, denoted as cost2...costN.
[0241] Step 4: If cost1 is the smallest of all rate distortion costs, use the TMMIP technique for the current block, set the TMMIP usage flag for the current block to true, and write it to the bitstream. If cost1 is not the smallest rate distortion cost, use another intra-prediction technique for the current block, set the TMMIP usage flag for the current block to false, and write it to the bitstream. Note that information such as flags or indices for other intra-prediction techniques is transmitted based on their definitions and is not described in detail here.
[0242] Step 5: The encoder determines the residual block of the current block based on the predicted block of the current block and the original block of the current block, and performs operations such as transformation and quantization, entropy coding, and loop filtering on the residual block of the current block. The specific process can be found in the related content above, and will not be described in detail here to avoid repetition.
[0243] The decoder-related scheme in this embodiment is described below.
[0244] The decoder parses block-level flags. If intra-mode is used for the current block, the decoder parses or retrieves sequence-level permission flags such as sps_tmmip_enable_flag. Sequence-level permission flags are used to indicate whether the current sequence is permitted to use template matching-based MIP mode derivation techniques. If all tmmip permission flags are true, it indicates that the decoder is currently permitted to use TMMIP techniques.
[0245] For example, the decoder process can be implemented as follows:
[0246] Step 1: If sps_tmmip_enable_flag is true, the decoder parses the TMMIP usage flag for the current block. Otherwise, the current decoding process does not need to decode the block-level TMMIP usage flag, and the block-level TMMIP usage flag is set to false by default. If the TMMIP usage flag for the current block is true, Step 2 is performed. Otherwise, Step 3 is performed.
[0247] Step 2: First, the decoder performs reconstructed sample padding on rows and columns adjacent to the outside of the template area. The padding process is the same as the padding method in the original intra-prediction process. For example, the decoder can traverse from the bottom left corner to the top right corner to pad. If all reconstructed samples are available, padding is performed sequentially with all available reconstructed samples. If all reconstructed samples are unavailable, padding is performed with the average value. If some reconstructed samples are available, padding is performed first with the available reconstructed samples, and for the remaining unavailable reconstructed samples, the decoder traverses from the bottom left corner to the top right corner until the first available reconstructed sample appears, and uses the first available reconstructed sample to pad for the previous unavailable position. Next, the decoder takes the reconstructed samples outside the padded template region as input and predicts the samples within the template region using the permitted MIP mode.
[0248] For example, there are 16 MIP modes permitted for use with 4x4 blocks. There are 8 MIP modes permitted for use with blocks whose width or height is equal to 4, or with 8x8 blocks. There are 6 MIP modes permitted for use with blocks of other sizes. In addition, the MIP transpose function can be used with blocks of any size, and the TMMIP prediction modes described above are the same as those used in MIP technology.
[0249] As an example, the specific prediction calculation process includes the following: First, the decoder performs half-downsampling on the reconstructed samples. For example, the decoder determines the downsampling step size based on the block size. Next, the decoder adjusts the splicing order of the upper and left downsampled reconstructed samples depending on whether transposition is required. If transposition is not required, the left downsampled reconstructed sample is spliced after the upper downsampled reconstructed sample, and the resulting vector is used as input. If transposition is required, the upper downsampled reconstructed sample is spliced after the left downsampled reconstructed sample, and the resulting vector is used as input. Next, the decoder obtains the MIP matrix coefficients using the traversed prediction mode as an index, and obtains the output vector by calculating the MIP matrix coefficients and the input. Finally, the decoder upsamples the output vector according to the number of samples in the output vector and the size of the current template. If upsampling is not required, the vectors are arranged sequentially according to the horizontal direction and output as prediction blocks in the template region. If upsampling is required, upsampling is performed first according to the horizontal direction, and then according to the vertical direction. Upsampling This process involves upsampling until the size matches the template size, and then outputting the predicted blocks within the template region.
[0250] Furthermore, the decoder uses DIMD technology to derive the optimal intra-prediction mode, i.e., the optimal DIMD mode. In DIMD technology, the gradient value of the reconstructed sample within the template region is calculated based on the Sobel operator, and the gradient value is transformed based on the angle value corresponding to different prediction modes to obtain the amplitude value under the corresponding prediction mode.
[0251] Next, the decoder calculates the strain cost using the prediction blocks of the template obtained by traversing each MIP mode and the reconstructed samples within the template, and records the optimal MIP mode and transpose information according to the principle of least cost. Furthermore, the decoder traverses all intra-prediction modes that are permitted to be used, calculates the amplitude value under each intra-prediction mode, and records the optimal DIMD prediction mode according to the principle of maximum amplitude.
[0252] Finally, based on the obtained optimal MIP mode and transpose information, the decoder optionally downsamples the reconstructed samples adjacent to the upper and left sides of the current block, splices them based on the transpose information to obtain the input vector, reads the matrix coefficients under the current mode using the MIP mode as the index, and then obtains the output vector by calculating the input vector and matrix coefficients. The decoder transposes the output based on the transpose information, upsamples the output vector based on the size of the current block and the number of samples in the output vector to obtain an output of the same size as the current block, which can be the optimal MIP prediction block for the current block. Furthermore, for the optimal DIMD prediction mode, the decoder obtains the corresponding interpolation filtering coefficients, performs interpolation filtering on the reconstructed samples adjacent to the upper and left sides to obtain prediction samples at all positions within the current block, and records them as the optimal DIMD prediction block. For each prediction sample, the decoder weights the optimal MIP prediction block and the optimal DIMD prediction block according to preset weights to obtain a new prediction block. The new prediction block is the prediction block for the current block.
[0253] Step 3: The decoder continues to analyze information such as the use flags or indices of other intra-prediction techniques and determines the final predicted block for the current block based on the analyzed information.
[0254] Step 4: The decoder analyzes the bitstream to obtain the frequency-domain residual block (also called frequency-domain residual information) of the current block, and then performs inverse quantization and inverse transform on the frequency-domain residual block of the current block to obtain the residual block (also called time-domain residual block or time-domain residual information) of the current block. Next, the decoder adds the predicted block of the current block to the residual block of the current block to obtain the reconstructed sample block.
[0255] Step 5: After techniques such as loop filtering are performed on all reconstruction sample blocks in the current image, the final reconstructed image is obtained.
[0256] Selectively, the reconstructed image may be used as a video output or as a reference for subsequent decoding.
[0257] In this embodiment, the calculation process for the optimal DIMD prediction block can be found in the above description of DIMD technology, and will not be detailed here to avoid duplication. Furthermore, the fusion weights of the optimal MIP prediction block and the optimal DIMD prediction block can be a preset value, for example, the optimal MIP prediction block accounts for 5 / 9 and the optimal DIMD prediction block accounts for 4 / 9. Of course, in other alternative embodiments, the fusion weights of the optimal MIP prediction block and the optimal DIMD prediction block may be other values, and this application is not specifically limited to such values.
[0258] While preferred embodiments of this application have been described in detail above with reference to the attached drawings, this application is not limited to the detailed contents of the above embodiments. Within the scope of the technical idea of this application, various simple modifications can be made to the technical proposal of this application, and any such simple modifications will fall within the scope of protection of this application. For example, each specific technical feature described in the above specific embodiments may be combined by any appropriate means, provided that they are not contradictory, and in order to avoid unnecessary redundancy, this application will not describe various possible combinations again. Furthermore, for example, between various different embodiments of this application, any combination should be considered disclosed in this application, as long as it does not contradict the idea of this application. It should be understood that in the various method embodiments of this application, the magnitude of the sequence number of each process does not indicate the execution order. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0259] The above describes in detail the method embodiment of this application. Hereinafter, the apparatus embodiment of this application will be described in detail with reference to Figures 9 to 11.
[0260] Figure 9 is a block diagram showing a decoder 500 according to an embodiment of this application.
[0261] As shown in Figure 9, the decoder 500 may include an analysis unit 510, a prediction unit 520, and a reconstruction unit 530. The analysis unit 510 is configured to analyze the bitstream and obtain the residual block of the current block in the current sequence. The prediction unit 520 is configured to determine the optimal MIP mode for predicting the current block based on the strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, the strain costs corresponding to multiple MIP modes include the strain costs obtained by using multiple MIP modes to predict samples in template regions adjacent to the current block. The prediction unit 520 is configured to determine a first intra prediction mode, which is determined based on distortion costs corresponding to a plurality of MIP modes and is an almost optimal MIP mode for predicting the current block, an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode, or an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode, and includes at least one of them. The prediction unit 520 is configured to predict the current block based on the optimal MIP mode and the first intra prediction mode to obtain a predicted block of the current block. The reconstruction unit 530 is configured to obtain a reconstructed block of the current block based on the residual block of the current block and the predicted block of the current block.
[0262] In some embodiments, specifically, the prediction unit 520 analyzes the bitstream of the current sequence to obtain a first flag, and when the first flag is used to indicate that it is permitted to predict an image block in the current sequence using the optimal MIP mode and the first intra prediction mode, it is configured to determine the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes.
[0263] In some embodiments, specifically, the prediction unit 520 analyzes the bitstream to obtain a second flag when the first flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra prediction mode, and when the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra prediction mode, it is configured to determine the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes.
[0264] In some embodiments, the analysis unit 510 further determines the arrangement order of a plurality of MIP modes based on the distortion costs corresponding to the plurality of MIP modes, determines the coding method used for the optimal MIP mode based on the arrangement order of the plurality of MIP modes, and is configured to decode the bitstream of the current sequence based on the coding method used for the optimal MIP mode to obtain the index of the optimal MIP mode.
[0265] In some embodiments, the codeword length of the coding method used for the first n MIP modes in the arrangement order is smaller than the codeword length of the coding method used for the MIP modes following the nth MIP mode in the arrangement order, and / or a variable length coding method is used for the first n MIP modes, and a truncated binary (TB) coding method is used for the MIP modes following the nth MIP mode.
[0266] In some embodiments, the prediction unit 520 specifically predicts the current block based on the optimal MIP mode to obtain a first predicted block, predicts the current block based on the first intra prediction mode to obtain a second predicted block, and is configured to weight the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the first intra prediction mode to obtain the predicted block of the current block.
[0267] In some embodiments, before the prediction unit 520 weights the first predicted block and the second predicted block based on the weight of the optimal MIP mode and the weight of the first intra prediction mode to obtain the predicted block of the current block, If the first intra-prediction mode includes an intra-prediction mode derived from a suboptimal MIP mode or TIMD mode, the weights of the optimal MIP mode and the first intra-prediction mode are determined based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra-prediction mode includes an intra-prediction mode derived from the DIMD mode, the system is further configured to set both the optimal MIP mode weight and the weight of the first intra-prediction mode to a preset value.
[0268] In some embodiments, the prediction unit 520 is specifically: The system is configured to predict samples within the template region based on a third flag and multiple MIP modes, and to obtain the strain costs corresponding to the multiple MIP modes under each state of the third flag, the third flag being used to indicate whether or not to transpose the input and output vectors corresponding to the MIP modes. The prediction unit 520 is configured to determine the optimal MIP mode based on the strain costs corresponding to multiple MIP modes under each state of the third flag.
[0269] In some embodiments, the prediction unit 520 is specifically: If the current block size is a preset size, it is configured to determine the optimal MIP mode based on the distortion cost corresponding to multiple MIP modes.
[0270] In some embodiments, the prediction unit 520 is specifically: If the image frame in which the current block is located is an I-frame, and the size of the current block is a preset size, the system is configured to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0271] In some embodiments, the prediction unit 520 is specifically: If the image frame in which the current block is located is a B-frame, the system is configured to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0272] In some embodiments, the prediction unit 520 determines the optimal MIP mode for predicting the current block based on the strain costs corresponding to multiple MIP modes, Get the MIP mode used for the adjacent block next to the current block, The configuration further determines which MIP mode is used for adjacent blocks, selecting from multiple MIP modes.
[0273] In some embodiments, the prediction unit 520 determines the optimal MIP mode for predicting the current block based on the strain costs corresponding to multiple MIP modes, Reconstruction sample padding is performed on the reference region adjacent to the outside of the template region to obtain the reference row and reference column of the template region. Using the reference row and reference column as input, the system predicts samples within the template region using each of the multiple MIP modes to obtain multiple prediction blocks corresponding to the multiple MIP modes. It is further configured to determine the strain costs corresponding to multiple MIP modes based on multiple prediction blocks and reconstruction blocks within the template region.
[0274] In some embodiments, the prediction unit 520 is specifically: The input vector is obtained by downsampling the reference row and reference column. By taking an input vector as input and traversing multiple MIP modes, samples within the template region are predicted, and output vectors corresponding to multiple MIP modes are obtained. The system is further configured to upsample the output vectors corresponding to multiple MIP modes to obtain prediction blocks corresponding to multiple MIP modes.
[0275] In some embodiments, the prediction unit 520 is specifically configured to determine an optimal MIP mode based on the sum of absolute differences (SATD) of differential transforms corresponding to a plurality of MIP modes within a template region.
[0276] FIG. 10 is a block diagram showing an encoder 600 according to an embodiment of the present application.
[0277] As shown in FIG. 10, the encoder 600 can include a prediction unit 610, a residual unit 620, and an encoding unit 630. The prediction unit 610 is configured to determine an optimal MIP mode for predicting a current block in a current sequence based on a distortion cost corresponding to an intra prediction (MIP) mode based on a plurality of matrices. The distortion costs corresponding to the plurality of MIP modes include distortion costs obtained by predicting samples within a template region adjacent to the current block using the plurality of MIP modes. The prediction unit 610 is configured to determine a first intra prediction mode, and the first intra prediction mode is determined based on the distortion costs corresponding to the plurality of MIP modes and includes at least one of a sub-optimal MIP mode for predicting the current block, an intra prediction mode derived from a decoder-side intra mode derivation (DIMD) mode, and an intra prediction mode derived from a template-based intra mode derivation (TIMD) mode. The prediction unit 610 is configured to predict the current block based on the optimal MIP mode and the first intra prediction mode to obtain a predicted block of the current block. The residual unit 620 is configured to obtain a residual block of the current block based on the predicted block of the current block and the original block of the current block. The encoding unit 630 is configured to encode the residual block of the current block to obtain a bitstream of the current sequence.
[0278] In some embodiments, the prediction unit 610 is specifically, Get the first flag, If the first flag is used to indicate that it is permitted to predict image blocks in the current sequence using the optimal MIP mode and the first intra-prediction mode, the system is configured to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes. The encoding unit 630 specifically, It is configured to encode the residual block of the current block and the first flag to obtain a bitstream.
[0279] Predicting the current block based on the optimal MIP mode and the first intra-prediction mode to obtain the predicted block for the current block includes the following: In some embodiments, the prediction unit 610 is specifically, If the first flag is used to indicate that it is permitted to predict the image block in the current sequence using the optimal MIP mode and the first intra-prediction mode, then predict the current block based on the optimal MIP mode and the first intra-prediction mode to obtain the first rate distortion cost. Predict the current block based on at least one intra-prediction mode to obtain at least one rate distortion cost, If the first rate distortion cost is less than or equal to the minimum of at least one rate distortion cost, the system is configured to determine the predicted block obtained by predicting the current block based on the optimal MIP mode and the first intra prediction mode as the predicted block of the current block. The encoding unit 630 specifically, It is configured to encode the residual block of the current block, the first flag, and the second flag to obtain a bitstream. If the first rate distortion cost is less than or equal to the minimum of at least one rate distortion cost, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode; if the first rate distortion cost is greater than the minimum of at least one rate distortion cost, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode.
[0280] In some embodiments, the encoding unit 630 specifically, Based on the distortion cost corresponding to multiple MIP modes, the arrangement order of multiple MIP modes is determined. Based on the arrangement order of multiple MIP modes, the coding scheme to be used for the optimal MIP mode is determined. It is configured to encode the residual blocks of the current block, encode the index of the optimal MIP mode based on the coding scheme used for the optimal MIP mode, and obtain a bitstream.
[0281] In some embodiments, the codeword length of the coding scheme used for the first n MIP modes in the sequence order is smaller than the codeword length of the coding scheme used for the MIP mode following the nth MIP mode in the sequence order, and / or The first n MIP modes use a variable-length coding scheme, and the MIP modes following the nth MIP mode use a truncated binary (TB) coding scheme.
[0282] In some embodiments, the prediction unit 610 is specifically, Based on the optimal MIP mode, predict the current block to obtain the first predicted block. Based on the first intra prediction mode, the current block is predicted to obtain the second predicted block. The system is configured to obtain the prediction block for the current block by weighting the first and second prediction blocks based on the optimal MIP mode weight and the first intra prediction mode weight.
[0283] In some embodiments, the prediction unit 610 weights the first and second prediction blocks based on the optimal MIP mode weights and the first intra-prediction mode weights to obtain the prediction block for the current block. If the first intra-prediction mode includes an intra-prediction mode derived from a suboptimal MIP mode or TIMD mode, the weights of the optimal MIP mode and the first intra-prediction mode are determined based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra-prediction mode includes an intra-prediction mode derived from the DIMD mode, the system is further configured to set both the optimal MIP mode weight and the weight of the first intra-prediction mode to a preset value.
[0284] In some embodiments, the prediction unit 610 is specifically configured to predict samples in a template region based on a third flag and a plurality of MIP modes to obtain strain costs corresponding to the plurality of MIP modes under each state of the third flag, the third flag being used to indicate whether or not to transpose the input and output vectors corresponding to the MIP modes. The prediction unit 610 is configured to determine the optimal MIP mode based on the strain costs corresponding to multiple MIP modes under each state of the third flag.
[0285] In some embodiments, the prediction unit 610 is specifically, If the current block size is a preset size, it is configured to determine the optimal MIP mode based on the distortion cost corresponding to multiple MIP modes.
[0286] In some embodiments, the prediction unit 610 is specifically, If the image frame in which the current block is located is an I-frame, and the size of the current block is a preset size, the system is configured to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0287] In some embodiments, the prediction unit 610 is specifically, If the image frame in which the current block is located is a B-frame, the system is configured to determine the optimal MIP mode based on the distortion costs corresponding to multiple MIP modes.
[0288] In some embodiments, the prediction unit 610 determines the optimal MIP mode for predicting the current block in the current sequence based on the strain costs corresponding to multiple MIP modes, Get the MIP mode used for the adjacent block next to the current block, The configuration further determines which MIP mode is used for adjacent blocks, selecting from multiple MIP modes.
[0289] In some embodiments, the prediction unit 610 determines the optimal MIP mode for predicting the current block in the current sequence based on the strain costs corresponding to multiple MIP modes. Reconstruction sample padding is performed on the reference region adjacent to the outside of the template region to obtain the reference row and reference column of the template region. Using the reference row and reference column as input, the system predicts samples within the template region using each of the multiple MIP modes to obtain multiple prediction blocks corresponding to the multiple MIP modes. It is configured to determine the strain costs corresponding to multiple MIP modes based on multiple prediction blocks and reconstruction blocks within the template region.
[0290] In some embodiments, the prediction unit 610 is specifically, The input vector is obtained by downsampling the reference row and reference column. By taking an input vector as input and traversing multiple MIP modes, samples within the template region are predicted, and output vectors corresponding to multiple MIP modes are obtained. The system is configured to upsample output vectors corresponding to multiple MIP modes to obtain prediction blocks corresponding to multiple MIP modes.
[0291] In some embodiments, the prediction unit 610 is specifically, It is configured to determine the optimal MIP mode based on the differential transformation absolute sum (SATD) corresponding to multiple MIP modes within the template region.
[0292] The apparatus embodiments can correspond to method embodiments, and similar descriptions can be found in the method embodiments. To avoid duplication, such descriptions are omitted here. Specifically, the decoder 500 shown in Figure 9 may correspond to an entity that performs method 300 in the embodiments of this application. Furthermore, the aforementioned and other operations and / or functions of each unit in the decoder 500 are used to realize the corresponding processes in each method, such as method 300. Similarly, the encoder 600 shown in Figure 10 may correspond to an entity that performs method 400 in the embodiments of this application. That is, the aforementioned and other operations and / or functions of each unit in the encoder 600 are used to realize the corresponding processes in each method, such as method 400.
[0293] Furthermore, each unit in the decoder 500 or encoder 600 according to the embodiment of this application may be integrated into one or several other units individually or in total, or some of these units may be further divided into several functionally smaller units. This allows similar operations to be achieved without affecting the realization of the technical effects of the embodiment of this application. The units are divided based on their logic functions. In practical applications, the function of one unit may be realized by several units, or the function of several units may be realized by one unit. In other embodiments of this application, the decoder 500 or encoder 600 may include other units, and in practical applications, these functions may be realized by the cooperation of the other units or by the cooperation of several units. According to another embodiment of this application, for example, a decoder 500 or encoder 600 according to an embodiment of this application is constructed by executing a computer program (including program code) capable of executing each step of the corresponding method in a general-purpose computer device such as a general-purpose computer equipped with processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM), thereby realizing the encoding method or decoding method according to an embodiment of this application. The computer program can be recorded, for example, on a computer-readable storage medium, mounted on an electronic device via the computer-readable storage medium, and operated within it, thereby realizing the corresponding method in the embodiment of this application.
[0294] In other words, the above-mentioned unit may be implemented in hardware form, in software form by instructions, or in combination of hardware and software. Specifically, each step of the method embodiment in the embodiments of this application can be completed by an integrated logic circuit in hardware or by instructions in software form in a processor. The steps of the method disclosed in the embodiments of this application can be executed and completed directly by a hardware decoding processor, or by a combination of hardware and software in a decoding processor. Optionally, the software can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium resides in memory. The processor reads information from memory and, together with the processor hardware, completes the steps of the method embodiment.
[0295] Figure 11 is a block diagram showing an electronic device 700 according to an embodiment of this application.
[0296] As shown in Figure 11, the electronic device 700 includes a processor 710 and a computer-readable storage medium 720. , transceiver 730 and It includes at least the following. Note that the processor 710 、 Computer-readable storage medium 720 , and transceiver 730 by bus or other means each otherIt may be connected. The computer-readable storage medium 720 is used to store a computer program 721 containing computer instructions. The processor 710 is used to execute the computer instructions stored in the computer-readable storage medium 720. The processor 710 is the computing core and control core of the electronic device 700 and is suitable for implementing one or more computer instructions, specifically, for implementing a corresponding process or corresponding function by loading and executing one or more computer instructions.
[0297] For example, the processor 710 may also be called a CPU. The processor 710 may include, but is not limited to, a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.
[0298] For example, the computer-readable storage medium 720 may be high-speed RAM memory, or non-volatile memory, such as at least one magnetic disk storage device. Selectively, the computer-readable storage medium 720 may be at least one computer-readable storage medium located away from the processor 710. Specifically, the computer-readable storage medium 720 includes, but is not limited to, volatile memory and / or non-volatile memory. Non-volatile memory may be ROM, programmable ROM (PROM), erasable programmable read-only memory (erasable PROM, EPROM), electrically erasable programmable read-only memory (electrically EPROM, EEPROM), or flash memory. Volatile memory may be RAM that functions as an external high-speed cache. As illustrative but not limited examples, various types of RAM are available, such as static random access memory (static RAM, SRAM), dynamic random access memory (dynamic RAM, DRAM), synchronous dynamic random access memory (synchronous DRAM, SDRAM), double data rate synchronous dynamic random access memory (double data rate SDRAM, DDRSDRAM), enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), synchronous link dynamic random access memory (synch link DRAM, SLDRAM), and direct rambus random access memory (direct rambus RAM, DRRAM).
[0299] In one embodiment, the electronic device 700 may be an encoder or encoding frame according to the embodiment of this application. The computer-readable storage medium 720 stores a first computer instruction. The processor 710 loads and executes the first computer instruction stored in the computer-readable storage medium 720 to realize the corresponding step in the encoding method according to the embodiment of this application. In other words, the first computer instruction in the computer-readable storage medium 720 is loaded and executed by the processor 710 to perform the corresponding step. To avoid redundancy, the explanation is omitted here.
[0300] In one embodiment, the electronic device 700 may be a decoder or decoding framework according to the embodiment of this application. The computer-readable storage medium 720 stores a second computer instruction. The processor 710 loads and executes the second computer instruction stored in the computer-readable storage medium 720 to realize the corresponding step in the decoding method according to the embodiment of this application. In other words, the second computer instruction in the computer-readable storage medium 720 is loaded and executed by the processor 710 to perform the corresponding step. To avoid redundancy, the explanation is omitted here.
[0301] According to another aspect of this application, a coding system is provided in an embodiment of this application. The coding system includes the encoder and decoder described above.
[0302] In another aspect of this application, embodiments of this application provide a computer-readable storage medium (Memory). The computer-readable storage medium is a storage device in the electronic device 700 and is used to store programs and data. For example, the computer-readable storage medium may be a computer-readable storage medium 720. To make it clear, the computer-readable storage medium 720 herein may include an internal storage medium in the electronic device 700, and of course may include an extended storage medium supported by the electronic device 700. The computer-readable storage medium provides a storage space in which the operating system of the electronic device 700 is stored. Furthermore, the storage space stores one or more computer instructions suitable for being loaded and executed by the processor 710, and these computer instructions may be one or more computer programs 721 (including program code).
[0303] According to another aspect of this application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions, which are stored on a computer-readable storage medium, and for example, the computer instructions may be computer program 721. In this case, electronic equipment 700 may be a computer, and the processor 710 reads a computer instruction from the computer-readable storage medium 720, and the processor 710 executes the computer instruction, thereby causing the computer to execute an encoding method or decoding method according to the various selectable methods described above.
[0304] In other words, when implemented by software, all or part of the above embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes described in the embodiments of this application are executed, or the functions described in the embodiments of this application are realized. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device. The computer instructions may be stored on a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, fiber optic cable, digital subscriber line (DSL), etc.) or wirelessly (e.g., infrared, radio, microwave, etc.).
[0305] In conjunction with the exemplary units and process steps described in the embodiments disclosed herein, it will be apparent to those skilled in the art that the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software will depend on the specific application of the invention and design limitations. Those skilled in the art may implement the described functions using different methods for each specific application, but such implementations should not be considered beyond the scope of this application.
[0306] Finally, the above is merely a specific embodiment of the present application, and the scope of protection of this application is not limited thereto. Any modification or substitution that a person skilled in the art could easily conceive within the scope of the art disclosed in this application should be included within the scope of protection of this application. Accordingly, the scope of protection of this application should be determined by the scope of protection of the claims.
Claims
1. A decoding method applied to a decoder, wherein the decoding method is This involves analyzing the bitstream to obtain the residual blocks of the current block, and Determining the optimal MIP mode for predicting the current block based on strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, wherein the strain costs corresponding to the multiple MIP modes include strain costs obtained by predicting samples in a template region adjacent to the current block using the multiple MIP modes. Determining a first intra-prediction mode, wherein the first intra-prediction mode is determined based on the strain costs corresponding to the plurality of MIP modes and includes at least one of the following for predicting the current block: a suboptimal MIP mode, an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode, and an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode. Based on the optimal MIP mode and the first intra prediction mode, predict the current block to obtain a predicted block for the current block. Based on the residual block of the current block and the predicted block of the current block, a reconstructed block of the current block is obtained. including, A decoding method characterized by the following features.
2. Based on the strain costs corresponding to the aforementioned multiple MIP modes, determining the optimal MIP mode for predicting the current block is: The process involves analyzing a bitstream to obtain a first flag, which is a sequence-level flag. When the first flag is used to indicate that it is permitted to predict an image block using the optimal MIP mode and the first intra-prediction mode, the optimal MIP mode is determined based on the distortion costs corresponding to the plurality of MIP modes, including, The decoding method according to feature 1.
3. If the first flag is used to indicate that it is permitted to predict an image block using the optimal MIP mode and the first intra-prediction mode, then determining the optimal MIP mode based on the distortion costs corresponding to the plurality of MIP modes is: If the first flag is used to indicate that it is permitted to predict image blocks using the optimal MIP mode and the first intra prediction mode, then the bitstream is analyzed to obtain the second flag, If the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode, then the optimal MIP mode is determined based on the strain costs corresponding to the plurality of MIP modes, including, The decoding method according to feature 2.
4. Predicting the current block based on the optimal MIP mode and the first intra prediction mode to obtain a predicted block for the current block is: Based on the aforementioned optimal MIP mode, the current block is predicted to obtain a first predicted block. Based on the first intra prediction mode, predict the current block to obtain a second predicted block, Based on the weights of the optimal MIP mode and the weights of the first intra prediction mode, the first prediction block and the second prediction block are weighted to obtain the prediction block of the current block. including, The decoding method according to feature 1.
5. Before obtaining the prediction block of the current block by weighting the first prediction block and the second prediction block based on the weights of the optimal MIP mode and the weights of the first intra prediction mode, the decoding method If the first intra-prediction mode includes an intra-prediction mode derived from the suboptimal MIP mode or the TIMD mode, the weights of the optimal MIP mode and the first intra-prediction mode are determined based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra prediction mode includes an intra prediction mode derived from the DIMD mode, then the weights of the optimal MIP mode and the weights of the first intra prediction mode are both set to preset values. Further including, The decoding method according to feature 4.
6. If the current block size is a preset size, determining the optimal MIP mode based on the strain cost corresponding to the multiple MIP modes is: If the image frame in which the current block is located is an I-frame and the size of the current block is the preset size, then the optimal MIP mode is determined based on the distortion cost corresponding to the plurality of MIP modes. including, The decoding method according to feature 1.
7. Before determining the optimal MIP mode for predicting the current block based on the strain costs corresponding to the plurality of MIP modes, the decoding method: Reconstruction sample padding is performed on the reference region adjacent to the template region to obtain the reference row and reference column of the template region. Using the aforementioned reference row and reference column as input, the system predicts samples within the template area using each of the multiple MIP modes to obtain multiple prediction blocks corresponding to the multiple MIP modes. Based on the plurality of prediction blocks and the reconstruction blocks within the template region, the strain cost corresponding to the plurality of MIP modes is determined. Further including, The decoding method according to feature 1.
8. An encoding method applied to an encoder, wherein the encoding method is Determining the optimal MIP mode for predicting the current block based on strain costs corresponding to multiple matrix-based intra-prediction (MIP) modes, wherein the strain costs corresponding to the multiple MIP modes include strain costs obtained by predicting samples in a template region adjacent to the current block using the multiple MIP modes. Determining a first intra-prediction mode, wherein the first intra-prediction mode is determined based on the strain costs corresponding to the plurality of MIP modes and includes at least one of the following for predicting the current block: a suboptimal MIP mode, an intra-prediction mode derived from a decoder-side intra-mode derivation (DIMD) mode, and an intra-prediction mode derived from a template-based intra-mode derivation (TIMD) mode. Based on the optimal MIP mode and the first intra prediction mode, predict the current block to obtain a predicted block for the current block. Based on the predicted block of the current block and the original block of the current block, the residual block of the current block is obtained. Encoding the residual block of the current block to obtain a bitstream, including, An encoding method characterized by the following features.
9. Based on the strain costs corresponding to the aforementioned multiple MIP modes, determining the optimal MIP mode for predicting the current block is: The first flag is obtained, and the first flag is a sequence-level flag. If the first flag is used to indicate that it is permitted to predict an image block using the optimal MIP mode and the first intra-prediction mode, the optimal MIP mode is determined based on the distortion costs corresponding to the plurality of MIP modes, Encoding the residual block of the current block to obtain a bitstream is: This includes encoding the residual block of the current block and the first flag to obtain the bitstream, The encoding method according to feature 8.
10. Predicting the current block based on the optimal MIP mode and the first intra prediction mode to obtain a predicted block for the current block is: If the first flag is used to indicate that it is permitted to predict an image block using the optimal MIP mode and the first intra-prediction mode, then predict the current block based on the optimal MIP mode and the first intra-prediction mode to obtain a first rate distortion cost, Predicting the current block based on at least one intra-prediction mode to obtain at least one rate distortion cost, If the first rate distortion cost is less than or equal to the minimum value of the at least one rate distortion cost, the predicted block obtained by predicting the current block based on the optimal MIP mode and the first intra prediction mode is determined to be the predicted block of the current block. Includes, Encoding the residual block of the current block and the first flag to obtain the bitstream is: This includes encoding the residual block of the current block, the first flag, and the second flag to obtain the bitstream, If the first rate distortion cost is less than or equal to the minimum of the at least one rate distortion cost, the second flag is used to indicate that it is permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode; if the first rate distortion cost is greater than the minimum of the at least one rate distortion cost, the second flag is used to indicate that it is not permitted to predict the current block using the optimal MIP mode and the first intra-prediction mode. The encoding method according to feature 9.
11. Predicting the current block based on the optimal MIP mode and the first intra prediction mode to obtain a predicted block for the current block is: Based on the aforementioned optimal MIP mode, the current block is predicted to obtain a first predicted block. Based on the first intra prediction mode, predict the current block to obtain a second predicted block, Based on the weights of the optimal MIP mode and the weights of the first intra prediction mode, the first prediction block and the second prediction block are weighted to obtain the prediction block of the current block. including, The encoding method according to feature 8.
12. Before obtaining the prediction block of the current block by weighting the first prediction block and the second prediction block based on the weights of the optimal MIP mode and the weights of the first intra prediction mode, the encoding method: If the first intra-prediction mode includes an intra-prediction mode derived from the suboptimal MIP mode or the TIMD mode, the weights of the optimal MIP mode and the first intra-prediction mode are determined based on the strain cost corresponding to the optimal MIP mode and the strain cost corresponding to the first intra-prediction mode. If the first intra prediction mode includes an intra prediction mode derived from the DIMD mode, then the weights of the optimal MIP mode and the weights of the first intra prediction mode are both set to preset values. Further including, The encoding method according to feature 11.
13. It is a decoder, It comprises a processor and memory in which computer programs are stored. When the computer program is executed by the processor, the computer program causes the processor to execute the decoding method described in any one of claims 1 to 7. A decoder characterized by the following features.
14. It is an encoder, It comprises a processor and memory in which computer programs are stored. When the computer program is executed by the processor, the computer program causes the processor to execute the encoding method described in any one of claims 8 to 12. An encoder characterized by the following features.
15. A method for transmitting a bitstream, The bitstream is generated by performing the encoding method described in any one of claims 8 to 12, The transmission of the aforementioned bitstream, including, A method for transmitting a bitstream characterized by the following.