Intra template matching prediction method, video encoding / decoding method, device and system
Intra template matching prediction methods improve video encoding and decoding by constructing candidate lists for reference blocks, addressing the challenge of high bandwidth in high-resolution video compression.
Patent Information
- Application Number
- JP2025538824
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2026-01-27
AI Technical Summary
Existing digital video compression standards face challenges in further reducing bandwidth and traffic pressure due to the increasing demand for high-resolution video, particularly in handling spatial and temporal redundancy within and between frames.
Intra template matching prediction methods are employed to construct a candidate list for video encoding and decoding, utilizing a search range to find reference blocks and calculate differences, determining N reference blocks and their order, and optimizing coding costs for improved efficiency.
Enhances coding efficiency by reducing redundancy through intra template matching prediction, leading to better compression performance and reduced bandwidth requirements.
Smart Images

Figure 2026502984000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE Embodiments of the present disclosure relate to, but are not limited to, video technology, and more particularly to intra template matching prediction methods, video encoding and decoding methods, devices, and systems. [Background technology]
[0002] Digital video compression technology primarily compresses massive amounts of digital video data for easier transmission and storage. Currently, most common video coding standards (e.g., H.266 / Versatile Video Coding (VVC)) employ a block-based hybrid coding framework. Each frame in a video is divided into square largest coding units (LCUs) of the same size (e.g., 128x128, 64x64, etc.). Each LCU can be divided into rectangular coding units (CUs) according to a set of rules. Coding units may be further divided into prediction units (PUs) and transform units (TUs). The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filters. The prediction module includes intra-prediction and inter-prediction for reducing or removing redundancy within a video. Inter-prediction involves motion estimation and motion compensation. Because there is strong correlation between adjacent pixels in a single frame of a video, video coding techniques use intra-prediction to remove spatial redundancy between adjacent pixels. Because there is strong similarity between adjacent frames in a video, video coding techniques use inter-prediction to remove temporal redundancy between adjacent frames, improving coding efficiency. Residual information from the predicted signal undergoes block-based transformation, quantization, and entropy coding to become a bitstream.
[0003] With the rapid growth of Internet video and the increasing demand for video resolution, although existing digital video compression standards can save a lot of video data, better digital video compression technologies are needed to further reduce the bandwidth and traffic pressure of digital video transmission. Summary of the Invention
[0004] The following is a summary of the subject matter described in detail herein, but is not intended to limit the scope of protection of the claims. [Means for solving the problem]
[0005] An embodiment of the present disclosure provides a method for constructing a candidate list for intra template matching prediction, the method including: determining a first search range for performing intra template matching prediction (intraTMP) on a current block; searching for a reference block template based on the first search range; calculating a difference between the searched reference block template and a current block template, where the reference block template has a one-to-one correspondence with a reference block; constructing a candidate list for intraTMP based on the difference; and determining N reference blocks in the candidate list and an order of the N reference blocks, where N≧2.
[0006] An embodiment of the present disclosure further provides a video decoding method, the method comprising: The method includes decoding an intra template matching prediction intraTMP mode usage flag of the current block, and if it is determined based on the intraTMP mode usage flag that the current block uses intraTMP mode, further decoding an intraTMP index of the current block, where the intraTMP index indicates the position of a reference block used by the current block in an intraTMP candidate list, constructing a candidate list, determining a reference block used by the current block based on the intraTMP index and the candidate list, and performing intra prediction on the current block based on the reference block used by the current block.
[0007] An embodiment of the present disclosure further provides a video encoding method, which includes: when it is determined that a current block allows use of a multi-candidate intra template matching prediction intraTMP mode, constructing an intraTMP candidate list according to any of the intraTMP candidate list construction methods of the present disclosure, where the candidate list includes N reference blocks, where N is greater than or equal to 2; calculating coding costs for predicting the current block based on the N reference blocks in the candidate list, and performing rate-distortion optimization of the current block using the smallest coding cost as the coding cost for the multi-candidate intraTMP mode; and when it is determined through rate-distortion optimization that the current block is to be intra-predicted using the multi-candidate intraTMP mode, encoding syntax elements related to the multi-candidate intraTMP mode for the current block.
[0008] An embodiment of the present disclosure further provides a candidate list construction device for intra template matching prediction, the device including a processor and a memory storing a computer program, the processor being capable of implementing the candidate list construction method for intra template matching prediction according to any embodiment of the present disclosure when executing the computer program.
[0009] An embodiment of the present disclosure further provides a video decoding device, the device including a processor and a memory storing a computer program, the processor being capable of implementing the video decoding method according to any embodiment of the present disclosure when executing the computer program.
[0010] An embodiment of the present disclosure further provides a video encoding device, the device including a processor and a memory storing a computer program, the processor being capable of implementing the video encoding method according to any of the embodiments of the present disclosure when executing the computer program.
[0011] An embodiment of the present disclosure further provides a video encoding and decoding system, which includes a video encoding device according to any embodiment of the present disclosure and a video decoding device according to any embodiment of the present disclosure.
[0012] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium having a computer program stored therein, the computer program being capable of implementing a method according to any of the embodiments of the present disclosure when executed by a processor.
[0013] An embodiment of the present disclosure further provides a computer program product, which includes a computer program that, when executed by a processor, is capable of implementing a method according to any embodiment of the present disclosure.
[0014] An embodiment of the present disclosure further provides a method for determining a search range of an intraTMP, which includes, when a current block allows use of the intraTMP mode, determining a first search distance in a width direction and a second search distance in a height direction of the first search range relative to a base point representing the position of the current block: A method including: calculating the product of the width of the current block and a first proportionality coefficient, and determining the larger value of this product and a minimum search distance in a predetermined width direction as the first search distance; and calculating the product of the height of the current block and a second proportionality coefficient, and determining the larger value of this product and a minimum search distance in a predetermined height direction as the second search distance, wherein the first proportionality coefficient and the second proportionality coefficient are equal or different; or A method including: determining the larger of the width of the current block and the minimum search distance in a predetermined width direction, and determining the product of the larger value and a first proportionality coefficient as the first search distance; and determining the larger of the height of the current block and the minimum search distance in a predetermined height direction, and determining the product of the larger value and a second proportionality coefficient as the second search distance, wherein the first proportionality coefficient and the second proportionality coefficient are equal or different; or a method including: obtaining the first search distance by multiplying a width of a current block by a corresponding first proportional coefficient, the first proportional coefficient being plural, and the larger the first proportional coefficient, the larger the width of the corresponding current block; and obtaining the second search distance by multiplying a height of the current block by a corresponding second proportional coefficient, the second proportional coefficient being plural, and the larger the second proportional coefficient, the larger the height of the corresponding current block; This includes determining by
[0015] Other aspects will become apparent after reading and understanding the drawings and detailed description. [Brief explanation of the drawings]
[0016] The drawings provide an understanding of the embodiments of the present disclosure, constitute a part of the specification, and are used to explain the technical means of the present disclosure together with the embodiments of the present disclosure, but are not intended to limit the technical means of the present disclosure.
[0017] [Figure 1A] FIG. 1 is a schematic diagram of an encoding and decoding system according to an embodiment of the present disclosure. [Figure 1B] FIG. 1 is a framework diagram of an encoding terminal according to one embodiment of the present disclosure. [Figure 1C] FIG. 1 is a framework diagram of a decoding terminal according to an embodiment of the present disclosure. [Figure 2A] 1 is a schematic diagram illustrating predicting a current block using an intra prediction method; [Figure 2B] 1 is a schematic diagram illustrating a current block predicted using a multi-reference row intra-prediction method; [Figure 3] FIG. 1 is a schematic diagram of a conventional intra mode used in a non-wide-angle mode in VVC. [Figure 4] FIG. 1 is a schematic diagram of a conventional intra mode used in a wide-angle mode in VVC. [Figure 5] FIG. 1 is a schematic diagram of the conventional intra mode used in AVS3. [Figure 6] FIG. 1 is a schematic diagram illustrating intra prediction based on IBC mode. [Figure 7] FIG. 1 is a schematic diagram illustrating inter prediction based on template matching technology. [Figure 8] FIG. 1 is a schematic diagram illustrating intra prediction based on the intraTMP mode. [Figure 9] 1 is a flowchart of a method for constructing a candidate list for intraTMP according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic diagram of setting a search distance according to one embodiment of the present disclosure. [Figure 11] 1 is a flowchart of a video decoding method according to an embodiment of the present disclosure. [Figure 12] 1 is a flowchart of a video encoding method according to an embodiment of the present disclosure. [Figure 13A] FIG. 10 is a schematic diagram of the position indicated by the BV during the first stage of searching according to one embodiment of the present disclosure. [Figure 13B] FIG. 10 is a schematic diagram illustrating determining a local search range in a second stage search based on a BV reserved in a first stage search, according to an embodiment of the present disclosure. [Figure 14] FIG. 2 is a module diagram of an intra-prediction device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018] While multiple embodiments are described in this disclosure, these descriptions are illustrative and not limiting, and those skilled in the art will recognize that there may be many more embodiments and implementations within the scope encompassed by the embodiments described in this disclosure.
[0019] In describing the present disclosure, terms such as "exemplary" or "for example" are used to express an example, illustration, or explanation. An embodiment described as "exemplary" or "for example" in the present disclosure should not be construed as being preferred or superior to other embodiments. In this specification, "and / or" describes a relationship between related objects and indicates that three types of relationships may exist. For example, "A and / or B" can represent three cases: when only A exists, when both A and B exist, and when only B exists. "Plural" refers to two or more cases. Furthermore, to clearly describe the technical means of the embodiments of the present disclosure, terms such as "first" and "second" are used to distinguish between identical or similar objects that have essentially the same function and operation. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and do not necessarily mean different things.
[0020] As used herein, "including any one or more of the following: Option 1, Option 2, ..." or "including any one or more of Option 1, Option 2, ..." means including any one of the listed options or any combination of multiple items of the listed options. For example, "including any one or more of the following: A, B" or "including any one or more of A and B" means including only A, including only B, or including both A and B. For further example, "including any one or more of the following: A, B, C" or "including any one or more of A, B, C" means including only A, including only B, including only C, including A and B, including A and C, including B and C, or including all of A, B, and C. The same analogy applies to cases with more options.
[0021] In describing representative illustrative embodiments, the specification may have previously expressed a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not depend on the particular order of steps described herein, the method or process should not be limited to the particular order of steps described. As one of ordinary skill in the art would readily understand, other order of steps are possible. Thus, the particular order of steps described in the specification should not be construed as a limitation on the scope of the claims. Furthermore, the scope of the method and / or process claims should not be limited to performing the steps in the order described; as one of ordinary skill in the art would readily understand, these orders can be varied while remaining within the spirit and scope of the embodiments of the present disclosure.
[0022] The intra prediction method and video encoding / decoding method according to the embodiments of the present disclosure may be applicable to various video encoding standards, such as H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), AVS (Audio Video coding Standard), and other standards established by MPEG (Moving Picture Experts Group), AOM (Alliance for Open Media), JVET (Joint Video Experts Team), extensions of these standards, or other custom standards.
[0023] FIG. 1A is a block diagram of a video encoding / decoding system applicable to an embodiment of the present disclosure. As shown in the drawing, the system includes an encoding end 1 and a decoding end 2. The encoding end 1 generates a bitstream. The decoding end 2 can decode the bitstream. The decoding end 2 can receive the bitstream from the encoding end 1 via a link 3. The link 3 includes one or more media or devices for transferring the bitstream from the encoding end 1 to the decoding end 2. For example, the link 3 includes one or more communication media that enable the encoding end 1 to directly transmit the bitstream to the decoding end 2. The encoding end 1 modulates the bitstream according to a communication standard and transmits the modulated bitstream to the decoding end 2. The one or more communication media may include wireless and / or wired communication media and may form part of a packet network. As another example, the bitstream may be output to a storage device via an output interface 15, and the decoding end 2 can read the stored data from the storage device by streaming or downloading.
[0024] As shown in the figure, the encoding terminal 1 includes a data source 11, a video encoding device 13, and an output interface 15. The data source 11 may include a video capture device (e.g., a camera), an archive containing previously captured data, a feed interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. The video encoding device 13, also referred to as a video encoding terminal, is configured to encode data from the data source 11 and output it to the output interface 15. The output interface 15 may include at least one of a conditioner, a modem, and a transmitter. The decoding terminal 2 includes an input interface 21, a video decoding device 23, and a display device 25. The input interface 21 includes at least one of a receiver and a modem. The input interface 21 can receive a bitstream via link 3 or from a storage device. The video decoding device 23, also referred to as a video decoding terminal, is configured to decode the received bitstream. The display device 25 is configured to display the decoded data. The display device 25 may be integrated with other devices in the decoding terminal 2 or may be installed independently. The display device 25 is optional for the decoding terminal. In other examples, the decoding end may include other devices or equipment that apply the decoded data.
[0025] 1B is a block diagram of an exemplary video encoding device applicable to embodiments of the present disclosure. As shown in the drawing, the video encoding device 10 includes a segmentation unit 101, a prediction unit 100, a residual generation unit 102, a transform processing unit 104, a quantization unit 106, an inverse quantization unit 108, an inverse transform unit 110, a reconstruction unit 112, a filter unit 113, a decoded image buffer 114, and an entropy encoding unit 115.
[0026] The division unit 101 cooperates with the prediction unit 100 to divide the received video data into slices, coding tree units (CTUs), or other relatively large units. The received video data may be a video sequence including video frames such as I frames, P frames, and B frames.
[0027] The prediction unit 100 divides the CTU into coding units (CUs) and performs intra-prediction coding or inter-prediction coding on the CUs. When performing intra-prediction and inter-prediction on a CU, the CU can be divided into one or more prediction units (PUs).
[0028] The prediction unit 100 includes an inter prediction unit 121 and an intra prediction unit 126 .
[0029] The inter prediction unit 121 is configured to perform inter prediction on the PU and generate prediction data for the PU. The prediction data includes a prediction block for the PU, motion information for the PU, and various syntax elements. The inter prediction unit 121 may include a motion estimation (ME) unit and a motion compensation (MC) unit. The motion estimation unit may be configured to perform motion estimation to generate a motion vector, and the motion compensation unit may be configured to obtain or generate a prediction block based on the motion vector.
[0030] The intra prediction unit 126 is configured to perform intra prediction on the PU and generate prediction data for the PU. The prediction data for the PU includes a prediction block and various syntax elements for the PU.
[0031] The residual generation unit 102 (indicated in the drawing by a circle with a plus sign after the division unit 101) is configured to subtract the predicted block of the PU obtained by dividing the CU from the original block of the CU to generate a residual block of the CU.
[0032] The transform processing unit 104 is configured to divide a CU into one or more transform units (TUs). The division of prediction units and transform units may be different. A residual block associated with a TU is a sub-block obtained by dividing the residual block of the CU. One or more transforms are applied to the residual block associated with the TU to generate a coefficient block associated with the TU.
[0033] The quantization unit 106 is configured to quantize the coefficients in the coefficient block based on a quantization parameter, and can change the degree of quantization of the coefficient block by adjusting the quantizer parameter (QP).
[0034] Inverse quantization unit 108 and inverse transform unit 110 are configured to apply inverse quantization and inverse transform, respectively, to the coefficient block to obtain a reconstructed residual block associated with the TU.
[0035] The reconstruction unit 112 (denoted in the drawing by a circle with a plus sign after the inverse transform processing unit 110) is configured to add the reconstructed residual block and the prediction block generated by the prediction unit 100 to generate a reconstructed image.
[0036] The filter unit 113 is configured to perform a loop filter on the reconstructed image.
[0037] The decoded image buffer 114 stores the loop-filtered reconstructed image. The intra prediction unit 126 can extract reference images of blocks adjacent to the current block from the decoded image buffer 114 to perform intra prediction. The inter prediction unit 121 can perform inter prediction on a PU of the current frame image using a reference image of a previous frame cached in the decoded image buffer 114.
[0038] The entropy coding unit 115 is configured to perform entropy coding operations on the received data (eg, syntax elements, quantized coefficient blocks, motion information, etc.) to generate a video bitstream.
[0039] In other examples, video encoding device 10 may include more, fewer, or different functional components than those in this example, such as omitting transform processing unit 104, inverse transform processing unit 110, etc.
[0040] 1C is a block diagram of an exemplary video decoding device applicable to embodiments of the present disclosure. As shown in the drawing, the video decoding device 15 includes an entropy decoding unit 150, an inverse quantization unit 154, an inverse transform processing unit 156, a prediction unit 152, a reconstruction unit 158, a filter unit 159, and a decoded image buffer 160.
[0041] The entropy decoding unit 150 is configured to perform entropy decoding on the received coded video bitstream to extract syntax elements, quantized coefficient blocks, motion information of PUs, etc. The prediction unit 152, the inverse quantization unit 154, the inverse transform processing unit 156, the reconstruction unit 158, and the filter unit 159 can all perform corresponding operations based on the syntax elements extracted from the bitstream.
[0042] Inverse quantization unit 154 is configured to perform inverse quantization on coefficient blocks associated with quantized TUs.
[0043] Inverse transform processing unit 156 is configured to apply one or more inverse transforms to the dequantized coefficient blocks to produce reconstructed residual blocks of the TUs.
[0044] The prediction unit 152 includes an inter prediction unit 162 and an intra prediction unit 164. If the current block uses intra prediction coding, the intra prediction unit 164 determines the intra prediction mode of the PU based on syntax elements decoded from the bitstream, and performs intra prediction in combination with reconstructed reference information neighboring the current block obtained from the decoded image buffer 160. If the current block uses inter prediction coding, the inter prediction unit 162 determines a reference block of the current block based on the motion information of the current block and corresponding syntax elements, and performs inter prediction on the reference block obtained from the decoded image buffer 160.
[0045] The reconstruction unit 158 (indicated in the drawing by a circle with a plus sign after the inverse transform processing unit 155) is configured to obtain a reconstructed image based on a reconstructed residual block associated with the TU and a predicted block of the current block generated by intra- or inter-prediction performed by the prediction unit 152.
[0046] The filter unit 159 is configured to perform a loop filter on the reconstructed image.
[0047] The decoded image buffer 160 is configured to store the loop-filtered reconstructed image to be used as a reference image for subsequent motion compensation, intra-prediction, inter-prediction, etc., and can also output the filtered reconstructed image as decoded video data for display on a display device.
[0048] In other embodiments, video decoder 15 may include more, fewer, or different functional components. For example, inverse transform processing unit 155 may be omitted in certain cases.
[0049] Based on the above-described video encoding device and video decoding device, the following basic encoding and decoding processes can be performed. At the encoding end, a frame image is divided into blocks or into multiple slices, which are then further divided into blocks. Slices within the same image can be processed in parallel. A current block is subjected to intra-prediction, inter-prediction, or other algorithm to generate a predicted block of the current block. The predicted block is subtracted from the current block's original block to obtain a residual block. The residual block is then transformed and quantized to obtain a quantized coefficient matrix. The quantized coefficient matrix is then entropy-coded to generate a bitstream. At the decoding end, intra-prediction or inter-prediction is performed on the current block to generate a predicted block of the current block. Meanwhile, the quantized coefficient matrix obtained by decoding the bitstream is subjected to inverse quantization and inverse transform to obtain a residual block. The predicted block and residual block are then added to obtain a reconstructed block. The reconstructed block forms a reconstructed image, and a loop filter is applied to the reconstructed image based on the image or block to obtain a decoded image. The encoding end also obtains a decoded image (also called a reconstructed image after loop filtering) through the same operations as the decoding end. The reconstructed image after the loop filter can be used as a reference frame for inter-prediction of a subsequent frame. Mode information and parameter information, such as block partition information, prediction, transform, quantization, entropy coding, and loop filter, determined by the encoding end can be written into a bitstream. The decoding end decodes the bitstream or analyzes it based on existing information to determine mode information and parameter information, such as block partition information, prediction, transform, quantization, entropy coding, and loop filter, used by the encoding end. This ensures that the decoded image obtained by the encoding end is the same as the decoded image obtained by the decoding end.
[0050] Although the above illustrates an example of a block-based hybrid coding framework, embodiments of the present disclosure are not limited thereto. As technology advances, one or more modules within this framework and one or more steps in this process may be replaced or optimized. Embodiments of the present disclosure relate to, but are not limited to, the above-mentioned intra prediction units and corresponding intra prediction methods at the encoding and decoding ends.
[0051] In this specification, the current block may be a coding unit (CU) or prediction unit (PU) currently being coded or decoded, or a block-level coding unit such as a sub-block obtained by dividing a CU or PU.
[0052] <Intra prediction> The intra prediction method predicts the current block using reconstructed pixels that have already been coded and decoded around the current block as reference pixels. For example, as shown in FIG. 2A, a 4x4 block in the diagram is the current block, and the pixels in one column to the left and one row above the current block are the reference pixels of the current block. Intra prediction predicts the current block using these reference pixels. These reference pixels may all have already been coded and decoded, but some may be unavailable. For example, if the current block is located at the leftmost edge of the entire frame, the reference pixels to the left of the current block may be unavailable. Or, if the lower left portion of the current block has not yet been coded and decoded when coding and decoding the current block, the reference pixels to the lower left may also be unavailable. If reference pixels are unavailable, they may be filled in using available reference pixels, a specific value, or a specific method, or no filling may be performed.
[0053] The multi-reference line (MRL) intra prediction method can improve coding efficiency by using more reference pixels. Figure 2B shows an example using four reference rows / columns.
[0054] <Conventional intra prediction mode> There are multiple intra prediction modes, and as technology advances and blocks become larger, the number of prediction modes continues to increase. For example, HEVC uses 35 intra prediction modes, including planar mode, DC mode, and 33 angle modes. As shown in Figure 3, VVC uses 67 intra prediction modes, including planar mode, DC mode, and 65 angle modes. In addition to these 67 modes, VVC also provides wide-angle modes, such as the dashed modes in Figure 4 (i.e., two ranges from -14 to -1 and from 67 to 80), for rectangular blocks with a relatively large difference between their length and width. These modes replace some of the standard modes. AVS3 uses 66 prediction modes, including DC mode, planar mode, bilinear mode, pulse-code modulation mode (PCM mode), and 62 angle modes, as shown in Figure 5.
[0055] <Inter prediction> Videos are composed of multiple images. To display videos smoothly, videos contain tens to hundreds of frames per second, such as 24, 30, 50, 60, or 120 frames per second. As a result, videos have significant temporal redundancy, or in other words, significant temporal correlation. Inter-frame prediction exploits this temporal correlation to improve compression efficiency. Inter-frame prediction often exploits temporal correlation by utilizing "motion." A very simple "motion" model is when an object is at a specific position in a picture corresponding to a particular point in time, but after a certain amount of time has passed, it translates to a different position in the picture corresponding to that point in time. This is the translation in video encoding and decoding. Inter-frame prediction uses motion information to represent "motion." Basic motion information includes information on a reference frame (or reference picture) and a motion vector (MV). The codec determines a reference image based on information about the reference image, and determines the coordinates of a reference block based on information about a motion vector and the coordinates of a current block. In the reference image, the reference block is determined using the coordinates of the reference block. Using the determined reference block as a predicted block is the most basic prediction method of inter prediction.
[0056] Not all motion in video is this simple. Even motion that can be considered translational can undergo subtle changes over time (such as slight deformations, brightness changes, and noise changes). Better prediction results can be achieved by using multiple reference blocks to predict the current block. For example, in currently commonly used bidirectional prediction, two reference blocks are used to predict the current block. The two reference blocks can be one forward reference block and one backward reference block. Later, it became possible to use both forward reference blocks or both backward reference blocks. Future video coding standards may support prediction using multiple reference blocks. One simple way to generate a prediction block using two reference blocks is to average the pixel values of corresponding positions in the two reference blocks to obtain the prediction block. To achieve better prediction results, weighted averaging, such as BCW (Bi-prediction with CU-level weighting), currently used in VVC, can also be used. GPM (Geometric Partitioning Mode) in VVC can also be considered a special type of bidirectional prediction. To use bidirectional prediction, it is of course necessary to find two reference blocks, which requires information on two sets of reference images and motion vectors.
[0057] <Intra Block Copy (IBC)> Intra Block Copy (IBC) technology can significantly improve the compression efficiency of screen content coding. Therefore, IBC mode is always used for screen content coding, from HEVC to VVC. Screen content is computer-generated, unlike camera-captured content. Screen content is noise-free, contains text, computer graphics, and other elements, and has clear boundaries. Screen content also contains a large amount of overlapping content. As shown in Figure 6, inter prediction uses a reference block in a reference image as the prediction block for the current block, but the reference image is not the current image. On the other hand, IBC mode applies the inter prediction method to intra prediction, finding a block from the already coded and decoded part of the current image (also known as the reconstructed part) and using it as the prediction block for the current block. IBC mode is also known as intra picture block compensation mode or current picture referencing (CPR) mode.
[0058] IBC mode uses a block vector (BV) to represent the positional difference between the current block and a reference block. This is similar to the motion vector (MV) used in inter-prediction. The encoding stage determines the best matching block for the current block using a block matching method within a search range and encodes the BV. IBC can be considered a type of intra-prediction method, but it can also be considered a separate type of prediction method independent of intra-prediction and inter-prediction.
[0059] Template Matching (TM) Template matching (TM) technology was first used in inter prediction, where it exploits the correlation between neighboring pixels to select a region surrounding the current block as a template. When encoding and decoding a current block, the left and upper sides of the current block have already been encoded and decoded according to the encoding order. In actual hardware implementations, it may not be possible to guarantee that the left and upper sides of the current block have already been decoded when decoding begins. For example, in HEVC, the prediction process for an inter block can be performed in parallel because neighboring reconstructed pixels are not required when generating a prediction block for an inter-coded block. On the other hand, for intra-coded blocks, the reconstructed pixels on the left and upper sides must be used as reference pixels. By appropriately adjusting the hardware design, the reconstructed pixels on the left and upper sides of the current block can be used. However, the reconstructed pixels on the right and lower sides of the current block are not available in the encoding order of current standards such as VVC.
[0060] As shown in FIG. 7, rectangular regions to the left and top of the current block are set as templates. The height of the left template is generally the same as the height of the current block, and the width of the top template is generally the same as the width of the current block, but may differ. The best matching position of the template is searched for in the reference frame, thereby determining the motion information (i.e., motion vector) of the current block. This process can be roughly described as starting from a starting position in a reference frame and searching within a predetermined surrounding range. Search rules such as the search range and search step width can be set in advance. Each time a position is moved to, the degree of match between the template corresponding to that position and the templates surrounding the current block is calculated. The degree of match can be measured by differences (e.g., sum of absolute difference (SAD), sum of absolute transformed difference (SATD), mean-square error (MSE), etc.), with smaller values of SAD, SATD, MSE, etc. indicating a higher degree of match. A cost is calculated based on the template's predicted block corresponding to the position and the template's reconstructed blocks around the current block. The motion information for the current block is determined based on the position with the highest matching degree of the searched template. By utilizing the correlation between neighboring pixels, the motion information suitable for the template may also be suitable for the current block.
[0061] Because template matching cannot be applied to all blocks, a means for determining whether the current block uses template matching can be used. For example, a control switch can be used to indicate whether the current block uses template matching. A traditional template matching technique is decoder-side motion vector derivation (DMVD). Both the encoding and decoding ends can derive motion information by searching using a template, or find better motion information based on existing motion information. This does not require the transmission of specific motion vectors or motion vector differences. Instead, the encoding and decoding ends perform searches according to the same rules, ensuring consistency between the encoding and decoding. While template matching can improve compression performance, it also requires searching at the decoding end, which increases the complexity of the decoding end.
[0062] <Intra-template matching prediction (intraTMP)> Intra template matching prediction (intraTMP) is a technology that combines IBC and TM. Just as applying TM to an inter frame reduces the coding overhead of the motion vector, applying TM to an IBC frame reduces the coding overhead of the block vector. As an example, without coding the block vector, the block with the highest matching degree found based on the TM is directly used as the prediction block for the intraTMP mode of the current block, and is included in the rate-distortion optimization to determine the intra prediction mode to be used for the current block.
[0063] As an example of intraTMP, as shown in FIG. 8, the inverted L-shaped area at the top left of the current block is the template for the current block. The local area R1 of the current CTU, the CTU area R2 at the top left of the current block, the CTU area R3 above the current block, and the CTU area R4 to the left of the current block are the reconstruction areas available for search. The reference block template should be searched within this available reconstruction area, but the actual search range may be smaller than this reconstruction area. However, this is merely an example, and in actual applications, the available reconstruction area may be different. In the example shown in the drawing, the search found the best matching block within area R2 (the reference block corresponding to the reference block template with the smallest difference from the current block template, i.e., the highest degree of matching). The reference block within area R2 in the drawing is the best matching block, and the shaded area surrounding the left and top of this reference block is the template for this reference block (also referred to as the reference block template corresponding to the reference block).
[0064] As mentioned above, IBC can significantly improve the compression efficiency of screen content encoding. One major reason for this is that screen content often contains many overlapping blocks, has clear boundaries, and, in terms of color, areas of the same color (luminance and chrominance) are connected. However, this situation rarely exists in camera-captured content. Camera-captured content inevitably contains noise, and while some areas of camera-captured content may appear uniform in color at first glance, there are slight variations in luminance and chrominance, and camera-captured content rarely has clear boundaries. Furthermore, due to factors such as viewing angles, it is difficult to find identical blocks in camera-captured content, but repeated textures do exist. Camera-captured content does indeed contain similar overlapping blocks (blocks with noise or slight variations in luminance).
[0065] IntraTMP determines the best matching block found by intra template matching as the final predicted block. That is, when decoding a current block, whether or not the current block uses intraTMP can be determined by decoding a single flag. If the current block uses intraTMP, the decoder searches for a best matching block within the reconstruction region using intra template matching and uses the reconstructed value of the best matching block as the predicted value of the current block. Because the decoder does not have the original value of the current block during the search, it can only use the block with the highest template match as the best matching block found by intraTMP. However, because the template is highly correlated with the current block but not the current block itself, the best matching block found by template matching may not necessarily be the best matching block for the actual current block. There is room for further improvement in coding efficiency in intraTMP mode.
[0066] In view of this, one embodiment of the present disclosure provides an intra prediction method for intraTMP. As shown in Figure 9, the method includes the following steps:
[0067] In step S110, a search range for performing intra template matching prediction intraTMP on the current block is determined.
[0068] In step S120, a reference block template is searched based on the search range, and the difference between the searched reference block template and the current block template is calculated. The reference block template has one-to-one correspondence with the reference block.
[0069] In step S130, a candidate list of intraTMP is constructed based on the difference, and N reference blocks in the candidate list and an order of the N reference blocks are determined, where N≧2.
[0070] The "reference block template corresponding to one reference block" referred to in this specification refers to the template of the reference block. As shown in FIG. 8, one reference block and its template are shown in the R2 region. The reference block template is represented by the shaded area in the drawing and is also referred to as the reference block template corresponding to the reference block. The size and shape of the reference block template are the same as those of the current block template, and the relative positional relationship between the reference block template and the reference block is also the same as that between the current block template and the current block. In the example shown in the drawing, the current block template is an L-shaped region surrounding the left and top sides of the current block, and the reference block template is an L-shaped region surrounding the left and top sides of the reference block. The embodiments of the present disclosure do not limit the number of rows and columns contained in the reference block template and the current block template. Furthermore, the template of a block may extend to the upper right and / or lower left sides of the block.
[0071] An embodiment of the present disclosure constructs a candidate list including multiple reference blocks using difference-based template matching. Any of the multiple reference blocks in the candidate list may be used as a reference block when intra-predicting a current block in intraTMP mode. The first option in this candidate list is the best matching block determined by template matching. However, using this best matching block to predict the current block may not necessarily result in the best coding efficiency. Using other reference blocks in the candidate list for prediction may result in better overall coding efficiency. By constructing the candidate list, a single index can be used to indicate the reference block with the highest coding efficiency. If the encoding end constructs the candidate list in a similar manner and finds the reference block indicated by the index based on the index to predict the current block, coding efficiency can be improved.
[0072] In one exemplary embodiment of the present disclosure, the search range is located in a reconstruction region of the current image, and the difference between the reference block template and the current block template is determined based on the SAD, SATD, or MSE between the reconstruction values of the reference block template and the reconstruction values of the current block template. In this embodiment, the difference between the reference block template and the current block template is expressed using the SAD, SATD, or MSE between the reconstruction values of the reference block template and the reconstruction values of the current block template, thereby reflecting the similarity between the two templates.
[0073] In one exemplary embodiment of the present disclosure, the order of the N reference blocks in the candidate list is determined based on the ascending order of the difference between the corresponding reference block templates. Because the templates of the reference blocks have a strong correlation with the reference block, if the reference block template with the highest similarity determined by template matching (i.e., calculating the template difference) is the reference block with the highest similarity to the current block, the corresponding reference block will also likely be the reference block with the highest similarity to the current block. Therefore, the embodiment of the present disclosure determines the order of the N reference blocks in the candidate list based on the ascending order of the difference between the corresponding reference block templates, thereby increasing the probability that a candidate reference block in the front will be selected and saving coding overhead due to shorter codewords in this case.
[0074] In one exemplary embodiment of the present disclosure, the reference block in the candidate list is represented by a block vector (BV) of the reference block, which is used to indicate the position of the reference block relative to the current block, and the BV corresponding to the reference block template is the BV of the reference block corresponding to the reference block template.
[0075] In this embodiment, the reference block in the candidate list is identified by its BV, i.e., it is the BV that is actually inserted into the candidate list. The position of the current block can be represented by a specified base point, which can be a pixel point in the current block. In this embodiment, the base point is the top left point (pixel point) of the current block, but the present disclosure is not limited thereto. The top right point, center point, or a point adjacent to the center point of the current block can also be used as the base point. In another example, a point in the current block template can be used as the base point, and if the relative position between the base point and the current block is constant and known, it can be used to position the current block. In this embodiment, assuming the coordinates of the base point are (50,50) and the coordinates of the top left point of a searched reference block are (120,120), the BV used to search for this reference block can be expressed as (70,70), that is, a position offset from the base point, which can be represented graphically as a vector pointing from the base point to the top left point of the reference block (see FIG. 8). The point obtained by adding the position offset amount represented by BV to the coordinates of the base point is the position indicated by this BV, which is the upper left point of the reference block in Fig. 8. For convenience of explanation, in this specification, the BV of a reference block corresponding to one reference block template is called the BV corresponding to the reference block template, and these two also have a one-to-one correspondence.
[0076] In one exemplary embodiment of the present disclosure, searching for a reference block template based on the first search range, calculating the difference between the searched reference block template and the current block template, and constructing the candidate list of intraTMP based on the difference includes: determining one BV group based on a first search step width and the first search range, where the position indicated by the BV group is within the first search range; searching for a corresponding reference block template based on the BV group; calculating the difference between the searched reference block template and the current block template; and populating the candidate list with BVs corresponding to the N reference block templates with the smallest difference.
[0077] In this embodiment, a reference block template is searched for based on a BV. As mentioned above, one BV can indicate the position of a reference block. For example, if the top left point of the current block is used as the base point, the BV can indicate the top left point of the reference block. Because the size and shape of the reference block are the same as those of the current block, the area occupied by the reference block, in other words, the reconstructed pixel points included in the reference block, can be determined based solely on the position indicated by one BV. Because the relative positions of the reference block template and the reference block are constant, the area occupied by the reference block template can also be determined by one BV. Therefore, a search for a reference block template can be realized using one group of BVs. To uniquely determine one reference block based on one BV, the BV of the reference block can be inserted into a candidate list as the reference block identifier.
[0078] In constructing a candidate list according to this embodiment, it is necessary to determine the N reference blocks in the candidate list and the order of the N reference blocks based on the differences between the searched reference block templates and the current block template. In one example, the N searched reference block templates are first inserted into the candidate list. Starting with the (N+1)th searched reference block template, the difference between the currently searched reference block template and the N reference block templates in the candidate list is compared. If the difference between this reference block template is less than the maximum difference between the N reference block templates in the candidate list, the candidate list is updated to delete the BV corresponding to this maximum difference, and the BV corresponding to this current reference block template is added to the candidate list. After the last searched reference block template is processed, the construction of the candidate list is completed. In the construction process, the N BVs in the candidate list are sorted in ascending order of their corresponding differences, which facilitates comparison. The difference corresponding to a BV is the difference between the reference block template corresponding to the BV. In another example, after all reference block templates have been searched, the N reference block templates with the smallest differences are inserted into the candidate list based on the differences between the reference block templates, thereby completing the construction of the candidate list.
[0079] In this specification, the difference of the reference block template refers to the difference between the reference block template and the current block template, and for convenience of explanation, is referred to as the difference of the reference block template.
[0080] In one exemplary embodiment of the present disclosure, searching for a reference block template based on the first search range, calculating a difference between the searched reference block template and a current block template, and constructing the candidate list of intraTMP based on the difference includes: determining a BV group based on a first search step width and the first search range, performing a search based on the BV group, calculating differences between the searched reference block template and a current block template, and populating the candidate list with BVs corresponding to the N reference block templates with the smallest differences; determining M second search ranges based on BVs corresponding to the M reference block templates with the smallest differences found in the first search; determining M BV groups based on a second search step width and the M second search ranges; searching for corresponding reference block templates in the M second search ranges based on the M BV groups; calculating the difference between the found reference block templates and the current block template; and updating the candidate list based on the difference.
[0081] Here, the second search step width is smaller than the first search step width, the second search range is smaller than the first search range, and the second search ranges do not overlap each other. M≧N.
[0082] In one example of this embodiment, the BV is expressed as a position offset amount relative to a base point, and the base point is a point (a pixel point or a partial pixel point) in the current block. The M second search ranges respectively cover positions indicated by the BVs corresponding to the M reference block templates, and the positions indicated by the BVs are determined based on the base point and the position offset amount.
[0083] This embodiment is a stepwise search method that performs two search stages, with the search step width of the next stage being smaller than that of the previous stage. The next stage includes multiple search ranges, each of which is a part of the search range of the previous stage. First, a first search stage is performed in a first search range using a large search step width, and N BVs are selected and added to a candidate list based on the magnitude of the differences between the searched reference block templates. In the second search stage, M second search ranges are determined based on the M BVs recorded after the first search stage. The second search ranges are searched using a smaller search step width, and the candidate list is updated based on the differences between the searched reference block templates. The stepwise search is a search method that gradually becomes finer, allowing for the rapid and accurate identification of reference block templates with high matching degrees within the reconstruction region to complete the construction of the candidate list.
[0084] In one example of this embodiment, after updating the candidate list based on the difference, the method further comprises: determining M' third search ranges based on BVs corresponding to the M' reference block templates with the smallest differences found in the second search, and determining M' BV groups based on a third search step width and the M' third search ranges; searching for corresponding reference block templates in the M" third search ranges based on the M' BV groups, respectively, calculating differences between the found reference block templates and the current block template, and updating the candidate list based on the differences.
[0085] Here, the third search step width is smaller than the second search step width, the third search range is smaller than the second search range, and they do not overlap each other, and M'≧N.
[0086] One BV group determined for each third search range is an integer pixel BV, or one BV group determined for each third search range is a partial pixel BV, and the reconstructed value of the reference block template corresponding to the partial pixel BV is obtained by interpolation.
[0087] This embodiment is a three-stage search method, and a more detailed third stage search is performed based on the second stage search. This allows for searching reference block templates at more positions, increasing the likelihood of finding a reference block template with a higher actual matching degree, and also increasing the probability that the reference block corresponding to the reference block template is closer to the current block. Therefore, the method of this embodiment can improve coding efficiency.
[0088] In one exemplary embodiment of the present disclosure, updating the candidate list based on the difference includes: The minimum difference d1 among the differences of the reference block templates searched in the same local search range is determined, and d1 <D N In this case, the candidate list is updated to D N and adding a BV corresponding to d1 to the candidate list.
[0089] where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating, and the difference corresponding to the BV refers to the difference of the reference block template corresponding to the BV. The local search range is the second search range or the third search range.
[0090] As described above, one local search range is determined based on a BV corresponding to one reference block template previously searched. The BV used to determine this local search range is also a BV within this local search range, and the reference block template previously searched based on this BV also belongs to the reference block templates searched in this local search range. In this specification, the reference block templates searched in the same local search range include not only reference block templates searched after this local search range is determined, but also reference blocks corresponding to the BV used to determine this local search range. Taking the second search range as an example, the reference block templates searched in the same second search range include reference block templates searched within this second search range in the second search stage and reference block templates (searched in the first search stage) corresponding to the BV used to determine this second search range. With respect to the third search range, the reference block templates searched in the same third search range include the reference block templates searched within this third search range in the third stage search and the reference block templates (searched in the first or second stage search) corresponding to the BV used to determine this third search range.
[0091] The update process for the candidate list according to this embodiment can be applied after the second or third search stage. In this embodiment, when updating, the BVs are limited to those corresponding to reference block templates searched from the same BV group. That is, only one BV of a reference block template searched from one local search range can be added to the candidate list at most. This local search range can be the second search range, the third search range, etc. In the second search stage, this method is applied to each BV group among the determined M BV groups. In the third search stage, this method is applied to each BV group among the determined M' BV groups. Furthermore, in any of the embodiments, when adding a new BV to the candidate list, the N BVs in the updated candidate list can be reordered based on the ascending order of the corresponding differences.
[0092] In this embodiment, the maximum number of BVs corresponding to reference block templates searched from one local search range that can be added to a candidate list is limited to one. This is because the differences between closely spaced reference block templates are generally similar, making it easy for BVs from multiple closely spaced reference block templates to be included in the candidate list. In this case, the reference blocks in the candidate list are overly concentrated at a specific location. If one reference block at that location is not highly similar to the current block, the candidate list will contain multiple reference blocks that are not highly similar to the current block. This results in reduced adaptability of the candidate list. By limiting the number, the locations of the reference blocks added to the candidate list are not too concentrated, and to some extent, this alleviates the situation where a block that closely matches the current block cannot be found in the candidate list due to differences in the texture characteristics of these reference blocks.
[0093] In one exemplary embodiment of the present disclosure, updating the candidate list based on the difference includes: Determine the smallest K differences among the differences of the reference block templates searched in the same local search range, and at least one of the K differences is D N If it is less, updating the candidate list.
[0094] The N BVs in the candidate list after updating are the N BVs with the smallest differences among the BVs corresponding to the K differences and the N BVs in the candidate list before updating, where K is a predetermined threshold, K≧2.
[0095] where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating. The difference corresponding to the BV refers to the difference of the reference block template corresponding to the BV. The local search range is the second search range or the third search range.
[0096] The difference between this embodiment and the previous embodiment is that the maximum number of BVs corresponding to reference block templates searched from one local search range that are added to a candidate list is limited to K, where K is an integer equal to or greater than 2. This predetermined threshold can be predetermined (i.e., set to a specific value by default) or can be set at the encoding end and then transmitted from the encoding end to the decoding end.
[0097] In one exemplary embodiment of the present disclosure, updating the candidate list based on the difference includes: For each reference block template searched, the difference of that reference block template is D N If it is smaller than D, update the candidate list. N and adding a BV corresponding to the reference block template to the candidate list.
[0098] where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating, and the difference corresponding to a BV refers to the difference of the reference block template corresponding to the BV.
[0099] In this embodiment, there is no limit on the number of BVs corresponding to reference block templates searched in one local search range that are added to the candidate list. The BVs of each reference block template searched in the local search range can be added to the candidate list if the corresponding difference is sufficiently small. In this embodiment, multiple BVs added to the candidate list may be concentrated in one local area, which may cause some problems in adaptability. However, if the reference block at that position has a relatively high degree of matching with the current block, it may be possible to find a reference block that is closest to the highest degree of matching. The update process for the candidate list in this embodiment includes both cases where the candidate list is not updated based on difference comparison and cases where the candidate list is updated based on difference comparison.
[0100] In this embodiment, the difference can be calculated for each reference block template, and then the difference comparison and update process can be performed immediately. Alternatively, the difference between the reference block templates searched from one local search range can be calculated, and then the difference comparison and update process can be performed for each of them. Alternatively, the difference between the reference block templates searched from all local search ranges can be calculated, and then the difference comparison and update process can be performed for each of them. The first processing method requires fewer cache resources.
[0101] In one exemplary embodiment of the present disclosure, the size of the first search range is determined based on the size of the current block.
[0102] In one example of this embodiment, the first search distance in the width direction and the second search distance in the height direction of the first search range relative to the base point representing the position of the current block are calculated in the following manner: a method including: calculating the product of the width of the current block and a first proportional coefficient, and determining the larger of the product and a predetermined minimum search distance in the width direction as the first search distance; and calculating the product of the height of the current block and a second proportional coefficient, and determining the larger of the product and a predetermined minimum search distance in the height direction as the second search distance, wherein the first proportional coefficient and the second proportional coefficient are equal to or different from each other; A method including: determining the larger of a width of a current block and a predetermined minimum search distance in a width direction, and determining the product of the larger of the width and a first proportionality coefficient as the first search distance; and determining the larger of a height of a current block and a predetermined minimum search distance in a height direction, and determining the product of the larger of the height and a second proportionality coefficient as the second search distance, wherein the first proportionality coefficient and the second proportionality coefficient are equal or different; and The search distance is determined by one of a method including: multiplying the width of the current block by a corresponding first proportional coefficient to obtain the first search distance; and multiplying the height of the current block by a corresponding second proportional coefficient to obtain the second search distance, wherein there are a plurality of first proportional coefficients, and the larger the first proportional coefficient, the larger the width of the corresponding current block; and there are a plurality of second proportional coefficients, and the larger the second proportional coefficient, the larger the height of the corresponding current block.
[0103] In this embodiment, the size of the search range is represented using search distances in the width and height directions. As shown in Figure 10, searchRangeWidth in the drawing represents the first search distance of the first search range in the width direction from the base point (the upper left point of the current block), and searchRangeHeight represents the second search distance of the first search range in the height direction from the base point. Of course, the size of the search range in this embodiment can also be represented in a different way, for example, by defining twice the searchRangeWidth in the drawing as the search distance in the width direction and twice the searchRangeHeight as the search distance in the height direction.
[0104] In this embodiment, different search ranges are adopted for different current blocks. Because the reference blocks have the same size as the current block, this method ensures that the number of reference blocks searched does not vary significantly according to the size of the current block, and ensures that there are a sufficient number of reference blocks available for matching, thereby ensuring the coding effect of the intraTMP mode.
[0105] In one exemplary embodiment of the present disclosure, determining a first search range of the intraTMP for the current block includes determining an area to be covered by the first search range based on a base point representing the position of the current block, a search distance relative to the base point, and a reconstruction area available at the time of the search.
[0106] In this embodiment, when determining the area actually covered by the first search range, as shown in FIGS. 8 and 10, the base point, the search distance (in two directions) relative to the base point, and the available reconstruction area during the search must be taken into consideration. This available search area is related to the set search direction. In the example shown in FIG. 10, the search (i.e., finding the reference block template at the corresponding position based on the BV) can be performed only on the left, top, upper left, lower left, and upper right sides of the current block, meaning that the reconstruction areas in these directions are available. In other examples, the search direction can be limited to the left, top, and upper left sides, and the present disclosure is not limited thereto. The available reconstruction area during the search can also be set directly. For example, in the example of FIG. 8, the reconstruction areas in the CTU to which the current block belongs and the CTUs above, left, and upper left of the current block are allowed to be used during the search. Furthermore, other restrictions can also be set. For example, it can be specified that the areas above and to the left of the current block in the CTU to which the current block belongs are unavailable.
[0107] After determining one BV group based on the search range and search step width, if a reference block sample searched based on the BV is not within the available reconstruction area, the reference block template can be discarded. Alternatively, when determining the BV, it is immediately determined whether a reference block template searched based on one BV is within the available reconstruction area, and if not, the BV is discarded. This ensures that the reference block template to be searched is within the available reconstruction area.
[0108] The embodiments of the present disclosure further provide a video decoding method, as shown in Figure 11, which includes the following steps:
[0109] In step S210, the intra template matching prediction intraTMP mode use flag of the current block is decoded.
[0110] In step S220, if it is determined based on the intraTMP mode usage flag that the current block uses intraTMP mode, the intraTMP index of the current block is further decoded, which indicates the position in the intraTMP candidate list of the reference block used by the current block.
[0111] In step S230, a candidate list is constructed, a reference block used by the current block is determined based on the intraTMP index and the candidate list, and intra prediction is performed on the current block based on the reference block used by the current block.
[0112] The intraTMP mode in this embodiment is a multi-candidate intraTMP mode. During decoding, a candidate list can be constructed, and a reference block used by the current block is determined based on the intraTMP index and the candidate list. The current block is intra-predicted based on the reference block used by the current block. Since the candidate list includes multiple reference blocks, the current block may find a reference block with a higher degree of matching during prediction, thereby improving coding efficiency.
[0113] In one exemplary embodiment of the present disclosure, the candidate list is constructed according to the intraTMP candidate list construction method according to any embodiment of the present disclosure. Note that when constructing a candidate list according to the intraTMP candidate list construction method according to any embodiment of the present disclosure, it is not necessary to construct a candidate list of length N; a candidate list of length less than N may be constructed to simplify processing. For example, if it is determined based on the intraTMP index that the reference block used by the current block is at the third position in the candidate list and N=5, the decoding end can construct a candidate list of length 3. The construction method is the same, only the length is different.
[0114] In one exemplary embodiment of the present disclosure, after decoding the intraTMP index of the current block, the method further comprises: If the intraTMP index indicates the first position in the candidate list, the candidate list is not constructed, and intra prediction is performed on the current block according to a single candidate intraTMP mode; and If the intraTMP index indicates a position other than the first position in the candidate list, constructing the candidate list and determining a reference block to be used by the current block based on the intraTMP index and the candidate list.
[0115] In this embodiment, if the intraTMP index indicates the first position in the candidate list, the reference block used by the current block can be found by simply using the single candidate intraTMP mode, so there is no need to build a candidate list, thereby reducing the complexity of decoding.
[0116] In one exemplary embodiment of the present disclosure, the method further includes decoding an intraTMP multi-candidate flag and determining whether use of a multi-candidate intraTMP mode is allowed based on the intraTMP multi-candidate flag, wherein the intraTMP multi-candidate flag is a sequence-level, picture-level, or slice-level flag.
[0117] After determining that the current block uses the intraTMP mode based on the intraTMP mode use flag, the method: If it is determined based on the intraTMP multi-candidate flag that the use of the multi-candidate intraTMP mode is permitted, further decoding the intraTMP index of the current block; If it is determined based on the intraTMP multi-candidate flag that the use of the multi-candidate intraTMP mode is not permitted, the method further includes skipping decoding for the intraTMP index of the current block and performing intra prediction for the current block according to the single-candidate intraTMP mode.
[0118] In this embodiment, a high-level intraTMP multi-candidate flag is used to indicate whether to allow the use of multi-candidate intraTMP mode. Thus, if the intraTMP multi-candidate flag indicates that the use of multi-candidate intraTMP mode is not allowed, when the intraTMP mode use flag is decoded and it is determined that the current block uses intraTMP mode, there is no need to decode the intraTMP index, and the current block can be intra-predicted directly according to the single-candidate intraTMP mode, thereby simplifying the decoder processing flow.
[0119] In one exemplary embodiment of the present disclosure, decoding the intraTMP index of the current block comprises: De-binarizing the intraTMP index using an analysis method that supports variable length coding, fixed length coding, truncated Unicode, or truncated binary code; or analyzing the value of the first binary symbol in the intraTMP index, and if the value indicates that the intraTMP index is coded using variable length coding or truncated Unicode, using an analysis method corresponding to variable length coding or truncated Unicode to de-binarize the intraTMP index; and if the value indicates that the intraTMP index is coded using fixed length coding or truncated binary code, using an analysis method corresponding to fixed length coding or truncated binary code to de-binarize the intraTMP index.
[0120] In this embodiment, the analysis method for de-binarizing the intraTMP index is determined based on the value of the first binary symbol in the intraTMP index. If the intraTMP index can use multiple encoding methods, decoding the intraTMP index can be easily and conveniently achieved.
[0121] In one exemplary embodiment of the present disclosure, before constructing the candidate list, the method further includes decoding an intraTMP mode search step size index, which is used to indicate an index of a search step size to be used among a plurality of candidate search step sizes.
[0122] When constructing the candidate list, a reference block template is searched for in the first search range using a search step width determined based on the search step width index.
[0123] In this embodiment, different search step widths can be used to perform intraTMP mode searches on the current block, which has better adaptability to different images.
[0124] The present disclosure further provides a video encoding method, as shown in Figure 12, which includes the following steps:
[0125] In step S310, if it is determined that the current block allows the use of the multi-candidate intra template matching prediction intraTMP mode, a candidate list for intraTMP is constructed according to the method according to any embodiment of the present disclosure.
[0126] In step S320, the coding costs for predicting the current block based on the N reference blocks in the candidate list are calculated, and the smallest coding cost among them is taken as the coding cost for the multi-candidate intraTMP mode to perform rate-distortion optimization of the current block.
[0127] In step S330, if the rate-distortion optimization determines that the current block is intra predicted using the multi-candidate intraTMP mode, the syntax elements related to the multi-candidate intraTMP mode of the current block are coded.
[0128] The intraTMP mode in this embodiment is a multi-candidate intraTMP mode. A candidate list is constructed during encoding. When the intraTMP mode is selected for rate-distortion optimization, the reference blocks used by the current block in the multi-candidate intraTMP mode are indicated by encoding syntax elements related to the multi-candidate intraTMP mode for the current block. Because the candidate list includes multiple reference blocks, the current block may find a reference block with a higher matching degree during prediction, thereby improving coding efficiency.
[0129] In one exemplary embodiment of the present disclosure, encoding a syntax element associated with a multi-candidate intraTMP mode of the current block comprises: encoding an intraTMP mode use flag of the current block to indicate that the current block uses intraTMP mode; and encoding an intraTMP index of the current block to indicate the position in the candidate list of the reference block with the lowest coding cost among the N reference blocks in the candidate list.
[0130] In this embodiment, by encoding the intraTMP mode usage flag and the intraTMP index, the decoding end can determine the reference block used by the current block based on these two syntax elements and perform intra prediction on the current block based on the reference block used by the current block.
[0131] In one exemplary embodiment of the present disclosure, encoding the intraTMP index of the current block comprises: implementing binarization of the intraTMP index using variable length coding, fixed length coding, truncated Unicode, or truncated binary code; or When the value of the intraTMP index is within a first numerical range, the intraTMP index is binarized by variable length coding or truncated Unicode, and when the value of the intraTMP index is within a second numerical range, the intraTMP index is binarized by fixed length coding or truncated binary code.
[0132] The value of the first numerical range is less than the value of the second numerical range.
[0133] In this embodiment, when encoding, variable length coding or truncated Unicode can be adopted when the value of the intraTMP index is relatively small, and fixed length coding or truncated binary code can be adopted when the codeword is short and the value of the intraTMP index is relatively large, thereby saving encoding overhead.
[0134] In one example embodiment of the present disclosure, determining that the current block is allowed to use multi-candidate intraTMP mode includes determining that the current block is allowed to use multi-candidate intraTMP mode if none of the conditions under which the use of multi-candidate intraTMP mode is not allowed is true, where the conditions under which the use of multi-candidate intraTMP mode is not allowed include a sequence-level, picture-level, or slice-level intraTMP multi-candidate flag indicating that the use of multi-candidate intraTMP mode is not allowed.
[0135] In one example of this embodiment, when encoding screen content, the intraTMP multi-candidate flag is set to a value indicating that the use of multi-candidate intraTMP mode is not permitted, and when encoding video captured by a camera, the intraTMP multi-candidate flag is set to a value indicating that the use of multi-candidate intraTMP mode is permitted.
[0136] This embodiment can accommodate different usage scenarios, and when the use of multi-candidate intraTMP mode is appropriate, the intraTMP multi-candidate flag can be coded to indicate that the use of multi-candidate intraTMP mode is permitted, thereby improving the coding effect; when the use of multi-candidate intraTMP mode is not appropriate, the intraTMP multi-candidate flag can be coded to indicate that the use of multi-candidate intraTMP mode is not permitted, thereby avoiding unnecessary coding complexity.
[0137] In an exemplary embodiment of the present disclosure, the method further includes, when performing intra prediction coding on the current block according to the single candidate intraTMP mode, encoding an intraTMP mode usage flag of the current block to indicate that the current block uses the intraTMP mode, and encoding an intraTMP index of the current block to indicate the first position in the candidate list of the reference block used by the current block. This embodiment does not require a separate flag to indicate whether the current block adopts the single candidate intraTMP mode or the multi-candidate intraTMP mode, and can reduce complexity at the decoding end by making the determination based on the intraTMP index.
[0138] An embodiment of the present disclosure further provides a multi-candidate intra template matching prediction (intraTMP) method, in which N candidates (N≧2) are set in the intraTMP, that is, a candidate list (denoted as intraTMPCandList[N]) of length N is set in the intraTMP.
[0139] The encoding end finds multiple reference block templates within the set search range according to the set search rule, calculates the difference between each of the multiple reference block templates and the current block template based on the reconstructed pixel values of the multiple reference block templates and the reconstructed pixel values of the current block template, and inserts the reference block position identifiers corresponding to the N reference block templates into the intraTMP candidate list in order of smallest difference.
[0140] The encoding end calculates the difference between each of the N reference blocks in the candidate list and the current block based on the reconstructed pixel values of the N reference blocks and the original pixel values of the current block, and determines the value of the intraTMP index based on the position in the candidate list of the reference block with the smallest difference. This reference block with the smallest difference is the reference block used by the current block in intraTMP mode, that is, the best matching block found. When the BV is inserted into the candidate list as the position identifier of the reference block, the BV located at the position indicated by the intraTMP index in the candidate list is also called the BV used by the current block in intraTMP mode.
[0141] If the encoding end performs rate-distortion optimization and selects intraTMP mode for the current block from multiple intra prediction modes (i.e., it decides that the current block will use intraTMP mode), after encoding a flag indicating that the current block will use intraTMP, it further encodes an intraTMP index to indicate the position in the candidate list of the reference block used by the current block.
[0142] Correspondingly, the decoding syntax is as follows:
[0143] intraTMPFlag if(intraTMPFlag) { intraTMPIndex} Here, intraTMPFlag is a flag indicating whether the current block uses the intraTMP mode, and intraTMPIndex is an intraTMP index indicating the position in the candidate list of the reference block used by the current block.
[0144] During decoding, if intraTMPFlag is true (e.g., 1), intraTMPIndex is further analyzed. The decoding end uses the same method to construct an intraTMP candidate list intraTMPCandList, finds the position identifier at the position indicated by intraTMPIndex in intraTMPCandList, and finds the corresponding reference block based on this position identifier. The reconstructed value of this reference block can be used as the predicted value of the current block.
[0145] In this embodiment, when constructing the intraTMPCandList, each time a BV is searched within the search range, the difference between the reference block template corresponding to the BV and the current block template is calculated. The reference block template is a block searched within the reconstruction region that has the same shape and size as the current block. The difference may be SAD, SATD, SSE, etc. When constructing the intraTMPCandList, the BVs corresponding to the searched reference block templates may be inserted into the intraTMPCandList in order of smallest difference. Alternatively, the searched reference block templates may be sorted in order of smallest corresponding difference, and the reference blocks corresponding to the first N reference block templates listed are set as the N reference blocks in the intraTMPCandList. Alternatively, only the N candidates with the smallest difference may be kept, and reference block templates whose order exceeds N may be directly discarded, thereby saving computational overhead.
[0146] Typically, blocks corresponding to adjacent BVs are relatively close to each other, especially when the BVs support sub-pixel precision such as 1 / 2, 1 / 4, 1 / 8, or 1 / 16 precision. The reference block template corresponding to the sub-pixel BV must be obtained by interpolation. When interpolating the intra template (intraTmp), the same filter as the inter-interpolation filter can be used, thereby reducing the memory required for additional filters. Alternatively, a simpler interpolation method can be used. While a 12-tap filter is used for inter-interpolation, in this embodiment, a filter with fewer taps, such as an 8-tap, 4-tap, or 2-tap filter, can be used to reduce the amount of calculation.
[0147] When BV supports sub-pixel accuracy, if the candidates are arranged based only on the difference between the reference block templates without any control, multiple candidates are likely to be concentrated in a very narrow range. In this embodiment, the following method for controlling the BV of candidates in the intraTMPCandList is provided to prevent them from being too concentrated.
[0148] The first method is as follows.
[0149] In the search process, all possible BVs are not searched sequentially. For example, the usual search order is from left to right and from top to bottom. Generally, integer pixel BVs can be searched sequentially. As shown in FIGS. 13A and 13B, if the currently searched BV is (x0, y0), the next BV is (x0+1, y0) (if the boundary of the search range has not been reached). In the first method, a sparse search is first performed. For example, for integer pixel BVs, if the currently searched BV is (x0, y0), the next BV is (x0+4, y0) (if the boundary of the search range has not been reached). Template matching (i.e., searching a reference block template and calculating the difference between the searched reference block template and the current block template) is performed once every predetermined number of pixels, and template matching is performed according to a set search step width. The search step width can be a preset value such as 2, 3, 4, 8, etc. Similar processing can be performed in the vertical direction.
[0150] First, find the N BVs with the smallest corresponding differences, and then search again for these N BVs with the smallest corresponding differences (the differences are also called cost or distortion cost) in a small local search range based on each BV to improve them. For example, in the first stage search, the search interval in the x and y directions is 4 pixels, and this local search range can be set to 4x4. If the difference of the reference block template searched in the local search range is small, the corresponding BV can replace the BV in the candidate list, and the N candidates in the candidate list can be reordered. As a result, there will be a certain distance between the BVs corresponding to the N reference blocks in the resulting candidate list.
[0151] As shown in FIG. 13A, a first-stage search is performed with a preset step size. The upper left point of the searched reference block is indicated by an "X" in the figure. After the search, three already-sorted BVs are found, and the upper left points of the reference block corresponding to these three BVs (i.e., the positions indicated by the BVs) are indicated by the "X" points in FIG. 13B. In this example, the horizontal search step size is 4, and the vertical search step size is also 4. In the second-stage search, a local search range is determined based on the three already-sorted BVs. For example, in this example, as shown in FIG. 13B, it is a 4x4 rectangular area covering the positions indicated by the three already-sorted BVs (small squares indicated by "X" in the figure). If the difference between the reference block templates searched from each 4x4 local search range is smaller than the difference between the corresponding sorted BVs, the BV corresponding to the newly searched reference block template can replace the corresponding sorted BV in the candidate list and participate in the sorting of the intraTMPCandList. Otherwise, the candidate list does not need to be updated.
[0152] In this embodiment, the horizontal and vertical sizes of the local search ranges are exactly the same as the search step width of the first stage, so that overlapping of the local search ranges can be avoided.
[0153] If subpixel accuracy is supported, a third-stage search can be performed after the second-stage search. The BVs used in the third-stage search are subpixel BVs. For example, a 1 / 2-pixel search is performed within a range of 1 pixel above, below, left, and right based on the position indicated by the integer pixel BV selected in the second-stage search, and a 1 / 2-pixel search is performed within a range of 1 pixel above, below, left, and right based on the integer pixel BV selected in the second stage. This results in four BVs, which are obtained by shifting the x-coordinates of the integer pixel BVs by ±1 / 2 pixels and the y-coordinates by ±1 / 2 pixels. In another example, four more BVs can be obtained by simultaneously shifting the x-coordinate and y-coordinate by ±1 / 2 pixels. In other words, four or eight BVs can be set and searched within this local search range. In other embodiments, subpixel BVs can be used in the second-stage search, or subpixel BVs can be used only in the fourth-stage search.
[0154] When determining one local search range based on one BV, the position indicated by this BV can be the center point or a point close to the center of this local search range, but is not limited to this. As shown in Figure 13B, the position indicated by the BV used in the first stage search can be located at a point close to the center within a 4x4 local search range (the coordinates in this local search range are written as (2,2)). Note that the position indicated by the BV used in the first stage search can also be set as the lower right point of the local search range to determine the local search range.
[0155] The construction of the candidate list is a process that must be performed at both the encoding end and the decoding end, thereby ensuring that the candidate list obtained at the encoding end matches the candidate list obtained at the decoding end.
[0156] In this embodiment, the number of items required for intraTMPCandList is N, and after the first stage of search, N BVs (written to the candidate list) are recorded. In the second stage of search, N local search ranges are determined based on the N BVs to continue the search. In another embodiment, more BVs can be recorded after the first stage of search (other BVs can be additionally saved in addition to the N BVs written to the candidate list), for example, M BVs (M>N, e.g., M=2N). This increases opportunities for refinement search and reduces areas suitable for the second stage of search that were overlooked in the sparse search of the first stage.
[0157] In this embodiment, a maximum number of BVs that can be reserved in each local search range can be set. That is, it is a threshold value for the maximum number of BVs that can be inserted into the intraTMPCandList in each local search range. This threshold value can be determined based on the length N of the candidate list and the size of the local search range. For example, when N is relatively small, each local search range can be set to reserve a relatively large number of BVs, thereby preventing the candidate reference blocks from concentrating too much. When N is relatively large (when the number of candidate reference blocks is large), more BVs can be reserved in each local search range, thereby ensuring coverage while maintaining a certain degree of precision.
[0158] To determine the number of BVs to be reserved in each local search range, one of the following methods can be used.
[0159] Method 1 At most, one BV from one local search range can be reserved in the intraTMPCandList.
[0160] Method 2 Any number of BVs from one improvement area can be reserved in the intraTMPCandList, meaning there is no limit to the maximum number of BVs that can be reserved in each improvement area. If the number of candidates is large enough, this setting can improve precision.
[0161] Method 3 A threshold value K is set so that the number of BVs to be retained from each local search range in intraTMPCandList is less than or equal to K. In each local search range, we first determine the K BVs with the smallest difference by sorting, and then try to add these K BVs to intraTMPCandList.
[0162] The encoder and decoder must perform the same search to ensure that the lists they build are the same. Typically, a larger search range means more BVs can be explored, increasing the possibilities, but also increasing the complexity. Therefore, setting a reasonable search range can balance performance and complexity.
[0163] intraTMP itself is an intra-block copy technique, which copies a block of the same size as the current block. That is, the larger the current block, the larger the area to be copied; the smaller the current block, the smaller the area to be copied. One way to do this is to set the search range relative to the block size. For example, the horizontal search range is set as searchRangeWidth = ratio * width, and the vertical search range is set as searchRangeHeight = ratio * height. Here, ratio is a multiple (e.g., 4, 5, 6, etc.), width is the width of the current block, and height is the height of the current block. However, the search range must not exceed the available reconstruction area. Considering that the current codec supports a minimum of 4x4 small blocks, taking a 4x4 small block as an example, and ignoring the maximum available area limit, if ratio is set to 5, then searchRangeWidth and searchRangeHeight will be 20. Since this range is very small, the ideal case for searching intraTMP is to find a texture that repeats the current block, so a threshold can be set to ensure that the minimum search range is not too small. Specifically, the size of the search range can be set in one of the following ways, and the size of this search range is represented by the search distance from the base point of the current block position.
[0164] Method 1 searchRangeWidth = max(ratio * width, thrLowerBoundary) searchRangeHeight = max(ratio * height, thrLowerBoundary) where thrLowerBoundary is the minimum search range (e.g., 64, 128, etc.).
[0165] Method 2 The following method can also be used:
[0166] searchRangeWidth = ratio * max(width, thrLowerBoundary) searchRangeHeight = ratio * max(height, thrLowerBoundary) where thrLowerBoundary is 16 or 32, etc.
[0167] where width and height are the width and height of the current block, ratio is the proportional coefficient for the width and height directions, searchRangeWidth and searchRangeHeight are the search distances for the width and height directions, and thrLowerBoundary is the minimum search distance for the width and height directions.
[0168] Method 3 In this method, a larger ratio is set for the small block. For example, if the width or height is less than 16, the corresponding ratio is 5 otherwise.
[0169] The search range can be set by the above method even for the single candidate intraTmp without requesting the multi-candidate to set the search range.
[0170] In one embodiment, a high-level control syntax can be set to control the search step size of the first step, for example, sps_intraTmp_search_step_idx of one sps (sequence parameter set) controls the search step size of the first step. If sps_intraTmp_search_step_idx is 0, the search step size is 3, and if sps_intraTmp_search_step_idx is 1, the search step size is 4. A larger search step size can be set for high-resolution video, and a smaller search step size can be set for low-resolution video.
[0171] In this embodiment, the intraTMPCandList is sorted, and from a statistical perspective, candidates at the front are more likely to be selected. Variable length coding can be set for binarization and de-binarization of the intraTMPIndex as follows, or truncated unary coding (TU) can be used.
[0172] [Table 1]
[0173] If the probability of each candidate reference block being selected is approximately the same, fixed-length coding or truncated binary coding can be used to realize the binarization of intraTMPCandList. In the above table, Bin index is the index of the binary symbol, where Bin index 0 indicates the first binary symbol, and Bin index 1 indicates the second binary symbol.
[0174] In a scenario where N is large, earlier candidates have higher probabilities, later candidates have lower probabilities, and later candidates are closer to each other in probability, it is possible to configure the codewords to be shorter for smaller intraTMPIndex values and longer for larger intraTMPIndex values. The reference blocks of several candidates in the candidate list can use the same code length, as shown in the following example:
[0175] [Table 2]
[0176] In this example, N is 15, and indices 3 to 6 use codewords of the same length, and indices 7 to 14 use codewords of the same length. The x in the table above can be obtained by encoding using truncated binary.
[0177] In this embodiment, high-level control syntax can be used to control whether to use the multi-candidate technique. If the multi-candidate technique is not used, an existing technique, i.e., a single-candidate method, can be used. As an example, one SPS (sequence parameter set) flag, such as sps_intra_tmp_multi_cand_enabled_flag, is used. If the value of sps_intra_tmp_multi_cand_enabled_flag is 1, the current sequence uses the intraTMP multi-candidate method; otherwise, the current sequence uses the intraTMP single-candidate method.
[0178] The corresponding syntax is as follows:
[0179] intra_tmp_flag If(sps_intra_tmp_multi_cand_enabled_flag && intra_tmp_flag) { intra_tmp_index }
[0180] An example of a usage scenario of the above high-level control syntax is to set sps_intra_tmp_multi_cand_enabled_flag to 1 for a sequence captured by a camera, and to set sps_intra_tmp_multi_cand_enabled_flag to 0 for a screen content sequence. Of course, picture-level or slice-level control can also be achieved using flags such as PPS (picture parameter set), picture header, slice header, etc.
[0181] By setting more candidates for intraTMP, the embodiments of the present disclosure can reduce the situation where the best matching block found by template matching is not ideal, thereby improving compression performance.
[0182] An embodiment of the present disclosure further provides a candidate list construction device for intra template matching prediction. As shown in Fig. 14, the device includes a processor 71 and a memory 73 storing a computer program, and when the processor 71 executes the computer program, the device can realize the candidate list construction method for intra template matching prediction according to any embodiment of the present disclosure.
[0183] An embodiment of the present disclosure further provides a video decoding device. Referring to Figure 14, the device includes a processor and a memory storing a computer program, and when the processor executes the computer program, the device can realize the video decoding method according to any embodiment of the present disclosure.
[0184] An embodiment of the present disclosure further provides a video encoding device, which, as shown in Figure 17, includes a processor and a memory storing a computer program, and when the processor executes the computer program, can realize the video encoding method according to any embodiment of the present disclosure.
[0185] The processor according to the above embodiments of the present disclosure may be a general-purpose processor (including a CPU, a network processor (NP), a microprocessor, etc.) or other conventional processor. The processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), discrete logic, other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or other equivalent integrated or discrete logic circuitry, or a combination thereof. That is, the processor according to the above embodiments may be any processor device or combination of devices that implements each method, step, and logical block diagram disclosed in the embodiments of the present disclosure. When parts of the embodiments of the present disclosure are implemented in software, software instructions can be stored in a suitable non-volatile computer-readable storage medium, and one or more processors can execute these instructions in hardware to implement the methods according to the embodiments of the present disclosure. As used herein, the term "processor" refers to the above structure or any other structure suitable for implementing the techniques described herein.
[0186] An embodiment of the present disclosure further provides a video encoding / decoding system including a video encoding device according to any one of the embodiments of the present disclosure and a video decoding device according to any one of the embodiments of the present disclosure.
[0187] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, can implement a method according to any embodiment of the present disclosure.
[0188] An embodiment of the present disclosure provides a computer program product including a computer program, which, when executed by a processor, can implement a method according to any embodiment of the present disclosure.
[0189] In one or more exemplary embodiments above, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium for transmitting a computer program from one place to another, such as a communication protocol. Thus, computer-readable media generally correspond to non-transitory tangible computer-readable storage media or communication media such as signals or carrier waves. Data storage media may be any available medium accessible by one or more computers or one or more processors to store and / or retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0190] By way of non-limiting example, such computer-readable storage media include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can be referred to as a computer-readable medium; for example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, microwave, etc., the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, microwave, etc., are included within the definition of medium. However, computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory (momentary) media, but refer to non-transitory, tangible storage media. As used herein, magnetic disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, Blu-ray disks, etc., where magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically using a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0191] In some aspects, the functionality described herein is provided in dedicated hardware and / or software modules configured for encoding and decoding, or in a combined codec. Alternatively, the techniques may be implemented entirely in one or more circuits or logic elements.
[0192] The technical solutions according to the embodiments of the present disclosure can be implemented in a wide variety of apparatuses or devices, including wireless mobile phones, integrated circuits (ICs), or sets of ICs (e.g., chipsets). The various components, modules, or units described in the embodiments of the present disclosure are described to emphasize functional aspects of apparatuses configured to perform the described techniques, but are not necessarily implemented by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.
Claims
1. A method for constructing a candidate list for intra template matching prediction, comprising: determining a first search range for performing intra template matching prediction (intraTMP) on a current block; Searching for a reference block template based on the first search range and calculating a difference between the searched reference block template and a current block template, wherein the reference block template has a one-to-one correspondence with the reference block; constructing a candidate list of intraTMP based on the difference, and determining N reference blocks in the candidate list and an order of the N reference blocks; N≧2, A method characterized by:
2. the first search range is located in a reconstruction region of the current image; The difference between the reference block template and the current block template is determined based on SAD, SATD, or MSE between the reconstructed values of the reference block template and the reconstructed values of the current block template.
2. The method of claim 1 .
3. The order of the N reference blocks in the candidate list is determined by the ascending order of the differences of the corresponding reference block templates.
2. The method of claim 1 .
4. The reference block in the candidate list is represented by a block vector (BV) of the reference block, where the BV of the reference block indicates the position of the reference block relative to the current block, and the BV corresponding to the reference block template is the BV of the reference block corresponding to the reference block template.
4. The method according to claim 1 or 3.
5. Searching for a reference block template based on the first search range, calculating a difference between the searched reference block template and a current block template, and constructing the candidate list of intraTMP based on the difference, determining a BV group based on a first search step width and the first search range, wherein a position indicated by the BV group is within the first search range; Searching for a corresponding reference block template based on the BV group; calculating differences between the searched reference block templates and the current block template, and populating the candidate list with BVs corresponding to the N reference block templates with the smallest differences; 5. The method of claim 4.
6. Searching for a reference block template based on the first search range, calculating a difference between the searched reference block template and a current block template, and constructing the candidate list of intraTMP based on the difference, determining a BV group based on a first search step width and the first search range, performing a search based on the BV group, calculating differences between the searched reference block template and a current block template, and populating the candidate list with BVs corresponding to the N reference block templates with the smallest differences; determining M second search ranges based on BVs corresponding to the M reference block templates with the smallest differences found in the first search; determining M BV groups based on a second search step width and the M second search ranges; searching for corresponding reference block templates in the M second search ranges based on the M BV groups; calculating differences between the found reference block templates and a current block template; and updating the candidate list based on the differences; the second search step width is smaller than the first search step width, the second search range is smaller than the first search range, and the second search ranges do not overlap each other, and M≧N; 5. The method of claim 4.
7. The BV is expressed as a position offset amount relative to a base point, the base point being a point in the current block; each of the M second search ranges covers a position indicated by a BV corresponding to the M reference block templates, and the position indicated by the BV is determined based on the base point and the position offset amount; 7. The method of claim 6.
8. After updating the candidate list based on the difference, the method further comprises: determining M' third search ranges based on BVs corresponding to the M' reference block templates with the smallest differences found in the second search, and determining M' BV groups based on a third search step width and the M' third search ranges; searching for corresponding reference block templates in M″ third search ranges based on the M′ BV groups, respectively, calculating a difference between the searched reference block templates and a current block template, and updating the candidate list based on the difference; wherein the third search step width is smaller than the second search step width, the third search range is smaller than the second search range and does not overlap each other; one BV group determined for each third search range is a BV of an integer pixel, or one BV group determined for each third search range is a BV of a partial pixel, and the reconstructed values of the reference block template corresponding to the BV of the partial pixel are obtained by interpolation; 7. The method of claim 6.
9. The updating of the candidate list based on the difference may include: The minimum difference d among the differences of the reference block templates searched in the same local search range 1 Determine d 1 <D N In this case, the candidate list is updated to D N Remove the BV corresponding to d from the candidate list, 1 adding a BV corresponding to where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating, where the difference corresponding to the BV refers to the difference of the reference block template corresponding to the BV, and the local search range is the second search range or the third search range; 9. The method according to claim 6 or 8.
10. The updating of the candidate list based on the difference may include: Determine the smallest K differences among the differences of the reference block templates searched in the same local search range, and at least one of the K differences is D N If so, updating the candidate list; The N BVs in the candidate list after updating are the N BVs with the smallest differences among the BVs corresponding to the K differences and the N BVs in the candidate list before updating, where K is a predetermined threshold, K≧2; where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating, where the difference corresponding to the BV refers to the difference of the reference block template corresponding to the BV, and the local search range is the second search range or the third search range; 9. The method according to claim 6 or 8.
11. The updating of the candidate list based on the difference may include: For each searched reference block template, the difference of the reference block template is D N If it is smaller than D, update the candidate list. N and adding a BV corresponding to the reference block template to the candidate list; where D N is the maximum difference among the differences corresponding to the N BVs in the candidate list before updating, and the difference corresponding to the BV refers to the difference of the reference block template corresponding to the BV; 9. The method according to claim 6 or 8.
12. K is determined based on one or more of the following parameters: N; and the size of the local search range.
11. The method of claim 10.
13. the size of the first search range is determined based on the size of the current block; 2. The method of claim 1 .
14. A first search distance in the width direction and a second search distance in the height direction of the first search range relative to a base point representing the position of the current block are a method including: calculating the product of the width of the current block and a first proportionality coefficient, and determining the larger value of either the product or a minimum search distance in a predetermined width direction as the first search distance; and calculating the product of the height of the current block and a second proportionality coefficient, and determining the larger value of either the product or a minimum search distance in a predetermined height direction as the second search distance, wherein the first proportionality coefficient and the second proportionality coefficient are equal to or different from each other; A method including: determining a larger value of a width of a current block and a minimum search distance in a predetermined width direction, and determining a product of the larger value and a first proportionality coefficient as the first search distance; and determining a larger value of a height of a current block and a minimum search distance in a predetermined height direction, and determining a product of the larger value and a second proportionality coefficient as the second search distance, wherein the first proportionality coefficient and the second proportionality coefficient are equal to or different from each other; and a method including: obtaining the first search distance by multiplying a width of the current block by a corresponding first proportional coefficient, the first proportional coefficient being plural, and the larger the first proportional coefficient, the larger the width of the corresponding current block; and obtaining the second search distance by multiplying a height of the current block by a corresponding second proportional coefficient, the second proportional coefficient being plural, and the larger the second proportional coefficient, the larger the height of the corresponding current block.
14. The method of claim 13.
15. A video decoding method, comprising: Decoding an intra template matching prediction (intraTMP) mode usage flag of the current block; if it is determined based on the intraTMP mode usage flag that the current block uses intraTMP mode, further decoding an intraTMP index of the current block, wherein the intraTMP index indicates a position in an intraTMP candidate list of a reference block used by the current block; constructing a candidate list; determining reference blocks to be used by a current block based on the intraTMP index and the candidate list; and performing intra prediction on the current block based on the reference blocks to be used by the current block. A method characterized by:
16. The candidate list is constructed by a method for constructing an intraTMP candidate list according to any one of claims 1 to 14.
16. The method of claim 15.
17. The method comprises: decoding an intraTMP multi-candidate flag; and determining whether use of a multi-candidate intraTMP mode is permitted based on the intraTMP multi-candidate flag, wherein the intraTMP multi-candidate flag is a sequence-level, image-level, or slice-level flag; After determining that the current block uses the intraTMP mode based on the intraTMP mode use flag, the method: If it is determined that the use of the multi-candidate intraTMP mode is permitted based on the intraTMP multi-candidate flag, further decoding the intraTMP index of the current block; If it is determined that the use of the multi-candidate intraTMP mode is not permitted based on the intraTMP multi-candidate flag, skip decoding the intraTMP index of the current block, and perform intra prediction on the current block according to the single-candidate intraTMP mode.
16. The method of claim 15.
18. After decoding the intraTMP index of the current block, the method comprises: If the intraTMP index indicates the first position in the candidate list, do not construct the candidate list, and perform intra prediction on the current block according to a single candidate intraTMP mode; If the intraTMP index indicates a position other than the first position in the candidate list, constructing the candidate list and determining a reference block to be used by the current block based on the intraTMP index and the candidate list.
16. The method of claim 15.
19. Decoding the intraTMP index of the current block comprises: De-binarizing the intraTMP index using an analysis method that supports variable length coding, fixed length coding, truncated Unicode, or truncated binary code; or analyzing the value of a first binary symbol in the intraTMP index, and if the value is either 0 or 1, using an analysis method corresponding to variable length coding or truncated Unicode to de-binarize the intraTMP index, and if the value is the other of 0 or 1, using an analysis method corresponding to fixed length coding or truncated binary code to de-binarize the intraTMP index; 16. The method of claim 15.
20. Before constructing the candidate list, the method comprises: further comprising decoding an intraTMP mode search step size index; the search step size index is used to indicate an index of a search step size to be used among a plurality of candidate search step sizes; When constructing the candidate list, a reference block template is searched for in the first search range using a search step width determined based on the search step width index.
17. The method of claim 16.
21. A video encoding method, comprising: If it is determined that the current block allows the use of a multi-candidate intra-template matching prediction (intraTMP) mode, constructing a candidate list for intraTMP according to the method of any one of claims 1 to 14, wherein the candidate list includes N reference blocks, where N≧2; Calculating coding costs for predicting the current block based on the N reference blocks in the candidate list, and performing rate-distortion optimization of the current block using the smallest coding cost as a multi-candidate intraTMP mode coding cost; and encoding syntax elements associated with the multi-candidate intraTMP mode of the current block if the rate-distortion optimization determines that the current block is intra predicted using the multi-candidate intraTMP mode. A method characterized by:
22. The encoding of the syntax elements related to the multi-candidate intraTMP mode of the current block as described above includes: encoding an intraTMP mode use flag of the current block to indicate that the current block uses intraTMP mode; and encoding an intraTMP index of the current block to indicate the position in the candidate list of a reference block with the lowest coding cost among the N reference blocks in the candidate list; 22. The method of claim 21 .
23. As mentioned above, encoding the intraTMP index of the current block is implementing binarization of the intraTMP index using variable length coding, fixed length coding, truncated Unicode, or truncated binary code; or When the value of the intraTMP index is within a first range of values, binarizing the intraTMP index using variable length coding or truncated Unicode, and when the value of the intraTMP index is within a second range of values, binarizing the intraTMP index using fixed length coding or truncated binary code; the value of the first numerical range is less than the value of the second numerical range; 23. The method of claim 22.
24. Determining that the current block allows use of multi-candidate intraTMP mode includes: determining that the current block is permitted to use the multi-candidate intraTMP mode when none of the conditions for not permitting the use of the multi-candidate intraTMP mode are satisfied; wherein the condition for disallowing use of multi-candidate intraTMP mode includes a sequence-level, image-level, or slice-level intraTMP multi-candidate flag indicating that use of multi-candidate intraTMP mode is disallowed.
22. The method of claim 21 .
25. setting the intraTMP multi-candidate flag to a value indicating that use of multi-candidate intraTMP mode is not permitted when encoding screen content; When encoding video captured by a camera, the intraTMP multi-candidate flag is set to a value indicating that use of the multi-candidate intraTMP mode is permitted.
25. The method of claim 24.
26. The method comprises: When performing intra prediction coding on the current block according to the intraTMP mode of the single candidate, encoding an intraTMP mode usage flag of the current block to indicate that the current block uses the intraTMP mode; and encoding an intraTMP index of the current block to indicate a first position in the candidate list of a reference block used by the current block; 22. The method of claim 21 .
27. A bitstream comprising: Generated by the video encoding method of any one of claims 21 to 26, A bitstream characterized in that
28. A candidate list construction device for intra template matching prediction, comprising: a processor and a memory in which a computer program is stored; The processor, when executing the computer program, is capable of implementing the candidate list construction method according to any one of claims 1 to 14. An apparatus characterized in that
29. A video decoding device, a processor and a memory in which a computer program is stored; The processor, when executing the computer program, is capable of implementing the video decoding method according to any one of claims 15 to 20. An apparatus characterized in that
30. A moving image encoding device, a processor and a memory in which a computer program is stored; The processor, when executing the computer program, is capable of implementing the video coding method according to any one of claims 21 to 26. An apparatus characterized in that
31. 1. A video encoding and decoding system, comprising: A video encoding device according to claim 20 and a video decoding device according to claim 29, A system characterized by:
32. 1. A non-transitory computer-readable storage medium, comprising: The computer-readable storage medium has stored thereon a computer program, which, when executed by a processor, is capable of implementing the method of any one of claims 1 to 26. A computer-readable storage medium comprising:
33. 1. A computer program product comprising: including computer programs, The computer program, when executed by a processor, is capable of implementing the method of any one of claims 1 to 26.
1. A computer program product comprising: