Template-based intra-mode derivation fusion with non-angled mode

Template-based intra-mode derivation enhances video encoding efficiency by autonomously determining optimal prediction angles, addressing inefficient intra-prediction in existing methods and reducing data transmission needs.

DE112024002874T5Pending Publication Date: 2026-05-13KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
KONINKLIJKE PHILIPS NV
Filing Date
2024-06-25
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing video encoding methods struggle with inefficient intra-prediction, particularly when motion-based prediction is not effective, leading to suboptimal compression and increased data transmission requirements.

Method used

Implementing template-based intra-mode derivation (TIMD) to autonomously determine the best prediction angles by analyzing neighboring blocks and forming a linear combination of multiple intra-prediction modes, including a non-angular mode, to enhance prediction accuracy and reduce data transmission.

Benefits of technology

Improves video encoding efficiency by optimizing intra-prediction, reducing data requirements, and enhancing compression rates through adaptive angle determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To better predict the pixels of a block in a video, we propose a method for predicting the pixel colors of an image block, comprising: - Determining that template-based intra-mode derivation is applicable to predicting the block; - Determining a set of intra-prediction modes and their respective costs; - Determining a first TIMD mode, which is the first of the set of lowest-cost intra-prediction modes, and a second TIMD mode, which is the second of the set of second-lowest-cost intra-prediction modes; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Checking whether the third TIMD mode differs from the first TIMD mode and whether the third TIMD mode differs from the second TIMD mode.- based on the successful verification, forming a final TIMD mode comprising a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - forming a prediction of the block based on a template of the block and the final TIMD mode.;
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF INVENTION

[0001] The invention relates to the predictive coding of blocks from previously decoded blocks of the same image using template-based intra-mode derivation technology, which can be used in video encoders and decoders. BACKGROUND OF THE INVENTION

[0002] Video encoding achieved great compression rates by exploiting the fact that objects typically move but don't necessarily change shape. Therefore, the same information already decoded for a previous frame can be reused to predict the pixels of that object in a current frame. For example, if the buildings in a cityscape undergo a simple panning motion, a global motion vector can be used to move the pixels from their previous position in the previous frame to the position of a currently decoded block. With regard to prediction correction, assuming no other changes occur, such as a change in lighting, these predicted pixels will be close enough to the values ​​of the original frame to be encoded to provide a faithful reconstruction.In the case of changes, prediction is still useful, since far fewer bits are needed to communicate the difference (also called residual) than to encode the entire pixel block itself.

[0003] However, motion-based prediction is not always the best prediction (e.g., at the beginning of a new scene).

[0004] In this case, the so-called intra-prediction can be used, in which a block at a certain position, which is currently being decoded and reconstructed, can be predicted from already decoded blocks of the same image (at higher positions in the image or to the left of the current block).

[0005] Linearly predictable features in the image can be utilized by employing a so-called angular intra-prediction. This is a prediction that forecasts a currently predicted pixel from a neighboring, already reconstructed pixel that can be retrieved at a specific angular direction. For example, a skyscraper might exhibit largely the same colors in vertical lines that follow the structure of the windows and the wall between them. Retrieving a neighboring pixel that is vertically higher (corresponding to angular direction 26 in, for example, HEVC) provides a good prediction for the color components of the current pixel.

[0006] In older MPEG codecs, the prediction direction or directions for a block were usually explicitly signaled.

[0007] With newer codecs, it has become clear that data processing can be more cost-effective than communication for some devices or communication systems. Consequently, one goal can be for the decoder to recognize vertical patterns on its own, eliminating the need for explicit instructions to use the vertical angle prediction mode. Only correction bits need to be sent if necessary, thus saving a few bits.

[0008] Template-based intra-mode derivation (TIMD) is a technique that allows the decoder to autonomously derive what should be a good prediction, such as an angle prediction of 45 degrees. In this approach, a template is created, for example, block width × height = 2 pixels above a block (the current block to be predicted, as well as all surrounding blocks that could be good candidates for the prediction). The predictability is then checked, for example, for an angle prediction of 45 degrees. The current block itself has not yet been decoded (the intra-prediction mode must first be set up), so these pixels cannot be used for the predictability check. However, another adjacent area can be used, located slightly further above the block and bordering the template; this second area is called the reference.If a diagonal pattern extends downward through the reference area into the template block (both for the currently decoded block and for a good candidate reference block on which the prediction of the block pixel color components can be based), it is likely that this pattern will continue downward from the template area into the block area and thus still be a good predictor. Different prediction directions can be tested, and a cost measurement can indicate which directions are the best predictors (e.g., downward rather than diagonally for the skyscraper).

[0009] It is desirable to continuously improve these possibilities. BRIEF SUMMARY OF THE INVENTION

[0010] The encoding of video images can be improved by using a method for predicting a block of pixel colors in an image, including: - Determine that the template-based intra-mode derivative is applicable for predicting the block; - Determining a variety of intra-forecast modes, and determining the respective costs for the variety of intra-forecast modes; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode (i.e., is not identical to it) and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming a final TIMD mode comprising a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode.

[0011] A non-angular mode is a mode that is not an angular mode.

[0012] The procedure may incur costs, each based on the differences between predicted samples obtained by applying the respective intra-prediction mode to samples of a template reference and reconstructed template samples.

[0013] The procedure for determining the third TIMD mode may depend on the condition that the first TIMD mode and the second TIMD mode are not non-angular intra-prediction modes.

[0014] The procedure, whereby the determination of the final TIMD mode depends on a third cost value of the non-angular intra-prediction mode being less than a threshold.

[0015] The procedure, whereby the threshold costs are based on the lowest of the costs of the first TIMD mode and the second TIMD mode.

[0016] The procedure, where the threshold costs are a multiplication of a scaling value by the lowest of the costs of the first TIMD mode and the second TIMD mode.

[0017] The method in which the linear combination involves multiplying each TIMD mode by a weighting that depends on its cost.

[0018] The method in which the weights are calculated as the numerator of the sum of all weights minus the respective weighting of the respective TIMD mode divided by a denominator that is a multiple of the sum of all weights.

[0019] The method in which the non-angular intra-prediction mode is a DC mode or a planar mode.

[0020] The method according to one of the preceding claims, wherein the non-angle-based intra-prediction mode is based on a reference block that is usable for the intra-block copy of an adjacent block of the block, or on a reference block that is usable for the intra-TMP prediction of an adjacent block of the block.

[0021] The procedure, wherein the non-angular intra-forecast mode is selected as the one with the lowest cost from a set of non-angular intra-forecast modes.

[0022] The invention may be useful in devices comprising one of the following circuits: a video decoder comprising a computing circuit connected to a data storage device, wherein the computing circuit is arranged to: - Receiving a bitstream containing encoded video data; - wherein the encoded video data includes an initial indication that a pixel block is encoded based on a template-based intra-mode derivation; - wherein the encoded video data includes a second indication that a final TIMD mode is based on a non-angle-dependent intra-prediction mode; - Determining a variety of costs for a variety of intra-forecast modes, where each cost factor for a given IPM is based on differences in a template of the block between the following: firstly predicted samples obtained by prediction from reference samples of the template using the respective IPM; and subsequently reconstructed samples of the original; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming the final TIMD mode, which comprises a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode.

[0023] A comprehensive video encoder: - Receiving an original video image that includes a block to be encoded using template-based intra-prediction coding; - Determining a variety of intra-forecast modes, and determining the respective costs for the variety of intra-forecast modes; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming a final TIMD mode comprising a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode. - Encoding in encoded video data a first indication that a pixel block is encoded based on a template-based intra-mode derivation, and a second indication that a final TIMD mode is based on a non-angle-based intra-prediction mode. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Some features are shown in the accompanying drawings as examples and are not exhaustive. In the drawings, identical reference symbols refer to similar elements. Fig. Figure 1 shows an example of a video encoding / decoding system in which embodiments of the present disclosure can be implemented. Fig. Figure 2 shows an example of an encoder in which embodiments of the present disclosure can be implemented. Fig. Figure 3 shows an example of a decoder in which embodiments of the present disclosure can be implemented. Fig. Figure 4 shows an example quadtree partitioning of a Coding Tree Block (CTB). Fig. Figure 5 shows an exemplary quadtree, which is based on the exemplary quadtree partitioning of the CTB in Fig. 4 corresponds to. Fig. Figure 6 shows examples of binary tree and ternary tree partitions. Fig. Figure 7 shows an example of a combined quadtree and multi-type tree partitioning of a CTB. Fig. Figure 8 shows an example tree, which uses the combined quadtree and multi-type tree partitioning of the in Fig. 7 CTB shown corresponds to this. Fig. Figure 9 shows an exemplary set of reference samples chosen for intraprediction of a current block. Fig. 10A and Fig. Figure 10B shows exemplary intraprediction modes. Fig. Figure 11 shows an example of a current block and associated reference samples. Fig. Figure 12 shows an example of the application of an intraprediction mode (e.g., an angle mode) to predict a current block. Fig. 13A shows an example of an interprediction performed for a current block in a current image. Fig. Figure 13B shows an example motion vector. Fig. Figure 14 shows an example of a bi-prediction performed for a current block. Fig. 15A shows examples of spatial candidate neighboring blocks relative to a currently coded block. Fig. Figure 15B shows exemplary positions of two temporal blocks located at the same point relative to a current block. Fig. Figure 16A shows an example of an intra-block copy (IBC). Fig. Figure 16B shows an example of a template-matching prediction (TMP) mode (e.g., Intra-TMP mode) for predicting or determining a current block according to some embodiments. Fig. Figure 17A shows an example of template-based intra-mode derivation (TIMD) for encoding a current block according to some embodiments. Fig. Figure 17B shows an example of template-based intra-mode derivation (TIMD) for encoding a current block using an adjacent block encoded with intra-block copy (IBC) or intra-TMP, according to some embodiments. Fig. Figure 18 shows an example of TIMD signaling for decoding a current block according to some embodiments. Fig. Figure 19 shows a flowchart of a procedure for applying a TIMD technique to a current block according to some embodiments. Fig. Figure 20 shows a flowchart of a procedure for applying a TIMD technique to a current block according to some embodiments. Fig. Figure 21 shows a flowchart of a procedure for applying a TIMD technique to a current block according to some embodiments. Fig. Figure 22 shows a flowchart of a method for applying a TIMD technique to decode a current block according to some embodiments. Fig. Figure 23 shows a flowchart of a procedure for applying a TIMD technique to encode a current block according to some embodiments. Fig. Figure 24 shows a flowchart of a procedure for applying a TIMD technique to a current block according to some embodiments. Fig. Figure 25 shows a block diagram of an example of a computer system in which embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0025] The following description presents numerous specific details to facilitate a comprehensive understanding of the revelation. However, it is obvious to those skilled in the art that the revelation, including its structures, systems, and procedures, can be implemented even without these specific details. The description and presentation provided here are the usual means used by experienced professionals in this field to communicate the content of their work to other professionals in the field as effectively as possible. In other instances, known procedures, processes, components, and logic have not been described in detail so as not to obscure aspects of the revelation unnecessarily.

[0026] References in the description to "an embodiment," "a specific form," "an exemplary embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or property, but that not every embodiment necessarily includes that particular feature, structure, or property. Furthermore, such expressions do not necessarily refer to the same embodiment. Moreover, when a particular feature, structure, or property is described in connection with an embodiment, it is assumed that it is within the knowledge of those skilled in the art to apply that feature, structure, or property in connection with other embodiments, regardless of whether these are explicitly described or not.

[0027] Furthermore, it should be noted that individual implementations can be described as a process, which is represented as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart can describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. Moreover, the order of the operations can be rearranged. A process terminates when its operations are completed, but it may still contain further steps not included in a diagram. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. If a process corresponds to a function, its termination can correspond to the function returning to the calling function or to the main function.

[0028] The term "computer-readable medium" includes, but is not limited to, portable and non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A computer-readable medium may include a non-transitory medium capable of storing data that does not contain carrier waves and / or volatile electronic signals that propagate wirelessly or via wired connections. Examples of non-transitory media include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, and storage devices.A computer-readable medium can store code and / or machine-executable instructions that may represent a procedure, function, subroutine, program, routine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be transferred by any suitable means, including memory release, message transmission, token transmission, network transmission, or the like.

[0029] Furthermore, embodiments can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments required to perform the necessary tasks (e.g., a computer program product) can be stored on a computer-readable or machine-readable medium. A processor (or processors) can then perform the necessary tasks.

[0030] A video sequence comprising multiple images / frames can be represented digitally for storage and / or transmission. Representing a video sequence digitally can require a large number of bits. Large amounts of data associated with video sequences can require significant resources for storage and / or transmission. Video encoding can be used to compress the size of a video sequence for more efficient storage and / or transmission. Video decoding can be used to decompress a compressed video sequence for viewing and / or other uses.

[0031] Fig. Figure 1 shows an example of a video encoding / decoding system 100 in which embodiments of the present disclosure can be implemented. The video encoding / decoding system 100 comprises a source device 102, a transmission medium 104, and a destination device 106. The source device 102 encodes a video sequence 108 into a bitstream 110 for more efficient storage and / or transmission. The source device 102 can store and / or send / transmit the bitstream 110 to / from the destination device 106 via the transmission medium 104. The destination device 106 decodes the bitstream 110 to display the video sequence 108. The destination device 106 can receive the bitstream 110 from the source device 102 via the transmission medium 104. The source device 102 and / or the destination device 106 can be any of a variety of different devices (e.g.,a desktop computer, laptop computer, tablet computer, smartphone, portable device, television, camera, video game console, set-top box, video streaming device, etc.).

[0032] Source device 102 can (e.g., for encoding video sequence 108 into bitstream 110) comprise one or more video sources 112, encoders 114, and / or output interfaces 116. Video source 112 can provide and / or generate a video sequence 108 based on a recording of a natural scene and / or a synthetically generated scene. A synthetically generated scene can be a scene consisting of computer-generated graphics and / or screen content. The video source 112 can comprise a video recording device (e.g., a video camera), a video archive containing previously recorded natural scenes and / or synthetically generated scenes, a video feed interface for receiving recorded natural scenes and / or synthetically generated scenes from a video content provider, and / or a processor for generating synthetic scenes.

[0033] A video sequence, such as video sequence 108, can comprise a series of images (also called frames). The impression of motion in a video sequence can be achieved by displaying images of the video sequence sequentially with a constant or variable time interval between the images. An image can comprise one or more sample arrays containing intensity values. The intensity values ​​can be captured (e.g., measured, determined, provided) at a series of regularly spaced locations within an image. A color image can (e.g., typically comprises) one luminance sample array and two chrominance sample arrays. The luminance sample array can include intensity values ​​representing the brightness (e.g., luma component, Y) of an image. The chrominance sample arrays can include intensity values ​​representing the blue and red components of an image (e.g., chroma components, Cb and Cr), respectively, separately from the brightness.Other color sample arrays may be based on other color schemes (e.g., a red-green-blue (RGB) scheme). A pixel in a color image can refer to / comprise / be associated with all intensity values ​​(e.g., luma component, chroma components) for a given position in the sample arrays used to represent color images (e.g., three sample arrays are used for one luma component or two chroma components). A monochrome image may contain a single luminance sample array. A pixel in a monochrome image can refer to / comprise / be associated with the intensity value (e.g., luma component) at a specific position in the single luminance sample array used to represent monochrome images.

[0034] Encoder 114 can encode video sequence 108 into bitstream 110. Encoder 114 can apply / use one or more prediction techniques (e.g., for encoding video sequence 108) to reduce redundant information in video sequence 108. Redundant information is information that can be predicted at a decoder and does not need to be transmitted to the decoder for accurate decoding of video sequence 108. For example, Encoder 114 can apply spatial prediction (e.g., intra-frame or intraprediction), temporal prediction (e.g., inter-frame or interprediction), inter-layer prediction, and / or other prediction techniques to reduce redundant information in video sequence 108. Encoder 114 can partition the images comprising video sequence 108, for example, into rectangular areas called blocks, before applying one or more prediction techniques.Coder 114 can then code a block using one or more prediction techniques.

[0035] For temporal prediction, encoder 114 can search for a block similar to the block being encoded in another frame (e.g., called the reference frame) of video sequence 108. The block identified during the search (e.g., called the prediction block) can then be used to predict the encoded block. For spatial prediction, encoder 114 can construct a prediction block based on data from reconstructed neighboring samples of the block to be encoded within the same frame of video sequence 108. A reconstructed sample is a sample that has been encoded and subsequently decoded. Encoder 114 can determine a prediction error (e.g., also called the residual) based on the difference between an encoded block and a prediction block. The prediction error can represent non-redundant information that can be sent / transmitted to a decoder for accurate decoding of video sequence 108.

[0036] Encoder 114 can apply a transformation to the prediction error (e.g., using a discrete cosine transform (DCT) or another transformation) to generate transformation coefficients. Encoder 114 can then form bitstream 110 based on these transformation coefficients and other information, which is used to determine prediction blocks based on prediction types, motion vectors, and / or prediction modes.

[0037] For example, encoder 114 can perform one or more quantization and entropy coding operations on the transformation coefficients and / or other information used to determine the prediction blocks before the bitstream 110 is formed. The quantization and / or entropy coding can further reduce the number of bits required to store and / or transmit the video sequence 108.

[0038] Output interface 116 can be configured to write and / or store the bitstream 110 on the transmission medium 104 for transmission to the destination device 106. Additionally or alternatively, output interface 116 can be configured to send / transmit, upload, and / or stream the bitstream 110 to the destination device 106 via the transmission medium 104. Output interface 116 can include a wired and / or wireless transmitter configured to send / transmit, upload, and / or stream the bitstream 110 according to one or more proprietary, open-source, and / or standardized communication protocols (e.g.,...).Standards of Digital Video Broadcasting (DBV), Advanced Television Systems Committee (ATSC), Integrated Services Digital Broadcasting (ISDB), Data Over Cable Service Interface Specification (DOCSIS), 3rd Generation Partnership Project (3GPP), Institute of Electrical and Electronics Engineers (IEEE), Internet Protocol (IP), Wireless Application Protocol (WAP) and / or any other communications protocol).

[0039] The transmission medium 104 can be wireless, wired, and / or computer-readable. For example, the transmission medium 104 can include one or more wires, cables, air interfaces, optical disks, flash memory, and / or magnetic storage. Additionally or alternatively, the transmission medium 104 can include one or more networks (e.g., the Internet) or file servers configured to store and / or send / transmit encoded video data.

[0040] Target device 106 can decode bitstream 110 for display as video sequence 108. Target device 106 can include one or more input interfaces 118, decoders 120, and / or video displays 122. Input interface 118 can be configured to read the bitstream 110 stored by source device 102 on the transmission medium 104. Additionally or alternatively, input interface 118 can be configured to receive, download, and / or stream the bitstream 110 from source device 102 via the transmission medium 104. Input interface 118 can include a wired and / or wireless receiver configured to receive, download, and / or stream the bitstream 110 according to one or more proprietary, open-source, standardized communication protocols, and / or any other communication protocol (e.g., as referenced herein).

[0041] Decoder 120 can decode video sequence 108 from encoded bitstream 110. Decoder 120 can generate prediction blocks for frames of video sequence 108 in a similar way to encoder 114 and determine the prediction errors for the blocks in order to decode video sequence 108. Decoder 120 can generate the prediction blocks using / based on prediction types, prediction modes, and / or motion vectors received in bitstream 110. Decoder 120 can determine the prediction errors using the transformation coefficients received in bitstream 110. Decoder 120 can determine the prediction errors by weighting the transformation basis functions using the transformation coefficients. Decoder 120 can combine the prediction blocks and the prediction errors to decode video sequence 108.Video sequence 108 on the destination device 106 can, but does not necessarily have to, be the same transmitted video sequence as the video sequence 108 transmitted by source device 102. Decoder 120 can decode a video sequence that approximates video sequence 108, for example due to lossy compression of video sequence 108 by encoder 114 and / or due to errors introduced into the encoded bitstream 110 during transmission to the destination device 106.

[0042] Video display 122 can display video sequence 108 to a user. Video display 122 can include a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display and / or any other display device suitable for displaying video sequence 108.

[0043] Video encoding / decoding system 100 is merely an example, and other video encoding / decoding systems besides video encoding / decoding system 100 and / or modified versions of video encoding / decoding system 100 can perform the procedures and processes described herein. For example, video encoding / decoding system 100 may include other components and / or arrangements. For example, the video source 112 may be located outside the source device 102. Likewise, video display 122 may be located outside the destination device 106 or may be omitted entirely (e.g., if video sequence 108 is intended for use by a machine and / or storage device). In one example, source device 102 may further include a video decoder, and destination device 106 may further include a video encoder.For example, source device 102 can be configured to further receive an encoded bitstream from destination device 106 to support bidirectional video transmission between the devices.

[0044] Encoder 114 and / or decoder 120 can operate according to one or more proprietary or industry-specific video coding standards. For example, Encoder 114 and / or decoder 120 can operate according to one or more proprietary, open-source, and / or standardized protocols (e.g., International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.263, ITU-T H.264, and Moving Picture Expert Group (MPEG)-4 Visual (also known as Advanced Video Coding (AVC)), ITU-T H.265 and MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)), ITU-T H.265 and MPEG-I Part 3 (also known as Versatile Video Coding (VVC)), the WebM VP8 and VP9 codecs, and / or AOMedia Video 1 (AV1), and / or any other video coding protocol).

[0045] Fig. Figure 2 shows an example encoder. The one in Fig. The encoder 200 shown can implement one or more of the processes described herein. Encoder 200 can encode a video sequence 202 into a bitstream 204 for more efficient storage and / or transmission. Encoder 200 can be used in the video encoding / decoding system 100 as described in Fig. 1 shown (e.g., as encoder 114) or implemented in any computer, communication, or electronic device (e.g., desktop computer, laptop computer, tablet computer, smartphone, portable device, television, camera, video game console, set-top box, video streaming device, etc.). Encoder 200 can comprise one or more of an inter-prediction unit 206, an intra-prediction unit 208, combiners 210 and 212, a transform and quantize unit (TR + Q) 214, an inverse transform and quantize unit (iTR + iQ) 216, an entropy coding unit 218, one or more filters 220, and / or a buffer 222.

[0046] Encoder 200 can partition images (e.g., frames) of video sequence 202 (e.g., comprising these) into blocks and encode video sequence 202 block by block. Encoder 200 can apply / perform a prediction technique on a block being encoded using either the Interprediction Unit 206 or the Intraprediction Unit 208. Interprediction Unit 206 can perform an interprediction by searching for a block similar to the block encoded in another, reconstructed image (e.g., a reference image) of video sequence 202. A reconstructed image is an image that has been encoded and subsequently decoded. The block identified during the search (e.g., referred to as the prediction block) can then be used to predict the block to be encoded in order to remove redundant information.Interprediction Unit 206 can exploit temporal redundancy or similarities in the scene content from frame to frame in video sequence 202 to determine the prediction block. For example, the scene content between the frames of video sequence 202 may be similar, apart from differences due to motion and / or affine transformation of the screen content over time.

[0047] Intraprediction Unit 208 can perform an intraprediction by forming a prediction block based on data from reconstructed neighboring samples of the block to be coded within the same frame of video sequence 202. A reconstructed sample is a sample that has been coded and subsequently decoded. Intraprediction Unit 208 can exploit spatial redundancy or similarities in scene content within a frame of video sequence 202 to determine the prediction block. For example, the texture of an area containing scene content in a frame may resemble the texture in the immediate vicinity of the area containing scene content in the same frame.

[0048] Encoder 210 can determine a prediction error (e.g., referred to as residual) based on the difference between a coded block and the predicted block. The prediction error can represent non-redundant information that can be sent / transmitted to a decoder for the accurate decoding of video sequence 202.

[0049] Transformation and quantization unit (TR + Q) 214 can transform and quantize the prediction error. Transformation and quantization unit 214 can transform the prediction error into transformation coefficients, for example, by applying a DCT to reduce correlated information in the prediction error. Transformation and quantization unit 214 can quantize the coefficients by mapping data of the transformation coefficients to a predefined set of representative values. Transformation and quantization unit 214 can quantize the coefficients to reduce irrelevant information in the bitstream 204. Irrelevant information is information that can be removed from the coefficients without causing any visible and / or perceptible distortion in the video sequence 202 after decoding (e.g., on a receiving device).

[0050] The entropy coding unit 218 can apply one or more entropy coding methods to the quantized transformation coefficients to further reduce the bit rate. For example, the entropy coding unit 218 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and / or syntax-based context-based binary arithmetic coding (SBAC). The entropy-coded coefficients can be packed to form the bitstream 204.

[0051] The inverse transformation and quantization unit (iTR + iQ) 216 can inversely quantize and inversely transform the quantized transformation coefficients to determine a reconstructed prediction error. The combiner 212 can combine the reconstructed prediction error with the prediction block to form a reconstructed block. The filter 220 can filter the reconstructed block, for example, using a deblocking filter and / or a sample adaptive offset (SAO) filter. The buffer 222 can store the reconstructed block for predicting one or more other blocks in the same and / or a different frame of the video sequence 202.

[0052] The Encoder 200 may also include an Encoder Control Unit. The Encoder Control Unit may be configured to control one or more units of the Encoder 200, as shown in Fig. Figure 2 shows the encoder control unit. The encoder control unit can control the one or more units of the encoder 200 so that the bitstream 204 can be generated in accordance with the requirements of one or more proprietary coding protocols, industry standards for video coding, and / or any other video coding protocol. For example, the encoder control unit can control the one or more units of the encoder 200 so that the bitstream 204 can be generated in accordance with one or more of ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, ​​AV1, and / or any other video coding standard / format.

[0053] The encoder control unit can be configured to attempt to minimize (or reduce) the bit rate of bitstream 204 and / or maximize (or increase) the reconstructed video quality (e.g., within the limitations of a proprietary coding protocol, an industry video coding standard, and / or any other video coding protocol). For example, the encoder control unit can be configured to attempt to minimize or reduce the bit rate of bitstream 204 so that the reconstructed video quality does not fall below a certain level / threshold, and / or to maximize or increase the reconstructed video quality so that the bit rate of bitstream 204 does not exceed a certain level / threshold.The encoder control unit can determine / control one or more of: the partitioning of the images of the video sequence 202 into blocks, whether a block is interpredicted by the interprediction unit 206 or intrapredicted by the intraprediction unit 208, a motion vector for the interprediction of a block, an intraprediction mode from a variety of intraprediction modes for the intraprediction of a block, the filtering performed by filter 220, and / or one or more transformation types and / or quantization parameters applied by the transformation and quantization unit 214. The encoder control unit can determine / control one or more of the above based on a rate distortion measure for a block or image being encoded.The encoder control unit can determine / control one or more of the above to reduce the rate distortion measure for a block or image being encoded.

[0054] The prediction type used to encode a block (intra- or inter-prediction), the block's prediction information (intra-prediction mode for intra-prediction, motion vector, etc.), and / or transformation and / or quantization parameters can be sent to the entropy coding unit 218 for further compression (e.g., to reduce the bit rate). For example, the entropy coding unit 218 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and / or syntax-based context-based binary arithmetic coding (SBAC) to achieve further compression. The prediction type, prediction information, and / or transformation and / or quantization parameters can be packed with the prediction error to form the bitstream 204.

[0055] Coder 200 is merely an example, and coders other than Coder 200 and / or modified versions of Coder 200 can perform the procedures and processes described herein. For example, Coder 200 may include other components and / or arrangements. One or more of the Fig. The two components shown can optionally be included in the encoder 200 (e.g., entropy coding unit 218 and / or filter 220).

[0056] Fig. Figure 3 shows an example decoder. One in Fig. Decoder 300 shown can implement one or more of the processes described herein. Decoder 300 can decode a bitstream 302 into a decoded video sequence 304 for display and / or other use. Decoder 300 can be used in the video encoding / decoding system 100. Fig. 1 and / or be implemented in a computer, communication, or electronic device (e.g., desktop computer, laptop computer, tablet computer, smartphone, portable device, television, camera, video game console, set-top box, and / or video streaming device). Decoder 300 may comprise an entropy decoding unit 306, an inverse transform and quantization unit (iTR + iQ) 308, a combiner 310, one or more filters 312, a buffer 314, an interprediction unit 316, and / or an intraprediction unit 318.

[0057] Decoder 300 can include a decoder control unit configured to control one or more Decoder 300 units. The decoder control unit can control the one or more Decoder 300 units to decode the 302 bitstream in accordance with the requirements of one or more proprietary coding protocols, video coding industry standards, and / or any other communication protocol. For example, the decoder control unit can control the one or more Decoder 300 units to decode the 302 bitstream in accordance with one or more of ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, ​​AV1, and / or any other video coding standard / format.

[0058] The decoder control unit can determine / control one or more of: whether a block is interpredicted by the interprediction unit 316 or intrapredicted by the intraprediction unit 318, a motion vector for the interprediction of a block, an intraprediction mode among several intraprediction modes for the intraprediction of a block, filtering performed by the filter(s) 312, and / or one or more inverse transformation types and / or inverse quantization parameters to be applied by the inverse transformation and quantization unit 308. One or more of the control parameters used by the decoder control unit can be packed in the bitstream 302.

[0059] The entropy decoding unit 306 can subject the bitstream 302 to entropy decoding. For example, the entropy decoding unit 306 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and syntax-based context-based binary arithmetic coding (SBAC) to decompress the prediction type used to encode a block (intra- or inter-prediction), the block's prediction information (intra-prediction mode for intra-prediction, motion vector, etc.), and transformation and quantization parameters. The inverse transformation and quantization unit 308 can inversely quantize and / or inversely transform the quantized transformation coefficients to determine a decoded prediction error. The combiner 310 can combine the decoded prediction error with the prediction block to form a decoded block.The forecast block can be generated by the intra-forecasting unit 318 or the inter-forecasting unit 316 (e.g., as above with respect to encoder 200 in ). Fig. (described in section 2). Filter 312 can filter the decoded block, for example, using a deblocking filter and / or a sample adaptive offset (SAO) filter. Buffer 314 can store the decoded block in bitstream 302 to predict one or more other blocks in the same and / or a different frame of the video sequence. The decoded video sequence 304 can be output by filter(s) 312, as described in Fig. 3 shown.

[0060] Decoder 300 is merely an example, and decoders other than Decoder 300 and / or modified versions of Decoder 300 can perform the procedures and processes described herein. For example, Decoder 300 may have different components and / or arrangements. One or more of the Fig. The 3 components shown can optionally be included in decoder 300 (e.g. entropy decoding unit 306 and / or filter 312).

[0061] Although in Fig. 2 and Fig. Not shown in Figure 3, both encoder 200 and decoder 300 can further include an intrablock copy unit in addition to the interprediction and intraprediction units. The intrablock copy unit can function similarly to an interprediction unit, but can predict blocks within the same image. For example, the intrablock copy unit can exploit repeating patterns that appear in the screen content. The screen content can include computer-generated text, graphics, animations, etc.

[0062] Video encoding and / or decoding can be performed block by block. The process of partitioning an image into blocks can be adaptive depending on the image content. For example, larger block partitions can be used in image areas with higher homogeneity to improve encoding efficiency.

[0063] An image (e.g., in HEVC or another encoding standard / format) can be divided into non-overlapping square blocks, which can be called Coding Tree Blocks (CTBs). The CTBs can contain samples from a sample array. A CTB can have a size of 2n x 2n samples, where n can be specified by a parameter of the encoding system. For example, n can be 4, 5, 6, or any other value. A CTB can also have any other size. A CTB can further subdivide into coding blocks (CBs) of half the vertical and half the horizontal size by recursive quadtree partitioning. The CTB can form the root of the quadtree. A CB that is not further subdivided by recursive quadtree partitioning can be called the leaf CB of the quadtree; otherwise, it is called the non-leaf CB of the quadtree. A CB can have a minimum size, which is specified by a parameter of the encoding system.For example, a CB can have a minimum size of 4x4, 8x8, 16x16, 32x32, 64x64 samples, or another minimum size. A CB can be further partitioned into one or more prediction blocks (PBs) to perform inter- and / or intra-prediction. A PB can be a rectangular block of samples to which the same prediction type / mode can be applied. For transformations, a CB can be partitioned into one or more transformation blocks (TBs). A TB can be a rectangular block of samples that can determine / specify an applied transformation size.

[0064] Fig. Figure 4 shows an example of a quadtree partitioning of a CTB 400. Fig. Figure 5 shows an exemplary Quadtree 500, which is based on the exemplary Quadtree partitioning of the CTB 400 in Fig. 4 corresponds. As in the examples of Fig. 4 and Fig. As shown in Figure 5, CTB 400 can first be partitioned into four CBs with half the vertical and half the horizontal size. Three of the resulting CBs from the first-level partitioning of CTB 400 are leaf CBs. The three leaf CBs from the first-level partitioning of CTB 400 are in Fig. 4 and Fig. 5 are each labeled 7, 8, and 9. The non-leaf CB of the first-level partitioning of CTB 400 is partitioned into four sub-CBs with half the vertical and half the horizontal size. Three of the resulting sub-CBs of the second-level partitioning of CTB 400 are leaf CBs. The three leaf CBs of the second-level partitioning of CTB 400 are in Fig. 4 and Fig. 5 are each designated 0, 5 and 6. Finally, the non-leaf CB of the second-level partitioning of CTB 400 is partitioned into four leaf CBs with half the vertical and half the horizontal size. The four leaf CBs are in Fig. 4 and Fig. 5 is labelled with 1, 2, 3 and 4 respectively.

[0065] The exemplary CTB 400 from Fig. 4 is partitioned into 10 leaf CBs, each labeled 0 through 9, but can be partitioned into other sets of leaf CBs. The 10 leaf CBs can correspond to 10 CB leaf nodes (e.g., 10 CB leaf nodes of Quadtree 500, as in Fig. 5 shown). In other examples, a CTB can be partitioned into a different number of leaf CBs. The resulting quadtree partitioning of CTB 400 can be scanned using a Z-scan (e.g., left to right, top to bottom) to establish the sequence order for encoding / decoding the CB leaf nodes. A numerical label (e.g., indicator, index) for each CB leaf node in Fig. 4 and Fig. 5 can correspond to the sequence order for encoding / decoding. For example, CB leaf node 0 can be encoded / decoded first and CB leaf node 9 last. Although in Fig. 4 and Fig. Not shown in Figure 5, each CB leaf node can contain one or more PBs and / or TBs.

[0066] An image in VVC (or any other encoding standard / format) can be partitioned in a similar way (as in HEVC). An image can first be partitioned into non-overlapping square CTBs. The CTBs can then be partitioned into CBs of half vertical and half horizontal size using recursive quadtree partitioning. A quadtree leaf node (e.g., in VVC) can be further partitioned into CBs of different sizes using binary or ternary tree partitioning (or any other partitioning method).

[0067] Fig. Figure 6 shows examples of binary and ternary tree partitions. A binary tree partition can divide a parent block into two halves, either vertically (602) or horizontally (604). The resulting partitions can be half the size of the parent block. In other examples, the resulting partitions can be smaller and / or larger than half the size of the parent block. A ternary tree partition can divide a parent block into three parts, either vertically (606) or horizontally (608). Fig. Figure 6 shows an example where the middle partition can be twice the size of the other two end partitions in the ternary tree partitions. In other examples, partitions can have different sizes relative to each other and to the parent block. Binary and ternary tree partitions are examples of multi-type tree partitioning. Multi-type tree partitioning can involve partitioning a parent block into other sets of smaller blocks. The block partitioning strategy (e.g., in VVC) can be described as a combination of quadtree and multi-type tree partitioning (quadtree + multi-type tree partitioning), since binary and / or ternary tree partitioning is added to quadtree partitioning.

[0068] Fig. Figure 7 shows an example of a combined quadtree and multi-type tree partitioning of a CTB 700. Fig. Figure 8 shows an example of a tree 800, which uses the combined quadtree and multi-type tree partitioning of the in Fig. The CTB 700 shown in section 7 corresponds to this. In both Fig. In both Figures 7 and 8, quadtree partitioning is represented by solid lines and multi-type tree partitioning by dashed lines. For simplicity, CTB 700 uses the same quadtree partitioning as the one in Figure 7. Fig. The CTB 400 described in section 4 is shown, and a description of the quadtree partitioning of CTB 700, which is similar to that for CTB 400, is omitted. The quadtree partitioning of CTB 700 is merely an example, and a CTB can also be quadtree-partitioned in ways other than the CTB 700. Additional multi-type tree partitions of CTB 700 can be shown relative to those in section 4. Fig. The three-sheet CBs shown in section 4 are to be performed. The three-sheet CBs in Fig. 4, which are in Fig. If leaf CBs 7 are shown as further subdivided, the leaf CBs can be 5, 8, and 9. The three leaf CBs can be further partitioned using one or more binary and / or ternary tree partitions.

[0069] The leaf CB 5 from Fig. 4 can be partitioned into two CBs based on a vertical binary tree partitioning. The two resulting CBs can be leaf CBs, which are in Fig. 7 and Fig. Sheet 8 is marked with 5 and 6 respectively. Sheet CB 8 of Fig. Tree 4 can be partitioned into three CBs based on a vertical ternary tree partitioning. Two of the three resulting CBs can be leaf CBs, which are divided into Fig. 7 and Fig. 8 are labeled 9 and 14, respectively. The remaining non-leaf CB can first be partitioned into two CBs based on a horizontal binary tree partitioning. One of the two CBs can be a leaf CB labeled 10. The other of the two CBs can be further partitioned into three CBs based on a vertical ternary tree partitioning. The resulting three CBs can be leaf CBs, which are in Fig. 7 and Fig. 8 are each labelled 11, 12 and 13. Sheet CB 9 of Fig. Tree 4 can be partitioned into three CBs based on a horizontal ternary tree partitioning. Two of the three CBs can be leaf CBs, which are divided into Fig. 7 and Fig. 8 are each labeled 15 and 19. The remaining non-leaf CB can be partitioned into three CBs based on further horizontal ternary tree partitioning. The resulting three CBs can be all leaf CBs that are in Fig. 7 and Fig. 8 are each labelled 16, 17 and 18.

[0070] In total, CTB 700 can be partitioned into 20 leaf CBs, each designated 0 to 19. These 20 leaf CBs can correspond to 20 leaf nodes (e.g., 20 leaf nodes of the in Fig. 8 tree 800 shown). The resulting combination of quadtree and multi-type tree partitioning of the CTB 700 can be scanned with a Z-scan (from left to right, top to bottom) to form the sequence order for encoding / decoding the CB leaf nodes. A numerical label for each CB leaf node in Fig. 7 and Fig. 8 can correspond to the sequence order for encoding / decoding, with CB leaf node 0 being encoded / decoded first and CB leaf node 19 being encoded / decoded last. Although in Fig. 7 and Fig. Not shown in Figure 8, it should be noted that each CB leaf knot can contain one or more PBs and / or TBs.

[0071] A coding standard / format (e.g., HEVC, VVC, or any coding standard / format) can define various units (e.g., in addition to specifying different blocks (e.g., CTBs, CBs, PBs, TBs)). Blocks can comprise a rectangular range of samples in a sample array. Units can include the combined sample blocks from the different sample arrays (e.g., luma and chroma sample arrays) that form an image, as well as syntax elements and prediction data of the blocks. A coding tree unit (CTU) can comprise the combined CTBs of the different sample arrays (i.e., the Y, Cb, and Cr pixel color components) and form a complete unit in a coded bitstream. A coding unit (CU) can comprise the combined CBs of the different sample arrays and the syntax structures used to encode the samples of the CBs.A prediction unit (PU) can comprise the combined prediction blocks (PBs) of the various sample arrays and syntax elements used to predict the PBs. A transformation unit (TU) can comprise transformation blocks (TBs) of the various sample arrays and syntax elements used to transform the TBs. The prediction block is a set of pixel positions for which a certain prediction would apply (e.g., a first block is predicted with motion compensation, and the adjacent block is intra-predicted), and a transformation unit would typically be a set of pixels (e.g., 8 horizontally by 8 vertically) to which a certain transformation is applied, usually to the residues after a prediction, such as a discrete cosine transform or a discrete sine transform.PBs and TBs (or units) do not need to be combined, as modern codecs have considerable flexibility in deciding what is best. For example, an 8x16 region might be predicted by a single motion-compensated prediction (e.g., if it is part of a large background making the same panning motion), but it might be advantageous to split it into two 8x8 transformation blocks, or vice versa.

[0072] A block can refer to CTB, CB, PB, TB, CTU, CU, PU, ​​and / or TU (e.g., in the context of HEVC, VVC, or any other encoding format / standard). A block can be used to reference similar data structures in the context of any video encoding format / standard / protocol. For example, a block can refer to a macroblock in the AVC standard, a macroblock or subblock in the VP8 encoding format, a superblock or subblock in the VP9 encoding format, and / or a superblock or subblock in the AV1 encoding format.

[0073] In intraprediction, samples of a block to be coded (e.g., also called the current block) can be predicted from samples in the column immediately adjacent to the leftmost column of the current block and samples in the row immediately adjacent to the topmost row of the current block. The samples from the immediately adjacent column and row can be collectively referred to as reference samples. Each sample of the current block can be predicted (e.g., in an intraprediction mode) by projecting the sample's position in the current block in a specific direction onto a point along the reference samples. The sample can be predicted by interpolation between the two nearest reference samples to the projection point if the projection does not fall directly on a reference sample. A prediction error (e.g.,(referred to as residual) can be determined for the current block based on differences between the predicted sample values ​​and the original sample values ​​of the current block.

[0074] Predicting samples and determining a prediction error based on the difference between the predicted samples and the original samples can be performed (e.g., at an encoder) for a variety of intraprediction modes (e.g., including non-directional intraprediction modes). The encoder can select any of the many intraprediction modes and the corresponding prediction error to code the current block. The encoder can send information about the selected prediction mode and the corresponding prediction error to a decoder to decode the current block. The decoder can decode the current block by predicting the samples of the current block using the intraprediction mode specified by the encoder and / or by combining the predicted samples with the prediction error.

[0075] Fig. Figure 9 shows an exemplary set of reference samples 902, which were determined for intraprediction of a current block 904. The current block 904 can correspond to a block that is being encoded and / or decoded. The current block 904 can correspond to block 3 of the partitioned CTB 700, as shown in Fig. Figure 7 shows that, as described herein, the numerical designations 0 to 19 of the blocks of the partitioned CTB 700 can correspond to the sequence order for encoding / decoding the blocks and can be used as such in the example of Fig. 9 can be used.

[0076] For the current block 904, with size wx h-samples, the reference samples 902 can comprise: 2 w-samples (or any other number of samples) from the row immediately adjacent to the top row of the current block 904, 2 h-samples (or any other number of samples) from the column immediately adjacent to the leftmost column of the current block 904, and the upper-left adjacent corner sample to the current block 904. The current block 904 can be square, so w = h = s. In other examples, a current block need not be square, so w ≠ h. Available samples from adjacent blocks of the current block 904 can be used to create the set of reference samples 902. There may be no samples available to create the set of reference samples 902, for example, if the samples are outside the image of the current block, the samples are part of another slice of the current block (e.g.,(if the concept of slices is used) and / or the samples belong to blocks that have been intercoded and a restricted intraprediction is specified. The intraprediction does not have to depend on the interpredicted blocks if, for example, a restricted intraprediction is specified.

[0077] Samples that may not be available for creating the Reference Samples 902 set may include samples in blocks that have not yet been encoded and reconstructed in an encoder and / or decoded in a decoder, based on the encoding / decoding sequence order. Restricting the inclusion of such samples in the Reference Samples 902 set may allow for the determination of identical prediction results for both the encoder and the decoder. In the example of Fig. 9. Samples from adjacent blocks 0, 1, and 2 may be available for creating reference samples 902, provided that these blocks are encoded and reconstructed in an encoder and decoded in a decoder before the current block 904 is encoded. For example, the samples from adjacent blocks 0, 1, and 2 may be available for creating reference samples 902 if no other issues exist (e.g., as mentioned above) that prevent the availability of the samples from adjacent blocks 0, 1, and 2. The portion of reference samples 902 from adjacent block 6 may not be available due to the encoding / decoding sequence order (e.g., because block 6 may not yet have been encoded and reconstructed in the encoder and / or decoded in the decoder based on the encoding / decoding sequence order).

[0078] In some examples, unavailable samples from reference samples 902 can be replaced by one or more of the available reference samples 902. For example, an unavailable reference sample can be replaced by the nearest available reference sample. The nearest available reference sample can be determined by proceeding clockwise through the reference samples 902 from the position of the unavailable reference. If no reference samples are available, the reference samples 902 can be populated, for example, with the mean of the dynamic range of the image to be encoded.

[0079] Reference samples 902 can be filtered based on the size of the currently coded block 904 and an applied intraprediction mode. Fig. Figure 9 shows an exemplary set of reference samples determined for intraprediction of a current block. Reference samples can also be determined in ways other than those described above. For example, in other cases, multiple reference lines can be used (e.g., in VVC).

[0080] Samples of the current block 904 can be intrapredicted based on reference samples 902, for example, based on (e.g., after) determining and (optionally) filtering reference samples 902. At least some (e.g., most) encoders / decoders can support a variety of intraprediction modes according to one or more video coding standards. For example, HEVC supports 35 intraprediction modes, including a planar mode, a direct current (DC) mode, and 33 angle modes. VVC supports 67 intraprediction modes, including a planar mode, a DC mode, and 65 angle modes. The planar and DC modes can be used to predict smooth and gradually changing image areas. Angle modes can be used to predict directional structures in areas of an image. Any number of intraprediction modes can be supported.

[0081] Fig. 10A and Fig. Figure 10B shows exemplary intraprediction modes. Fig. Figure 10A shows 35 intraprediction modes, such as those supported by HEVC. These 35 intraprediction modes can be identified by the indices 0 through 34. Prediction mode 0 corresponds to planar mode. Prediction mode 1 corresponds to DC mode. Prediction modes 2 through 34 correspond to angular modes. Prediction modes 2 through 18 can be referred to as horizontal prediction modes because the primary source of the prediction is in a horizontal direction. Prediction modes 19 through 34 can be referred to as vertical prediction modes because the primary source of the prediction is in a vertical direction.

[0082] Fig. Figure 10B shows 67 intraforecast modes, such as those supported by VVC. These 67 intraforecast modes can be identified by the indices 0 through 66. Forecast mode 0 corresponds to planar mode. Forecast mode 1 corresponds to DC mode. Forecast modes 2 through 66 correspond to angular modes. Forecast modes 2 through 34 can be described as horizontal forecast modes because the primary forecast source is in a horizontal direction. Forecast modes 35 through 66 can be described as vertical forecast modes because the primary forecast source is in a vertical direction.

[0083] Some of the in Fig. The intraprediction modes shown in 10B can be adaptively replaced by wide-angle directions, since blocks in VVC do not have to be squares.

[0084] Fig. Figure 11 shows a current block 904 and corresponding reference samples 902 from Fig. 9. To further describe how intraforecast modes are applied to determine a forecast (e.g., a forecast block) of the current block 904, shows Fig. 11. The current block 904 and reference samples 902 are placed in a two-dimensional x, y plane, where a sample can be referenced as p[x][y]. To simplify the prediction process, reference samples 902 can be placed in two one-dimensional arrays. The reference samples 902 above the current block 904 can be placed in the one-dimensional array ref1[x]: ref1[x]=p[−1+x][−1],(x≥0).

[0085] The reference samples 902 to the left of the current block 904 can be placed in the one-dimensional array ref2[y]: ref2[y]=[−1][−1+y],(y≥0).

[0086] The prediction process can involve determining a predicted sample p[x][y] (e.g., a predicted value) at a location [x][y] in the current block 904. For planar mode, a sample at the location [x][y] in the current block 904 can be predicted by determining / calculating the mean of two interpolated values. The first of the two interpolated values ​​can be based on a horizontal linear interpolation at the location [x][y] in the current block 904. The second interpolated value can be based on a vertical linear interpolation at the location [x][y] in the current block 904. The predicted sample p[x][y] in the current block 904 can be determined / calculated as follows: p[x][y]=12⋅s(h[x][y]+v[x][y]+s), where h[x][y]=(s−x−1)⋅ref2[y]+(x+1)⋅ref1[s] the horizontal linear interpolation at the position [x] [y] in the current block can be 904 and v[x][y]=(s−y−1)⋅ref1[x]+(y+1)⋅ref2[s] The vertical linear interpolation at position [x] [y] in the current block 904 can be s. s can be equal to a side length (e.g., a number of samples on a side) of the current block 904.

[0087] For DC mode, a sample at a position [x] [y] in the current block 904 can be predicted by the mean of the reference samples 902. The predicted sample p [x] [y] in the current block 904 can be determined / calculated as follows: p[x][y]=12⋅s(∑x=0s−1ref1[x]+∑y=0s−1ref2[y]).

[0088] For angular modes, a sample at a location [x][y] in the current block 904 can be predicted by projecting the location [x][y] in a direction specified by a given angular mode to a point on the horizontal or vertical line of samples that include reference samples 902. The sample at location [x][y] can be predicted by interpolation between the two nearest reference samples to the projection point if the projection does not fall directly on a reference sample. The direction specified by the angular mode can be indicated by an angle φ defined relative to the y-axis for vertical prediction modes (e.g., modes 19 to 34 in HEVC and modes 35 to 66 in VVC). The direction specified by the angular mode can be indicated by an angle φ defined relative to the x-axis for horizontal prediction modes (e.g., modes 2 to 18 in HEVC and modes 2 to 34 in VVC).

[0089] Fig. Figure 12 shows an example of the application of an intraprediction mode (e.g., an angular mode such as the vertical prediction mode 906) to predict a current block 904. Fig. Figure 12 shows in particular the prediction of a sample at a location [x] [y] in the current block 904 for a vertical prediction mode 906. The vertical prediction mode 906 can be given by an angle φ with respect to the vertical axis. The location [x] [y] in the current block 904, in vertical prediction modes, can be projected onto a point (e.g., called the projection point) on the horizontal line of the reference samples ref1[x]. The reference samples 902 are in Fig. 12 are only partially shown for the sake of simplicity. As in Fig. As shown in Figure 12, the projection point on the horizontal line of the reference samples ref1[x] does not necessarily correspond exactly to a reference sample. A predicted sample p[x][y] in the current block 904 can be determined / calculated by linear interpolation between the two reference samples, for example, if the projection point lies at a fractional sample position between two reference samples. The predicted sample p[x][y] can be determined / calculated as follows: p[x][y]=(1−if)⋅ref1[x+ii+1]+if⋅ref1[x+ii+2]. i i can be the integer part of the horizontal displacement of the projection point relative to the position [x] [y]. i i can be determined / calculated as a function of the tangent of the angle φ of the vertical prediction mode 906 as follows: ii=⌊(y+1)⋅tanφ⌋. i fcan be the fraction of the horizontal displacement of the projection point relative to the position [x] [y] and can be determined / calculated as follows: if=((y+1)⋅tanφ)−⌊(y+1)⋅tanφ⌋, where ⌊⋅⌋ The integer floor function is...

[0090] For horizontal prediction modes, a position [x][y] of a sample in the current block 904 can be projected onto the vertical line of the reference samples ref2 [y]. A predicted sample p [x] [y] for horizontal prediction modes can be determined / calculated as follows: p[x][y]=(1−if)⋅ref2[y+ii+1]+if⋅ref2[y+ii+2]. i i can be the integer part of the vertical displacement of the projection point relative to the position [x] [y]. i i can be determined / calculated as a function of the tangent of the angle φ of the horizontal prediction mode as follows: ii=⌊(x+1)⋅tanφ⌋. i fcan be the fraction of the vertical displacement of the projection point relative to the position [x] [y]. i f can be determined / calculated as follows: if=((x+1)⋅tanφ)−⌊(x+1)⋅tanφ⌋, where ⌊⋅⌋ The integer floor function is...

[0091] The interpolation functions given by equations (7) and (10) can be implemented by an encoder and / or a decoder (e.g. encoder 200 in Fig. 2 and / or decoder 300 in Fig. 3) be implemented. The interpolation functions can be implemented using finite impulse response (FIR) filters. For example, the interpolation functions can be implemented as a set of two-tap FIR filters. The coefficients of the two-tap FIR filters can be expressed by (1-i f ) or i fThe predicted sample p[x][y] can be calculated for angle intraprediction with a predefined measure of sample accuracy (e.g., 1 / 32 sample accuracy or accuracy defined by another metric). For a sample accuracy of 1 / 32, the set of two-tap FIR interpolation filters can include up to 32 different two-tap FIR interpolation filters—one for each of the 32 possible values ​​of the fraction of the projected displacement i. f Other levels of sample accuracy can be used in other examples.

[0092] In some examples, FIR filters can be used to predict chroma samples and / or luma samples. For example, the two-tap interpolation FIR filter can be used to predict chroma samples, and the same and / or a different interpolation technique / filter can be used for luma samples. For example, a four-tap FIR filter can be used to determine a predicted value of a luma sample. The coefficients of the four-tap FIR filter can be determined using i f (e.g., similar to the two-tap FIR filter). For a sample accuracy of 1 / 32, a set of 32 different four-tap FIR filters can include up to 32 different four-tap FIR filters—one for each of the 32 possible values ​​of the fraction of the projected displacement i. fIn other examples, different levels of sample accuracy can be used. The set of four-tap FIR filters can be stored and referenced in a lookup table (LUT) based on i f . A predicted sample p [x] [y] for vertical prediction modes can be determined based on the four-tap FIR filter as follows: p[x][y]=∑i=03fT[i]⋅ref1[x+iIdx+i], where fT[i], i = 0...3, can be the filter coefficients, and Idx is an integer shift. A predicted sample p[x][y] for horizontal prediction modes can be determined based on the four-tap FIR filter as follows: p[x][y]=∑i=03fT[i]⋅ref2[y+iIdx+i].

[0093] Additional reference samples can be determined / created by projecting the position [x] [y] of a sample to be predicted in the current block 904 onto a negative x-coordinate. For example, the position [x] [y] of a sample can be projected onto a negative x-coordinate when negative vertical prediction angles φ are used. The additional reference samples can be determined / created by projecting the reference samples in ref2[y] onto the horizontal line of reference samples 902 from the vertical line of reference samples 902 using the negative vertical prediction angle φ. Additional reference samples can also be determined / created, for example, by projecting the position [x] [y] of a sample to be predicted in the current block 904 onto a negative y-coordinate. For example, the position [x] [y] of a sample can be projected onto a negative y-coordinate when negative horizontal prediction angles φ are used.The additional reference samples can be determined / created by projecting the reference samples in ref1[x] onto the horizontal line of reference samples 902 to the vertical line of reference samples 902 using the negative horizontal prediction angle φ.

[0094] An encoder can determine / predict samples of a currently coded block (e.g., current block 904) for a variety of intraprediction modes (e.g., using one or more of the functions described herein). For example, an encoder can determine / predict samples of a current block for each of the 35 intraprediction modes in HEVC and / or 67 intraprediction modes in VVC. For each applied intraprediction mode, the encoder can determine a corresponding prediction error for the current block based on a difference (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), or sum of absolute transformed differences (SATD)) between the prediction samples determined for that intraprediction mode and the original samples of the current block. The encoder can then determine / select one of the intraprediction modes to code the current block based on the determined prediction errors.For example, the encoder can determine / select one of the intraprediction modes that yields the smallest prediction error for the current block. In some examples, the encoder can determine / select the intraprediction mode for encoding the current block based on a rate distortion measure (e.g., Lagrangian rate distortion cost) determined using the prediction errors. The encoder can then send information about the determined / selected intraprediction mode and the corresponding prediction error (e.g., residual) to a decoder to decode the current block.

[0095] A decoder can determine / predict samples of a currently decoded block (e.g., current block 904) for an intraprediction mode. For example, a decoder can receive a specification of an intraprediction mode (e.g., an angle intraprediction mode) from an encoder for a current block. The decoder can create a set of reference samples and perform an intraprediction based on the intraprediction mode specified by the encoder for the current block in a manner similar to what is described above for the encoder. The decoder can add predicted values ​​of the samples (e.g., determined based on the intraprediction mode) of the current block to a residual of the current block to reconstruct the current block. In some examples, a decoder for a current block does not need to receive a specification of an angle intraprediction mode from an encoder.Instead, the decoder can determine an intraprediction mode by other decoder-side means.

[0096] While various examples herein correspond to intraprediction modes in HEVC and VVC, the methods, devices and systems described herein may also be applied to / used for other intraprediction modes (e.g., as used in other video coding standards / formats such as VP8, VP9, ​​AV1, etc.).

[0097] Intraprediction can exploit correlations between spatially adjacent samples in the same frame of a video sequence to perform video compression. Interprediction is another encoding tool that can be used for video compression. Interprediction can exploit correlations in the time domain between sample blocks in different frames of a video sequence. For example, an object may be visible across multiple frames of a video sequence. The object may be moving (e.g., through translation and / or affine motion) or remain stationary across the multiple frames. A current block of samples in a currently encoded frame may be associated with / share a corresponding block of samples in a previously decoded frame. The corresponding block of samples can accurately predict the current block of samples.The corresponding block of samples may be shifted from the current block of samples, for example, due to movement of the object represented in both blocks across their respective images. The previously decoded image may be a reference image. The corresponding block of samples in the reference image may be a reference block for motion-compensated prediction. An encoder may use a block mapping technique to estimate the displacement (or movement) of the object and / or to determine the reference block in the reference image.

[0098] Similar to intraprediction, an encoder can determine a difference between a current block and a prediction for that current block. For example, an encoder might determine a difference based on / after determining / generating a prediction for the current block (e.g., using interprediction). The difference could be a prediction error (e.g., a residual). The encoder can store and / or transmit (e.g., signal) the prediction error and / or other related prediction information in / over a bitstream. The prediction error and / or other related prediction information can be used for decoding and / or other forms of utilization. A decoder can decode the current block by predicting the samples of the current block (e.g., using the related prediction information) and combining the predicted samples with the prediction error.

[0099] Fig. Figure 13A shows an example of an inter-prediction. The inter-prediction can be performed for a current block 1300 in a current image 1302 that is being coded. An encoder (e.g., encoder 200 as in Fig. (as shown in Figure 2) can perform an interprediction to determine and / or generate a reference block 1304 in a reference image 1306. Reference block 1304 can be used to predict the current block 1300. Reference images (e.g., reference image 1306) can be previously decoded images available at the encoder and / or a decoder. The availability of a previously decoded image can depend on whether the previously decoded image is available in a decoded image buffer at the time the current block 1300 is encoded and / or decoded. The encoder can search the one or more reference images 1306 for a block (e.g., a candidate reference block) that is similar (or substantially similar) to the current block 1300. The encoder can determine the best-matching block from the blocks (e.g., candidate reference blocks) examined during the search.The best-fitting block may be reference block 1304. The coder may determine that reference block 1304 is the best-fitting reference block based on one or more cost criteria. One or more cost criteria may include a rate distortion criterion (e.g., Lagrange rate distortion cost). One or more cost criteria may be based on a difference (e.g., SSD, SAD, and / or SATD) between prediction samples of reference block 1304 and original samples of the current block 1300.

[0100] The encoder can search for reference block 1304 within a reference region (e.g., a search area 1308). The reference region (e.g., a search area 1308) can be centered around a merged block (or position) 1310 of the current block 1300 in the reference image 1306. The merged block 1310 can occupy the same position in reference image 1306 as the current block 1300 in the current image 1302. The reference region (e.g., search area 1308) can extend at least partially outside the reference image 1306. A constant bounding extension can be used, for example, when the reference region (e.g., search area 1308) extends outside the reference image 1306. The constant bounding extension can be used so that values ​​of the samples in a row or column of reference image 1306 that are immediately adjacent to a section of the reference region (e.g.,Search area 1308, which extends beyond reference image 1306, can be used for sample positions outside of reference image 1306. A subset of potential positions or all potential positions within the reference region (e.g., search area 1308) can be searched for reference block 1304. The coder can use one or more search implementations to determine and / or generate reference block 1304. For example, the coder can determine a set of candidate search positions based on motion information from neighboring blocks (e.g., a motion vector 1312) to the current block 1300.

[0101] During interprediction, the encoder can search for one or more reference images to determine and / or generate the most suitable reference block. The reference images searched for by the encoder can be included (e.g., added) to one or more reference image lists. For example, in HEVC and VVC (and / or in one or more other communication protocols), two reference image lists can be used (e.g., a reference image list 0 and a reference image list 1). A reference image list can contain one or more images. Reference image 1306 of reference block 1304 can be specified by a reference index that points to a reference image list containing reference image 1306.

[0102] Fig. Figure 13B shows an example motion vector. A displacement between reference block 1304 and current block 1300 can be interpreted as an estimate of the motion between reference block 1304 and current block 1300 across their respective frames. The displacement can be represented by a motion vector 1312. For example, motion vector 1312 can be specified by a horizontal component (MVx) and a vertical component (MVy) relative to the position of current block 1300. A motion vector (e.g., motion vector 1312) can have fractional or integer resolution. A motion vector with fractional resolution can point between two samples in a reference frame to allow a better estimation of the motion of current block 1300. For example, a motion vector can have a resolution of 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, or any other fractional sample resolution.Interpolation between the two samples at integer positions can generate a reference block and the corresponding samples at fractional positions, for example, when a motion vector refers to a non-integer sample in the reference image. The interpolation can be performed using a filter with two or more taps.

[0103] The encoder can determine a difference (e.g., a corresponding sample-to-sample difference) between reference block 1304 and the current block 1300. The encoder can determine this difference based on / after reference block 1304 has been determined and / or generated using an interprediction for the current block 1300. The difference can be a prediction error (e.g., a residual). The encoder can store and / or transmit (e.g., signal) the prediction error and / or related motion information in / over a bitstream. The prediction error and / or related motion information can be used for decoding (e.g., decoding the current block 1300) and / or other purposes. The motion information can include the motion vector 1312 and a reference indicator / index.The reference indicator can specify reference image 1306 in a reference image list. In other examples, the motion information can include a specification of the motion vector 1312 and / or a specification of the reference indicator / index. The reference indicator can specify reference image 1306 in the reference image list that includes reference image 1306. A decoder can decode the current block 1300 by determining and / or generating the reference block 1304, which can correspond to / form a prediction of the current block 1300 (e.g., can be considered as such). The decoder can determine and / or generate the reference block 1304, for example, based on the associated motion information. The decoder can decode the current block 1300 based on the combination of the prediction (e.g., a reference block) with the prediction error (e.g., a residual block).

[0104] The interprediction, as in Fig. As shown in Figure 13A, a prediction for the current block 1300 can be made using a reference image 1306 as the source. An inter-prediction based on the prediction of a current block using a single image can be called a uni-prediction.

[0105] Interprediction of a current block using bi-prediction can be based on two images (e.g., the prediction source can be derived from the two images). Bi-prediction can be useful, for example, when a video sequence includes fast movements, camera pans, zooms, and / or scene changes. Bi-prediction can also be useful for capturing scene fade-outs or transitions from one scene to another, effectively displaying two images simultaneously at different intensity levels.

[0106] To perform interprediction (e.g., at an encoder and / or a decoder), a uni-prediction and / or bi-prediction may be available / used. The specific type of interprediction (e.g., uni-prediction and / or bi-prediction) performed may depend on the slice type of the current block. For example, for P-slices, only uni-prediction may be available / used to perform interprediction. For B-slices, either uni-prediction or bi-prediction may be available / used to perform interprediction. An encoder may determine and / or generate a reference block for predicting a current block, for example, from a reference image list 0, if the encoder uses uni-prediction.An encoder can determine and / or generate a first reference block to predict a current block from a reference image list 0 and determine and / or generate a second reference block to predict the current block from a reference image list 1, for example, if the encoder uses a bi-prediction.

[0107] Fig. Figure 14 shows an example of a Bi prediction. Two reference blocks, 1402 and 1404, can be used to predict a current block, 1400. Reference block 1402 can be located in a reference image from either reference image list 0 or reference image list 1. Reference block 1404 can be located in another reference image from either reference image list 0 or reference image list 1. As shown in Fig. As shown in Figure 14, reference block 1402 can be located in a first image that precedes (e.g., temporally) the current image of block 1400, and reference block 1404 can be located in a second image that follows (e.g., temporally) the current image of block 1400. The first image can precede the current image with respect to the Picture Order Count (POC). The second image can follow the current image with respect to the POC. In other examples, the reference images can both precede or follow the current image with respect to the POC. A POC can be / specify an order in which images are output (e.g., from a decoded image buffer). A POC can be / specify an order in which images should generally be displayed. Output images do not necessarily have to be displayed but can be subject to other processing and / or use (e.g., transcoding).The two reference blocks determined and / or generated using the Bi prediction function can correspond to the same reference image (e.g., be contained within it). The reference image can be included in both reference image list 0 and reference image list 1, for example, if the two reference blocks correspond to the same reference image.

[0108] A configurable weighting and / or offset value can be applied to one or more interprediction reference blocks. An encoder can enable the use of a weighted prediction using a flag in an image parameter set (PPS). The encoder can send / signal the weighting and / or offset parameters in a slice segment header for the current block 1400. Different weighting and / or offset parameters can be sent / signaled for luma and / or chroma components.

[0109] The encoder can determine and / or generate reference blocks 1402 and 1404 for the current block 1400 using interprediction. The encoder can determine a difference between the current block 1400 and each of the reference blocks 1402 and 1404. This difference can be a prediction error or a residual. The encoder can store and / or transmit / signal the prediction errors and / or their associated movement information in a bitstream. The prediction errors and their associated movement information can be used for decoding and / or other purposes.

[0110] The motion information for reference block 1402 can include a motion vector 1406 and / or a reference indicator / index. The reference indicator can specify a reference image of reference block 1402 in a reference image list. In some examples, the motion information for reference block 1402 can include a specification of the motion vector 1406 and / or a specification of the reference index. The reference index can specify the reference image of reference block 1402 in the reference image list.

[0111] The motion information for reference block 1404 can include a motion vector 1408 and / or a reference index / indicator. The reference indicator can specify a reference image of reference block 1404 in a reference image list. The motion information for reference block 1404 can include a specification of the motion vector 1408 and / or a specification of the reference index. The reference index can specify the reference image of reference block 1404 in the reference image list.

[0112] A decoder can decode the current block 1400 by determining and / or generating reference blocks 1402 and 1404. The decoder can determine and / or generate reference blocks 1402 and 1404, for example, based on their respective associated motion information. Reference blocks 1402 and 1404 can correspond to / constitute the prediction of the current block 1400 (e.g., be considered as such) (e.g., be used to generate a prediction block). The decoder can decode the current block 1400 based on the combination of the prediction and the prediction errors.

[0113] Motion information can be predictively encoded before it is stored and / or transmitted / signaled in / over a bitstream (e.g., in HEVC, VVC, and / or other video coding standards / formats / protocols). The motion information for a current block can be predictively encoded based on motion information from one or more blocks adjacent to the current block. The motion information of the adjacent block(s) can often correlate with the motion information of the current block, since the motion of an object represented in the current block is often identical (or similar) to the motion of objects in the adjacent block(s). Motion information prediction techniques (such as those in HEVC and VVC) can include extended motion vector prediction (AMVP) and / or block merging between predictions (e.g., merge mode).

[0114] A coder (e.g., coder 200, as in Fig. (as shown in Figure 2) can encode a motion vector. The encoder can encode the motion vector (e.g., using AMVP) as the difference between a motion vector of a currently coded block and a motion vector predictor (MVP). An encoder can determine / select the MVP from a list of MVP candidates. The MVP candidates can be previously decoded motion vectors of adjacent blocks in the current image of the current block and / or blocks at or near the merged position of the current block in other reference images. The encoder and / or a decoder can mutually generate and / or determine the list of MVP candidates.

[0115] The encoder can determine / select an MVP from the list of MVP candidates. The encoder can then send / signal an indication of the selected MVP and / or a movement vector difference (MVD) in / over a bitstream. The encoder can indicate the selected MVP in the bitstream using an index / indicator. The index can indicate the selected MVP in the list of MVP candidates. The MVD can be determined / calculated based on a difference between the movement vector of the current block and the selected MVP. For example, for a movement vector (which includes, for instance, a horizontal component (MVx) and a vertical component (MVy)) that indicates a position relative to a position of the currently coded block, the MVD can be represented by two components: MVx and MVy. x and MVD y MVD x and MVD y can be determined / calculated as follows: MVDx=MVx−MVPx, MVDy=MVy−MVPy.

[0116] MVDx and MVDy can each represent horizontal and vertical components of the MVD. MVPx and MVPy can each represent horizontal and vertical components of the MVP.

[0117] A decoder (e.g., decoder 300 as in Fig. (as shown in Figure 3) can decode the motion vector by adding the MVD to the MVP specified in / over the bitstream. The decoder can decode the current block by determining and / or generating the reference block. For example, the decoder can determine and / or generate the reference block based on the decoded motion vector. The reference block can correspond to / form the prediction of the current block (e.g., a prediction block) (e.g., be considered as such). The decoder can decode the current block by combining the prediction with the prediction error.

[0118] The list of MVP candidates (e.g., in HEVC, VVC, and / or one or more other communication protocols) for AMVP can include two or more candidates (e.g., candidates A and B). Candidates A and B can include: up to two (or any other set of) spatial MVP candidates determined / derived from five (or any other set of) spatially neighboring blocks of a currently encoded block; one (or any other set of) temporal MVP candidate determined / derived from two (or any other set of) temporally located blocks (e.g., when the two spatial MVP candidates are unavailable or identical); and / or zero-movement MVP candidates (e.g., when one or both of the spatial MVP candidates or temporal MVP candidates are unavailable).The list of MVP candidates can be made up of other sets of spatial MVP candidates, spatial neighboring blocks, temporal MVP candidates and / or temporal blocks located in the same position.

[0119] Fig. Figure 15A shows exemplary spatial candidate neighbor blocks for a current block. For example, five (or any other number) spatial candidate neighbor blocks can be located relative to a currently coded block 1500. The five spatial candidate neighbor blocks can be A0, A1, B0, B1, and B2. Fig. 15B shows temporally colocated blocks for the current block. For example, two (or any other set) temporally colocated blocks can be located relative to the currently encoded block 1500. The two temporally colocated blocks can be C0 and C1. The two temporally colocated blocks can be in one or more reference frames, which may differ from the current frame of the current block 1500.

[0120] A coder (e.g., coder 200 as in Fig. (as shown in Figure 2) can encode a motion vector using a merge of prediction blocks (e.g., a merge mode). For example, the encoder (e.g., using the merge mode) can reuse the same motion information from an adjacent block (e.g., one of the adjacent blocks A0, A1, B0, B1, and B2) to interpredict a current block. Similarly, the encoder (e.g., using the merge mode) can reuse the same motion information from a temporal block at the same location (e.g., one of the temporal blocks at the same location C0 and C1) to interpredict a current block. No MVD needs to be sent (e.g., specified, signaled) for the current block because the same motion information can be used for the current block as for an adjacent block or a temporal block at the same location (e.g., one of the temporal blocks C0 and C1).(at the encoder and / or a decoder). Signaling overhead for sending / signaling the movement information of the current block can be reduced because the MVD for the current block does not need to be specified. The encoder and / or the decoder can mutually generate a candidate list of movement information from adjacent blocks or blocks in the same temporal position as the current block (e.g., in a manner similar to AMVP). The encoder can determine whether to use (e.g., adopt) movement information from an adjacent block or a block in the same temporal position in the candidate list to predict movement information of the currently coded block. The encoder can signal / send a reference to the specific movement information from the candidate list in / over a bitstream. For example, the encoder can signal / send an indicator / index.The index can specify particular movement information in the list of candidate movement information. The coder can signal / send the index to indicate the specific movement information.

[0121] A list of candidate movement information for the merge mode (e.g., in HEVC, VVC, or other encoding formats / standards / protocols) can include: up to four (or any other set of) spatial merge candidates derived / determined from five (or any other set of) spatial neighboring blocks (e.g., as in Fig. 15A shown); a (or any other set of) temporal merge candidate derived from two (or any other set of) temporal blocks located at the same position (e.g., as in Fig. 15B shown); and / or additional merge candidates, including bi-predictor candidates and candidates with zero motion vectors. In some examples, the spatially adjacent blocks and the temporally co-located blocks used for the merge mode may be the same as the spatially adjacent blocks and the temporally co-located blocks used for AMVP.

[0122] Interprediction can be performed in ways and variations other than those described here. For example, motion information prediction techniques other than AMVP and merge mode can be used. While various examples herein correspond to interprediction modes as used in HEVC and VVC, the methods, devices, and systems described herein can also be applied to / used for other interprediction modes (e.g., as used for other video coding standards / formats such as VP8, VP9, ​​AV1, etc.). History-based motion vector prediction (HMVP), combined intra- / interprediction mode (CIIP), and / or a merge mode with motion vector difference (MMVD) (e.g., as described in VVC) can be performed / used and are within the scope of protection of this disclosure.

[0123] A block mapping operation (or block mapping technique) can be applied / used (e.g., in interprediction) to determine a reference block in a different frame than that of a currently encoded (e.g., encoded and / or decoded) block. A block mapping operation can also be applied / used to determine a reference block in the same frame as the currently encoded block. The reference block in the same frame as the current block, determined by block mapping, often cannot accurately predict the current block (e.g., in camera-recorded videos). Prediction accuracy for videos with screen content may not be affected as much if, for example, a reference block in the same frame as the current block is used for encoding. Screen content videos can include, for example, computer-generated text, graphics, animations, etc. Screen content videos can (e.g.,can include frequently repeated patterns (e.g., repeated patterns of text and / or graphics) within the same image.

[0124] Using a reference block (e.g., as determined by block mapping) in the same image as a currently encoded block can enable efficient compression for videos containing screen content.

[0125] A prediction technique (e.g., in HEVC, VVC, and / or other coding standards / formats / protocols) can be used to exploit the correlation between sample blocks within the same image (e.g., videos with screen content). The prediction technique can be intra-block copying (IBC) or current picture referencing (CPR). An encoder can apply / use a block mapping technique (e.g., similar to inter-prediction) to determine a displacement vector (e.g., a block vector (BV)). The BV can specify a relative position of a reference block (e.g., according to an intra-block compensated prediction) that best matches the current block from a position of the current block. For example, the relative position of the reference block can be the relative position of an upper-left corner (or any other point / pattern) of the reference block.The BV can specify a relative shift from the current block to the reference block that best matches the current block. The encoder can determine the best-matching reference block from the blocks checked during a search (e.g., in a manner similar to interprediction). The encoder can determine that a reference block is the best match based on one or more cost criteria. The one or more cost criteria can include a rate distortion criterion (e.g., Lagrange rate distortion cost). The one or more cost criteria can be based, for example, on one or more differences (e.g., an SSD, a SAD, a SATD, and / or a difference determined by a hash function) between the predicted samples of the reference block and the original samples of the current block. A reference block can be derived from previously decoded sample blocks (e.g.,The reference block may contain / comprise decoded blocks of samples from the current image before they are processed by in-loop filtering operations (e.g., deblocking and / or SAO filtering).

[0126] Fig. Figure 16A shows an example of an intra-block copy (e.g., an IBC mode). The in Fig. The example shown in Figure 16A may correspond to the screen content. The rectangular parts / sections with arrows starting at their boundaries may be the currently coded blocks. The rectangular parts / sections to which the arrows point may be the reference blocks for predicting the respective current blocks.

[0127] Using IBC, a reference block can be determined and / or generated for a current block. The encoder can determine a difference (e.g., a corresponding sample-to-sample difference) between the reference block and the current block. The difference can be a prediction error or a residual. The encoder can store and / or transmit / signal the prediction error and / or associated prediction information in / over a bitstream. The prediction error and / or associated prediction information can be used for decoding and / or other forms of utilization. The prediction information can include a bitstream variable (BV). The prediction information can include a hint about the BV. A decoder (e.g., Decoder 300 as in Fig. (as shown in Figure 3) can decode the current block by determining and / or generating the reference block. The decoder can determine and / or generate the current block, for example, based on the prediction information (e.g., the BV). The reference block can correspond to / form the prediction (e.g., a prediction block) of the current block (e.g., be considered as such). The decoder can decode the current block by combining the prediction (e.g., prediction block) with the prediction error (e.g., residual or residual block).

[0128] A block variable (BV) can be predictively encoded (e.g., in HEVC, VVC, and / or other encoding standards / formats / protocols) before it is stored and / or sent / signaled in / over a bitstream. For example, the BV for a current block can be predictively encoded based on a BV of one or more blocks adjacent to the current block. For example, an encoder can predictively encode a BV using merge mode (e.g., in a manner similar to that described here for inter-prediction), AMVP (e.g., as described here for inter-prediction), or a technique similar to AMVP. An AMVP-like technique could be BV prediction and differential encoding (or AMVP for inter-block).

[0129] A coder (e.g., coder 200 as in Fig. (as shown in Figure 2), which performs BV prediction and coding, can encode a BV as the difference between the BV of a currently coded block and a block vector predictor (BVP). An encoder can select / determine the BVP from a list of BVP candidates. The candidate BVPs may include / correspond to previously decoded BVs of adjacent blocks in the current image of the current block. The encoder and / or a decoder can mutually generate or determine the list of BVP candidates.

[0130] The encoder can send / signal a specification of the selected BVP and a block vector difference (BVD) in / over a bitstream. The encoder can specify the selected BVP in the bitstream using an index / indicator. The index can indicate (e.g., reference) the selected BVP in the list of BVP candidates. The BVD can be determined / calculated based on a difference between a BV of the current block and the selected BVP. For example, for a BV (e.g., represented by a horizontal component (BVx) and a vertical component (BVy)) that specifies a position relative to a position of the currently coded block, the BVD can be represented by two components: BVx and BVy. x and BVD y BVD x and BVD y can be determined / calculated as follows: BVDx=BVx−BVPx, BVDy=BVy−BVPy,

[0131] BVDx and BVDy can each represent horizontal and vertical components of the BVD. BVPx and BVPy can each represent horizontal and vertical components of the BVP. A decoder (e.g., decoder 300 as in Fig. (as shown in Figure 3) can decode the block variable (BV) by adding the block variable distortion (BVD) to the block variable specified in / over the bitstream. The decoder can decode the current block by determining and / or generating the reference block. For example, the decoder can determine and / or generate the reference block based on the decoded BV. The reference block can correspond to / form the prediction (e.g., a prediction block) of the current block (e.g., be considered as such). The decoder can decode the current block by combining the prediction (e.g., the prediction block) with the prediction error (e.g., residual or residual block).

[0132] The same block value (BV) can be used for the current block as for an adjacent block, and a block value document (BVD) does not need to be signaled / sent separately for the current block, as is the case in merge mode. A block value candidate (BVP) that can correspond to a decoded BV of the adjacent block can itself be used as a BV for the current block. Not sending the BVD reduces signaling overhead.

[0133] A list of BVP candidates (e.g., in HEVC, VVC, and / or another coding standard / format / protocol) can include two (or more) candidates. The candidates can include candidates A and B. Candidates A and B can include: up to two (or any other set of) spatial candidate BVPs determined / derived from five (or any other set of) spatially neighboring blocks of a currently encoded block; and / or one or more of the two most recent (or any set of) encoded BVs (e.g., when spatially neighboring candidates are not available). Spatially neighboring candidates may not be available if, for example, adjacent blocks are encoded using intra- or inter-prediction.The positions of the spatial candidate neighbor blocks relative to a current block, encoded using IBC, can be illustrated in a similar way to the spatial candidate neighbor blocks used to encode motion vectors in interprediction (e.g., as in ). Fig. 15A). For example, five spatial candidate neighboring blocks of an actual block coded using IBC can each be designated A0, A1, B0, B1 and B2, as shown in Fig. 15A shown.

[0134] Accordingly, a block value (BV) can be determined for the current block in IBC mode (e.g., a merge mode for IBC or using AMVP). The BV specifies a displacement from the current block to a position of a reference block in the same image (e.g., frame) as the current block. Therefore, IBC works similarly to the one in Fig. Inter-prediction mode described in 13A-B.

[0135] Intra-prediction match prediction (e.g., Intra-TMP mode) is another intra-prediction mode that, similar to IBC, determines a prediction block from the reconstructed portion of a current frame to determine (at the decoder) or predict (at the encoder) a current block within the same current frame. Unlike IBC, where a reference block is determined by the encoder and signaled to the decoder to be used as the prediction block, Intra-TMP does not require such signaling. Instead, both the encoder and the decoder can mutually perform template matching operations between a template of the current block (e.g.,The algorithm performs a search of candidate reference blocks (using an L-shaped template, a template that includes only the top portion, or a template that includes only the left portion) and templates of candidate reference blocks within a search area (within the reconstructed portion) to determine a reference block to use as the prediction block. The candidate reference block templates may match the current template in shape, size, and / or orientation (e.g., be identical). The reference block may be determined based on a template, with the reference block being identified as the one most similar to the current template. In some examples, the similarity may be represented or determined based on the cost of template matching (e.g., SAD, SSD, SATD, SSE, etc.).

[0136] For example, within a search area (determined, for instance, based on the size of the current block), the encoder can search the reconstructed portion of the current frame for the template most similar to the current template and use the reference block (corresponding to the most similar template) as the prediction block. Based on the reference block used as the prediction block, the encoder can send a signal to the decoder indicating the use of IntraTMP mode. Upon decoding the IntraTMP mode indication, the decoder can perform the same IntraTMP operations as the encoder to determine the same reference block to be used as the prediction block.

[0137] Fig. Figure 16B shows an example of a TMP mode (e.g., an intra-TMP mode) for predicting or determining a current block 1600 based on template matching, which a decoder can do according to some embodiments. The current block 1600 comprises a rectangular block of samples (to be decoded by prediction) in an image or video frame of the current image 1602, which is to be encoded by the encoder or decoded by the decoder. To perform a TMP (template-based prediction) to determine a reference block (RB) 1610 for the current block 1600, an encoder (e.g., the encoder or the decoder) can determine or create a current template 1608 of the current block 1600. The encoder can determine or create the current template 1608 based on samples in a reconstructed region of the current image 1602.This can, for example, include a region above the block to be predicted, which has the width of the block and a height of, for example, 2 pixels, and includes a number of pixels that have already been reconstructed in this template (as parts of previously reconstructed blocks or units), and whose template can therefore be used for local similarity tests with the template of another block, with this reference block being a good candidate for intra-prediction.Adding a previously reconstructed template region (usually, but not necessarily, rectangular and of the same height as the block) allows for the creation of a combined template that is better suited to finding similar structures. There is likely to be a correlation between the pixel color structure within the block and immediately in front of (above and to the left of) the block. Therefore, the probability that the content of the reference block for prediction (1610) will have a similar structure also increases, thus providing a good, or at least reasonable, prediction for the current block. For example, the current template might include 1608 samples in the reconstructed region adjacent to the samples of the current block 1600. For instance, the current template might include 1608 samples in the reconstructed region to the left and / or above the current block 1600.Block vector (BV) 1630 indicates a shift from the current block 1600 to the specified block RB 1610.

[0138] After determining or creating the current template 1608 of the current block 1600, the coder can search a TMP search area 1606 (also referred to herein as the TMP reference area) for a reference template 1612 of a reference block (RB) (e.g., RB 1610) that best matches or is most similar to the current template 1608 of the current block 1600. For example, the coder can determine a multitude of candidate templates corresponding to a multitude of respective candidate reference blocks (RBs) 1614 ( Fig. Figure 16B shows three matching examples) from which the reference template 1612 and the reference block (RB) 1610 can be determined (after they have been ordered according to a criterion of matching, also called cost, such as a sum of differences or a transformation-based difference, etc.). As in Fig. As shown in 16B, the candidate templates of candidate reference blocks 1614 match the current template 1608 in shape, alignment and size.

[0139] In some examples, the coder can search in TMP search area 1606 for a candidate template of a candidate RB that best matches the current template 1608 by determining the cost between the current template 1608 and each of the candidate templates of candidate reference blocks 1614 in TMP search area 1606. For example, the cost can be based on a difference (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), sum of absolutely transformed differences (SATD), or a difference determined based on a hash function) between a candidate template of a candidate RB and the current template 1608. In the Fig. In the example illustrated in Figure 16B, reference template 1612 is determined by RB 1610 to be the best match for current template 1608 (e.g., based on the cost difference between reference template 1612 and current template 1608). A block vector (BV) can specify the displacement of an RB (e.g., RB 1610) relative to a current block (e.g., current block 1600).

[0140] In some examples, the TMP search area 1606 encompasses a section of a reconstructed region of the current image 1602. The TMP search area 1606 specifies the regions in which the encoder or decoder can search for candidate templates (e.g., candidate templates of candidate reference blocks 1614) to determine the reference template 1612 and the corresponding RB 1610. In some examples, the TMP search area 1606 may include regions 1606A, 1606B, 1606C, and 1606D. Relative to the current block 1600, region 1606A (R1) can include a section of the current CTU, region 1606C (R2) can include a section of the upper left CTU, region 1606B (R3) can include a section of the aforementioned CTU, and region 1606D (R4) can include a section of the left CTU. The CTUs are the result of image partitioning operations, which were described in more detail above.For example, an encoder or decoder can search for a matching template within TMP search area 1606. For example, reference template 1612 can be determined from RB 1610 to best match the current template 1608 of the current block 1600 based on the SAD cost or other costs, as described above. The decoder can use RB 1610 to predict the current block 1600, as described above.

[0141] In some examples, the dimensions of the TMP search area 1606 (denoted as View_area_w, Search_area_h) can be set proportionally to the dimensions of the current block 1600 (denoted as BlkW, BlkH) to obtain a fixed number of cost comparisons (e.g., SAD or other difference comparisons) per pixel. More precisely, the dimensions of the TMP search area 1606 can be calculated as follows: SearchRange_w=a∗BlkW SearchRange_h=a∗BlkH

[0142] As shown above, “a” (or alpha) can be a constant that controls a trade-off between gain and complexity for the encoder or decoder. For example, “a” can be equal to 5. In Fig. Furthermore, it should be noted in section 16B that the dimensions of the TMP search area 1606 are illustrated as examples and are not exhaustive. In practical implementation, for example, the dimensions of the regions may vary and / or one or more regions may not exist. Fig. In the example illustrated in Figure 16B, sections of the reconstructed region located directly above and directly to the left of the current block 1600 may not be available for predictions or determinations and can therefore be excluded from the TMP search area 1620. This could be because, for example, a reconstructed region in these sections would overlap with the current block 1600, resulting in an invalid position for predicting or determining the current block 1600. A similar restriction could also be based on the unavailability of samples due to the encoding or decoding sequence, or on the samples being outside the TMP search area or the current image 1602.

[0143] In some embodiments, while searching and performing TMP in the TMP search area 1606, the encoder can determine (e.g., generate) a candidate list of up to a predetermined maximum (e.g., 19) of template-matching block vectors (BVs) ordered in ascending order according to template-matching costs (e.g., SAD, SATD, SSD, SSE, etc.). Template-matching block vectors are block vectors that correspond to (e.g., indicate) the position of the relevant candidate reference blocks whose templates were compared to the current template. In some examples, the encoder can generate the candidate list within each TMP search area 1606A-D. Accordingly, the template-matching BV whose template has the lowest cost can be selected as the predictor (e.g., a single predictor) that specifies the prediction block (e.g.,the reference block, which is specified by the template-matching BV in relation to the current block).

[0144] In some examples, to speed up the template matching process, not every template of potential candidate reference blocks in the TMP search area 1606 is searched and compared to the current template. Instead, positions of potential candidate reference blocks can be underscanned by a factor (e.g., a factor of 3) in the vertical and / or horizontal direction. In some examples, after determining the best-matching template and the corresponding candidate reference block, a refinement process can be performed. In the refinement process, a second template matching search for the best match can be performed with a reduced region based on the factor to scan positions that were skipped in the first template matching search.

[0145] In some examples, the following modes can be applied to further enhance Intra-TMP: multi-predictor fusion, sub-pel precision, and linear filter model. In multi-predictor fusion mode, several predictors are blended to select and derive the final prediction block. The blend weights can be calculated either from the template-matching costs of each predictor or using a weighting derivation method based on a Wiener filter. In sub-pel precision mode, when using a single predictor, sub-pel precision can be applied at ½-pel, ¼-pel, and ¾-pel precision, each with eight possible directions. In linear filter model mode, a linear filter can be derived between the reference template and the current template and applied to the reference block. This mode can be used for a single predictor when sub-pel precision is not employed.

[0146] After determining the reference template 1612 from RB 1610, the encoder can use RB 1610 to encode the current block 1600. For example, an encoder can determine a difference (e.g., a corresponding sample-to-sample difference) between the current block 1600 and RB 1610. This difference can be referred to as a prediction error or a residual. The encoder can store the prediction error or residual in a bitstream and / or signal it for decoding by a decoder.

[0147] To perform the TMP to determine the current block 1600, a decoder can perform the same (or reverse) operations as the encoder, as described above. Fig. As described in Section 16B, if the decoder receives information from the encoder that TMP is being used to predict the current block 1600 (e.g., via a flag), the decoder can similarly determine or construct the current template 1608 of the current block 1600. After determining or constructing the current template 1608, the decoder can further search the TMP search area 1606 for a reference template of an RB that best matches the current template 1608. For example, from the candidate reference blocks 1614, the decoder can determine reference template 1612 of reference block 1610 as the one that best matches the current template 1608. After determining reference template 1612 of RB 1610, the decoder can use RB 1610 (corresponding to reference template 1612) to determine the current block 1600.For example, the decoder can combine the residual received from the encoder with RB 1610 to reconstruct the current block 1600. Unlike with IBC, BV 1630, which specifies RB 1610, may not need to be provided by the encoder to the decoder in order for the decoder to determine RB 1610.

[0148] In some examples, IntraTMP can be enabled for the current block 1600 based on the size of that block. For instance, the encoder and decoder can mutually determine whether IntraTMP is enabled based on whether the size of the current block is less than or equal to a size threshold (e.g., 64). In some examples, the size threshold can be configurable, e.g., signaled from the encoder to the decoder.

[0149] Fig. Figure 17A shows an example of template-based intra-mode derivation (TIMD) for encoding a current block according to some embodiments. TIMD is a type of intra-prediction where an intra-prediction mode (IPM) can be mutually determined (e.g., independently derived) by an encoder (e.g., encoder 200) and a decoder (e.g., decoder 300), so that the IPM determined (e.g., selected) by the encoder does not need to be signaled to the decoder (because the encoder knows that the decoder can ultimately draw the correct conclusion from the received compressed data). Therefore, the signal bandwidth is reduced, and TIMD can be considered a type of decoder-side intra-mode derivation process.

[0150] As in Fig. As shown in Figure 17A, for TIMD, a video encoder (e.g., encoder 200 or encoder 300) can determine a template 1704 (= 1704A and 1704B) for the current block 1702. Template 1704 can comprise one or more regions of samples within a reconstructed area 1708 of a current image (or frames) of the current block 1702. The one or more regions can include reconstructed samples that are adjacent to (e.g., bordering) the current block 1702. In some examples, the one or more regions of template 1704 can comprise a template area 1704A to the left of the current block 1702 (e.g., a left template) and a template area 1704B above the current block 1702 (e.g., an upper template). Or only one of them. In some examples, template 1704 may include a region that lies above and to the left of the current block 1702 (e.g., the region enclosed by reference to template 1706 and template areas 1704A-B).Template areas 1704A and 1704B can have a thickness (e.g., width or height) of R1 and R2 samples, respectively. For example, R1 and / or R2 can have 2 samples, 4 samples, 8 samples, etc. Template area 1704A can have a height of N samples, which can correspond to the height of the current block 1702. Template area 1704B can have a width of M samples, which can correspond to the width of the current block 1702.

[0151] In TIMD, template 1706 of the reference is used to derive a template predictor for template 1704. Predicting how a pattern continues before template 1704, i.e., from 1706 to 1704, can be a good predictor of how it continues from template 1704 to a block to be reconstructed. For example, the video encoder can determine (e.g., select and retain) the reference of template 1706 as a region of samples (within the reconstructed region 1708) that is adjacent to (e.g., bordering) template 1704. For example, the reference of template 1706 might include reconstructed samples to the left and above template 1704. The reference of template 1706 might have a thickness to the left of template region 1704A containing the R1 samples and a thickness above template region 1704B containing the R2 samples. For example, R1 and / or R2 can be 1 sample, 2 samples, 4 samples, etc.In some examples, the reference of template 1706 may include an upper region whose width is greater than that of template area 1704B (e.g., a width greater than or equal to twice the width of template area 1704B, 2 (M + L1) samples, 2 (M + L1) + R1 samples, etc.). In some examples, the reference of template 1706 may include a left region whose height is greater than that of template area 1704A (e.g., a height greater than or equal to twice the height of template area 1704A, 2 (N + L1) samples, 2 (N + L1) + R2 samples, etc.).

[0152] In some examples, a TIMD mode predictor can be determined using a list of candidate intra-prediction modes (IPMs). The list might include IPMs from a list of most probable modes (MPMs). In some examples, one or more candidate IPMs generated from a DC mode, a planar mode, a horizontal and / or vertical DC mode (summing or averaging the values ​​in only one horizontal template above or only one vertical template to the left), a horizontal and / or vertical planar mode, and / or prediction modes generated from a template of the reference block (e.g., prediction block) of adjacent blocks (e.g., bordering the current block) encoded with IBC or IntraTMP modes can be added to the candidate IPM list. The cost (e.g., SSD, SSE, SAD, or SATD) for each candidate IPM in the list can be determined based on differences between reconstructed samples in template 1704 (i.e.,The best possible reconstruction that a decoder would perform for these pixels after completing the entire decoding process in template 1704) and predicted samples of template 1704, generated based on reference template 1706 and using a candidate IPM, are determined (i.e., this gives an indication of whether at least one of the candidate predictors would be well-suited to predict the samples in 1704 based on the already reconstructed neighboring samples in 1706, since a good predictor leads to a low cost or difference measure). For example, the video encoder can determine the predicted samples by applying the candidate IPM to reference samples of template 1706.

[0153] In some examples, a first IPM and a second IPM are selected from the list with the lowest cost (determined for candidate IPMs in the list) to determine (e.g., derive) a first TIMD mode and a second TIMD mode. In one example, the first and second TIMD modes are determined as the first and second IPMs, respectively. For instance, there might be a diagonal striping pattern extending from the vertical portion of 1706 into the region of 1704A, but if these diagonal stripes continue further, they could also be accurately predicted by pixels in the horizontal area above 1706 all the way to 1704A. The mean of these two angle predictions could then be calculated.

[0154] In some examples, the first and second TIMD modes can be determined by refining the first and second IPMs. For instance, an angle mode range can be extended from a first region of the list (e.g., 67 modes) to a second region (e.g., 131 modes), and the cost of the two adjacent modes (i.e., + / - 1 mode) of each selected IPM can be determined. For example, the first TIMD mode can be determined as the IPM with the lowest cost among the costs of the first IPM and its two adjacent modes in the second region. The second TIMD mode can be determined as the IPM with the lowest cost among the costs of the second IPM and its two adjacent modes in the second region.

[0155] In some embodiments, a TIMD mode predictor can be determined based on the combination (e.g., by mixing or fusion) of the first TIMD mode and the second TIMD mode. For example, the TIMD mode predictor can be a linear combination of the first and second TIMD modes. The weighting of the first and second TIMD modes in the linear combination can be determined based on the costs of the first and second TIMD modes, respectively. For example, a weighting for a TIMD mode can be determined to be inversely proportional to the cost of that TIMD mode. The TIMD mode predictor can be applied to template 1704 to determine a prediction block for the current block 1702. An encoder can generate a residual (e.g., a prediction error) based on a difference between the prediction block and the current block 1702.A decoder can reconstruct the current block 1702 based on the mutually generated prediction block and the residual received by the encoder in a bitstream.

[0156] Fig. Figure 17B shows an example of template-based intra-mode derivation (TIMD) for encoding a current block using an adjacent block coded with intra-block copy (IBC) or intra-template matching prediction (Intra-TMP) according to some embodiments. For example, in addition to one or more candidate IPMs derived from templates 1704A-B of the current block 1702 according to Fig. 17A were determined, one or more further candidate IPMs using an adjacent block 1710 coded in IBC mode (as in Fig. 16A) or an adjacent block 1720 encoded in Intra-TMP mode (as described in Fig. 16B), can be derived. The neighboring blocks 1710 and 1720 can, for example, adjoin the current block 1702, like blocks that are described in Fig. 15A are described.

[0157] As shown, the adjacent block 1710 may have previously been encoded using a reference block 1714, specified by the block vector 1712, which was determined (e.g., derived) in IBC mode, as in Fig. As described in section 16A, block vector 1712 indicates a displacement from the adjacent block 1710 to its reference block 1714 (this reference block could potentially be found, i.e., it can be found independently in the decoder in exactly the same way as in the encoder, based on the match / cost of its adjacent template above and to the left, as explained). In IBC mode, reference block 1714 can be used as a prediction block to encode (i.e., to encode and decode) the adjacent block 1710. A candidate IPM for TIMD can be derived based on a position of reference block 1714, which is used as a prediction block to encode the adjacent block 1710. The candidate IPM can be stored as a block vector indicating a displacement from the current block 1702 to reference block 1714.

[0158] A coder (e.g., coder and / or coder) can determine a template for the candidate IPM that matches template 1704 of the current block 1702 in shape, size, and / or orientation. For example, the coder can determine template area 1716A to the left of reference block 1714 (e.g., a left template used to determine or predict the adjacent block 1710) and template area 1704B above reference block 1714 (e.g., an upper template used to determine or predict the adjacent block 1710), each containing samples used to generate a prediction of template area 1704A and template area 1704B, respectively. The coder can determine the cost of the first IPM based on the cost of template matching between template 1716 (e.g., including template ranges 1716A and 1716B) and template 1704 (e.g., including templates 1704A-B).

[0159] As shown, the adjacent block 1720 may have been previously encoded using a reference block 1724 specified by block vector 1722 (e.g., Intra-TMP-BV), which was determined (e.g., derived) in Intra-TMP mode, as in Fig. Described in section 16B, block vector 1722 specifies a shift from the adjacent block 1720 to the reference block 1724. In Intra-TMP mode, reference block 1724 can be used as a prediction block, which is used for encoding (e.g., encoding and decoding) the adjacent block 1720. As described above in Fig. As explained in section 16B, templates of candidate RBs can be searched in a TMP search area and compared to template 1728 of the adjacent block 1720, and template 1730 (corresponding to reference block 1724) can be determined as the most similar (i.e., most matching) template 1728. In intra-TMP modes, the coder determines templates (of candidate RBs) that match template 1728 in size, shape, and / or orientation. A candidate IPM for TIMD can be derived based on the position of reference block 1724, which is used as a prediction block for coding the adjacent block 1720. The candidate IPM can be stored as a block vector indicating a shift from the current block 1702 to reference block 1724.

[0160] For this candidate IPM, the encoder (e.g., encoder and / or decoder) can determine a template that matches the shape, size, and / or orientation of template 1704 of the current block 1702. For example, the encoder can determine template area 1726A to the left of reference block 1724 (e.g., a left template used to determine or predict the adjacent block 1720) and template area 1724B above reference block 1724 (e.g., an upper template used to determine or predict the adjacent block 1720), each containing samples used to generate a prediction of template area 1704A and template area 1704B, respectively. Templates 1726 and 1730 may not be identical. The coder can calculate costs for the IPM based on the cost of template matching between template 1726 (e.g., including template ranges 1726A and 1726B) and template 1704 (e.g.,including templates 1704A-B).

[0161] In some examples, several adjacent blocks (e.g., adjacent blocks 1710 or 1720) encoded with IBC or IntraTMP modes are added to or appended to the current block 1702. Costs (e.g., SAD, SSE, or SATD) can be determined based on differences between reconstructed samples in template 1704 and predicted samples in template 1704, which were generated from templates determined (e.g., generated) from the positions of reference blocks (e.g., prediction blocks) used to determine or predict adjacent blocks.

[0162] Fig. Figure 18 shows an example of TIMD signaling for decoding a current block according to some embodiments. At block 1802, a decoder receives (e.g., from a bitstream) an indication of whether TIMD is applied to generate a prediction block for encoding the current block. As explained above, TIMD is a type of intraprediction where a template of the current block in the reconstructed area of ​​the image is used to derive an intraprediction mode for encoding the current block. For example, the indication could be a flag showing whether TIMD is enabled (e.g., being applied).

[0163] At block 1804, the decoder determines whether TIMD should be applied based on the information received at block 1802.

[0164] At block 1810, the decoder determines a TIMD mode predictor based on a variety of intra-prediction modes (IPMs) based on the indication that TIMD is applied (e.g., enabled or selected), without additional signaling from an encoder specifying a particular intra-prediction mode. As above in Fig. 17A and Fig. As described in Figure 17B, the TIMD mode predictor can be determined based on identifying two IPMs from the plurality of IPMs that have the lowest costs determined for the plurality of IPMs. Since the encoder and decoder mutually (i.e., independently and identically) determine the TIMD mode predictor, the signaling of each specific intra-prediction mode (IPM) from the bitstream can be omitted. In some embodiments, the plurality of IPMs does not include candidate IPMs, which are determined (e.g., derived) based on reference blocks (e.g., prediction blocks) used to encode adjacent blocks of the current block, as in Figure 17B. Fig. described in section 17B. At block 1812, the decoder generates a prediction block based on the specified TIMD mode predictor (e.g., as in Fig. 17A and Fig. 17B described).

[0165] At block 1806, the decoder receives an indication of an intra-prediction mode (IPM) from the bitstream (e.g., parses it), based on the indication that TIMD is not applied (e.g., disabled or not selected). The indication can comprise a variety of bits representing an index that specifies the IPM from a list of IPMs (e.g., an MPM list). At block 1808, the decoder generates a prediction block based on the specified IPM (e.g., as in Fig. 17A and Fig. (described in 17B). At block 1814, the decoder reconstructs the current block based on the prediction block and a residual, which may have been received from a bitstream, for example.

[0166] In existing technologies, TIMD has been introduced as a decoder-side technique where an intra-prediction of a current block can be performed to determine a prediction block for reconstructing the current block, without requiring explicit signaling of specific intra-prediction modes in the bitstream. Furthermore, a TIMD mode predictor can be generated as a linear combination of two or more intra-prediction modes (IPMs) from a list of IPMs, selecting the one with the lowest cost (e.g., SATD or SAD) among the IPMs in the list. The list of IPMs can include one or more angular IPMs, a DC mode, and / or one or more planar modes (e.g., a vertical planar mode or a horizontal planar mode). The linear combination of multiple IPMs is also referred to as TIMD-IPM fusion or TIMD-IPM mixing.

[0167] Embodiments of the present disclosure relate to a method for improving the mixing (e.g., combining or fusion) of a plurality of TIMD modes derived from a plurality of intraprediction modes to generate a TIMD mode predictor. More specifically, a video encoder (e.g., an encoder and / or decoder) can condition at least one TIMD mode (e.g., two TIMD modes derived from the plurality of intraprediction modes) with a TIMD mode corresponding to a non-angular intraprediction mode (IPM). NA ) corresponds, combine, For example, IPM can NA a DC mode or a planar mode. The video encoder can determine (e.g., derive) weights of a linear combination of the TIMD mode and at least one TIMD mode based on the cost of the TIMD modes in the linear combination. For example, a weight for a TIMD mode (e.g., a derived IPM or an IPM NA) inversely proportional to the costs of the TIMD mode.

[0168] In some examples, the video encoder can determine whether the TIMD mode, which is an IPM NA corresponds to, is to be combined, using at least one TIMD mode to generate the TIMD mode predictor, based on whether the at least one TIMD mode corresponds to the IPM NA includes and / or at least one IPM NA include (e.g., whether at least one TIMD mode consists exclusively of angled IPMs). For example, if at least one TIMD mode does not include non-angled intra-modes and / or the IPM NA Including, the video encoder can determine at least one TIMD mode with the IPM NA to combine. In other examples, the video encoder may additionally consider one or more criteria to determine whether the TIMD mode is at least compatible with the IPM. NAto be combined. By conditionally combining / mixing / merging the at least one TIMD mode with the TIMD mode, such that the TIMD mode predictor more likely includes a fusion / combination / mixing of intra-modes including a non-angular intra-prediction mode (IPM), a more accurate intra-prediction of the current block can be achieved.

[0169] In some examples, one or more of the criteria may affect the comparison of IPM costs. NA include a threshold (e.g., a fixed value or a value determined from the cost of at least one TIMD mode). For example, the video encoder can determine the IPM NA -to combine data, with at least one TIMD mode based on costs below the threshold.

[0170] These and other features of the present disclosure are described in more detail below.

[0171] Fig. Figure 19 shows a flowchart 1900 of a method for applying a TIMD technique to a current block according to some embodiments. The method of flowchart 1900 can be implemented by a video encoder (e.g. encoder 200 in Fig. 2 or decoder 300 in Fig. 3) be implemented. The flowchart procedure can be performed reciprocally by an encoder and a decoder, so that an intra-prediction mode does not need to be signaled by the encoder to the decoder in order to reconstruct the current block using intra-prediction techniques.

[0172] At block 1902, the video encoder, based on the application of TIMD (e.g., enabled) to the current block, determines a multitude of costs for a multitude of intra-prediction odis (IPMs). That is, multiple predictions, e.g., along a few angles, can be selected for the set, and the cost of each prediction can be determined based on the template and its surrounding / preceding reference. The cost of each IPM among the multitude of IPMs is based on the differences between: predicted samples of a template generated from the IPM applied to reference samples of the template; or generated from the template of the reference block (e.g., prediction block) of a neighboring block (e.g., adjacent to the current block) encoded in IBC or Intra-TMP mode; and reconstructed samples of the template (i.e., samples that have already been fully decoded).Examples of the current block template and the template reference are given with reference to . Fig. 17A is shown and described. Examples of the current block template and the reference block template, used to determine a prediction block of an adjacent block, are given with reference to Fig. 17B shown and described.

[0173] In some examples, the video encoder can be a decoder that determines the manifold costs based on receiving (e.g., parsing and decoding) a bitstream that indicates that TIMD is being applied (e.g., enabled) to the current block (i.e., that the current block is to be decoded using the TIMD method). For example, an encoder can determine that TIMD is being used to encode the current block and signal this in the bitstream to the decoder. In some examples, the indication can be a syntax element (e.g., a flag like `timd_flag`) that is enabled and signaled at the block level (e.g., per block) in the bitstream.

[0174] At block 1903, the video encoder determines a TIMD mode predictor based on the multitude of IPMs. Block 1903 can include blocks 1904 to 1920.

[0175] At block 1904, the video encoder determines a first TIMD mode and a second TIMD mode, based on a first IPM and a second IPM, from the plurality of IPMs with the lowest cost among the plurality. In some embodiments, the first IPM and the second IPM can be determined from a second plurality of IPMs. The second plurality of IPMs can be a subset of the plurality of IPMs and can exclude one or more candidate IPMs determined from reference blocks (e.g., prediction blocks) used to encode adjacent blocks of the current block.

[0176] At block 1906, the video encoder determines whether the second TIMD mode should be selected to determine the TIMD mode predictor.

[0177] At block 1916, the video encoder determines the TIMD mode predictor based on the first TIMD mode, since the second TIMD mode was not selected. For example, the TIMD mode predictor can be determined as the first TIMD mode.

[0178] At block 1908, the video encoder determines, based on the selected second TIMD mode (e.g., to be considered when blending with the first TIMD mode), whether to select a third TIMD mode based on a non-angular IPM. As explained above, the encoder and decoder operate independently and mutually determine (i.e., the decoder and encoder in the same way) whether to select the third TIMD mode, thus eliminating the need for additional signaling in the bitstream to conditionally select the third TIMD mode. In some embodiments, the video encoder determines the third TIMD mode based on the non-angular IPM from a variety of non-angular IPMs. For example, the third TIMD mode can be determined before block 1908, such as before block 1906 or block 1904.

[0179] In some examples, the video encoder generates a list of non-angular IPMs, which represents the multitude of non-angular IPMs (IPM). NA The non-angular IPMs can include a DC mode and / or a planar mode. For example, the non-angular IPMs can include a vertical DC mode, a horizontal DC mode, a vertical planar mode, and / or a horizontal planar mode. In some examples, the non-angular IPMs can include a candidate IPM generated from a template of a reference block (e.g., used to determine a prediction block) of an adjacent block to the current block, coded with IBC or Intra-TMP prediction modes, as shown above. Fig. 17B described.

[0180] In some examples, the number of candidate IPMs in the list can be based on the size of the current block. For example, a planar mode (and / or directed planar modes) can be included in the list based on an area of ​​the current block (i.e., a product of the width and height of the current block) that is greater than a number of samples (e.g., 32) (or less than the number of samples). Similarly, a DC mode (and / or DC directional modes) can be included in the list based on an area that is greater than a second number (or less than a second number).

[0181] In some examples, the third TIMD mode can be determined as the non-angular IPM with the lowest cost among the multitude of non-angular IPMs. Costs (e.g., SSE, SAD, SATD, etc.) of the non-angular IPMs can be determined in a similar manner to the costs of the IPMs in Block 1902. For example, the video encoder can calculate the costs based on differences between reconstructed template samples and predicted template samples generated from the template reference samples using a corresponding non-angular IPM in the list. In some embodiments, the costs of the multitude of non-angular IPMs can be determined by the video encoder prior to Block 2008 of the Fig. 20 are calculated as described below, wherein the video encoder determines whether at least one non-angular IPM exists among at least two different IPMs from available adjacent blocks. In these embodiments, the video encoder, acting as a TIMD mode predictor, can select a non-angular IPM as the most cost-effective IPM.

[0182] In some examples, the costs assigned to a non-angular IPM in the list can be adjusted based on the type of non-angular IPM. For instance, the costs can be scaled by a factor f (e.g., greater than 1) based on the assumption that the non-angular IPM is of a first type, so other types of non-angular IPMs might have the lowest costs. For example, f can be set to 2 for the DC mode, so the adjusted costs for the DC mode are more likely to be higher than the costs for the planar mode, which may lead to a greater likelihood of selecting the planar mode.

[0183] For example, the planar mode may be preferred to the DC mode if the following condition (19) is met: costPlanar <f.costDC where f is a scaling factor greater than 1 (e.g., 2). Each type of non-angular IPM in the list can be assigned a corresponding scaling factor, which can be different or the same.

[0184] In some embodiments, the costs for non-angular modes may have already been calculated prior to or as part of Block 1902, so the video encoder may not need to recalculate the costs. For example, the Block 1902 IPMs may include candidates from an MPM list, which may include one or more non-angular IPMs.

[0185] In some examples, the costs of non-directional, non-angular IPMs are first compared. Then, the costs of one or more directional modes (e.g., horizontal and / or vertical) of the lower-cost non-angular IPM can be further calculated to select the non-angular IPM for block 1908. For example, if the list of non-angular IPM candidates includes {DC, planar} and the planar mode has the lower cost, the costs of horizontal and vertical planar modes can be further determined / calculated. For instance, a scaling factor (e.g., greater than 1) can be applied to the costs of the directional modes to favor the non-directional modes.

[0186] In some examples, the video encoder determines whether the third TIMD mode is selected based on one or more conditions. Examples of different conditions are described below. For instance, the video encoder determines whether the third TIMD mode (IPM) is selected. NA ) should be selected based on whether both the first TIMD mode (IPM1) and the second TIMD mode (IPM2) are not the non-angled IPMs (e.g., both the first TIMD mode and the second TIMD mode are different from the non-angled IPM).

[0187] In other words, the third TIMD mode can be selected if the following condition (20) is determined: IPM1!=IPMNAi&&IPM2!=IPMNAi where: IPM1 is the selected first TIMD mode, IPM2 is the selected second TIMD mode (different from IPM1), and IPM NAThe i-th non-angular intra-forecast mode in the non-angular IPM list. IPM1 and IPM2 may have the lowest cost (and / or the lowest adjusted cost) of the multitude of IPMs. In some examples, this condition can be useful so that certain non-angular modes are not overweighted, since if the selected IPM1 and / or IPM2 are already non-angular modes, activating a third IPM is NA It's as if the weighting of the non-angled IPM1 and / or IPM2 is distorted. In other examples, the AND ("&&") condition could be a different logical operator and, for example, replaced by an OR ("||") condition.

[0188] In some examples, non-angular directional modes are considered different from angular non-directional modes. If IPM1 is an angular mode and IPM2 is the DC mode, but IPM NAIf the mixing process is planar (or a planar directional mode), it can be performed using IPM. NA enclose, since IPM NA can be considered a mode that differs from IPM1 and IPM2.

[0189] In some examples, one or more non-angular modes generated from templates of reference blocks (e.g., for determining prediction blocks) of blocks adjacent to the current block can be designated as IPM1, IPM2, and / or IPM NA The modes can be selected based on the lowest cost (e.g., template customization costs). For example, IPM1 could be such a non-angular mode, IPM2 could be an angular mode, and IPM NA It can be another (different from IPM1, for example) non-angular mode, generated from the template of the reference block (e.g., prediction block) of an adjacent block to the current block.

[0190] In some examples, the video encoder may determine whether to select the third TIMD mode based on the first TIMD mode and the second TIMD mode, which does not include the non-angular IPM. In some further examples, the determination of whether to select the third TIMD may be based on whether the first TIMD mode and the second TIMD mode include a non-angular TIMD mode (e.g., one that includes the non-angular TIMD mode).

[0191] In some examples, the selection of the third TIMD mode may also be based on a cost comparison of the specified / selected non-angular IPM (costIPM). NA) against a cost threshold. For example, the cost threshold can be a fixed value. In some examples, the threshold costs can be based on one or more costs from the multitude of IPMs, such as the first IPM and / or the second IPM. In one example, the threshold cost calculation involves scaling the costs of a selected TIMD mode (e.g., the first TIMD mode with the lowest costs) by a scaling factor a greater than 1. For example, the third TIMD mode, the IPM NA includes, for which the mixture is selected in the fusion process if the following condition (21) is met: CostsIPMNA <a.KonstenIPM1 where: a is greater than 1 (e.g., a can be a value in the range of 1 to 2); Cost IPM NA a cost factor (e.g. SATD, SSE, SAD, etc.) of IPM NAis; and CostIPM1 is the cost of the first TIMD mode. In some examples, the first TIMD mode may have the lowest cost. In other examples, the first TIMD mode may have the second-lowest cost, and / or CostIPM NA may refer to a combination (e.g., an average, mean, etc.) of the costs of one or more of the lowest costs of IPMs among the multitude of IPMs.

[0192] In some examples, conditionally mixing the third TIMD to determine the TIMD mode predictor using the cost threshold can allow the selective addition of the non-angular mode, resulting in a reduced prediction error between predicted template samples and reconstructed template samples. Reducing this prediction error should also decrease the prediction error between current block samples and a prediction block for the current block generated based on applying the TIMD mode predictor to the current block template.

[0193] In some examples, the video encoder can determine, separately or in addition to one or more of the conditions described above, whether to select the third TIMD mode for generating the TIMD mode predictor based on the current block size, the current block image resolution, and / or the template size. For example, a range of current block sizes (or a maximum or minimum size threshold) can be defined for which the third TIMD mode (including a non-angular IPM) is selected. For example, smaller current blocks (e.g., 4x4, 4x8...) tend to show a stronger spatial correlation with reconstructed or predicted template samples, as the prediction improves the closer the samples are to those used to generate the prediction. Conversely, a larger current block (e.g., 32x32, 16x32...) may show a higher spatial correlation with reconstructed or predicted template samples.) show a lower correlation between its lower right samples and the template prediction samples (which are mostly located at the top, left and / or top left of the current block).

[0194] In some examples, when the third TIMD mode is selected or enabled (i.e., the selected IPM) NA (intended for use in the fusion / mixing process), the video encoder can perform a position-dependent prediction combination process (PDPC). In the PDPC process, the video encoder can generate a linear combination of the prediction and reference samples to smooth the discontinuity of the boundaries, which can be applied to the non-angular prediction samples.

[0195] At block 1918, the video encoder determines the TIMD mode predictor, which comprises a linear combination of the first and second TIMD modes, based on the fact that the third TIMD mode was not selected. Block 1918 can include block 1920. At block 1920, the video encoder determines the respective weights of the first and second TIMD modes in the linear combination, based on their costs.

[0196] In some examples, each TIMD mode's weight can be determined to be inversely proportional to the mode's cost, such that lower-cost modes (e.g., SSE, SSD, or SADT) are assigned a higher weight. In other words, the lower-cost TIMD mode has a higher weight than the higher-cost TIMD mode. For example, a first weight (Weight1) for the first TIMD mode with the first cost (CostMode1) and a second weight (Weight2) for the second TIMD mode with the second cost (CostMode2) can be determined as follows: Weighting 1 = Cost Mode 2, Cost Mode 1 + Cost Mode 2 Weighting2 = 1 − Weighting1 = Cost Mode 1, Cost Mode 1 + Cost Mode 2

[0197] In block 1910, the video encoder determines the TIMD mode predictor based on the selection of the third TIMD mode. This predictor is a linear combination of the first, second, and third TIMD modes. Block 1910 can include blocks 1912 and 1914.

[0198] At block 1912, the video encoder determines a third weighting for the third TIMD mode in the linear combination of block 1910. In some examples, the third weighting may be based on the cost of the third TIMD mode (e.g., depend on or vary from it), as described in more detail below. At block 1914, the video encoder determines the respective weightings of the first and second TIMD modes in the linear combination of block 1910, based on their costs.

[0199] In some embodiments, weights (e.g., first, second, and third weights) of the TIMD modes (e.g., first, second, and third TIMD modes) selected to determine the TIMD mode predictor can be determined (e.g., calculated) based on the costs of the TIMD modes. For example, a weight for a TIMD mode can be determined such that it is inversely proportional to the cost of the TIMD mode and / or proportional to the difference between the sum of the costs and the cost of the TIMD mode. The weights of the TIMD modes can be determined, for example, according to the following equation (24): wi=∑i=1nCostIPMi−CostIPMi(n−1)×∑i=1nCostIPMi

[0200] w i is the weight assigned to the i-th IPM of the i-th TIMD mode. CostIPM iwi are the costs assigned to (or determined by) the i-th IPM of the i-th TIMD mode. For example, w1 represents the first weight of the first TIMD mode, w2 represents the second weight of the second TIMD mode, and w3 represents the third weight of the third TIMD mode. n (e.g., 3, 4, etc.) represents the number of TIMD modes to be combined (i.e., in a linear combination with the weights) to generate the TIMD mode predictor.

[0201] In some embodiments, the third weight (w3), corresponding to the third TIMD mode comprising a non-angular IPM, can be determined in block 1912 before the first and second weights of the first and second TIMD modes, respectively, are determined in block 1914. The third weight can be determined based on equation (24), where n is the number of three TIMD modes and i is the number 3.

[0202] In some examples, the video encoder can adjust (e.g., clip) the third weight to keep it within a weighting range, for example, if the third weight is below or above a minimum or maximum value within the range. For example, the third weight can be adjusted according to the following equation (25): w3˜=min(b,max(a,w3)) b can be the maximum value of the range [a, b] and is greater than a, which can be the minimum value of the range [a, b]. For example, a = 64 / 12 and b = 64 / 3.

[0203] In some examples, the video encoder can adjust the third weight after applying a transformation to the third weight. For example, the third weight can be adjusted (e.g., clipped) so that it lies within a weight range according to the following equation (26): w3˜=min(b,max(a,f(w3))) f(x) can be a linear transformation function of the weighting. For example, f(x)=x2−1.

[0204] In some examples, the first weight (w1) and / or the second weight (w2) can be determined based on the third weight, which may be an adjusted third weight. For example, the first weight and the second weight can be determined according to the following equations (27) and (28): w1=∑i=1n−1CostIPMi−CostIPM1∑i=1n−1CostIPMi w2=1−(w3˜+w1)

[0205] In some embodiments, the third weighting can be a weighting value (W). NA), which is determined based on the size of the current block or the image resolution. In some examples, a larger current block size (e.g., the area of ​​the current block or the product of its width and height) may be assigned a higher weight value. For example, the third weight may be set to a first weight value (e.g., 64 / 8) based on the current block size being less than or equal to a first value (e.g., 32). Similarly, the third weight may be set to a second weight value (e.g., 64 / 6) based on the current block size being larger than the first value. In some examples, the third weight may be set to zero because the size is smaller than a second value (e.g., 16).

[0206] Returning to blocks 1910 and 1918, the TIMD mode predictor comprises a linear combination of a number of selected TIMD modes (e.g., the first and second TIMD modes are selected in block 1918, and the first, second, and third TIMD modes are selected in block 1910). The linear combination includes the weights assigned to the TIMD modes, as described above for blocks 1912, 1914, and 1920. For example, the TIMD mode predictor (IPM) TIMD ) according to the following linear combination (29): IPPMTIMD=∑i=1nwi×IPMi w i The weight can be for an i-th TIMD mode that includes an i-th IPM. As described above, the weight (w3) for the third TIMD mode can be set to 0 if it is determined that the third TIMD mode was not selected in block 1908.

[0207] At block 1922, the video encoder generates a prediction block for the current block based on the TIMD mode predictor. For example, the video encoder can apply the TIMD mode predictor to the template of the current block to generate the prediction block for the current block (e.g., as above with reference to Fig. (described in sections 10 to 12 and 17).

[0208] In some examples, if the video encoder is a decoder, the decoder can reconstruct the current block based on the prediction block and a residual (e.g., a prediction error or a residual block). For example, the decoder can receive (e.g., decode) the residual signal of a bitstream.

[0209] In some examples, if the video encoder is an encoder, the encoder can determine (e.g., generate) the residual based on a difference between the prediction block and the current block. The encoder can then encode this determined residual in the bitstream.

[0210] Fig. Figure 20 shows a flowchart 2000 of a method for applying a TIMD technique to a current block according to some embodiments. The method of flowchart 2000 can be implemented by a video encoder (e.g. encoder 200 in Fig. 2 or decoder 300 in Fig. 3) be implemented. The flowchart procedure can be performed reciprocally by an encoder and a decoder, so that an intra-prediction mode does not need to be signaled by the encoder to the decoder to reconstruct the current block using intra-prediction techniques. In some examples, one or more steps of flowchart 2000 can be performed before block 1903 of the Fig. 19 will be carried out.

[0211] For block 2002, the video encoder determines the availability of one or more adjacent blocks for intraprediction of the current block. For example, the one or more adjacent blocks may be spatial candidate blocks contiguous with the current block, as described above with reference to FIG. 15A-B. Examples of these adjacent blocks are a block to the left, above, above-left, below-left, and above-right of the current block. In some examples, based on the availability of one or more adjacent blocks, the video encoder may determine intraprediction modes of one or more of the available adjacent blocks.

[0212] As above in relation to Fig. As described in section 17A, the video encoder can determine a template of the current block and a template reference to perform TIMD. For example, the template can include template regions derived from a reconstructed region of an image of the current block and defined relative to a position of the current block. The template regions can include a left template region adjacent to the current block and an upper template region adjacent to the current block. For example, the template reference can include samples of regions from the reconstructed region defined relative to the position of the current block and / or a position of the template. In some examples, template regions and the template reference can further depend on which of the adjacent blocks are available.

[0213] At block 2012, the video encoder determines a TIMD mode predictor based on a planar mode, based on the unavailability of one or more adjacent blocks (e.g., based on the fact that none of the adjacent blocks of the current block at block 2004 are available). For example, the TIMD mode predictor can be determined as a planar mode.

[0214] For block 2006, the video encoder determines one or more intra-prediction modes (IPMs) based on the availability of one or more adjacent blocks in block 2004. These IPMs are associated with one or more available adjacent blocks. In some examples, the determination of the one or more IPMs may be based on the application of TIMD to encode the current block. In other examples, the one or more IPMs may be determined from a variety of candidate IPMs, excluding candidate IPMs determined (e.g., derived) from reference blocks (e.g., used to determine prediction blocks) of adjacent blocks encoded using IBC or intra-TMP modes.

[0215] In block 2008, the video encoder determines whether one or more IPMs are non-angle intrapredictor modes (IPMs). In block 2014, based on whether one or more IPMs are non-angle IPMs, the video encoder determines the TIMD mode predictor based on a non-angle IPM from a variety of non-angle IPMs. Examples of non-angle IPMs are shown above. Fig. 17A, Fig. 17B and Fig. 19 described.

[0216] In one example, the video encoder determines the TIMD mode based on the non-angular IPM in response to the fact that at least one of the one or more IPMs is a non-angular IPM of at least two distinct IPMs. In some examples, the video encoder selects the lowest-cost non-angular IPM (e.g., SAD, SATD, SSE, etc.) from a variety of options as the TIMD mode predictor.

[0217] In some examples, the video encoder can determine whether the number of non-angle IPMs is greater than or equal to a threshold. For example, the threshold can be set based on the number of available neighboring blocks of the current block and / or the number of different IPMs of available neighboring blocks. For example, if at least 3 neighboring blocks are available, the threshold can be set to 2, where at least 2 of the IPM modes of the available neighboring blocks must be different. If, for example, only 1 or 2 neighboring blocks are available, the threshold can be set to 1.

[0218] In Block 2010, the video encoder determines a variety of costs for a variety of IPMs, which include one or more IPMs (e.g., referenced in Block 2006 and Block 2008). In some examples, Block 2010 may correspond to Block 1902 from Fig. 19, and the video encoder may perform the same or similar operations.

[0219] Fig. Figure 21 shows a flowchart 2100 of a method for applying a TIMD technique to a current block according to some embodiments. The method of flowchart 2100 can be performed by a video encoder (e.g. encoder 200 in Fig. 2 or decoder 300 in Fig. 3) be implemented. The flowchart procedure can be performed reciprocally by an encoder and a decoder, so that an intra-prediction mode does not need to be signaled by the encoder to the decoder to reconstruct the current block using intra-prediction techniques. In some examples, Flowchart 2100 shows more detailed operations that apply to blocks 1902 and 1904 of Fig. 19 correspond.

[0220] At block 2102, the video encoder, based on the application of TIMD for the current block, determines a variety of costs from a variety of intraprediction modes (IPMs). In some examples, block 2102 corresponds to block 1902 from Fig. 19, and the video encoder can perform the same or similar operations. Block 2102 can include block 2104.

[0221] At block 2104, the video encoder generates the multitude of IPMs, which includes candidate IPMs from a list of most probable modes (MPM). In some examples, the multitude of IPMs may also include a DC mode, a planar mode, a vertical planar mode, a horizontal planar mode, a vertical DC mode, a horizontal DC mode, or a combination thereof. In some examples, the multitude of IPMs may also include candidate IPMs generated from a template of the reference block (e.g., used to determine the prediction block) of a block adjacent to the current block. The cost of each IPM in the MPM list is then determined and corresponds to the respective cost of the multitude.

[0222] In some examples, the MPM list can be generated for each block for use in an intra-prediction mode. For example, to signal an intra-prediction mode (IPM), the encoder can encode an index into the MPM list that specifies a selected IPM. Using the MPM list reduces the number of bits involved in signaling the indices of the selected IPMs. In some examples, the MPM list contains 22 IPM candidates. For instance, the MPM list includes an initial section (e.g., 6 IPMs) of candidate IPMs, referred to as the primary MPM list. The primary MPMs are the planar intra-prediction mode, an IPM from a left-adjacent block, an IPM from an above-adjacent block, an IPM from a bottom-left-adjacent block, an IPM from an above-right-adjacent block, and an IPM from an above-left-adjacent block, as in the example above. Fig. Described in 15A, the MPM list may include a second section (e.g., the next 16 candidate IPMs) of candidate IPMs, referred to as the secondary MPM list. The secondary MPM list includes IPMs derived by offsets from the IPMs in the primary MPM list. In some examples, one or more decoder-side intra-mode derivation (DIMD) modes may be added to the final MPM list after the primary MPMs and before the secondary MPMs. In some examples, other IPMs not included in the MPM list are included in a separate non-MPM list.

[0223] At block 2106, the video encoder determines a first TIMD mode and a second TIMD mode, based on a first IPM and a second IPM, from the set of IPMs with the lowest cost among the set of costs. In some examples, block 2106 corresponds to block 1904 from Fig. 19, and the video encoder can perform the same or similar operations. Block 2106 can include blocks 2108 to 2114.

[0224] At block 2108, the video encoder determines whether wide-angle IPMs are available for intra-prediction.

[0225] At block 2110, based on the unavailability of the wide-angle IPMs, the video encoder determines that the first TIMD mode includes the first IPM (i.e., the first TIMD mode is determined as the first IPM) and that the second TIMD mode includes the second IPM (i.e., the second TIMD mode is determined as the second IPM).

[0226] In some examples, based on the availability of the wide-angle IPM, the number of available angle IPMs can be extended from a first range (e.g., 67 modes) to a second range (e.g., 131 modes), and each of the first and second IPMs can be refined (e.g., adjusted) to determine the first TIMD mode and the second TIMD mode.

[0227] At block 2112, the video encoder determines the first TIMD mode based on the availability of wide-angle IPMs as one of the following: the first IPM and one or more wide-angle IPMs adjacent to the first IPM. For example, the one or more IPMs could be the two adjacent IPM modes (e.g., + / - 1 IPM) next to the first IPM.

[0228] In block 2114, the video encoder determines the second TIMD mode based on the availability of wide-angle IPMs as one of the following: the second IPM and one or more wide-angle IPMs adjacent to the second IPM. Blocks 2112 and 2114 can be performed in any order. For example, one or more IPMs can be determined as the two adjacent IPM modes (e.g., + / - 1 IPM) to the second IPM.

[0229] Fig. Figure 22 shows a flowchart 2200 of a method for applying a TIMD technique to decode a current block according to some embodiments. The method of flowchart 2300 can be implemented by a decoder (e.g., decoder 300 in Fig. 3) be implemented. Unlike Flowchart 1900, which shows an example of how a video encoder determines whether to select the third TIMD mode to be combined / mixed / fused with the first and second TIMD modes, Flowchart 2200 shows an example in which an indication is signaled / received in a bitstream to indicate whether the third TIMD mode is selected. Many operations in Flowchart 2200 are the same as or similar to those in Flowchart 1900. A block in Flowchart 2200 that corresponds to a block in Flowchart 1900 indicates that the video encoder can perform the same or similar operations.

[0230] At block 2201, the video encoder receives from a bitstream a first indication that TIMD is applied for the current block, and a second indication of a TIMD mode predictor, which includes a TIMD mode based on a non-angular IPM.

[0231] In block 2202, the video encoder determines a variety of costs for a variety of intraprediction modes (IPMs) based on the initial input. The costs can be calculated similarly to those in block 1902. Fig. 19 described and determined.

[0232] At block 2203, the video encoder determines a TIMD predictor based on the multitude of IPMs. Block 2203 can include blocks 2204 to 2220.

[0233] At block 2204, the video encoder determines a first TIMD mode and a second TIMD mode, based on a first IPM and a second IPM, from the set of IPMs with the lowest cost among the set of costs. In some examples, block 2204 corresponds to block 1904. Fig. 19. At block 2206, the video encoder determines whether the second TIMD mode should be selected. In some examples, block 2206 corresponds to block 1906 from Fig. 19. At block 2216, the video encoder determines the TIMD mode predictor based on the first TIMD mode, given that the second TIMD mode has been selected. In some examples, block 2216 corresponds to block 1916. Fig. 19.

[0234] In block 2208, the video encoder determines whether to select a third TIMD mode based on the second specification. For example, based on the second specification, which indicates that the TIMD mode predictor includes a TIMD mode based on a non-angular IPM, the video encoder can select the third TIMD mode, which includes a first non-angular IPM, from a variety of non-angular IPMs. Examples of non-angular IPMs are given above in block 1908. Fig. 19 described. In contrast, in block 1908 of Fig. 19 The encoder and the decoder mutually inquire whether the third TIMD mode should be selected, and no indication is signaled by the encoder to be parsed by the decoder.

[0235] At block 2210, the video encoder determines, based on the second indication that the third TIMD mode is applied (e.g., enabled or selected), that the TIMD predictor comprises a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode. In some examples, the video encoder may interpret the linear combination of block 2210 similarly to the linear combination of block 1910. Fig. 19 determine. In some examples, the video encoder determines a third weighting of the third TIMD mode at block 2212, similar to that determined with respect to block 1912 by Fig. 19. In some examples, the video encoder determines the respective weights of the first TIMD mode and the second TIMD mode in block 2214, similar to the context of block 1914 of Fig. 19 described.

[0236] At block 2218, based on the second piece of information indicating that the third TIMD mode is not applied (e.g., disabled or not selected), the video encoder determines that the TIMD predictor comprises a linear combination of the first TIMD mode and the second TIMD mode, similar to the one used with respect to block 1918. Fig. 19 was described. Similar to Block 1920 by Fig. 19 The video encoder determines the weighting of the first TIMD mode and the second TIMD mode based on the costs of the first TIMD mode and the second TIMD mode, respectively.

[0237] At block 2222, the video encoder generates a prediction block for the current block based on the TIMD predictor. In some examples, block 2222 corresponds to block 1922 of Fig. 19.

[0238] Fig. Figure 23 shows a flowchart 2300 of a method for applying a TIMD technique to encode a current block according to some embodiments. The method of flowchart 2300 can be implemented by an encoder (e.g., encoder 200 in Fig. 2) be implemented.

[0239] At block 2302, the coder, based on the TIMD application for the current block, determines a variety of costs for a variety of intraprediction modes (IPMs). The cost for each IPM of the variety of IPMs is determined based on the differences between: predicted samples, a template of the current block generated from reference samples of the template and using the IPM; and reconstructed samples of the template. In some examples, block 2302 might correspond to block 1902 of Fig. 19.

[0240] In block 2304, the encoder determines a first TIMD mode and a second TIMD mode, based on a first IPM and a second IPM, from the set of IPMs with the lowest cost among the set of costs. In some examples, block 2304 may correspond to block 1904 of Fig. 19.

[0241] In block 2306, the coder determines that a third TIMD mode encompasses a variety of non-angular IPMs (e.g., not specified). For example, the third TIMD mode may be set to the lowest-cost non-angular IPM determined for the non-angular IPMs. Examples of determining the third TIMD mode are given with reference to block 1906. Fig. 19 described.

[0242] At block 2308, the coder selects one of the first and second TIMD mode predictors as the TIMD mode predictor based on: the first TIMD cost of a first TIMD mode predictor that includes a first linear combination of the first TIMD mode and the second TIMD mode; and the second TIMD cost of a second TIMD mode predictor that includes a second linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode. In other words, the TIMD costs between the first TIMD mode predictor, which is associated with the regular TIMD mode, and the second TIMD mode predictor, which is associated with mixing with a non-angular IPM, can be compared to select more effective intra-modes. In some examples, the TIMD costs may be the costs of a tariff distortion (RD costs) (e.g., an RD audit).Performing the RD check may, for example, involve calculating the bit rate of the encoded block data taking into account the proposed / selected candidate mode (i.e., entropy-coded quantized residuals with associated mode signaling) and calculating the distortion (e.g., using PSNR) that is representative of the associated reconstructed block; subsequently, the two costs are combined (e.g., by adding or accumulating) (e.g., using the Lagrange multiplier technique) to obtain a value that represents the rate-distortion trade-off for the current block encoded with the selected candidate mode.

[0243] At block 2310, the encoder encodes in a bitstream a first indication (e.g., `timd_flag` syntax element) of the TIMD applied to the current block; and a second indication (e.g., `timd_non-angular_fusion_flag` syntax element) of whether the selected TIMD mode predictor is the first or second TIMD mode predictor. In some examples, the signaling of the second indication may depend on the first indication, which specifies that TIMD is applied to the current block.

[0244] In some examples, the second specification may include a syntax element (e.g., a flag) that indicates which of the two TIMD mode predictors leads to lower TIMD costs (e.g., lower Lagrange costs resulting from performing the RD check).

[0245] In some examples, the encoder can further encode one or more values ​​that are associated with parameters for the third TIMD mode. For example, one or more indicators can include a third value (e.g., `timd_non_angular_fusion_idc` syntax element) that signals the non-angular IPM selected by the encoder, which is used to determine the third TIMD mode. Explicitly signaling this non-angular IPM can reduce the computational overhead in the decoder, since the IPM NAIt is no longer necessary to determine the cost (e.g., SSE, SAD, SSD, SATD, etc.) for each candidate in the non-angular IPM list based on the current block template (which would be signaled directly in the bitstream). For example, a value of 0 in the third indicator can indicate that the regular TIMD procedure is used during the fusion process (e.g., up to two IPMs are blended, but excluding the non-angular IPM from a non-angular IPM list). A value of 1 in the third indicator can indicate that DC mode is used as the third TIMD mode to generate the TIMD mode predictor, and a value of 2 in the third indicator can indicate that planar mode is used as the third TIMD mode to generate the TIMD mode predictor. In some examples, the third indicator can include one or more other values ​​to indicate, for example, whether a horizontal and / or vertical planar or DC mode is used.In some examples, the coder can signal another indicator (e.g., a flag or a syntax element) to indicate that a direction mode is being used (e.g., setting timd_non_ang_directional_flag to 0 if not used, and to 1 if used) and which direction is selected (e.g., setting timd_non_ang_direction_flag to 0 for horizontal direction and to 1 for vertical direction).

[0246] Fig. Figure 24 shows a flowchart 2400 of a method for applying a TIMD technique to a current block according to some embodiments. The method of flowchart 2400 can be performed by a video encoder (e.g. encoder 200 in Fig. 2 or decoder 300 in Fig. 3) be implemented. The flowchart procedure can be performed reciprocally by an encoder and a decoder, so that an intra-prediction mode does not need to be signaled by the encoder to the decoder in order to reconstruct the current block using intra-prediction techniques.

[0247] At block 2402, the video encoder determines a multitude of costs for a multitude of intra-prediction modes (IPMs) based on the template-based intra-mode derivation (TIMD) applied to the current block. The cost of each IPM from the multitude of IPMs can be determined based on the differences between: predicted samples, a template of the current block generated from reference samples of the template and using the IPM; and reconstructed samples of the template. For example, the template may include one or more template regions containing reconstructed samples that are adjacent to or near the current block, as shown above. Fig. 17A and Fig. As described in section 17B, the template can, for example, include a left template area to the left of the current block and / or a top template area above the current block.

[0248] In some examples, if the video encoder is an encoder, the encoder can signal in a bitstream that TIMD is being applied to the current block. In some examples, if the video encoder is a decoder, the decoder can receive from the bitstream (e.g., parse and decode) the signal that TIMD is being applied to the current block.

[0249] At block 2404, the video encoder determines a first TIMD mode and a second TIMD mode, based on a first IPM and a second IPM, from the multitude of IPMs with the lowest cost among the multitude. For example, block 2404 might correspond to block 1904 of Fig. 19. In some examples, the video encoder determines a variety of TIMD modes, each with a cost below a threshold determined based on the lowest-cost TIMD mode. For example, more than two TIMD modes can be determined and selected.

[0250] At block 2406, the video encoder determines a third TIMD mode based on a non-angular IPM from a variety of non-angular IPMs. In some examples, the third TIMD mode may be as above with respect to block 1908. Fig. 19 can be determined. In some examples, the third TIMD mode includes the non-angular IPM.

[0251] In some examples, the variety of non-angled IPMs includes a direct current (DC) mode and a planar mode. In one example, the variety of non-angled IPMs further includes a horizontal planar mode, a vertical planar mode, a horizontal DC mode, and / or a vertical DC mode.

[0252] In some examples, the multitude of non-angular IPMs may further include one or more candidate IPMs generated from templates of reference blocks of neighboring blocks (e.g., adjacent to the current block) encoded with IBC or IntraTMP modes.

[0253] In some examples, determining the third TIMD mode involves determining a second set of costs for the set of non-angular IPMs, where the cost for each non-angular IPM of the set of non-angular IPMs is based on the differences between: predicted template samples generated from the non-angular IPM applied to the template reference samples; and the reconstructed template samples.

[0254] In some examples, the third TIMD mode is determined to be the non-angular IPM with the lowest cost of the second set of costs.

[0255] In some examples, the cost for each non-angular IPM is further based on a scaling factor that is assigned to a type of non-angular IPM.

[0256] At block 2408, the video encoder determines, based on the check that neither the first RIMD mode nor the second TIMD mode includes the non-angular IPM (i.e., the non-angular third IPM is planar, and neither the first nor the second IPM is a planar intra-prediction mode), that a TIMD mode predictor includes a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode. For example, block 2408 might correspond to block 1910 of Fig. 19.

[0257] In some examples, the determination of the TIMD mode predictor that includes the linear combination is further based on both the first TIMD mode and the second TIMD mode that does not include non-angular IPMs.

[0258] In some examples, the determination of the TIMD mode predictor, which includes the linear combination, is further based on the non-angular IPM cost being less than or equal to a threshold. For example, the cost threshold for the video encoder can be determined based on the lowest costs of the first TIMD mode and the second TIMD mode. The video encoder can, for instance, determine the threshold cost to include the lowest cost multiplied by a scaling value.

[0259] In some examples, the determination of the TIMD mode predictor, which includes the linear combination, is further based on the size of the current block or the size of an image of the current block. In one example, the linear combination, which includes the third IPM, is further based on the current block size being greater than or equal to a first threshold size and / or the image size being greater than or equal to a second threshold size.

[0260] In some examples, the multitude of costs includes the sum of squared differences (SSD), the sum of absolute differences (SAD), or the sum of absolute transformed differences (SATD).

[0261] In some examples, the linear combination includes a first weight for the first TIMD mode, a second weight for the second TIMD mode, and a third weight for the third TIMD mode. For example, the first, second, and third weights can each be determined as inversely proportional to the respective costs of the first, second, and third TIMD modes. Alternatively, the first, second, and third weights can each be determined as proportional to the difference between the sum of the costs and the respective costs of the first, second, and third TIMD modes.

[0262] In some examples, the third weighting of the third TIMD mode can be determined as being proportional to the difference between the sum of the first, second and third costs and the third cost of the third TIMD mode.

[0263] In some examples, the third weighting includes a weight value based on the size of the current block or the image resolution of the current block.

[0264] In some examples, if the third weighting is outside a weighting range, it can be adjusted to fall within the weighting range.

[0265] At block 2410, the video encoder generates a prediction block for the current block based on the TIMD mode predictor and the template. For example, block 2410 might correspond to block 1922 of Fig. 19.

[0266] In some examples, if the video encoder is an encoder, it can determine a residual (e.g., a prediction error or a residual block) based on (e.g., a difference) between the prediction block and the current block, and encode the residual in the bitstream. In other examples, if the video encoder is a decoder, it can decode the residual of the current block from the bitstream and reconstruct the current block based on the prediction block and the residual (e.g., by combining or adding them).

[0267] Embodiments of the present disclosure can be implemented in hardware using analog and / or digital circuits, in software, by executing instructions through one or more general-purpose or specialized processors, or as a combination of hardware and software. Consequently, embodiments of the disclosure can be implemented in the environment of a computer system or other processing system (e.g., an IC in a television, mobile phone, set-top box, etc.). An example of such a computer system 2500 is described in Fig. 25 shown. Blocks depicted in the preceding figures, such as the blocks in Fig. 1, Fig. 2 and Fig. 3. can be executed on one or more computer systems 2500. Furthermore, each of the steps of the flowcharts presented in this disclosure can be implemented on one or more computer systems 2500.

[0268] Computer system 2500 includes one or more processors, such as processor 2504. Processor 2504 can be, for example, a special-purpose processor, a general-purpose processor, a microprocessor, or a digital signal processor. Processor 2504 can be connected to a communication infrastructure 2502 (for example, a bus or network). Computer system 2500 can also include main memory 2506, such as random-access memory (RAM), and secondary memory 2508.

[0269] The secondary storage device 2508 can, for example, include a hard disk drive 2510 and / or a removable storage device 2512, which is a magnetic tape drive, an optical drive, or the like. The removable storage device 2512 can read from and / or write to a removable storage unit 2516 in a known manner. The removable storage unit 2516 is a magnetic tape, an optical disk, or the like, which is read from and written to by the removable storage drive 2512. As is known to those skilled in the art in the relevant field(s), the removable storage unit 2516 includes a computer-usable storage medium on which computer software and / or data is / are stored.

[0270] In alternative implementations, the secondary storage 2508 may include other similar means to enable the loading of computer programs or other instructions into the computer system 2500. These means may include, for example, a removable storage unit 2518 and an interface 2514. Examples of such means may include a program cartridge and a cartridge interface (such as those found in video game devices), a removable memory chip (such as an EPROM or PROM) and an associated socket, a USB flash drive and a USB connector, and other removable storage units 2518 and interfaces 2514 that enable the transfer of software and data from the removable storage unit 2518 to the computer system 2500.

[0271] The Computer System 2500 can also include a Communication Interface 2520. The Communication Interface 2520 enables the transfer of software and data between the Computer System 2500 and external devices. Examples of a Communication Interface 2520 include a modem, a network interface (such as an Ethernet card), a communication port, etc. Software and data transmitted via the Communication Interface 2520 are in the form of signals, which can be electronic, electromagnetic, optical, or other signals that can be received by the Communication Interface 2520. These signals are provided to the Communication Interface 2520 via a Communication Path 2522. The Communication Path 2522 transmits signals and can be implemented using wires or cables, fiber optics, a telephone line, a mobile phone connection, an RF connection, and other communication channels.

[0272] As used herein, the terms “computer program medium” and “computer-readable medium” refer to tangible storage media, such as removable storage units 2516 and 2518, or a hard disk installed in the hard disk drive 2510. These computer program products are means of providing software to computer system 2500. Computer programs (also called computer control logic) can be stored in main memory 2506 and / or secondary memory 2508. Computer programs can also be received via communication interface 2520. When executed, these computer programs enable computer system 2500 to implement the present disclosure as discussed herein. In particular, when executed, the computer programs enable processor 2504 to implement the processes of the present disclosure, such as all the methods described herein.Accordingly, such computer programs represent controls of the Computer System 2500.

[0273] In another embodiment, features of the disclosure can be implemented in hardware, for example, using hardware components such as application-specific integrated circuits (ASICs) and gate arrays. The implementation of a hardware state machine to perform the functions described herein is also obvious to those skilled in the art.

Claims

[1] Method for predicting a block of pixel colors of an image, comprising: - Determine that template-based intra-mode derivation (TIMD) is applicable for predicting the block; - Determining a variety of intra-forecast modes, and determining the respective costs for the variety of intra-forecast modes; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming a final TIMD mode comprising a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode. [2] Method according to claim 1, wherein the respective costs are based on differences between predicted samples obtained by applying the respective intra-prediction mode to samples of a reference of the template and reconstructed samples of the template. [3] Method according to claim 1, wherein the determination of the third TIMD mode depends on the condition that the first TIMD mode and the second TIMD mode are not non-angular intra-prediction modes. [4] Method according to claim 1, wherein the determination of the final TIMD mode depends on the third costs of the non-angular intra-prediction mode being less than a threshold. [5] Method according to claim 4, wherein the threshold costs are based on the lowest of the costs of the first TIMD mode and the second TIMD mode. [6] Method according to claim 4, wherein the threshold costs are a product of a scaling value and the lowest of the costs of the first TIMD mode and the second TIMD mode. [7] Method according to any of the preceding claims, wherein the linear combination comprises multiplying each TIMD mode by a weighting that depends on its cost. [8] Method according to claim 7, wherein the weights are calculated as the numerator of the sum of all weights less the respective weight of the respective TIMD mode divided by a denominator that is a multiple of the sum of all weights. [9] Method according to any of the preceding claims, wherein the non-angular intra-prediction mode is a DC mode or a planar mode. [10] Method according to any of the preceding claims, wherein the non-angle-based intra-prediction mode is based on a reference block that is usable for the intra-block copy of an adjacent block of the block, or on a reference block that is usable for the intra-TMP prediction of an adjacent block of the block. [11] Method according to claim 1, wherein the non-angular intra-prediction mode is selected as the one with the lowest cost from a set of non-angular intra-prediction modes. [12] Video decoder comprising a computing circuit connected to a data storage device, wherein the computing circuit is arranged to: - Receiving a bitstream containing encoded video data; - wherein the encoded video data includes an initial indication that a pixel block is encoded based on a template-based intra-mode derivation; - wherein the encoded video data includes a second indication that a final TIMD mode is based on a non-angle-dependent intra-prediction mode; - Determining a variety of costs for a variety of intra-forecast modes, where each cost factor for a given IPM is based on differences in a template of the block between the following: predicted samples obtained by prediction from reference samples of the template using the respective IPM; and reconstructed samples of the original; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming the final TIMD mode, which comprises a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode. [13] Video encoders, comprehensive: - Receiving an original video image that includes a block to be encoded using template-based intra-prediction coding; - Determining a variety of intra-forecast modes, and determining the respective costs for the variety of intra-forecast modes; - Determining a first TIMD mode, which is the first of the multitude of intra-forecast modes with the lowest costs, and a second TIMD mode, which is the second of the multitude of intra-forecast modes with the second-lowest costs; - Determining a third TIMD mode, which is a non-angle-dependent intra-prediction mode; - Check if the third TIMD mode differs from the first TIMD mode and if the third TIMD mode differs from the second TIMD mode; - based on the successful verification, forming a final TIMD mode comprising a linear combination of the first TIMD mode, the second TIMD mode, and the third TIMD mode; and - Forming a prediction of the block based on a template of the block and the final TIMD mode. - Encoding in encoded video data a first indication that a pixel block is encoded based on a template-based intra-mode derivation, and a second indication that a final TIMD mode is based on a non-angle-based intra-prediction mode.