Template-based intra mode derivation and directional sample-by-sample fusion

By using template-based intra-frame pattern derivation and directional sample-by-sample fusion techniques, the problem of low efficiency in removing redundant information in video coding is solved, achieving more efficient video coding and decoding results.

CN121970314APending Publication Date: 2026-05-01OFINNO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
OFINNO LLC
Filing Date
2024-07-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video coding technologies suffer from low efficiency in removing redundant information when processing video sequences, especially in intra-frame prediction and inter-frame prediction, resulting in poor coding efficiency and decoding quality.

Method used

By employing template-based intra-frame pattern derivation (TIMD) and directional per-sample fusion techniques, redundant information in video sequences is reduced by performing intra-frame prediction pattern derivation and directional per-sample fusion on video blocks.

Benefits of technology

It improves the efficiency of video encoding and decoding quality, reduces the amount of data in the bitstream, and enhances the efficiency of video transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970314A_ABST
    Figure CN121970314A_ABST
Patent Text Reader

Abstract

A codec determines a plurality of costs applied to a plurality of intra prediction modes (IPMs) for predicting a template of a block based on template-based intra mode derivation (TIMD) applied to the block; a TIMD mode is determined based on an IPM of the plurality of IPMs. For each of the TIMD modes: a weight is determined based on a cost of the TIMD mode, and a directionality of the TIMD mode is determined based on comparing a first sub-cost for a cost of a first region of the template with a second sub-cost for a cost of a second region of the template. A TIMD mode predictor is determined based on a linear combination of the TIMD modes having weights adjusted according to the directionality of the TIMD modes. Based on the TIMD mode prediction value, the codec generates a prediction block for encoding the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 525,936, filed July 10, 2023, and U.S. Provisional Application No. 63 / 542,931, filed October 6, 2023, all of which are incorporated herein by reference in their entirety. Attached Figure Description

[0002] In the accompanying drawings, some features are shown by way of example and without limitation. In the drawings, the same numerals indicate the same elements.

[0003] Figure 1 An example video encoding / decoding system in which embodiments of the present disclosure can be implemented is shown.

[0004] Figure 2 An example encoder in which an embodiment of this disclosure can be implemented is shown.

[0005] Figure 3 An example decoder in which an embodiment of this disclosure can be implemented is shown.

[0006] Figure 4 An example quadtree partitioning of the Coding Tree Block (CTB) is shown.

[0007] Figure 5 It shows the relationship with Figure 4 The CTB example quadtree partitioning corresponds to the example quadtree.

[0008] Figure 6 Examples of binary tree partitioning and ternary tree partitioning are shown.

[0009] Figure 7 Examples of quadtree partitioning and multi-type tree partitioning with CTB combinations are shown.

[0010] Figure 8 It shows the relationship with Figure 7 The example trees shown correspond to the quadtree partitioning and multi-type tree partitioning of the CTB combination.

[0011] Figure 9 An example set of reference samples for intra-frame prediction determination for the current block is shown.

[0012] Figure 10A and Figure 10B An example intra-frame prediction mode is shown.

[0013] Figure 11 An example of the current block and its corresponding reference sample is shown.

[0014] Figure 12An example of applying an intra-frame prediction mode (e.g., angle mode) to predict the current block is shown.

[0015] Figure 13A An example of performing inter-frame prediction for the current block in the current image is shown.

[0016] Figure 13B An example motion vector is shown.

[0017] Figure 14 An example of bidirectional prediction performed on the current block is shown.

[0018] Figure 15A Example spatial candidate neighboring blocks are shown relative to the current block being encoded.

[0019] Figure 15B Example positions of two time-juxtaposed blocks relative to the current block are shown.

[0020] Figure 16 An example of intra-block copy (IBC) is shown.

[0021] Figure 17A An example of decoder-side intra-frame mode derivation (DIMD) for encoding the current block is shown according to some implementation schemes.

[0022] Figure 17B An example of template-based intra-frame mode derivation (TIMD) for encoding the current block is shown according to some implementation schemes.

[0023] Figure 18 An example is shown of signaling TIMD for decoding the current block, according to some implementation schemes.

[0024] Figure 19A A flowchart is shown as an example method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes.

[0025] Figure 19B A flowchart is shown as an example method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes.

[0026] Figure 20 A flowchart is shown as an example method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes.

[0027] Figure 21 A flowchart is shown as an example method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes.

[0028] Figure 22A block diagram of an example computer system in which embodiments of the present disclosure may be implemented is shown. Detailed Implementation

[0029] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure, including its structures, systems, and methods, can be practiced without these specific details. The descriptions and representations herein are common means used by those skilled in the art to most effectively communicate the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuits have not been described in detail to avoid unnecessarily obscuring aspects of this disclosure.

[0030] References to "an embodiment," "an embodiment," "an example embodiment," etc., in this specification indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is believed that the combination of other embodiments affects that feature, structure, or characteristic in a way that is within the knowledge of someone skilled in the art, whether explicitly stated or not.

[0031] Furthermore, note that each implementation can be described as a process depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. While a flowchart can describe operations as a sequential process, many of these operations can be performed in parallel or simultaneously. Moreover, the order of operations can be rearranged. When an operation of a process is completed, the process terminates, but there may be additional steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, the termination of the process can correspond to the function returning to the calling function or the main function.

[0032] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, including, or carrying instructions and / or data. Computer-readable media may include non-transitory media on which data can be stored but does not include carrier waves and / or transient electronic signals propagating via wireless or wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as optical discs (CDs) or digital multifunction discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0033] Furthermore, the implementation scheme can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) that perform the necessary tasks can be stored on a computer-readable or machine-readable medium. The processor can perform the necessary tasks.

[0034] Video sequences comprising multiple images / frames can be represented digitally for storage and / or transmission. Representing a video sequence digitally may require a large number of bits. The large data size associated with a video sequence may require significant resources for storage and / or transmission. Video encoding can be used to compress the size of the video sequence for more efficient storage and / or transmission. Video decoding can be used to decompress the compressed video sequence for display and / or other forms of consumption.

[0035] Figure 1An example video encoding / decoding system 100 in which embodiments of the present disclosure may be implemented is shown. The video encoding / decoding system 100 includes a source device 102, a transmission medium 104, and a target device 106. The source device 102 encodes a video sequence 108 into a bitstream 110 for more efficient storage and / or transmission. The source device 102 may store the bitstream 110 and / or may send / transmit the bitstream to the target device 106 via the transmission medium 104. The target device 106 decodes the bitstream 110 to display the video sequence 108. The target device 106 may receive the bitstream 110 from the source device 102 via the transmission medium 104. The source device 102 and / or the target device 106 may be any of a plurality of different devices (e.g., desktop computer, laptop computer, tablet computer, smartphone, wearable device, television, camera, video game console, set-top box, video streaming device, etc.).

[0036] Source device 102 may (e.g., for encoding video sequence 108 into bitstream 110) include one or more of video source 112, encoder 114, and / or output interface 116. Video source 112 may provide and / or generate video sequence 108 based on the capture of natural scenes and / or synthetically generated scenes. Synthetically generated scenes may be scenes that include computer-generated graphics and / or screen content. Video source 112 may include video capture equipment (e.g., a camera), video archives including previously captured natural scenes and / or synthetically generated scenes, a video feed interface for receiving captured natural scenes and / or synthetically generated scenes from a video content provider, and / or a processor for generating synthetic scenes.

[0037] A video sequence, such as video sequence 108, may include a series of pictures (also called frames). The video sequence may achieve a motion impression based on continuously presenting pictures of the video sequence using constant or variable time intervals between pictures. A picture may include one or more sample arrays of intensity values. Intensity values ​​can be acquired (e.g., measured, determined, provided) at a series of regularly spaced locations within the picture. A color picture may include (e.g., typically includes) a luminance sample array and two chrominance sample arrays. The luminance sample array may include intensity values ​​representing the luminance of the picture (e.g., the luminance component Y). The chrominance sample arrays may include intensity values ​​representing the blue and red components of the picture, separate from the luminance (e.g., chrominance components Cb and Cr). Other color picture sample arrays are possible based on different color schemes (e.g., red, green, blue (RGB) color schemes). For a given location in a sample array used to represent a color picture (e.g., three sample arrays for one luminance component and two chrominance components respectively), a pixel in the color picture may refer to / include / be associated with all intensity values ​​(e.g., luminance component, chrominance component). A monochrome picture may include a single luminance sample array. A pixel in a monochrome image can refer to / include / be associated with an intensity value (e.g., a luminance component) at a given location in a single luminance sample array used to represent the monochrome image.

[0038] Encoder 114 can encode video sequence 108 into bitstream 110. Encoder 114 can (e.g., for encoding video sequence 108) apply / use one or more prediction techniques to reduce redundant information in video sequence 108. Redundant information is information that can be predicted at the decoder and does not need to be transmitted to the decoder for accurate decoding of video sequence 108. For example, encoder 114 can apply spatial prediction (e.g., intra-frame prediction or intra-frame prediction), temporal prediction (e.g., inter-frame prediction or inter-frame prediction), inter-layer prediction, and / or other prediction techniques to reduce redundant information in video sequence 108. For example, before applying one or more prediction techniques, encoder 114 can divide the image including video sequence 108 into rectangular regions called blocks. Encoder 114 can then encode the blocks using one or more prediction techniques.

[0039] For temporal prediction, encoder 114 can search for blocks similar to the block being encoded in another picture (e.g., referred to as a reference picture) of video sequence 108. The block being encoded can then be predicted using the blocks identified during the search (e.g., referred to as prediction blocks). For spatial prediction, encoder 114 can form prediction blocks based on data from reconstructed neighboring samples of the block to be encoded within the same picture of video sequence 108. Reconstructed samples refer to samples that are encoded and then decoded. Encoder 114 can determine prediction error (e.g., also referred to as residual) based on the difference between the block being encoded and the prediction block. The prediction error can represent non-redundant information that can be sent / transmitted to the decoder for accurate decoding of video sequence 108.

[0040] Encoder 114 may apply a transform (e.g., using Discrete Cosine Transform (DCT) or any other transform) to the prediction error to generate transform coefficients. Encoder 114 may form bitstream 110 based on the transform coefficients and other information used to determine prediction blocks using / based on prediction type, motion vector, and / or prediction mode. Encoder 114 may perform one or more of quantization and entropy coding of the transform coefficients and / or other information used to determine prediction blocks, for example, before forming bitstream 110. Quantization and / or entropy coding may further reduce the number of bits required to store and / or transmit video sequence 108.

[0041] Output interface 116 may be configured to write and / or store bit stream 110 onto transmission medium 104 for transmission to target device 106. Alternatively, output interface 116 may be configured to send / transmit, upload, and / or stream bit stream 110 to target device 106 via transmission medium 104. Output interface 116 may include a wired and / or wireless transmitter configured to send / transmit, upload, and / or stream bit stream 110 according to one or more proprietary, open-source, and / or standardized communication protocols, such as Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Cable Data Service Interface Specification (DOCSIS) standards, 3rd Generation Partnership Project (3GPP) standards, Institute of Electrical and Electronics Engineers (IEEE) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and / or any other communication protocols.

[0042] The transmission medium 104 may include wireless, wired, and / or computer-readable media. For example, the transmission medium 104 may include one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic storage. Alternatively, the transmission medium 104 may include one or more networks (e.g., the Internet) or file servers configured to store and / or send / transmit encoded video data.

[0043] Target device 106 can decode bitstream 110 into video sequence 108 for display. Target device 106 may include one or more of input interface 118, decoder 120, and / or video display 122. Input interface 118 may be configured to be read by source device 102 from bitstream 110 stored on transmission medium 104. Alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 may include a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary, open-source, standardized communication protocols and / or any other communication protocols such as those cited herein.

[0044] Decoder 120 can decode video sequence 108 from encoded bitstream 110. Decoder 120 can generate prediction blocks for images of video sequence 108 in a manner similar to encoder 114, and determine the prediction error for each block, for example, to decode video sequence 108. Decoder 120 can generate prediction blocks using / based on prediction type, prediction mode, and / or motion vectors received in bitstream 110. Decoder 120 can determine the prediction error using transform coefficients received in bitstream 110. Decoder 120 can determine the prediction error by weighting the transform basis function using the transform coefficients. Decoder 120 can combine prediction blocks and prediction errors to decode video sequence 108. Video sequence 108 at target device 106 may or may not be the same video sequence as the transmitted video sequence (such as video sequence 108 transmitted by source device 102). Decoder 120 can decode video sequences that approximate video sequence 108, for example, due to lossy compression of video sequence 108 by encoder 114 and / or errors introduced into the encoded bitstream 110 during transmission to target device 106.

[0045] The video display 122 can display the video sequence 108 to a user. The video display 122 may include a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display, and / or any other display device suitable for displaying the video sequence 108.

[0046] The video encoding / decoding system 100 is merely an example, and video encoding / decoding systems different from and / or modified versions of the video encoding / decoding system 100 may similarly perform the methods and processes described herein. For example, the video encoding / decoding system 100 may include other components and / or arrangements. For example, the video source 112 may be external to the source device 102. Similarly, the video display 122 may be external to the target device 106 or omitted entirely, e.g., where the video sequence 108 is intended to be consumed by a machine and / or storage device. In this example, the source device 102 may also include a video decoder, and the target device 106 may also include a video encoder. For example, the source device 102 may be configured to further receive encoded bitstreams from the target device 106 to support bidirectional video transmission between the devices.

[0047] Encoder 114 and / or decoder 120 may operate according to one or more proprietary or industry video coding standards. For example, encoder 114 and / or decoder 120 may operate according to one or more proprietary, open-source and / or standardized protocols (such as ITU-T H.263, ITU-T H.264 and Moving Picture Experts Group (MPEG)-4 video (also known as Advanced Video Coding (AVC)), ITU-T H.265 and MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)), ITU-T H.265 and MPEG-I Part 3 (also known as Universal Video Coding (VVC)), WebM VP8 and VP9 codecs and / or AOMedia Video 1 (AV1) and / or any other video coding protocol).

[0048] Figure 2 An example encoder is shown. (e.g.) Figure 2 The encoder 200 shown can implement one or more of the processes described herein. Encoder 200 can encode video sequence 202 into bit stream 204 for more efficient storage and / or transmission. Encoder 200 can, for example... Figure 1 Implemented in the video encoding / decoding system 100 shown (e.g., as encoder 114), or in any computing, communication, or electronic device (e.g., desktop computer, laptop computer, tablet computer, smartphone, wearable device, television, camera, video game console, set-top box, video streaming device, etc.). Encoder 200 may include one or more of the following: inter-frame prediction unit 206, intra-frame prediction unit 208, combiners 210 and 212, transform and quantization unit (TR+Q) 214, inverse transform and quantization unit (iTR+iQ) 216, entropy coding unit 218, one or more filters 220, and / or buffers 222.

[0049] Encoder 200 can divide (e.g., include) images (e.g., frames) of video sequence 202 into blocks and encode video sequence 202 on a block-by-block basis. Encoder 200 can perform / apply prediction techniques on the blocks being encoded using inter-frame prediction unit 206 or intra-frame prediction unit 208. Inter-frame prediction unit 206 can perform inter-frame prediction by searching for blocks similar to the blocks being encoded in another reconstructed image (e.g., a reference image) of video sequence 202. A reconstructed image refers to an image that is encoded and then decoded. The blocks being encoded can then be predicted using blocks determined during the search (e.g., referred to as prediction blocks) to remove redundant information. Inter-frame prediction unit 206 can utilize temporal redundancy or similarity in the scene content of each image in video sequence 202 to determine prediction blocks. For example, the scene content between images of video sequence 202 can be similar, except for differences due to motion and / or affine transformations of screen content over time.

[0050] Intra-prediction unit 208 can perform intra-prediction by forming prediction blocks based on data from reconstructed neighboring samples of the block to be encoded within the same image from video sequence 202. Reconstructed samples refer to samples that are encoded and then decoded. Intra-prediction unit 208 can utilize spatial redundancy or similarity in scene content within images of video sequence 202 to determine prediction blocks. For example, the texture of a region of scene content in an image may be similar to the texture of the immediately surrounding region of scene content in the same image.

[0051] Combiner 210 can determine the prediction error (e.g., referred to as the residual) based on the difference between the block being encoded and the predicted block. The prediction error can represent non-redundant information that can be sent / transmitted to the decoder for accurate decoding of the video sequence 202.

[0052] Transform and quantization unit (TR + Q) 214 can transform and quantize the prediction error. Transform and quantization unit 214 can reduce relevant information in the prediction error by applying, for example, DCT, thereby transforming the prediction error into transform coefficients. Transform and quantization unit 214 can quantize the coefficients by mapping the data of the transform coefficients to a set of predefined representative values. Transform and quantization unit 214 can quantize the coefficients to reduce irrelevant information in bitstream 204. Irrelevant information refers to information that can be removed from the coefficients without producing visible and / or perceptible distortion in the video sequence 202 after decoding (e.g., at the receiving device).

[0053] Entropy coding unit 218 can apply one or more entropy coding methods to the quantized transform coefficients to further reduce the bit rate. For example, entropy coding unit 218 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and / or syntax-based context-based binary arithmetic coding (SBAC). The entropy-coded coefficients can be packed to form bit stream 204.

[0054] The inverse transform and quantization unit (iTR + iQ) 216 can perform inverse quantization and inverse transform on the quantized transform coefficients to determine the reconstruction prediction error. The combiner 212 can combine the reconstruction prediction error with the prediction block to form a reconstruction block. The filter 220 can filter the reconstruction block, for example, using a deblocking filter and / or a sample adaptive offset (SAO) filter. The buffer 222 can store the reconstruction block for prediction of one or more other blocks in the same and / or different images of the video sequence 202.

[0055] The encoder 200 may also include an encoder control unit. The encoder control unit can be configured to control, for example... Figure 2 The encoder 200 shown includes one or more units. The encoder control unit can control one or more units of the encoder 200 to generate bitstream 204 according to one or more proprietary coding protocols, industry video coding standards, and / or any other video coding protocols. For example, the encoder control unit can control one or more units of the encoder 200 to generate bitstream 204 according to one or more of ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, ​​AV1, and / or any other video coding standards / formats.

[0056] The encoder control unit can be configured to attempt to minimize (or reduce) the bit rate of bitstream 204 and / or (e.g., within the constraints of proprietary coding protocols, industry video coding standards, and / or any other video coding protocols) maximize (or improve) the reconstructed video quality. For example, the encoder control unit can be configured to attempt to minimize or reduce the bit rate of bitstream 204 such that the reconstructed video quality is not lower than a certain level / threshold, and / or maximize or improve the reconstructed video quality such that the bit rate of bitstream 204 does not exceed a certain level / threshold. The encoder control unit can determine / control one or more of the following: dividing the images of video sequence 202 into blocks, whether the blocks are inter-frame predicted by inter-frame prediction unit 206 or intra-frame predicted by intra-frame prediction unit 208, motion vectors for inter-frame prediction of blocks, intra-frame prediction mode among multiple intra-frame prediction modes for intra-frame prediction of blocks, filtering performed by filter 220, and / or one or more transform types and / or quantization parameters applied by transform and quantization unit 214. The encoder control unit can determine / control one or more of the above-mentioned factors based on the rate-distortion metric for the block or image being encoded. The encoder control unit can determine / control one or more of the above-mentioned factors to reduce the rate-distortion metric for the block or image being encoded.

[0057] The prediction type (intra-frame or inter-frame prediction) used to encode the block, the block's prediction information (intra-frame prediction mode (if intra-frame), motion vectors, etc.), and / or transform parameters and / or quantization parameters can be sent to entropy coding unit 218 for further compression (e.g., to reduce the bit rate). For example, entropy coding unit 218 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and syntax-based context-based binary arithmetic coding (SBAC) to achieve further compression. The prediction type, prediction information, and / or transform and / or quantization parameters can be packaged together with the prediction error to form bitstream 204.

[0058] Encoder 200 is merely an example, and encoders and / or modified versions of encoder 200, different from encoder 200, may perform the methods and processes described herein. For example, encoder 200 may include other components and / or arrangements. Figure 2 One or more of the components shown may optionally be included in encoder 200 (e.g., entropy coding unit 218 and / or filter 220).

[0059] Figure 3 An example decoder is shown. (e.g.) Figure 3The decoder 300 shown can implement one or more of the processes described herein. Decoder 300 can decode bitstream 302 into a decoded video sequence 304 for display and / or some other form of consumption. Figure 1 The video encoding / decoding system 100 is implemented in and / or in computing, communication, or electronic devices (e.g., desktop computers, laptop computers, tablet computers, smartphones, wearable devices, televisions, cameras, video game consoles, set-top boxes, and / or video streaming devices). The decoder 300 may include an entropy decoding unit 306, an inverse transform and quantization (iTR+iQ) unit 308, a combiner 310, one or more filters 312, a buffer 314, an inter-frame prediction unit 316, and / or an intra-frame prediction unit 318.

[0060] Decoder 300 may include a decoder control unit configured to control one or more units of decoder 300. The decoder control unit can control one or more units of decoder 300 to decode bitstream 302 according to the requirements of one or more proprietary coding protocols, industry video coding standards, and / or any other communication protocols. For example, the decoder control unit can control one or more units of decoder 300 to decode bitstream 302 according to one or more of ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, ​​AV1, and / or any other video coding standards / formats.

[0061] The decoder control unit can determine / control one or more of the following: whether the block is inter-predicted by the inter-frame prediction unit 316 or intra-frame predicted by the intra-frame prediction unit 318; the motion vector for the inter-frame prediction of the block; the intra-frame prediction mode among multiple intra-frame prediction modes for the intra-frame prediction of the block; the filtering performed by the filter 312; and / or one or more inverse transform types and / or inverse quantization parameters to be applied by the inverse transform and quantization unit 308. One or more of the control parameters used by the decoder control unit can be packed in the bit stream 302.

[0062] Entropy decoding unit 306 can perform entropy decoding on bitstream 302. For example, entropy decoding unit 306 can apply context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), and syntax-based context-based binary arithmetic coding (SBAC) to decompress the prediction type (intra-frame or inter-frame prediction), prediction information of the block (intra-frame prediction mode, motion vectors, etc. in the case of intra-frame prediction), and transform and quantization parameters used to encode the block. Inverse transform and quantization unit 308 can perform inverse quantization and / or inverse transform on the quantized transform coefficients to determine the decoding prediction error. Combiner 310 can combine the decoding prediction error with the prediction block to form a decoded block. The prediction block can be generated by intra-frame prediction unit 318 or inter-frame prediction unit 316 (e.g., as described above regarding...). Figure 2 (As described in encoder 200). Filter 312 can filter the decoded blocks, for example, using a deblocking filter and / or a sample adaptive offset (SAO) filter. Buffer 314 can store the decoded blocks for prediction of one or more other blocks in the same and / or different images of the video sequence in bitstream 302. The decoded video sequence 304 can be output from filter 312, as... Figure 3 As shown.

[0063] Decoder 300 is merely an example, and decoders and / or modified versions of decoder 300, different from decoder 300, may perform the methods and processes described herein. For example, decoder 300 may have other components and / or arrangements. Figure 3 One or more of the components shown may optionally be included in the decoder 300 (e.g., entropy decoding unit 306 and / or filter 312).

[0064] Although Figure 2 and Figure 3 Not shown in the diagram, but in addition to the inter-frame prediction unit and the intra-frame prediction unit, each of the encoder 200 and decoder 300 may also include an intra-frame block copying unit. The intra-frame block copying unit may perform / operate similarly to the inter-frame prediction unit, but may predict blocks within the same image. For example, the intra-frame block copying unit may utilize repeating patterns appearing in the screen content. The screen content may include computer-generated text, graphics, animations, etc.

[0065] Video encoding and / or decoding can be performed on a block-by-block basis. The process of dividing an image into blocks based on its content can be adaptive. For example, larger block partitions can be used in image regions with a high level of homogeneity to improve encoding efficiency.

[0066] Images (e.g., in HEVC or any other coding standard / format) can be divided into non-overlapping square blocks, which can be called coding tree blocks (CTBs). CTBs can include samples of a sample array. A CTB can have a size of 2nx2n samples, where n can be specified by parameters of the coding system. For example, n can be 4, 5, 6, or any other value. A CTB can have any other size. A CTB can be further divided into coding blocks (CBs) of semi-vertical and semi-horizontal sizes via recursive quadtree partitioning. A CTB can form the root of a quadtree. CBs not further partitioned as part of a recursive quadtree partition can be called leaf CBs of the quadtree; otherwise, they can be called non-leaf CBs of the quadtree. A CB can have a minimum size specified by parameters of the coding system. For example, a CB can have a minimum size of 4×4, 8×8, 16×16, 32×32, 64×64 samples, or any other minimum size. A CB can be further divided into one or more prediction blocks (PBs) for performing inter-frame prediction and / or intra-frame prediction. A PB can be a rectangular sample block to which the same prediction type / mode can be applied. For a transformation, the CB can be divided into one or more transform blocks (TBs). A TB can be a rectangular block of samples that can determine / indicate the size of the applied transformation.

[0067] Figure 4 An example quadtree partition of CTB 400 is shown. Figure 5 It shows the relationship with Figure 4 The example quadtree partitioning of CTB 400 corresponds to the example quadtree 500. For example... Figure 4 and Figure 5 As shown in the example, CTB 400 can first be divided into four leaf CBs of semi-vertical and semi-horizontal size. Three of the leaf CBs resulting from the first-level division of CTB 400 are leaf CBs. The three leaf CBs of the first-level division of CTB 400 are... Figure 4 and Figure 5 The sub-CBs are labeled 7, 8, and 9 respectively. The non-leaf CBs in the first-level partition of CTB 400 are divided into four sub-CBs of semi-vertical and semi-horizontal size. Three of the sub-CBs obtained from the second-level partition of CTB 400 are leaf CBs. The three leaf CBs in the second-level partition of CTB 400 are... Figure 4 and Figure 5 The numbers were labeled 0, 5, and 6 respectively. Finally, the non-leaf CBs of the second-level division of CTB400 were divided into four leaf CBs of semi-vertical and semi-horizontal size. The four leaf CBs were... Figure 4 and Figure 5 They are labeled as 1, 2, 3 and 4 respectively.

[0068] Figure 4The example CTB 400 is divided into 10 leaf CBs labeled 0 to 9, but it can be divided into other numbers of leaf CBs. The 10 leaf CBs can be combined with 10 CB leaf nodes (e.g., as shown in the example). Figure 5 The quadtree shown corresponds to the 10 leaf nodes of the CB (Central Branch Tree) in the quadtree 500. In other examples, the CTB can be divided into different numbers of leaf CBs. The resulting quadtree partition of the CTB 400 can be scanned using a z-scan (e.g., from left to right, from top to bottom) to form a sequence order for encoding / decoding the leaf nodes of the CBs. Figure 4 and Figure 5 The numerical label (e.g., indicator, index) of each CB leaf node can correspond to the sequence order used for encoding / decoding. For example, CB leaf node 0 can be encoded / decoded first, and CB leaf node 9 can be encoded / decoded last. Although not in Figure 4 and Figure 5 As shown in the figure, each CB leaf node may include one or more PBs and / or TBs.

[0069] Images in VVC (or any other encoding standard / format) can be partitioned in a similar manner (such as HEVC). The image can first be partitioned into non-overlapping square CTBs. Then, a recursive quadtree partitioning method can be used to divide the CTB into CBs of semi-vertical and semi-horizontal sizes. Quadtree leaf nodes (e.g., in VVC) can be further partitioned into CBs of unequal sizes using binary or ternary tree partitioning (or any other partitioning method).

[0070] Figure 6 Example binary tree partitioning and ternary tree partitioning are shown. A binary tree partition can divide the parent block in half along the vertical direction 602 or the horizontal direction 604. The resulting partition may be half the size of the parent block. In other examples, the resulting partition may correspond to a size less than and / or greater than half the size of the parent block. A ternary tree partition can divide the parent block into three parts along the vertical direction 606 or the horizontal direction 608. Figure 6 An example is shown where the middle partition can be twice the size of the other two terminal partitions in a ternary tree partition. In other examples, the partitions can be different sizes relative to each other and relative to the parent block. Binary tree partitions and ternary tree partitions are examples of multi-type tree partitions. Multi-type tree partitions can include dividing the parent block into an additional number of smaller blocks. (For example, in VVC) A block partitioning strategy can be called a combination of quadtree partitioning and multi-type tree partitioning (quadritree + multi-type tree partitioning) because it adds binary tree partitioning and / or ternary tree partitioning to the quadtree partitioning.

[0071] Figure 7 Examples of combined quadtree partitioning and multi-type tree partitioning of the CTB 700 are shown. Figure 8 It shows the relationship with Figure 7 The example tree 800 shows the combination of quadtree partitioning and multi-type tree partitioning of CTB 700. Figure 7 and Figure 8 In the diagram, quadtree partitions are shown with solid lines, and multi-type tree partitions are shown with dashed lines. For ease of explanation, the CTB 700 is shown as having the same characteristics as... Figure 4 The quadtree partitioning described herein is the same as that of CTB 400, and descriptions of quadtree partitioning in CTB 700, which is similar to that of CTB 400, are omitted. The quadtree partitioning of CTB 700 is merely an example, and CTB can perform quadtree partitioning in ways different from CTB 700. Additional multi-type tree partitioning of CTB 700 can be performed relative to... Figure 4 The three leaves CB shown are used for this purpose. Figure 4 In Figure 7 The three leaf CBs shown as being further divided can be leaves CB 5, 8, and 9. The three leaf CBs can be further divided using one or more binary tree partitions and / or ternary tree partitions.

[0072] Figure 4 Leaf CB 5 can be partitioned into two CBs based on a vertical binary tree partition. The two resulting CBs can be in... Figure 7 and Figure 8 Leaves CB are marked as 5 and 6 respectively. Figure 4 Leaf CB 8 can be partitioned into three CBs based on a vertical ternary tree partition. Two of the three resulting CBs can be from... Figure 7 and Figure 8 The leaf branches (CBs) are labeled 9 and 14 respectively. The remaining non-leaf CBs can be first partitioned into two CBs based on horizontal binary tree partitioning. One of the two CBs can be the leaf CB labeled 10. The other CB can be further partitioned into three CBs based on vertical ternary tree partitioning. The resulting three CBs can be... Figure 7 and Figure 8 Leaves CB are labeled 11, 12 and 13 respectively. Figure 4 Leaf CB 9 can be partitioned into three CBs based on a horizontal ternary tree partition. Two of the three CBs can be in... Figure 7 and Figure 8 The leaf CBs are labeled 15 and 19 respectively. The remaining non-leaf CBs can be partitioned into three CBs based on another horizontal ternary tree partition. All three resulting CBs can be... Figure 7 and Figure 8 Leaves CB are marked as 16, 17 and 18 respectively.

[0073] In summary, the CTB 700 can be divided into 20 leaf CBs labeled 0-19. These 20 leaf CBs can be associated with 20 leaf nodes (e.g., ...). Figure 8 The 20 leaf nodes of the tree 800 shown correspond to this. A z-scan (from left to right, from top to bottom) can be used to scan the resulting combination of quadtree partitions and multi-type tree partitions of the CTB 700 to form a sequence order for encoding / decoding the CB leaf nodes. Figure 7 and Figure 8 The numerical label of each CB leaf node can correspond to the sequence order used for encoding / decoding, where CB leaf node 0 is encoded / decoded first, and CB leaf node 19 is encoded / decoded last. Although Figure 7 and Figure 8 Not shown, but it should be noted that each CB leaf node may include one or more PBs and / or TBs.

[0074] Encoding standards / formats (e.g., HEVC, VVC, or any other encoding standard / format) can define various units (e.g., in addition to specifying various blocks (e.g., CTB, CB, PB, TB)). A block can include a rectangular area of ​​samples in a sample array. A unit can include juxtaposed sample blocks from different sample arrays forming an image (e.g., a luma sample array and a chroma sample array), as well as the block's syntax elements and prediction data. A coding tree unit (CTU) can include juxtaposed CTBs of different sample arrays and can form a complete entity in the encoded bitstream. A coding unit (CU) can include juxtaposed CBs of different sample arrays and a syntax structure for encoding samples of the CBs. A prediction unit (PU) can include juxtaposed PBs of different sample arrays and syntax elements for predicting the PBs. A transform unit (TU) can include TBs of different sample arrays and syntax elements for transforming the TBs.

[0075] A block can refer to any of CTB, CB, PB, TB, CTU, CU, PU, ​​and / or TU (e.g., in the context of HEVC, VVC, or any other encoding format / standard). In the context of any video encoding format / standard / protocol, a block can be used to refer to a similar data structure. For example, a block can refer to a macroblock in the AVC standard, a macroblock or subblock in VP8 encoding format, a superblock or subblock in VP9 encoding format, and / or a superblock or subblock in AV1 encoding format.

[0076] In intra-frame prediction, samples of the block to be encoded (e.g., also called the current block) can be predicted based on samples from the leftmost column immediately adjacent to the current block and samples from the topmost row immediately adjacent to the current block. Samples from the adjacent columns and rows can be collectively referred to as reference samples. Each sample of the current block can be predicted (e.g., in intra-frame prediction mode) by projecting the location of the sample in the current block in a given direction onto a point along the reference samples. If the projection does not fall directly on a reference sample, the sample can be predicted by interpolating between the two nearest reference samples of the projection point. The prediction error (e.g., called the residual) of the current block can be determined based on the difference between the predicted sample values ​​and the original sample values.

[0077] Predicted samples can be performed for multiple different intra-prediction modes (e.g., including non-directional intra-prediction modes) (e.g., at the encoder), and prediction errors can be determined based on the differences between the predicted samples and the original samples. The encoder can select one of the multiple intra-prediction modes and its corresponding prediction error to encode the current block. The encoder can send an indication of the selected prediction mode and its corresponding prediction error to the decoder to decode the current block. The decoder can decode the current block by predicting samples of the current block using the intra-prediction mode indicated by the encoder and / or by combining the predicted samples with the prediction error.

[0078] Figure 9 An example set of reference samples 902 for intra-frame prediction determination of the current block 904 is shown. The current block 904 may correspond to a block being encoded and / or decoded. The current block 904 may correspond to, for example, Figure 7 This corresponds to block 3 of the CTB 700 partition shown. As mentioned above, the numerical labels 0-19 of the partitioned CTB 700 blocks can correspond to the sequence order of encoding / decoding the blocks, and can be... Figure 9 This is how it is used in the examples.

[0079] For a current block 904 of size w × h samples, reference samples 902 may include: 2w samples (or any other number of samples) from the top row immediately adjacent to the current block 904, 2h samples (or any other number of samples) from the leftmost column immediately adjacent to the current block 904, and the sample adjacent to the top-left corner of the current block 904. The current block 904 may be square, such that w = h = s. In other examples, the current block need not be square, such that w ≠ h. Available samples from neighboring blocks of the current block 904 can be used to construct the set of reference samples 902. For example, if a sample is located outside the image of the current block, is part of a different slice of the current block (e.g., if the concept of slices is used), and / or belongs to a block that has already been inter-coded and indicates constrained intra-prediction, then the sample may not be available to construct the set of reference samples 902. For example, if constrained intra-prediction is indicated, intra-prediction may not depend on inter-prediction blocks.

[0080] Samples that may not be suitable for constructing the set of reference samples 902 may include samples within a block that have not yet been encoded and reconstructed at the encoder and / or decoded at the decoder based on the sequence order used for encoding / decoding. Restricting the inclusion of such samples in the set of reference samples 902 allows for the determination of the same prediction at both the encoder and decoder. Figure 9 In the example, samples from neighboring blocks 0, 1, and 2 can be used to construct reference sample 902, assuming these blocks are encoded and reconstructed at the encoder and decoded at the decoder before the current block 904 is encoded. For example, if no other issues (e.g., as mentioned above) prevent the availability of samples from neighboring blocks 0, 1, and 2, then samples from neighboring blocks 0, 1, and 2 can be used to construct reference sample 902. Due to the sequence order used for encoding / decoding, a portion of reference sample 902 from neighboring block 6 may be unavailable (e.g., because block 6 may not have been encoded and reconstructed at the encoder and / or decoded at the decoder based on the sequence order used for encoding / decoding).

[0081] In some examples, unavailable samples from reference sample 902 can be filled using one or more available reference samples 902. For example, an unavailable reference sample can be filled using the nearest available reference sample. The nearest available reference sample can be determined by moving clockwise through reference sample 902 from the location of the unavailable reference. For example, if no reference sample is available, reference sample 902 can be filled using the median value of the dynamic range of the image being encoded.

[0082] The reference sample 902 can be filtered based on the size of the current block 904 being encoded and the applied intra-prediction mode. Figure 9An example of determining reference samples for intra-frame prediction of a block is shown. Reference samples can be determined in ways different from those described above. For example, in other cases (such as in VVC), multiple reference lines may be used.

[0083] Intra-frame prediction of samples in the current block 904 can be performed based on reference sample 902 (e.g., determination and (optionally) filtering based on reference sample 902, e.g., after such determination and filtering). At least some (e.g., most) encoders / decoders can support multiple intra-frame prediction modes according to one or more video coding standards. For example, HEVC supports 35 intra-frame prediction modes, including planar mode, DC mode, and 33 angular modes. VVC supports 67 intra-frame prediction modes, including planar mode, DC mode, and 65 angular modes. Planar mode and DC mode can be used to predict smooth regions and gradually changing regions of an image. Angular mode can be used to predict directional structures in regions of an image. Any number of intra-frame prediction modes can be supported.

[0084] Figure 10A and Figure 10B An example intra-frame prediction mode is shown. Figure 10A The diagram illustrates 35 intra-frame prediction modes supported by HEVC. These 35 intra-frame prediction modes can be indicated / identified by indices 0 to 34. Prediction mode 0 can correspond to a planar mode. Prediction mode 1 can correspond to a DC mode. Prediction modes 2-34 can correspond to angular modes. Prediction modes 2-18 can be referred to as horizontal prediction modes because the primary source of prediction is in the horizontal direction. Prediction modes 19-34 can be referred to as vertical prediction modes because the primary source of prediction is in the vertical direction.

[0085] Figure 10B The diagram illustrates 67 intra-frame prediction modes supported by VVC. These 67 intra-frame prediction modes can be indicated / identified by indices 0 to 66. Prediction mode 0 can correspond to a planar mode. Prediction mode 1 corresponds to a DC mode. Prediction modes 2-66 can correspond to angular modes. Prediction modes 2-34 can be referred to as horizontal prediction modes because the primary source of prediction is in the horizontal direction. Prediction modes 35-66 can be referred to as vertical prediction modes because the primary source of prediction is in the vertical direction. Figure 10B Some of the intra prediction modes described in the document can be adaptively replaced by wide-angle orientation because blocks in VVC do not need to be square.

[0086] Figure 11 Showing from Figure 9 The current block 904 and the corresponding reference sample 902. To further describe how the intra-frame prediction mode is applied to determine the prediction (e.g., the predicted block) of the current block 904, Figure 11The current block 904 and reference sample 902 are shown in a two-dimensional x, y plane, where the sample can be referred to as To simplify the prediction process, reference sample 902 can be placed in two one-dimensional arrays. Reference sample 902 above the current block 904 can be placed in a one-dimensional array. middle:

[0087] The reference sample 902 to the left of the current block 904 can be placed in a one-dimensional array. middle:

[0088] The prediction process may include determining the position in the current block 904. Predicted samples at the location (e.g., predicted value). For planar patterns, the position in the current block 904 can be predicted by determining / calculating the average of two interpolated values. The sample at that location. The first of the two interpolated values ​​can be based on the sample at location 904 in the current block. Horizontal linear interpolation at position 904. The second interpolated value of the two interpolated values ​​can be based on the value at position 904 in the current block. Vertical linear interpolation at the current location. Predicted samples in block 904. It can be determined / calculated as follows:

[0089] in

[0090] It could be in the current block 904 at position Horizontal linear interpolation at the location, and

[0091] It can be the position in the current block 904. Vertical linear interpolation at the location. It can be equal to the length of the edge of the current block 904 (e.g., the number of samples on the edge).

[0092] For DC mode, the position in the current block 904 can be predicted by referring to the average value of sample 902. The sample at this location. The predicted sample in the current block 904. It can be determined / calculated as follows:

[0093] For angle patterns, the position can be determined by positioning the position in the direction specified by the given angle pattern. Projecting points onto the horizontal or vertical lines of a sample including reference sample 902 to predict the position in the current block 904. The sample at the location. If the projection does not fall directly on the reference sample, the location at the projection point can be predicted by interpolation between the two nearest reference samples. The direction specified by the angle mode can be given by an angle φ relative to the y-axis used for vertical prediction modes (e.g., modes 19-34 in HEVC and modes 35-66 in VVC). The direction specified by the angle mode can be given by an angle φ relative to the x-axis used for horizontal prediction modes (e.g., modes 2-18 in HEVC and modes 2-34 in VVC).

[0094] Figure 12 An example of applying an intra-frame prediction mode (e.g., an angle mode, such as vertical prediction mode 906) to predict the current block 904 is shown. Figure 12 Specifically, the position within the current block 904 for vertical prediction mode 906 is shown. The prediction of the sample at that location. Vertical prediction mode 906 can be given by the angle φ relative to the vertical axis. In vertical prediction mode, the position in the current block 904... It can be projected onto the reference sample Points on the horizontal line (e.g., referred to as projection points). For ease of illustration, reference sample 902 is... Figure 12 Only a portion is shown. For example... Figure 12 As shown in the figure, the reference sample The projection point on the horizontal line may not lie exactly on the reference sample. For example, if the projection point falls at the fractional sample location between two reference samples, the predicted sample in the current block 904 can be determined / calculated by performing linear interpolation between the two reference samples. Predicted Samples It can be determined / calculated as:

[0095] It can be the projection point relative to the position. The integer part of the horizontal displacement. It can be determined / calculated as a function of the tangent of the angle φ in the vertical prediction pattern 906:

[0096] It can be the projection point relative to the position. The fractional part of the horizontal displacement can be determined / calculated as:

[0097] in It is a floor function.

[0098] For the horizontal prediction mode, the position of the sample in the current block 904 It can be projected onto the reference sample On the vertical line. Prediction samples used for horizontal prediction patterns. It can be determined / calculated as:

[0099] It can be the projection point relative to the position. The integer part of the vertical displacement. It can be determined / calculated as a function of the tangent of the angle φ in the horizontal prediction model:

[0100] It can be the projection point relative to the position. The fractional part of the vertical displacement. It can be determined / calculated as:

[0101] in It is a floor function.

[0102] The interpolation functions given by equations (7) and (10) can be derived from the encoder and / or decoder (e.g., Figure 2 encoder 200 and / or Figure 3 The decoder (300) in the code is used for implementation. The interpolation function can be implemented using a finite impulse response (FIR) filter. For example, the interpolation function can be implemented as a set of two-tap FIR filters. The coefficients of the two-tap FIR filters can be obtained from... and Given that, in intra-frame angular prediction, the predicted samples can be calculated using a predefined level of sample accuracy (e.g., 1 / 32 sample accuracy or accuracy defined by any other metric). For 1 / 32 sample accuracy, the set of two-tap FIR interpolation filters can include up to 32 different two-tap FIR interpolation filters—used for projected shift. Each of the 32 possible values ​​for the decimal part. In other examples, different levels of sample accuracy may be used.

[0103] In some examples, FIR filters can be used to predict chroma samples and / or luminance samples. For instance, a two-tap interpolation FIR filter can be used to predict chroma samples, and the same and / or different interpolation techniques / filters can be used for luminance samples. For example, a four-tap FIR filter can be used to determine the predicted value for a luminance sample. This can be based on... To determine the coefficients of the four-tap FIR filter (similar to a two-tap FIR filter), the set of 32 different four-tap FIR filters can be used for 1 / 32 sample accuracy. Each of the 32 possible values ​​for the fractional part. In other examples, different levels of sample accuracy can be used. The set of four-tap FIR filters can be stored in a lookup table (LUT) and based on... For reference. For vertical prediction patterns, prediction samples. It can be determined based on a four-tap FIR filter as follows:

[0104] in These can be filter coefficients, and It is an integer displacement. For horizontal prediction modes, the predicted samples... It can be determined based on a four-tap FIR filter as follows:

[0105] If we want to predict the location of the sample in the current block 904 Projected onto a negative x-coordinate, supplementary reference samples can be determined / constructed. For example, if a negative vertical prediction angle φ is used, the position of the sample... It can be projected onto a negative x-coordinate. The vertical line of reference sample 902 can be projected using a negative vertical prediction angle φ. The reference sample in the middle is projected onto the horizontal line of reference sample 902 to determine / construct supplementary reference samples. For example, if the position of the sample in the current block 904 is to be predicted... Projected onto a negative y-coordinate, supplementary reference samples can be similarly determined / constructed. For example, if a negative horizontal prediction angle φ is used, the position of the sample... It can be projected onto a negative y-coordinate. The reference sample 902 can be projected onto the horizontal line by using a negative horizontal prediction angle φ. The reference sample in the middle is projected onto the vertical line of reference sample 902 to determine / construct the supplementary reference sample.

[0106] The encoder can (e.g., using one or more of the functions described herein) determine / predict samples of the current block being encoded (e.g., current block 904) against multiple intra prediction modes. For example, the encoder can determine / predict samples of the current block against each of the 35 intra prediction modes in HEVC and / or the 67 intra prediction modes in VVC. For each applied intra prediction mode, the encoder can determine the corresponding prediction error for the current block based on the difference between the predicted samples determined for the intra prediction mode and the original samples of the current block (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), or sum of absolute transform differences (SATD)). The encoder can determine / select one of the intra prediction modes to encode the current block based on the determined prediction error. For example, the encoder can determine / select one of the intra prediction modes that produces the minimum prediction error for the current block. In some examples, the encoder can determine / select the intra prediction mode to encode the current block based on a rate-distortion metric (e.g., Lagrangian rate-distortion cost) determined using the prediction error. The encoder can send an indication of the determined / selected intra-prediction mode and its corresponding prediction error (e.g., residual) to the decoder to decode the current block.

[0107] The decoder can determine / predict samples of the current block being decoded (e.g., current block 904) for an intra-prediction mode. For example, the decoder can receive an indication of the intra-prediction mode (e.g., an angular intra-prediction mode) from the encoder used for the current block. The decoder can construct a reference sample set and perform intra-prediction based on the intra-prediction mode indicated by the encoder used for the current block in a manner similar to that described above for the encoder. The decoder can add the predicted values ​​of the samples of the current block (e.g., determined based on the intra-prediction mode) to the residual of the current block to reconstruct the current block. In some examples, the decoder does not need to receive an indication of the angular intra-prediction mode from the encoder used for the current block. Instead, the decoder can determine the intra-prediction mode through other decoder-side means.

[0108] While the various examples in this paper correspond to intra-prediction modes in HEVC and VVC, the methods, devices, and systems described herein can be applied to / used for other intra-prediction modes (e.g., as used in other video coding standards / formats such as VP8, VP9, ​​AV1, etc.).

[0109] Intra-frame prediction can perform video compression by leveraging the correlation between spatially adjacent samples in the same frame of a video sequence. Inter-frame prediction is another coding tool that can be used to perform video compression. Inter-frame prediction can leverage the temporal correlation between sample blocks in different frames of a video sequence. For example, an object can be seen in multiple frames of a video sequence. The object may move (e.g., by some kind of translation and / or affine motion) or remain stationary in the multiple frames. The current sample block in the current frame being encoded may have a corresponding sample block in a previously decoded frame or be associated with that corresponding sample block. The corresponding sample block can accurately predict the current sample block. For example, the corresponding sample block may be shifted from the current sample block due to the movement of the object represented in the corresponding frames of those frames. The previously decoded frame can be a reference frame. The corresponding sample block in the reference frame can be a reference block for motion compensation prediction. The encoder can use block matching techniques to estimate the displacement (or motion) of the object and / or determine the reference block in the reference frame.

[0110] Similar to intra-frame prediction, the encoder can determine the difference between the current block and the prediction for the current block. The encoder can determine / generate the prediction for the current block, for example, based on (e.g., using inter-frame prediction), or determine the difference after determining / generating that prediction. This difference may be a prediction error (e.g., a residual). The encoder can store and / or transmit (e.g., signal) the prediction error and / or other relevant prediction information in / via the bitstream. The prediction error and / or other relevant prediction information can be used for decoding and / or other forms of overhead. The decoder can (e.g., by using relevant prediction information) predict samples of the current block and combine the predicted samples with the prediction error to decode the current block.

[0111] Figure 13A An example of inter-frame prediction is shown. Inter-frame prediction can be performed for the current block 1300 in the current image 1302 being encoded. The encoder (e.g., such as...) Figure 2The encoder 200 shown can perform inter-frame prediction to determine and / or generate a reference block 1304 in a reference picture 1306. Reference block 1304 can be used to predict the current block 1300. The reference picture (e.g., reference picture 1306) can be a previously decoded picture available at the encoder and / or decoder. The availability of the previously decoded picture can depend on / based on whether the previously decoded picture is available in the decoded picture buffer when the current block 1300 is encoded and / or decoded. The encoder can search for blocks (e.g., candidate reference blocks) similar to (or substantially similar to) the current block 1300 in one or more reference pictures 1306. The encoder can determine the best-matching block from the blocks (e.g., candidate reference blocks) tested during the search process. The best-matching block can be reference block 1304. The encoder can determine that reference block 1304 is the best-matching reference block based on one or more cost criteria. One or more cost criteria can include rate-distortion criteria (e.g., Lagrange rate-distortion cost). One or more cost criteria may be based on the difference between the predicted sample of reference block 1304 and the original sample of the current block 1300 (e.g., SSD, SAD, and / or SATD).

[0112] The encoder can search for reference block 1304 within a reference region (e.g., search range 1308). The reference region (e.g., search range 1308) can be located around a juxtaposed block (or location) 1310 of the current block 1300 in reference image 1306. Juxtaposed block 1310 can have the same location in reference image 1306 as the current block 1300 in current image 1302. The reference region (e.g., search range 1308) can extend at least partially beyond reference image 1306. For example, if the reference region (e.g., search range 1308) extends beyond reference image 1306, constant boundary extension can be used. Constant boundary extension can be used such that the values ​​of samples in rows or columns of reference image 1306 immediately adjacent to a portion of the reference region (e.g., search range 1308) extending beyond reference image 1306 are available at sample locations outside reference image 1306. Reference block 1304 can be searched within a subset or all of the potential locations within the reference region (e.g., search range 1308). The encoder may use one or more search implementations to determine and / or generate reference block 1304. For example, the encoder may determine a set of candidate search locations based on motion information (e.g., motion vector 1312) of the neighboring blocks of the current block 1300.

[0113] The encoder may search for one or more reference images during inter-frame prediction to determine and / or generate the best-matching reference block. The reference images searched by the encoder may be included in one or more reference image lists (e.g., added to those lists). For example, in HEVC and VVC (and / or in one or more other communication protocols), two reference image lists (e.g., reference image list 0 and reference image list 1) may be used. A reference image list may include one or more images. Reference image 1306 of reference block 1304 may be indicated by a reference index pointing to the reference image list that includes reference image 1306.

[0114] Figure 13B An example motion vector is shown. The displacement between reference block 1304 and current block 1300 can be interpreted as an estimate of the motion between reference block 1304 and current block 1300 in their respective images. This displacement can be represented by motion vector 1312. For example, motion vector 1312 can be indicated by a horizontal component (MVx) and a vertical component (Mvy) relative to the position of current block 1300. Motion vectors (e.g., motion vector 1312) can have fractional or integer resolution. Motion vectors with fractional resolution can be pointed between two samples in the reference image to provide a better estimate of the motion of current block 1300. For example, motion vectors can have a fractional sample resolution of 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, or any other fractional sample resolution. Interpolation between two samples at integer locations can be used to generate a reference block and a corresponding sample at its fractional location, for example, when the motion vector points to a non-integer sample value in the reference image. Interpolation can be performed by a filter with two or more taps.

[0115] The encoder can determine the difference (e.g., the corresponding sample-by-sample difference) between reference block 1304 and current block 1300. The encoder can determine the difference between reference block 1304 and current block 1300, for example, based on determining and / or generating reference block 1304 for current block 1300 using inter-frame prediction, or after determining and / or generating the reference block. This difference may be a prediction error (e.g., a residual). The encoder can store and / or transmit (e.g., signal) the prediction error and / or associated motion information in / via the bit stream. The prediction error and / or associated motion information can be used for decoding (e.g., decoding current block 1300) and / or other forms of consumption. Motion information may include motion vector 1312 and a reference indicator / index. The reference indicator may indicate reference image 1306 in a list of reference images. In other examples, motion information may include an indication of motion vector 1312 and / or an indication of a reference indicator / index. The reference indicator may indicate reference image 1306 in a list of reference images that includes reference image 1306. The decoder can decode the current block 1300 by determining and / or generating a reference block 1304, which can correspond to / form (e.g., be considered) the prediction of the current block 1300. The decoder can determine and / or generate the reference block 1304, for example, based on relevant motion information. The decoder can decode the current block 1300 by combining the prediction (e.g., the reference block) with the prediction error (e.g., the residual block).

[0116] like Figure 13A As shown, inter-frame prediction can be performed using a reference image 1306 as a prediction source for the current block 1300. Inter-frame prediction based on predictions for the current block using a single image can be referred to as unidirectional prediction.

[0117] Inter-frame prediction of the current block using bidirectional prediction can be based on two images (e.g., the prediction source can come from two images). For example, bidirectional prediction can be useful if the video sequence includes fast motion, camera panning, zooming, and / or scene changes. Bidirectional prediction can be used to capture a fade-out of a scene or a fade-out from one scene to another, where the two images can be effectively displayed simultaneously at different intensity levels.

[0118] One or both of unidirectional and bidirectional prediction are available for / can be used (e.g., at the encoder and / or at the decoder) to perform inter-frame prediction. The specific type of inter-frame prediction (e.g., unidirectional and / or bidirectional prediction) performed may depend on the slice type of the current block. For example, for a P slice, only unidirectional prediction is available for / can be used to perform inter-frame prediction. For a B slice, either unidirectional or bidirectional prediction is available for / can be used to perform inter-frame prediction. For example, if the encoder uses unidirectional prediction, the encoder can determine and / or generate a reference block for predicting the current block from reference picture list 0. For example, if the encoder uses bidirectional prediction, the encoder can determine and / or generate a first reference block for predicting the current block from reference picture list 0, and determine and / or generate a second reference block for predicting the current block from reference picture list 1.

[0119] Figure 14 An example of bidirectional prediction is shown. Two reference blocks, 1402 and 1404, can be used to predict the current block 1400. Reference block 1402 can be in the reference image of either reference image list 0 or reference image list 1. Reference block 1404 can be in the reference image of the other one of reference image list 0 or reference image list 1. Figure 14 As shown, reference block 1402 may (e.g., temporally) be in a first picture preceding the current picture of current block 1400, and reference block 1404 may (e.g., temporally) be in a second picture following the current picture of current block 1400. With regard to picture order count (POC), the first picture may precede the current picture. With regard to POC, the second picture may follow the current picture. In other examples, with regard to POC, the reference pictures may all precede or all follow the current picture. The POC may be / indicate the order in which pictures (e.g., from a decoded picture buffer) are output. The POC may be / indicate the order in which pictures are generally intended to be displayed. The output pictures are not necessarily displayed, but may undergo different processing and / or consumption (e.g., code conversion). Two reference blocks used for bidirectional prediction determination and / or generation may correspond to the same reference picture (e.g., be included in the same reference picture). For example, if two reference blocks correspond to the same reference picture, the reference picture may be included in both reference picture list 0 and reference picture list 1.

[0120] Configurable weights and / or offset values ​​can be applied to one or more inter-frame prediction reference blocks. The encoder can use flags in the Picture Parameter Set (PPS) to implement weighted prediction. The encoder can send / signal the weights and / or offset parameters in the slice header used for the current block 1400. Different weights and / or offset parameters can be sent / signaled for the luma and / or chroma components.

[0121] The encoder can use inter-frame prediction to determine and / or generate reference blocks 1402 and 1404 for the current block 1400. The encoder can determine the difference between the current block 1400 and each of the reference blocks 1402 and 1404. These differences can be prediction errors or residuals. The encoder can store and / or transmit / signalize the prediction errors and / or their corresponding motion information in / via the bit stream. The prediction errors and their corresponding motion information can be used for decoding and / or other forms of overhead.

[0122] Motion information for reference block 1402 may include motion vector 1406 and / or reference indicator / index. The reference indicator may indicate a reference image of reference block 1402 in a list of reference images. In some examples, motion information for reference block 1402 may include an indication of motion vector 1406 and / or an indication of reference index. The reference index may indicate a reference image of reference block 1402 in a list of reference images.

[0123] Motion information for reference block 1404 may include motion vector 1408 and / or reference index / indicator. The reference indicator may indicate a reference image of reference block 1404 in a list of reference images. Motion information for reference block 1404 may include an indication of motion vector 1408 and / or an indication of reference index. The reference index may indicate a reference image of reference block 1404 in a list of reference images.

[0124] The decoder can decode the current block 1400 by determining and / or generating reference blocks 1402 and 1404. The decoder can determine and / or generate reference blocks 1402 and 1404, for example, based on corresponding relevant motion information for the reference blocks 1402 and 1404. Reference blocks 1402 and 1404 can correspond to / form (e.g., be considered as) a prediction of the current block 1400 (e.g., used to generate a prediction block). The decoder can decode the current block 1400 based on combining the prediction with the prediction error.

[0125] For example, motion information can be predictively encoded (e.g., in HEVC, VVC, and / or other video coding standards / formats / protocols) before being stored in / via the bitstream and / or transmitted / notified by signaling. Motion information for the current block can be predictively encoded based on motion information from one or more blocks adjacent to it. Motion information from adjacent blocks is often correlated with the motion information of the current block because the motion of objects represented in the current block is typically the same (or similar) to the motion of objects in adjacent blocks. Motion information prediction techniques (such as those in HEVC and VVC) can include Advanced Motion Vector Prediction (AMVP) and / or inter-frame prediction block merging (e.g., merge modes).

[0126] Encoder (e.g., such as) Figure 2 The encoder 200 shown can encode motion vectors. The encoder can (e.g., using an AMVP) encode motion vectors as the difference between the motion vector of the current block being encoded and the predicted motion vector value (MVP). The encoder can determine / select MVPs from a list of candidate MVPs. Candidate MVPs can be previously decoded motion vectors of neighboring blocks in the current image of the current block and / or blocks at or near the juxtaposition of the current block in other reference images, or correspondences with these previously decoded motion vectors. The encoder and / or decoder can reciprocally generate and / or determine the candidate MVP list.

[0127] The encoder can determine / select an MVP from a list of candidate MVPs. The encoder can then send / signal an indication of the selected MVP and / or Motion Vector Difference (MVD) in / via the bit stream. The encoder can use an index / indicator to indicate the selected MVP in the bit stream. The index can indicate the selected MVP in the candidate MVP list. The MVD can be determined / calculated based on the difference between the motion vector of the current block and the selected MVP. For example, for a motion vector indicating the location relative to the current block being encoded (e.g., including horizontal (MVx) and vertical (MVy) components), the MVD can be composed of two components. and express.

[0128] and It can be determined / calculated as:

[0129]

[0130] MVDx and MVDy can represent the horizontal and vertical components of MVD, respectively. MVPx and MVPy can represent the horizontal and vertical components of MVP, respectively.

[0131] Decoder (e.g.) Figure 3 The decoder 300 shown can decode motion vectors by adding MVD to the bitstream / via the MVP indicated by the bitstream. The decoder can decode the current block by determining and / or generating a reference block. The decoder can determine and / or generate the reference block, for example, based on the decoded motion vectors. The reference block can correspond to / form (e.g., be considered as) the prediction of the current block. The decoder can decode the current block by combining the prediction with the prediction error.

[0132] The candidate MVP list for AMVP (e.g., in HEVC, VVC, and / or one or more other communication protocols) may include two or more candidates (e.g., candidates A and B). Candidates A and B may include: up to two (or any other number) spatial candidate MVPs determined / derived from five (or any other number) spatial neighboring blocks of the current block being encoded; one (or any other number) temporal candidate MVP determined / derived from two (or any other number) temporally juxtaposed blocks (e.g., in the case where neither spatial candidate MVP is available or they are identical); and / or a zero motion vector candidate MVP (e.g., in the case where one or both spatial or temporal candidate MVPs are unavailable). Additional numbers of spatial candidate MVPs, spatial neighboring blocks, temporal candidate MVPs, and / or temporally juxtaposed blocks may be available in the candidate MVP list.

[0133] Figure 15A Example spatial candidate neighbor blocks are shown for the current block. For example, five (or any other number) spatial candidate neighbor blocks can be located relative to the current block 1500 being encoded. The five spatial candidate neighbor blocks can be A0, A1, B0, B1, and B2. Figure 15B The diagram illustrates time-juxtaposed blocks used for the current block. For example, two (or any other number) time-juxtaposed blocks can be positioned relative to the current block 1500 being encoded. The two time-juxtaposed blocks can be C0 and C1. The two time-juxtaposed blocks can be within one or more reference images, which may differ from the current image of the current block 1500.

[0134] Encoder (e.g., such as) Figure 2The encoder 200 shown can encode motion vectors using inter-frame prediction block merging (e.g., merging mode). For example, the encoder (e.g., using merging mode) can reuse the same motion information from neighboring blocks (e.g., one of neighboring blocks A0, A1, B0, B1, and B2) for inter-frame prediction of the current block. Similarly, the encoder (e.g., using merging mode) can reuse the same motion information from temporally juxtaposed blocks (e.g., one of temporally juxtaposed blocks C0 and C1) for inter-frame prediction of the current block. There is no need to send (e.g., indicate, signal) MVD for the current block because the same motion information as that from neighboring blocks or temporally juxtaposed blocks is available for the current block (e.g., at the encoder and / or decoder). Because there is no need to indicate MVD for the current block, the signaling overhead for sending / signaling motion information for the current block can be reduced. The encoder and / or decoder can reciprocally generate a candidate list of motion information based on the neighboring blocks or temporally juxtaposed blocks of the current block (e.g., in a manner similar to AMVP). The encoder can determine the motion information of the current block being encoded by using (e.g., inheriting) motion information from a neighboring block or a temporally juxtaposed block in a candidate list. The encoder can signal / send indications of the determined motion information from the candidate list in / via the bit stream. For example, the encoder can signal / send indicators / indexes. Indexes can indicate the determined motion information in the candidate motion information list. The encoder can signal / send indexes to indicate the determined motion information.

[0135] The list of candidate motion information for merge modes (e.g., in HEVC, VVC, or any other encoding format / standard / protocol) can include: from (e.g., such as Figure 15A As shown, up to four (or any other number) spatial merge candidates are derived / determined from five (or any other number) spatially adjacent blocks; from (e.g., such as...) Figure 15B The two (or any other number of) temporal merging candidates derived from the temporal juxtaposition blocks shown; and / or additional merging candidates including bidirectional prediction candidates and zero motion vector candidates. In some examples, the spatial neighbor blocks and temporal juxtaposition blocks used for merging patterns can be the same as those used for AMVP.

[0136] Inter-frame prediction can be performed in ways and variations different from those described herein. For example, motion information prediction techniques different from AMVP and merge modes can be used. While the various examples in this document correspond to inter-frame prediction modes such as those used in HEVC and VVC, the methods, apparatus, and systems described herein can be applied to / used for other inter-frame prediction modes (e.g., for other video coding standards / formats such as VP8, VP9, ​​AV1, etc.). History-based motion vector prediction (HMVP), combined intra / inter-frame prediction modes (CIIP), and / or (e.g., merge modes with motion vector differences as described in VVC) can be performed / used and are all within the scope of this disclosure.

[0137] Block matching operations (or techniques) can be applied / used (e.g., in inter-frame prediction) to determine reference blocks in images different from the current block being encoded (e.g., encoded and / or decoded). Block matching operations can also be applied / used to determine reference blocks in images identical to the current block being encoded. Reference blocks in images identical to the current block, as determined by block matching, may often inaccurately predict (e.g., for camera-captured video) the current block. For example, if reference blocks in images identical to the current block are used for encoding, prediction accuracy for screen content video may not be similarly affected. Screen content video can include, for example, computer-generated text, graphics, animations, etc. Screen content video can include (e.g., can often include) repeating patterns within the same images (e.g., repeating patterns of text and / or graphics). Using reference blocks (e.g., as determined by block matching) in images identical to the current block being encoded can provide efficient compression for screen content video.

[0138] Prediction techniques can be used (e.g., in HEVC, VVC, and / or any other coding standard / format / protocol) to leverage the correlation between sample blocks within the same frame (e.g., in screen content video). Prediction techniques can be Intra-Block Copy (IBC) or Current Picture Reference (CPR). The encoder can apply / use block matching techniques (e.g., similar to inter-frame prediction) to determine displacement vectors (e.g., block vectors (BV)). The BV can indicate the relative positioning (e.g., predicted based on intra-frame block compensation) of the reference block that best matches the current block relative to the current block. For example, the relative positioning of the reference block could be the relative positioning of the top-left corner (or any other point / sample) of the reference block. The BV can indicate the relative displacement from the current block to the reference block that best matches the current block. The encoder can determine the best-matching reference block from the blocks tested during the search process (e.g., in a manner similar to that used for inter-frame prediction). The encoder can determine that the reference block is the best-matching reference block based on one or more cost criteria. One or more cost criteria can include rate-distortion criteria (e.g., Lagrange rate-distortion cost). One or more cost criteria may be based on one or more differences (e.g., SSD, SAD, SATD, and / or differences determined based on a hash function) between, for example, the predicted samples of a reference block and the original samples of the current block. The reference block may correspond to or include previously decoded sample blocks (e.g., reconstructed samples) of the current image. The reference block may include decoded sample blocks of the current image prior to processing by in-loop filtering operations (e.g., deblocking and / or SAO filtering).

[0139] Figure 16 An example of IBC (e.g., IBC mode) is shown. Figure 16 The examples shown can be correlated with screen content. A rectangular section / segment with an arrow starting at its boundary can be the current block being encoded. The rectangular section / segment pointed to by the arrow can be a reference block used to predict the corresponding current block.

[0140] IBC can be used to determine and / or generate a reference block for the current block. The encoder can determine the difference (e.g., the corresponding sample-by-sample difference) between the reference block and the current block. This difference can be a prediction error or a residual. The encoder can store and / or transmit / signal the prediction error and / or related prediction information in / via the bit stream. The prediction error and / or related prediction information can be used for decoding and / or other forms of consumption. The prediction information may include a BV. The prediction information may include an indication of the BV. The decoder (e.g., such as...) Figure 3The decoder 300 shown can decode the current block by determining and / or generating a reference block. The decoder can determine and / or generate the current block, for example, based on prediction information (e.g., BV). The reference block can correspond to / form (e.g., be considered as) the prediction of the current block (e.g., the prediction block). The decoder can decode the current block by combining the prediction (e.g., the prediction block) with the prediction error (e.g., the residual or the residual block).

[0141] Before being stored and / or transmitted in / via the bit stream / notified by signaling, the BV can be predictively encoded (e.g., in HEVC, VVC, and / or any other encoding standard / format / protocol). For example, the BV for the current block can be predictively encoded based on the BVs of one or more blocks adjacent to the current block. For example, the encoder can predictively encode the BV using a merging mode (e.g., in a manner similar to that described herein for inter-frame prediction), (e.g., AMVP as described herein for inter-frame prediction), or a technique similar to AMVP. A technique similar to AMVP could be BV prediction and differential coding (or AMVP for IBC).

[0142] Encoders that perform BV prediction and encoding (e.g., such as...) Figure 2 The encoder 200 shown can encode the BV as the difference between the BV of the current block being encoded and the block vector prediction value (BVP). The encoder can select / determine a BVP from a list of candidate BVPs. Candidate BVPs can include previously decoded BVs of neighboring blocks in the current picture of the current block, or correspond to those previously decoded BVs. The encoder and / or decoder can reciprocally generate or determine the list of candidate BVPs.

[0143] The encoder can send / signal an indication of the selected BVP and block vector difference (BVD) in / via the bit stream. The encoder can use an index / indicator to indicate the selected BVP in the bit stream. The index can indicate (e.g., point to) the selected BVP in a list of candidate BVPs. The BVD can be determined / calculated based on the difference between the current block's BV and the selected BVP. For example, for a BV indicating the location relative to the current block being encoded (e.g., represented by a horizontal component (BVx) and a vertical component (BVy)), the BVD can be composed of two components. and express. and It can be determined / calculated as:

[0144]

[0145] BVDx and BVDy can represent the horizontal and vertical components of BVD, respectively. BVPx and BVPy can represent the horizontal and vertical components of BVP, respectively. Decoder (e.g.) Figure 3 The decoder 300 shown can decode a BV by adding a BVD to the bitstream / via a BVP indicated by the bitstream. The decoder can decode the current block by determining and / or generating a reference block. The decoder can determine and / or generate the reference block, for example, based on the decoded BV. The reference block can correspond to / form (e.g., be considered as) the prediction of the current block. The decoder can decode the current block by combining the prediction (e.g., the prediction block) with the prediction error (e.g., a residual or a residual block).

[0146] A BV identical to that of an adjacent block can be used for the current block, and there is no need for separate signaling / sending of a BVD for the current block, as is the case in merge mode. A BVP (in the candidate BVP) corresponding to the decoded BV of an adjacent block can itself be used as the BV for the current block. Not sending a BVD reduces signaling overhead.

[0147] The candidate BVP list (e.g., in HEVC, VVC, and / or any other coding standard / format / protocol) can include two (or more) candidates. Candidates can include candidates A and B. Candidates A and B can include: up to two (or any other number) spatial candidate BVPs determined / derived from five (or any other number) spatial neighboring blocks of the current block being encoded; and / or (e.g., in the case where spatial neighboring candidates are unavailable) one or more of the last two (or any other number) encoded BVs. For example, if intra-frame prediction or inter-frame prediction is used to encode neighboring blocks, spatial neighboring candidates may not be available. The position of spatial candidate neighboring blocks relative to the current block using IBC encoding can be compared with the position used in (e.g., as...) Figure 15A The spatial candidate neighbor blocks, which encode motion vectors in inter-frame prediction as shown in the example, are illustrated in a similar manner. For example, as... Figure 15A As shown, the five spatial candidate neighboring blocks of the current block using IBC encoding can be represented as A0, A1, B0, B1 and B2, respectively.

[0148] In the prior art, various decoder-side techniques have been introduced to perform intra-block prediction of the current block to determine the prediction block for reconstructing the current block without explicitly signaling any specific intra-block prediction mode in the bitstream. Such techniques are enabled by reciprocally (e.g., independently and identically) deriving the intra-block mode for encoding the current block based on the encoder and decoder using previously encoded / decoded samples (e.g., reconstructed samples). If the encoder and decoder reciprocally determine / derive the same intra-block mode, the signaling of the intra-block mode can be omitted. These decoder-side techniques include decoder-side intra-block mode derivation (DIMD) and template-based intra-block mode derivation (TIMD).

[0149] Figure 17A An example of decoder-side intra-frame mode derivation (DIMD) for encoding the current block is shown according to some implementations. The DIMD method performs texture gradient processing to derive multiple (e.g., two) DIMD modes. In some examples, a fusion / mixing scheme may be applied to DIMD such that DIMD prediction 1722 may include a combination (e.g., fusion or mixing) of DIMD modes 1716 and 1718A through 1718B. For example, DIMD modes 1718A and 1718B may be optimal modes (e.g., lowest cost) and combined with DIMD mode 1716, which may be a planar mode. DIMD prediction 1722 may include a weighted average (e.g., a linear combination) with weights 1728A through 1728B and 1726 determined respectively for DIMD modes 1718A through 1718B and 1716. DIMD prediction 1722 may be applied to a reference template 1712 (e.g., the current template) to determine a predicted block 1724 for the current block 1710. The selection of DIMD is signaled in the bitstream used for intra-coded blocks using a flag. At the decoder, if the DIMD flag is true, the intra-prediction mode is derived during reconstruction using the same previously encoded neighboring pixels. If not, the intra-prediction mode is resolved from the bitstream as in classic intra-coded modes.

[0150] In some examples, to derive the DIMD mode (and DIMD prediction value 1722) for the current block 1710, a reference template 1712 can be determined, and a reference sample set (e.g., neighboring pixels) can be selected. Gradient analysis is then performed on this reference sample set, as described above. Figures 11 to 12The gradient analysis described herein. For prescriptive purposes, these pixels should be the decoded / reconstructed pixels of the image. The principal angular direction (corresponding to the IPM) for the template (e.g., reference template 1712) can be determined, and it is assumed that this principal angular direction is very likely to be the same as the principal angular direction of the current block to be predicted. In some examples, gradient analysis can be performed using a simple 3x3 Sobel gradient filter defined by the following matrix convolved with the template:

[0151] For each selected reference sample (e.g., a pixel) in template 1712, each of these two matrices can be multiplied pointwise with a 3x3 Sobel filter window, and the results summed. The 3x3 Sobel filter window can be centered on the current reference sample and consists of the current reference sample's eight direct neighbors. Therefore, two values ​​corresponding to the gradient intensity at the current sample are obtained in the horizontal and vertical directions respectively (from the...). (multiplication) and (from and) (multiplication) The angle can be calculated for the window. The calculated angle can correspond to one of the angle IPMs (e.g., one of 65 angle IPMs) (e.g., converted to that angle IPM), and the associated amplitude (i.e., the amplitude for window positioning). This can be added to a gradient histogram (HoG) 1720 indexed by the corresponding IPM. After the magnitude and angle of each window location in template 1712 are processed, each entry in the resulting HoG 1720 represents the cumulative magnitude for the corresponding IPM. Therefore, HoG 1720 can be determined using the magnitudes corresponding to the IPM counts of the gradient analysis performed across each sample of reference template 1712.

[0152] In some examples, multiple IPMs can be determined for the mixing / fusion process, such as DIMD modes 1718A to 1718B. In some examples, the planar mode can have a fixed weight (e.g., ¼), and the remaining weights can be adjusted based on a subtraction of the fixed weight (e.g., 1 - ¼ = ¾).

[0153] In some examples, position-dependent DIMD patterns are introduced to adjust the weights 1728A to 1728B of the angle IPM in the DIMD. During DIMD IPM derivation, a process known as “position-dependent” determination is applied to determine which template region each selected orientation pattern in the selected orientation patterns originates from. To determine whether a particular reference sample in the template contributes to the inference of a particular DIMD pattern, three separate regions are considered within the DIMD template: the region above the current block (ABOVE), the region to the left of the current block (LEFT), and the upper-left region of the current block (ABOVE-LEFT). For example, the left region may include one or more columns of samples to the left of the current block. For example, the top region may include one or more rows of samples above the current block. For example, the upper-left region may include one or more samples located both above and to the left of the current block. Gradient calculations are performed separately for the samples in each region, resulting in three histograms respectively. , and For example, for directional patterns , Indicates direction The cumulative magnitude of all samples in the region ABOVE, and and This corresponds similarly to the LEFT and ABOVE-LEFT regions. It should be noted that, relative to regular DIMD, the template region can be expanded by one sample in the upper left and one sample in the lower right.

[0154] The complete gradient histogram for the entire template can then be calculated as the sum of three separate histograms. As in regular DIMD, the two orientation modes with the largest and second largest cumulative amplitudes in the histograms are selected as the primary DIMD modes. and secondary DIMD mode .

[0155] Additionally, histogram and Can be used to determine and / or Whether it depends on a specific template region ABOVE or LEFT. Specifically, for what is represented as an indication of The location correlation indicator can be defined as:

[0156] Therefore, the positional correlation of DIMD for each DIMD mode can be determined based on the analysis of the HoG peak amplitude of the selected angle IPM corresponding to the DIMD mode.

[0157] In some implementations, the fusion / mixing scheme used to determine the DIMD prediction value 1722 can be adjusted based on location-dependent DIMD patterns. For example, mixing is performed to combine the primary DIMD prediction value... and secondary DIMD predictions With plane prediction value Fusion. No DIMD mode was determined to be location-dependent (e.g., this means indicating...). In such cases, uniform mixing should be applied. Uniform weighting. , and It is derived based on the relative magnitude of the patterns in the histogram, and the final DIMD prediction can be calculated as:

[0158] Otherwise, if at least one DIMD pattern in the DIMD patterns is inferred to be location-related, for example, if the indication of location correlation is not zero, then a sample-based mixture is used. Different weights are used for each location. The predicted values ​​at each location are mixed.

[0159] if (For example, the indication of location relevance indicates the vertical or horizontal direction), then according to Calculate the predicted value Sample-based weights This makes the average weight used within the block approximately equal to the uniform weight. This allows for the use of higher weights in the portions of the block closer to the ABOVE or LEFT regions. The corresponding values ​​can be determined and predefined. Fixed range of maximum deviation (for example, usually) ). The higher the value, the greater the change in weight within the block. This is especially true for blocks of different sizes. The block:

[0160] if and According to The weights are calculated for the two predicted values ​​as shown above. .

[0161] Conversely, if and Then used for The weights are calculated as follows:

[0162] Finally, for plane prediction values The weights were then calculated as follows:

[0163] The final location-related DIMD prediction was then calculated as follows:

[0164] Ultimately, the planar pattern can be systematically applied to the mixing process with fixed weights (e.g., ¼ or 21 / 64).

[0165] Figure 17B An example of template-based intra-mode derivation (TIMD) for encoding the current block, according to some implementation schemes, is shown. TIMD is a type of intra-prediction where the intra-prediction mode (IPM) can be reciprocally determined (e.g., derived identically and independently) by the encoder (e.g., encoder 200) and the decoder (e.g., decoder 300), such that the IPM determined (e.g., selected) by the encoder does not need to be signaled to the decoder. Therefore, signaling bandwidth is reduced, and TIMD can be viewed as a type of decoder-side intra-mode derivation process.

[0166] like Figure 17B As shown, for TIMD, a video codec (e.g., encoder 200 or decoder 300) can determine a template 1704 for the current block 1702. Template 1704 can include one or more regions of samples in the reconstruction region 1708 of the current picture (or frame) of the current block 1702. One or more regions can include reconstruction samples adjacent to (e.g., near) the current block 1702. In some examples, one or more regions of template 1704 can include template region 1704A to the left of the current block 1702 (e.g., left template) and template region 1704B above the current block 1702 (e.g., top template). In some examples, template 1704 can include regions above and to the left of the current block 1702 (e.g., the region surrounded by the reference template 1706 and template regions 1704A to 1704B and referred to as the top-left template). The thicknesses (e.g., width and height, respectively) of template regions 1704A and 1704B can be R1 and R2 samples, respectively. For example, R1 and / or R2 can be 2 samples, 4 samples, 8 samples, etc. Template region 1704A can have a height of N samples, which can be the height of the current block 1702. Template region 1704B can have a width of M samples, which can be the width of the current block 1702. Therefore, template 1704A can include L1 columns of reconstructed samples to the left of the current block 1702 (e.g., adjacent to the current block). Similarly, template 1704B can include L2 rows of reconstructed samples above the current block 1702 (e.g., adjacent to the current block).

[0167] In TIMD, the reference for template 1706 is used to derive template predictions for template 1704. The video codec can determine (e.g., select and obtain) the reference for template 1706 as sample regions (within reconstruction region 1708) adjacent to (e.g., near) template 1704. For example, the reference for template 1706 may include reconstructed samples to the left and above template 1704. The reference for template 1706 may have a thickness to the left of template region 1704A with R1 samples and a thickness above template region 1704B with R2 samples. For example, R1 and / or R2 may be 1 sample, 2 samples, 4 samples, etc., and R1 may be different from R2. Therefore, the reference for template 1706 may include a number of R1 column reconstructed samples to the left of template 1704A and a number of R2 row reconstructed samples above template 1704A. In some examples, the reference to template 1706 may include an upper region with a width greater than that of template region 1704B (e.g., a width greater than or equal to twice the width of template region 1704B, 2(M+L1) samples, 2(M+L1)+R1 samples, etc.). In some examples, the reference to template 1706 may include a left-side region with a height greater than that of template region 1704A (e.g., a height greater than or equal to twice the height of template region 1704A, 2(N+L1) samples, 2(N+L1)+R2 samples, etc.).

[0168] In some examples, a list of candidate intra-prediction modes (IPMs) can be used to determine the TIMD mode prediction value. For example, this list may include IPMs from a list of most probable modes (MPMs). In some examples, one or more of DC modes, planar modes, horizontal DC modes, and / or vertical DC modes, or horizontal and / or vertical plane modes, may be added to the candidate IPM list. The cost (e.g., SSE, SAD, or SATD) for each candidate IPM in the list can be determined based on the difference between the reconstructed samples in template 1704 and the predicted samples of template 1704 generated using the candidate IPMs and a reference to template 1706. For example, a video codec can determine the predicted samples of template 1704 by applying the candidate IPMs to samples referenced to template 1706, similar to how intra-prediction modes can be applied to template 1704 to predict the current block 1702 in regular intra-prediction.

[0169] In some examples, a first IPM and a second IPM with the lowest cost (determined for candidate IPMs in the list) are selected from a list to determine (e.g., derive) a first TIMD pattern and a second TIMD pattern. In the example, the first TIMD pattern and the second TIMD pattern are determined as the first IPM and the second IPM, respectively.

[0170] In some examples, the first TIMD pattern and the second TIMD pattern can be determined by refining the first IPM and the second IPM. For example, the angular pattern range can be expanded from a first range (e.g., 67 patterns) to a second range (e.g., 131 patterns) of the list, and the cost of the two neighboring patterns (i.e., + / - 1 patterns) of each selected IPM can be determined. For example, the first TIMD pattern can be determined as the IPM with the minimum cost among the costs of the first IPM in the second range and its two neighboring patterns. The second TIMD pattern can be determined as the IPM with the minimum cost among the costs of the second IPM in the second range and its two neighboring patterns.

[0171] In some implementations, TIMD pattern predictions can be determined based on combining (e.g., mixing or fusing) a first TIMD pattern and a second TIMD pattern. For example, the TIMD pattern predictions can be a linear combination of the first and second TIMD patterns. The weights of each of the first and second TIMD patterns in the linear combination can be determined based on their respective costs. For example, for each TIMD pattern in the first and second TIMD patterns, the weights can be determined to be inversely proportional to the cost of the TIMD pattern. The TIMD pattern predictions can be applied to template 1704 to determine the prediction block used to predict the current block 1702. The encoder can generate a residual (e.g., prediction error) based on the difference between the prediction block and the current block 1702. The decoder can reconstruct the current block 1702 based on the reciprocally / identically generated prediction blocks and the residuals in the bitstream received from the encoder. For example, the decoder can combine / add the residuals obtained from the bitstream with the prediction blocks.

[0172] Figure 18 An example of signaling TIMD for decoding the current block is shown according to some implementations. At box 1802, the decoder (e.g., from the bitstream) receives an indication of whether TIMD is applied to generate a predicted block for encoding the current block. As explained above, TIMD is a type of intra-prediction where the template of the current block in the reconstructed region of the image is used to deduce the intra-prediction mode for encoding the current block. For example, the indication could be a flag indicating whether TIMD is enabled (e.g., applied). For example, the flag could be a single bit.

[0173] At box 1804, the decoder determines whether to apply TIMD based on the indication received at box 1802.

[0174] At box 1810, based on an indication of the TIMD being applied (e.g., enabled or selected), the decoder determines (e.g., derives) the TIMD mode prediction value based on multiple intra-prediction modes (IPMs), for example, without additional signaling from the encoder indicating a specific intra-prediction mode. (As stated above...) Figure 17B As described, TIMD mode prediction values ​​can be determined based on identifying two IPMs with the lowest cost among those determined for the multiple IPMs. Because the encoder and decoder reciprocally (e.g., independently and identically) determine the TIMD mode prediction values, signaling for any particular intra-prediction mode (IPM) can be omitted from the bitstream. At block 1812, the decoder generates a prediction block (e.g., as shown in the diagram) based on the determined TIMD mode prediction values. Figure 17B (As described in the text).

[0175] At box 1806, based on an indication that TIMD is not applied (e.g., disabled or not selected), the decoder receives (e.g., parses) an indication of an intra-prediction mode (IPM) from the bitstream. This indication may include multiple bits representing the index of an IPM in a list of specified IPMs (e.g., an MPM list). At box 1808, the decoder generates a prediction block (e.g., as indicated by the IPM) based on the indicated IPM. Figure 17B (As described in the document). At box 1814, the decoder reconstructs the current block based on the predicted block and, for example, the residual received from the bitstream.

[0176] In existing technologies, TIMD is introduced as a decoder-side technique where intra-frame prediction of the current block can be performed to determine the prediction block used to reconstruct the current block without explicitly signaling any particular intra-frame prediction mode in the bitstream. Furthermore, TIMD mode prediction values ​​can be generated as a linear combination of two or more IPMs in a list of intra-frame prediction modes (IPMs) that have the lowest cost among the IPMs in the list (e.g., SATD or SAD). The IPM list can include one or more angular IPMs, DC modes, and / or one or more planar modes (e.g., vertical plane modes or horizontal plane modes). Linear combinations of multiple IPMs are also referred to as TIMD IPM fusion or TIMD IPM mixing.

[0177] Generally, the TIMD pattern is selected to minimize the cost of the template region used for the current block, and through spatial proximity, the TIMD pattern is considered a potentially good candidate for predicting samples for the current block. However, the farther the current sample in the current block is from the template region, the less likely there is a spatial correlation between the template region and the current block. Therefore, applying the same TIMD weights to derive predicted samples may be inaccurate and may produce large prediction errors. Some location-dependent weights have been introduced as a HoG amplitude-dependent DIMD technique, which is not suitable for TIMD.

[0178] Embodiments of this disclosure relate to an improved method for mixing (e.g., combining or fusing) multiple TIMD modes derived from multiple intra-prediction modes to generate TIMD mode prediction values ​​that take into account the spatial proximity of the predicted samples to samples from which the TIMD modes are determined. Specifically, a video codec (e.g., an encoder and / or decoder) can determine the orientation (e.g., horizontal, vertical, or diagonal) of each TIMD mode, such that the weights of the TIMD modes can be adjusted based on the position of the predicted samples. The video codec can determine (e.g., derive) the weights of a linear combination of TIMD modes. For example, the weights used for the TIMD modes (e.g., derived IPMs or non-angular intra-prediction mode IPMs). NA The weights can be inversely proportional to the cost used for TIMD mode. Subsequently, in some examples, the weights can be adjusted based on the maximum weight value (e.g., range) to be used in the per-sample mixing. In some examples, the maximum weight value is not fixed and can be determined by the video codec based on the size of the current block and / or the image resolution of the current block (e.g., set or adjusted).

[0179] In some examples, determining multiple TIMD modes may include combining at least one TIMD mode (e.g., two TIMD modes derived from multiple intra-prediction modes) with a non-angular intra-prediction mode (IPM). NA The corresponding TIMD patterns can be conditionally combined. For example, IPM NA This can include DC mode or planar mode. The video codec can determine (e.g., derive) the weights of the TIMD modes and the linear combination of at least one TIMD mode based on the cost of the TIMD modes in the linear combination. For example, the weights used for the TIMD modes (e.g., derived IPM or IPM). NA It can be inversely proportional to the cost used for this TIMD mode.

[0180] In some examples, the video codec can be based on whether at least one TIMD mode includes IPM. NA And / or includes at least one IPM NA(e.g., whether at least one TIMD pattern is an angle IPM) to determine whether it will be associated with IPM. NA The corresponding TIMD mode is combined with at least one TIMD mode to generate TIMD mode prediction values. For example, if at least one TIMD mode does not include any non-angular intra-frame modes and / or does not include IPM. NA Then the video codec can determine to combine at least one TIMD mode with IPM. NA Combining them. In other examples, the video codec may further consider one or more criteria to determine whether to combine at least the TIMD mode with IPM. NA Combining is then performed. In some examples, the orientation of TIMD patterns can be determined based on the cost difference (e.g., SATD, SSE, SAD, etc.) calculated for the regions used for the current template. For example, a video codec can determine the ratio of a first cost for a first region of the template to a second cost for a second region of the template. For example, the first region could be the upper template, and the second region could be the left template. By dynamically adjusting the weights of one or more TIMD patterns used to determine the predicted samples during the TIMD fusion process, the locational relevance of the predicted samples or their spatial proximity to the template can be considered. Therefore, mixing intra-frame predictions can better suit the signal characteristics of the block to be predicted.

[0181] These and other features of this disclosure are further described below.

[0182] Figure 19A A flowchart 1900A illustrates a method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes. The method in flowchart 1900A can be derived from a video codec (e.g., Figure 2 encoder 200 or Figure 3 The decoder (300) in the flowchart is implemented. The flowchart method can be performed reciprocally (e.g., independently and identically) by the encoder and decoder, so that it is not necessary for the encoder to signal the intra-prediction mode to the decoder to reconstruct the current block using intra-prediction techniques.

[0183] At box 1902, the video codec determines multiple costs for multiple intra-prediction modes (IPMs) based on the TIMD being applied to the current block (e.g., enabled for the current block). Each cost for a given IPM among the multiple IPMs is based on the difference between: the predicted sample of the template generated from the IPM applied to the reference sample for the template; and the reconstructed sample of the template. Regarding... Figure 17B Examples of templates for the current block (e.g., templates 1704A to 1704B) and references to templates (e.g., references to template 1706) are shown and described.

[0184] In some examples, a video codec can be a decoder that determines multiple costs based on an indication received from the bitstream (e.g., parsing and decoding) of a TIMD that is being applied to the current block (e.g., enabled for that current block). For example, the encoder can determine that a TIMD is being applied to encode the current block and signal that indication to the decoder in the bitstream. In some examples, the indication can be a syntax element (e.g., a flag, such as) that is activated in the bitstream at the block level (e.g., per block) and signaled. ).

[0185] At box 1903, the video codec determines the TIMD mode prediction value based on multiple IPMs. Box 1903 may include boxes 1904 through 1920.

[0186] At box 1904, the video codec determines the first TIMD mode and the second TIMD mode based on the first IPM and the second IPM, which have the lowest cost among multiple IPMs.

[0187] At box 1906, the video codec determines whether to select the second TIMD mode to determine the TIMD mode prediction value.

[0188] At box 1916, since the second TIMD mode is not selected, the video codec determines the TIMD mode prediction value based on the first TIMD mode. For example, the TIMD mode prediction value can be determined as the first TIMD mode.

[0189] At box 1908, based on the selection of a second TIMD mode (e.g., to be considered when mixed with the first TIMD mode), the video codec determines whether to apply positional correlation weighting. As explained above, the encoder and decoder can independently and reciprocally / equally determine whether to apply positional correlation weighting, so that no additional signaling is required in the bitstream to conditionally enable / select positional correlation weighting. In some implementations, the video codec determines whether to apply positional correlation weighting based on the orientation determined for each of the selected first and second TIMD modes. For example, if the orientation is determined to be none, positional correlation weighting can be disabled (e.g., not enabled or not selected) and the positional correlation state can be set to 0 (indicating no positional correlation or diagonal direction).

[0190] In some implementations, the video codec can be based on the size of the current block (e.g., the height of the current block). and width The product of these factors determines whether location-dependent weighting is applied. For example, a video codec can use the size of the current block as a factor of a size threshold (e.g., set to 64) to determine whether location-dependent weighting is applied. Compare them.

[0191] if ,but

[0192] For example, location-related weighting can be disabled based on whether the size is less than or equal to a size threshold (e.g., In an alternative implementation with the same result, location-related weighting can be disabled based on a size less than a size threshold plus one (e.g., if...). ,but ).

[0193] In some examples, instead of using a size threshold, a height threshold can be used. ) and width threshold ( The positional relevance weighting is compared with the height and width of the current block to determine whether positional relevance weighting is enabled or disabled. For example, regardless of the positional relevance status (e.g., orientation) determined for the first TIMD mode and / or the second TIMD mode, if the width or height of the current block is less than (or equal to) a width threshold and a height threshold, respectively, the positional relevance status of the associated TIMD mode can be set to 0 (e.g., indicating no orientational relevance weighting or diagonal).

[0194] In some examples, the determination of whether to enable or disable location-related weighting depends on the location-related state (indicating orientation) determined for the selected TIMD mode (e.g., the first TIMD mode or the second TIMD mode) and the length of the current block corresponding to that orientation. For example, if the width of the current block ( ) below the threshold (e.g. And the selected TIMD mode ( i-th The location-related state of the pattern ) is determined as a level (e.g. If the positional correlation status is inaccurate, the video codec can determine (e.g., by...) Setting it to 0 disables positional dependency for the selected TIMD mode. Similarly, if the height of the current block ( ) below the threshold (e.g. And the location-related status of the selected TIMD mode ( ) is determined to be vertical (e.g. If the location-related state is not accurate, the video codec can determine (e.g., by...) Setting it to 0 disables location dependence for the selected TIMD mode. These settings are shown below: if

[0195] if

[0196] In some examples, the size threshold, height threshold, and / or width threshold can be predetermined (e.g., predefined or preconfigured) at the encoder and decoder. In other examples, the encoder can signal the values ​​of the size threshold, height threshold, and / or width threshold to the decoder in the bitstream.

[0197] At box 1918, with location-related weighting not selected / applied / enabled, the video codec determines the TIMD mode prediction value, which includes a linear combination of the first and second TIMD modes. Box 1918 may include box 1920. At box 1920, the video codec determines the corresponding weights of the first and second TIMD modes in the linear combination based on the costs of the first and second TIMD modes.

[0198] In some examples, each weight of a TIMD pattern can be determined to be inversely proportional to the cost of the TIMD pattern, such that a lower cost (e.g., SSE, SSD, or SADT) is associated with a higher weight. In other words, a TIMD pattern with a lower cost will have a higher weight than another TIMD pattern with a higher cost. For example, a first TIMD pattern with a first cost... The first weight of ) ) and for a second TIMD mode with a second cost ( The second weight of ) It can be determined as follows:

[0199]

[0200] At box 1910, based on location-related weighting being selected / applied / enabled, the video codec determines a TIMD mode prediction value comprising a linear combination of a first TIMD mode and a second TIMD mode, and the linear combination includes dynamic weights. Box 1910 may include boxes 1912 to 1915.

[0201] In some implementations, the weights (e.g., first weight and second weight) of the TIMD patterns selected to determine the TIMD pattern prediction value can be determined (e.g., calculated) based on the cost of the TIMD patterns. For example, the weights for the TIMD patterns can be determined to be inversely proportional to the cost of the TIMD patterns and / or proportional to the difference between the sum of the costs and the cost of the TIMD patterns. For example, the weights for the TIMD patterns can be determined according to the following equation (24):

[0202] It is the weight associated with the i-th IPM of the i-th TIMD pattern. costIPM i This is the cost associated with the i-th IPM of the i-th TIMD mode (e.g., determined for that i-th IPM). For example, This represents the first weight of the first TIMD mode, and This represents the second weight of the second TIMD mode. (e.g., 3, 4, etc.) indicates the number of TIMD patterns to be combined (i.e., in a linear combination with weights) to generate TIMD pattern predictions.

[0203] It should be noted that additional weights and orientations can be determined based on the additional TIMD patterns being mixed. One or more TIMD patterns can be selected and determined / selected from the MPM list or other patterns (e.g., DC, horizontal, vertical, and extended IPM for wide-angle IPM).

[0204] In some implementations, the orientation (e.g., first orientation and second orientation) of a TIMD pattern (e.g., a first TIMD pattern and a second TIMD pattern) is determined based on the cost of the TIMD pattern. For example, orientation may include a level indicating that predicted samples closer to the left boundary of the current block are likely to be spatially more relevant to the parameters (e.g., weights) used for the determined TIMD pattern; a vertical indicating that predicted samples closer to the upper boundary of the current block are likely to be spatially more relevant to the parameters (e.g., weights) used for the determined TIMD pattern; no orientation; and / or a diagonal indicating a combination of horizontal and vertical. In the example, no orientation may correspond to and indicate a diagonal direction.

[0205] For a selected TIMD pattern (e.g., a candidate that minimizes the cost to be computed in the template region), the orientation (indicating positional relevance) can be determined based on which part of the current template contributes most to the selection of the TIMD pattern. There can be as many positional relevance orientations as regions of the current template. For example, the current template may include a first region (e.g., template region 1704B) and a second region (e.g., template region 1704A). If the upper template is determined to be the most influential template for TIMD pattern selection, the positional relevance state / indicator associated with the selected TIMD pattern is set to "vertical," and per-sample vertical fusion is operated to mix this TIMD pattern. If the left template is determined to be the most influential template for TIMD pattern selection, the positional relevance state associated with the selected TIMD pattern is set to "horizontal," and per-sample horizontal fusion is operated to mix this TIMD pattern. If the influence of the templates in the selection of the TIMD pattern is balanced, the associated positional relevance state for the TIMD pattern can be set to "non-positional," and neither the horizontal per-sample process nor the vertical per-sample process is operated (e.g., meaning that rule-based block-by-block weighted mixing is implemented). In other words, per-sample diagonal fusion will be performed to blend this TIMD pattern. Therefore, the location-related state associated with the selected TIMD pattern affects the orientation of the per-sample TIMD fusion process. In some examples, the first and second regions do not overlap. In other examples, the first and second regions may overlap.

[0206] In some examples, the orientation of a TIMD pattern can be determined based on the availability of template samples and / or template regions. For example, only one template region might be available at the image boundary. For instance, if only a first region (e.g., the left template) exists or is available, a first orientation corresponding to that first region can be determined (e.g., positional correlation is set to 2 to indicate horizontal). Similarly, if only a second region (e.g., the upper template) exists or is available, a second orientation corresponding to that second region can be determined (e.g., positional correlation is set to 1 to indicate vertical). In some examples, when positional correlation is determined to be disabled (e.g., off), the value can be set to, for example, 0, which can indicate diagonal direction.

[0207] Therefore, the orientation (e.g., location-related status / indication) of a TIMD pattern can be determined based on one or more available template regions as follows:

[0208] In some examples, if only one component of the cost calculation (e.g., SATD, SSE, SAD, etc.) is available, location dependence can be disabled; for example, orientation associated with the selected TIMD pattern (e.g., location dependence status) can be disabled. ).

[0209] In some examples, the impact can be determined by comparing a first cost in a first region with a second cost in a second region. Costs can be calculated in the same way as TIMD costs, such as using SATD, SSE, or SAD. In some examples, a lower cost indicates a larger impact. In some examples, the ratio between the first and second costs can be calculated to determine which region has a larger impact, thus determining the orientation of the TIMD pattern.

[0210] For example, the orientation of the TIMD pattern can be determined as follows for a first cost for a first region (e.g., the left template region / area) and a second cost for a second region (e.g., the upper template region / area):

[0211]

[0212]

[0213] For illustrative purposes, costs are represented as SATD costs, but other costs are possible. It is the SATD cost associated with the selected TIMD pattern and calculated in the second region (e.g., the template above). It is the SATD cost associated with the selected TIMD pattern and calculated in the first region (e.g., the left template). It is the position-related parameter value associated with the i-th selected TIMD pattern (e.g. i Belongs to [0;2]). A value of 0 can indicate no positional dependence; a value of 1 can indicate the vertical positional dependence of the i-th selected TIMD pattern; and a value of 2 can indicate the horizontal positional dependence of the i-th selected TIMD pattern. It is a scaling factor and can, for example, be a value less than 1. For example, It can be a value included in the range [1 / 10; 1 / 6], such as 1 / 8.

[0214] In some examples, each of the first and second costs (e.g., represented as SATD costs in the examples below) can be normalized based on (e.g., for) the number of samples present in the corresponding template (e.g., quantity) before being compared to determine the orientation of the TIMD pattern (e.g., location-related status / indication):

[0215]

[0216]

[0217] For example, Represents the normalized cost (e.g., SATD) of the template region x (e.g., where x is the first region (e.g., the left region) or the second region (e.g., the upper region)). and It can be determined (e.g., calculated) as follows:

[0218]

[0219]

[0220] It is the height of the second template area (e.g., the template above). It is the width of the first template area (e.g., the left template). It is the height of the current block, and It is the width of the current block. It is the scaling factor.

[0221] For example, these parameters can be determined as follows:

[0222]

[0223]

[0224] In some examples, in the case of integer implementations, positive values ​​can be chosen. (For example However, for the specific implementation of floating-point operations, this positive value can be set to 1.

[0225] In some implementations, the ratio between costs can be calculated between template SATD and total SATD (cumulative across all templates) as follows:

[0226]

[0227]

[0228] For example, It can be in the range of [0.5; 1]. This can indicate, for example, the cumulative SATD within the first template region and the second template region (e.g., both the upper and left templates). In these examples, the logic here is shifted because it is assumed that if the majority of the total SATD cost comes from the template (e.g., the left template), then the positional correlation is towards the other template (e.g., the upper template, thus a vertical positional correlation state), since the minimum SATD cost implies a better prediction.

[0229] In some examples, when the candidate TIMD pattern is a non-angular pattern (e.g., DC or planar pattern), the associated positional relevance state (e.g., orientation) can be set to 0 (e.g., indicating no positional relevance) because these “flat” patterns are not expected to exhibit significant differences in relative template cost (e.g., SATD) when selected. However, in other examples, the positional relevance state can be determined (e.g., derived) for candidate non-angular TIMD patterns, similar to how the positional relevance state is determined for angular TIMD patterns, as described above.

[0230] In some examples, the orientation of a non-angular pattern can be derived / determined based on multiple oriented non-angular patterns (e.g., oriented DC, oriented plane) corresponding to the non-angular pattern. For example, the orientation can be determined to correspond to the direction of the lowest cost (e.g., SATD cost) of the oriented non-angular pattern (e.g., among the vertical, horizontal, and original patterns). For example, template (e.g., SATD) costs can be calculated for the original pattern, the horizontal pattern, and the vertical plane pattern. If the horizontal plane pattern has the lowest SATD cost among the three plane patterns, the positional relevance or orientation for the plane pattern can be set to horizontal. If the vertical plane pattern has the lowest SATD cost among the three plane patterns, the positional relevance or orientation for the plane pattern can be set to vertical. Otherwise, if the original plane pattern has the lowest SATD cost, the orientation can be set to diagonal (or possibly deactivated for per-sample fusion). Note that although the oriented non-angular pattern has the lowest SATD cost, the regular original non-angular pattern can be used during the fusion process. As noted above, examples regarding SATD cost have been described, but other types of costs, such as SAD, SSE, etc., can be used instead.

[0231] At box 1912, the video codec determines the first weight and first orientation of the first TIMD mode, as explained above. At box 1914, the video codec determines the second weight and second orientation of the second TIMD mode, as explained above.

[0232] At box 1915, the video codec determines the weight adjustment range (e.g. The weights of position-related TIMD modes are adjusted within this weight adjustment range. In some examples, the weight adjustment range (e.g., maximum weight value or maximum weight deviation) for the sample-based DIMD fusion process can be a fixed value, such as 10 ( In some examples, the weight adjustment range can be determined based on the current block size and / or the number of samples in the image (e.g., image resolution). In other words, the encoder and decoder can dynamically and reciprocally / equally determine the weight adjustment range for different images and / or blocks of different sizes.

[0233] Advantageously, the larger the selected current block, the less variation in the predicted value samples can be expected (e.g., flatter predicted values), meaning that reducing the sample-based weight range can increase the accuracy of the mixed predicted values. As an example, the weight adjustment range can be determined based on the size of the current block as follows:

[0234] in: It is the height of the current block (CU); It is the width of the current block (CU); This indicates the current block (CU) size threshold; values ​​exceeding this threshold are considered indicative of a problem. Then it changes. For example, . It is a fixed value, for example . It is necessary to, for example The offset value applied to a fixed value under the condition.

[0235] Furthermore, the same reasoning applies to the number of samples in an image, where for larger image resolutions, the variation in TIMD weights across blocks (CUs) should be reduced. As an example, the weight adjustment range can be determined based on the image size (e.g., the resolution of the image for the current block) as follows:

[0236] in: This is the height of the current image; This is the width of the current image; This represents a threshold for image size (e.g., characterized by the number of brightness samples). Above this threshold, Then it changes. For example, . It is necessary to, for example The offset value applied to a fixed value under the condition.

[0237] In some examples, the weight adjustment range It can be determined (e.g., adjusted or adapted) based on the current block size and the number of samples in the image, as follows:

[0238] in: It is the height of the current block; This is the width of the current block; This is the height of the current image; This is the width of the current image; This indicates the current block size threshold; values ​​exceeding this threshold are considered invalid. Then it changes. For example, . This represents the image size threshold (characterized by the number of brightness samples). Above this image size threshold, Then it changes. For example, . It is a fixed value, for example . It is necessary to, for example The offset value applied to a fixed value under the condition.

[0239] Returning to boxes 1910 and 1918, the TIMD pattern predictions comprise a linear combination of multiple selected TIMD patterns (e.g., selecting a first TIMD pattern and a second TIMD pattern at box 1918 or 1910). The linear combination includes weights determined for each TIMD pattern, as described above with respect to boxes 1910 and 1918. For example, the TIMD pattern predictions can be determined based on the following linear combination (29) ( IPM TIMD ):

[0240] These can be weights used for the i-th TIMD mode, including the i-th IPM. As described above, the weights (based on a determined weight adjustment range (e.g., maximum weight deviation / value) and the orientation determined for the i-th TIMD mode) are weighted accordingly. () can be location-related.

[0241] At box 1922, the video codec generates a prediction block for the current block based on the TIMD mode prediction value. For example, the video codec can apply the TIMD mode prediction value to the template of the current block to generate a prediction block for the current block (e.g., as described above regarding Figures 10 to 1922). Figure 12 and Figure 17B (As described).

[0242] In some examples, the TIMD predictions generated and applied per-sample can vary for each sample, based on the locational relevance of the predicted values ​​for the determined / selected / enabled TIMD pattern. Specifically, the TIMD predictions depend on the location / position of the reference sample and the locational relevance parameter (which determines the direction of the per-sample mixing) varies for each reference sample. For example, the locational relevance weights can depend on the TIMD pattern predictions to be used ( IPM TIMD The location of the generated predicted samples changes.

[0243] In some examples, when the video codec is a decoder, the decoder can reconstruct the current block based on the predicted block and the residual (e.g., prediction error or residual block). For example, the decoder can receive the residual from the bitstream (e.g., decode that residual).

[0244] In some examples, when the video codec is an encoder, the encoder can determine (e.g., generate) a residual based on the difference between the predicted block and the current block. The encoder can then encode the determined residual into a bitstream.

[0245] Figure 19B A flowchart 1900B illustrates a method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes. The method in flowchart 1900B can be derived from a video codec (e.g., Figure 2 encoder 200 or Figure 3 The decoder (300) in the flowchart is implemented. The flowchart method can be performed reciprocally / identically by the encoder and decoder, so that it is not necessary for the encoder to signal the intra-prediction mode to the decoder to reconstruct the current block using intra-prediction techniques. In some examples, many of the operations shown in flowchart 1900B can be performed similarly to the operations in flowchart 1900A.

[0246] At block 1932, the video codec determines multiple costs for multiple intra-prediction modes (IPMs) based on the TIMD being applied to the current block (e.g., enabled for the current block). In some examples, each cost for each IPM may include a first cost of the IPM applied to a first portion of the current template (e.g., to predict the first portion) and a second cost of the IPM applied to a second portion of the current template (e.g., to predict the second portion). In some examples, the current template includes a first portion and a second portion. In examples, the first portion and the second portion may not overlap. For example, the first portion and the second portion may respectively include samples to the left of the current block and samples above the current block. In another example, the first portion and the second portion may overlap. Each of the first cost and the second cost may be calculated similarly as described in block 1902.

[0247] At box 1933, the video codec determines the TIMD mode prediction value based on multiple IPMs. Block 1903 may include blocks 1933 through 1950.

[0248] At box 1934, the video codec determines multiple TIMD modes with the lowest cost among multiple costs from multiple IPMs based on a first plurality of IPMs. In some examples, box 1934 may correspond to box 1904. At box 1934, it may be based, for example, on having The lowest cost The number of IPMs determines more than two TIMD patterns.

[0249] At box 1936, the video codec determines the weight and orientation of each corresponding TIMD mode among multiple TIMD modes. The orientation of each TIMD mode can be determined based on a comparison of a first cost and a second cost for each TIMD mode. For example, the orientation can be associated with the smaller cost. Alternatively, the orientation can be determined based on the difference or ratio between the first and second costs. In some examples, additional details for calculating the orientation are provided with respect to boxes 1912 and 1914.

[0250] At box 1938, the video codec determines whether to select a second TIMD mode to determine the TIMD mode prediction value. In some examples, box 1938 may correspond to box 1906.

[0251] At box 1940, since the second TIMD mode is not selected, the video codec determines the TIMD mode prediction value based on the first TIMD mode. For example, the TIMD mode prediction value can be determined as the first TIMD mode. In some examples, box 1940 may correspond to box 1916.

[0252] At box 1942, the video codec generates a directional TIMD prediction value for each orientation, as a linear combination of one or more directional TIMDs among a plurality of TIMDs.

[0253] At box 1944, based on the selection of a second TIMD mode (e.g., to be considered when mixed with the first TIMD mode), the video codec determines whether to apply position-dependent weighting. In some examples, box 1944 may correspond to box 1908.

[0254] At box 1946, without applying location-related weighting, the video codec determines that the TIMD mode prediction value includes the sum of the directional TIMD prediction values ​​determined at box 1942. In some examples, the resulting TIMD mode prediction value at 1946 is equivalent to the TIMD mode prediction value determined at box 1918.

[0255] At box 1948, the video codec determines the maximum value (or range) of the weight adjustments to be applied to the location-related weighting. In some examples, box 1948 corresponds to box 1915. In some examples, the weight adjustment range can be a fixed value. In some examples, the weight adjustment range can be determined based on the size of the current block and / or the image size / resolution, as described above regarding box 1915.

[0256] At box 1950, the video codec determines the TIMD mode prediction value based on the directional TIMD prediction value adjusted according to the position of the maximum value and the reference sample. The TIMD mode prediction value is applied sample-by-sample to the reference sample (e.g., the sample of the current template) to derive the prediction sample for the prediction block used for the current block, as explained above with respect to box 1910.

[0257] At box 1952, the video codec generates a prediction block for the current block based on the TIMD mode prediction value. In some examples, box 1952 may correspond to box 1922.

[0258] Figure 20 A flowchart 2000 illustrates an example method for utilizing TIMD technology with directed per-sample fusion for current block applications, according to some implementation schemes. The method of flowchart 2000 can be derived from a video codec (e.g., Figure 2 encoder 200 or Figure 3 The decoder 300 in the flowchart is implemented. The flowchart method can be performed reciprocally / equally by the encoder and decoder, so that it is not necessary for the encoder to signal the intra-prediction mode to the decoder to reconstruct the current block using intra-prediction techniques. In some examples, one or more operations of flowchart 2000 can be implemented in the following ways: Figure 19A Before 1903 or in the box Figure 19BThe box was executed before 1933.

[0259] At box 2002, the video codec determines the availability of one or more neighboring blocks for intra-frame prediction of the current block. For example, one or more neighboring blocks could be spatial candidate blocks adjacent to the current block, as described above. Figure 15A As described. Examples of these neighboring blocks can include blocks to the left, above, upper left, lower left, and upper right of the current block. In some examples, the video codec can determine the intra-prediction mode of one or more available neighboring blocks based on one or more of the available neighboring blocks.

[0260] As mentioned above Figure 17B As described, in order to perform TIMD, the video codec can determine a template and a reference to the template for the current block. For example, the template may include a template region defined by the positioning of a reconstructed region of the image from the current block relative to the current block. The template region may include a template region to the left of the current block and a template region above the current block. For example, the reference to the template may include a sample of regions defined by the positioning of the reconstructed region relative to the current block and / or the positioning of the template. In some examples, the template region and the template reference may be further based on which neighboring blocks are available.

[0261] At box 2012, based on the fact that one or more neighboring blocks are unavailable at box 2004 (e.g., based on the fact that none of the neighboring blocks of the current block are available), the video codec determines the TIMD mode prediction value based on the planar mode. For example, the TIMD mode prediction value can be determined as a planar mode.

[0262] At box 2006, based on the availability of one or more neighboring blocks at box 2004, the video codec determines one or more intra-prediction modes (IPMs) associated with the one or more available neighboring blocks. In some examples, the determination of one or more IPMs may be based on the TIMD applied to the encoding of the current block.

[0263] At box 2008, the video codec determines whether one or more IPMs are not angular intra-prediction modes (IPMs). At box 2014, based on the fact that one or more IPMs are not angular IPMs, the video codec determines the TIMD mode prediction value based on the non-angular IPMs among a plurality of non-angular IPMs. (The above text is incomplete and requires further context.) Figure 17B and Figures 19A to 19B An example of a non-angular IPM is described in the document.

[0264] In the example, the video codec determines the TIMD mode in response to (e.g., based on) at least one of one or more IPMs being a non-angled IPM among at least two different IPMs. In some examples, the video codec may select the non-angled IPM with the lowest cost among multiple non-angled IPMs (e.g., SAD, SATD, SSE, etc.) as the TIMD mode prediction.

[0265] In some examples, the video codec can determine whether the number of one or more IPMs that are not angled modes is greater than or equal to a threshold number. For example, the threshold number can be set based on the number of available neighboring blocks for the current block and / or the number of different IPMs among the available neighboring blocks. For example, if at least 3 neighboring blocks are available, the threshold number can be set to 2, where at least 2 modes of the IPMs in the available neighboring blocks are different. For example, if only 1 or 2 neighboring blocks are available, the threshold number can be set to 1.

[0266] At box 2010, the video codec determines multiple costs for multiple IPMs, including one or more IPMs (as referenced in boxes 2006 and 2008). In some examples, box 2010 may be related to... Figure 19A The corresponding box is 1902, and the video codec can perform the same or similar operations.

[0267] Figure 21 A flowchart 2100 illustrates an example method for utilizing TIMD technology with directed per-sample fusion for a current block application, according to some implementation schemes. The method of flowchart 2100 can be derived from a video codec (e.g., Figure 2 encoder 200 or Figure 3 The decoder 300 in the flowchart is implemented. The flowchart method can be performed reciprocally / identically by the encoder and decoder, so that it is not necessary for the encoder to signal the intra-prediction mode to the decoder to reconstruct the current block using intra-prediction techniques. In some examples, flowchart 2100 shows the implementation of the same method as the encoder. Figure 19A Boxes 1902 and 1904 and / or Figure 19B More detailed operations correspond to boxes 1932 and 1934.

[0268] At box 2102, the video codec determines multiple costs for multiple intra-prediction modes (IPMs) based on the TIMD being applied to the current block. In some examples, box 2102 is related to... Figure 19A Corresponding to box 1902, and the video codec can perform the same or similar operations. Box 2102 may include box 2104.

[0269] At box 2104, the video codec generates multiple IPMs, including candidate IPMs from a list of most probable modes (MPMs). In some examples, the multiple IPMs may also include DC mode, planar mode, vertical planar mode, horizontal planar mode, vertical DC mode, horizontal DC mode, or combinations thereof. Subsequently, a cost is determined for each IPM in the MPM list, and this cost corresponds to each of the multiple costs.

[0270] In some examples, an MPM list can be generated for each block for use in intra-prediction modes. For instance, to signal the intra-prediction mode (IPM), the encoder can encode the index indicating the selected IPM into the MPM list. Using an MPM list reduces the number of bits involved in the signaling index of the selected IPM. In some examples, the MPM list includes 22 candidate IPMs. For example, the MPM list includes a first part (e.g., 6 IPMs) of the candidate IPMs, referred to as the main MPM list. The main MPMs are the planar intra-prediction mode, the IPM from the left adjacent block, the IPM from the upper adjacent block, the IPM from the lower left adjacent block, the IPM from the upper right adjacent block, and the IPM from the upper left adjacent block, as per [reference to...]. Figure 15A As described. The MPM list may include a second part, known as the minor MPM list (e.g., the next 16 candidate IPMs), of candidate IPMs. The minor MPM list includes or consists of IPMs derived from offsets of IPMs in the major MPM list. In some examples, one or more decoder-side intra-frame mode derivation (DIMD) modes may be added to the MPM list after the major MPMs and before the minor MPMs in the final MPM list. In some examples, other IPMs not included in the MPM list are included in a separate non-MPM list.

[0271] At box 2106, the video codec determines the first TIMD mode and the second TIMD mode based on a first IPM and a second IPM having the lowest cost among multiple IPMs. In some examples, box 2106 is related to... Figure 19A Box 1904 and / or Figure 19B This corresponds to box 1934, and the video codec can perform the same or similar operations. Box 2106 may include boxes 2108 to 2120.

[0272] At box 2108, the video codec determines whether the wide-angle IPM is available for intra-frame prediction. In some examples, as shown by the dashed lines, the wide-angle IPM is not configured, and boxes 2108, 2112, and 2114 can be omitted from flowchart 2100.

[0273] At box 2116, based on the fact that the wide-angle IPM is unavailable or unused, the video codec determines that the first TIMD mode includes the first IPM (e.g., the first TIMD mode is determined to be the first IPM) and the second TIMD mode includes the second IPM (e.g., the second TIMD mode is determined to be the second IPM).

[0274] In some examples, based on the available wide-angle IPMs, multiple available angle IPMs can be extended from a first range (e.g., 67 modes) to a second range (e.g., 131 modes), and each of the first IPM and the second IPM can be refined (e.g., adjusted) to determine the first TIMD mode and the second TIMD mode.

[0275] At box 2112, based on the availability of a wide-angle IPM, the video codec determines the first TIMD mode as one of the following IPMs: the first IPM and one or more IPMs adjacent to the first IPM in the wide-angle IPM. For example, one or more IPMs can be determined as two adjacent IPM modes (e.g., + / -1 IPM) adjacent to the first IPM.

[0276] At box 2114, based on the availability of the wide-angle IPM, the video codec determines the second TIMD mode as one of the following IPMs: the second IPM and one or more IPMs adjacent to the second IPM in the wide-angle IPM. Boxes 2112 and 2114 can be performed in any order. For example, one or more IPMs can be determined as two adjacent IPM modes (e.g., + / -1 IPM) adjacent to the second IPM.

[0277] At box 2118, the video codec determines a first orientation of a first TIMD mode having a first cost based on: a sub-cost of the first cost for a first region of the current template; and a sub-cost of the first cost for a second region of the current template. At box 2120, the video codec similarly determines a second orientation of a second TIMD mode having a second cost based on: a sub-cost of the second cost for a first region of the current template; and a sub-cost of the second cost for a second region of the current template. In some examples, the orientation for each TIMD mode can be determined as described with respect to boxes 1942 and 1912-1914.

[0278] In some implementations, dynamic weights can be advantageously determined for TIMD modes mixed (e.g., combined or fused) during TIMD fusion. For example, the decoder may receive a bitstream including encoded video data and an indication of enabling / applying / selecting TIMD for the current block. Based on enabling TIMD for the current block, the decoder may determine the current template for the current block, where the current template includes a first region and a second region of reconstructed samples. For example, the first and second regions may correspond to samples to the left and above the current template, respectively.

[0279] In some examples, the decoder can determine multiple TIMD patterns (e.g., 2 or 3) to be combined to generate TIMD pattern predictions. For example, the number or limit of TIMD patterns can be set to a predefined number. N Each TIMD can be selected based on the cost calculated for each of the multiple IPMs applied to generate the predicted sample for the reference sample used for the current template (i.e., the current template sample). In some examples, the weight and orientation (e.g., location-related state) of each selected TIMD pattern can be determined based on comparing a first cost for a first region with a second cost for a second region. The orientation can indicate or correspond to the direction in the per-sample oriented prediction fusion process.

[0280] In some examples, the weights of the TIMD pattern used in the TIMD pattern prediction can be adjusted based on the sample's location, according to a weight adjustment range (e.g., maximum weight bias / value). In some examples, the weight adjustment range can be a fixed value. In some implementations, the weight adjustment range can be determined based on the current block size and / or image size / resolution.

[0281] In some examples, TIMD pattern predictions can be determined based on the mixing weights and orientation of each TIMD pattern in the TIMD pattern prediction. The decoder can then use the TIMD pattern predictions to generate prediction blocks for decoding the current block.

[0282] The embodiments of this disclosure can be implemented in hardware using analog and / or digital circuitry, in software, by executing instructions by one or more general-purpose or special-purpose processors, or as a combination of hardware and software. Therefore, embodiments of this disclosure can be implemented in a computer system or other processing system environment. An example of such a computer system 2200 is... Figure 22 As shown in the figures above. The blocks depicted in the figures above (such as...) Figure 1 , Figure 2 and Figure 3The blocks in the flowchart described in this disclosure can be executed on one or more computer systems 2200. Furthermore, each step in the flowchart described in this disclosure can be implemented on one or more computer systems 2200.

[0283] Computer system 2200 includes one or more processors, such as processor 2204. Processor 2204 may be, for example, a dedicated processor, a general-purpose processor, a microprocessor, or a digital signal processor. Processor 2204 may be connected to communication infrastructure 2202 (e.g., a bus or network). Computer system 2200 may also include main memory 2206, such as random access memory (RAM), and may also include secondary memory 2208.

[0284] Secondary storage 2208 may include, for example, a hard disk drive 2210 and / or a removable storage drive 2212, representing a magnetic tape drive, an optical disc drive, etc. The removable storage drive 2212 can read from and / or write to the removable storage unit 2216 in a well-known manner. The removable storage unit 2216 represents a magnetic tape, optical disc, etc., read from and written to by the removable storage drive 2212. Those skilled in the art will understand that the removable storage unit 2216 includes a computer-usable storage medium in which computer software and / or data are stored.

[0285] In alternative implementations, secondary storage 2208 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 2200. Such means may include, for example, removable storage unit 2218 and interface 2214. Examples of such means may include program boxes and box interfaces (such as those found in video game devices), removable storage chips (such as EPROM or PROM) and associated sockets, thumb drives and USB ports, and other removable storage units 2218 and interfaces 2214 that allow software and data to be transferred from removable storage unit 2218 to computer system 2200.

[0286] Computer system 2200 may also include communication interface 2220. Communication interface 2220 allows software and data to be transferred between computer system 2200 and external devices. Examples of communication interface 2220 may include modems, network interfaces (such as Ethernet cards), communication ports, etc. Software and data transmitted via communication interface 2220 are in the form of signals, which may be electronic signals, electromagnetic signals, optical signals, or other signals that can be received by communication interface 2220. These signals are provided to communication interface 2220 via communication path 2222. Communication path 2222 carries signals and may be implemented using wires or cables, optical fibers, telephone lines, cellular telephone links, RF links, and other communication channels.

[0287] As used herein, the terms "computer program medium" and "computer-readable medium" are used to refer to tangible storage media, such as removable storage units 2216 and 2218 or a hard disk installed in hard disk drive 2210. These computer program products are means for providing software to computer system 2200. The computer program (also referred to as computer control logic) may be stored in main memory 2206 and / or secondary storage 2208. The computer program may also be received via communication interface 2220. When executed, such a computer program enables computer system 2200 to implement the disclosures discussed herein. Specifically, when executed, the computer program enables processor 2204 to implement the processes of this disclosure, such as any of the methods described herein. Therefore, such a computer program represents a controller for computer system 2200.

[0288] In another embodiment, the features of this disclosure can be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. It will also be apparent to those skilled in the art that a hardware state machine is implemented to perform the functions described herein.

Claims

1. A method, the method comprising: Based on template-based intra-mode derivation (TIMD) applied to the current block, multiple costs of multiple intra-prediction modes (IPMs) applied to predict the template of the current block are determined. The TIMD mode is determined based on the IPM among the multiple IPMs; For each TIMD pattern in the TIMD patterns: The weights are determined based on the cost of the TIMD pattern; and The orientation of the TIMD pattern is determined by comparing the following: The first sub-cost of the cost for the first region of the template; and The cost is a second sub-cost for the second region of the template; as well as The TIMD pattern prediction value is determined based on a linear combination of the TIMD patterns having weights adjusted according to the orientation of the TIMD pattern. as well as A prediction block for the current block is generated based on the TIMD pattern prediction value.

2. The method of claim 1, wherein the IPM has the lowest cost among the plurality of costs.

3. The method of any one of claims 1 to 2, wherein the cost of each of the plurality of IPMs is based on the difference between the following: The current block's template is derived from a reference sample of the template and a predicted sample generated using the IPM; and The reconstructed sample of the template.

4. The method of any one of claims 1 to 3, wherein the cost includes the sum of absolute transformation differences (SATD) cost, the sum of squared errors (SSE) cost, or the sum of absolute differences (SAD) cost.

5. The method of any one of claims 3 to 4, wherein the reference of the template includes a reconstructed sample above or to the left of the template.

6. The method of claim 5, wherein the reference of the template includes a first number of rows of reconstructed samples above the template and a second number of columns of reconstructed samples to the left of the template.

7. The method of any one of claims 1 to 6, wherein the first region is associated with a first orientation, and the second region is associated with a second orientation, and wherein the orientation is selected from at least the first orientation or the second orientation.

8. The method of any one of claims 1 to 7, wherein the first region includes the reconstruction sample to the left of the current block, and the second region includes the reconstruction sample above the current block.

9. The method of any one of claims 7 to 8, wherein the first orientation is horizontal and the second orientation is vertical.

10. The method of any one of claims 7 to 9, wherein the orientation is further selected from the first orientation, the second orientation, and an orientation indicating a diagonal orientation without direction.

11. The method of any one of claims 7 to 10, wherein the orientation corresponds to the first region based on the fact that the first sub-cost is less than the second sub-cost for the second region.

12. The method of claim 11, wherein the orientation corresponds to the first region based on the first sub-cost being less than the second sub-cost multiplied by a scaling factor.

13. The method of claim 12, wherein the scaling factor is less than 1.

14. The method of any one of claims 1 to 13, wherein, prior to the comparison, the first sub-cost and the second sub-cost are respectively normalized to a first normalized sub-cost and a second normalized sub-cost.

15. The method of claim 14, wherein the first sub-cost is normalized to the first normalized sub-cost based on the number of samples in the first region, and wherein the second sub-cost is normalized to the second normalized sub-cost based on the number of samples in the second region.

16. The method of any one of claims 1 to 15, wherein the cost is equal to the sum of the first sub-cost and the second sub-cost.

17. The method of any one of claims 1 to 16, further comprising: The orientation of each TIMD pattern in the TIMD pattern is determined based on the width and height of the current block.

18. The method of claim 17, wherein determining the orientation of each TIMD pattern in the TIMD patterns is based on the current block size being greater than or equal to a size threshold, wherein the size is based on the width and the height.

19. The method of any one of claims 17 to 18, wherein the orientation of each TIMD pattern in the TIMD patterns is determined based on: The width of the current block is greater than or equal to a width threshold; and The height of the current block is greater than or equal to a height threshold.

20. The method of any one of claims 1 to 19, wherein the weights are further determined based on the sum of the costs of the TIMD modes.

21. The method of claim 20, wherein the weights used for the TIMD mode are inversely proportional to the cost of the TIMD mode.

22. The method of any one of claims 1 to 21, further comprising: The weight adjustment range for adjusting the weights used to adjust the TIMD mode is determined based on the orientation.

23. The method of claim 22, wherein the weight adjustment range is a predetermined value.

24. The method of any one of claims 22 to 23, wherein the weight adjustment range is determined based on the size of the current block.

25. The method of any one of claims 22 to 24, wherein the weight adjustment range is determined based on the resolution of the image including the current block.

26. The method of any one of claims 1 to 25, wherein generating the prediction block comprises applying the TIMD pattern prediction value to a reference sample in the template to determine a prediction sample in the prediction block.

27. The method of claim 26, wherein the weights of the TIMD patterns in the linear combination are adjusted based on the location of the reference sample and the orientation of the TIMD patterns.

28. The method of any one of claims 1 to 27, wherein the TIMD mode comprises a first TIMD mode and a second TIMD mode based on a first IPM and a second IPM having the lowest cost among the plurality of costs, respectively.

29. The method of claim 28, wherein determining the first TIMD mode based on the first IPM comprises: The first TIMD pattern is selected from the first IPM and the first neighboring IPMs of the first IPM in a second plurality of IPMs that are different from the plurality of IPMs, wherein the selection is based on the cost calculated for the first IPM and the two neighboring IPMs.

30. The method of claim 29, wherein determining the second TIMD mode based on the second IPM comprises: The second TIMD mode is selected from the second IPM and the second neighboring IPM of the second IPM in the second plurality of IPMs, wherein the selection is based on the cost calculated for the second IPM and the second neighboring IPM.

31. The method of any one of claims 1 to 30, the method further comprising receiving from a bitstream a first indication that an indication TIMD is applied to the current block, wherein the plurality of costs are determined based on the first indication.

32. The method of any one of claims 1 to 31, the method further comprising receiving from a bitstream a second indication indicating that a position-related weighting is applied to the TIMD modes, wherein determining the orientation of each TIMD mode in the TIMD modes is further based on the second indication.

33. The method of any one of claims 1 to 32, further comprising: Decode the residual of the current block from the bit stream; as well as The current block is reconstructed based on the predicted block and the residual.

34. The method of claim 33, wherein the reconstruction includes combining the prediction block with the residual.

35. The method of any one of claims 1 to 30, the method further comprising signaling in the bit stream a first indication that TIMD is applied to the current block, wherein the determination of the plurality of costs is based on the first indication.

36. The method of any one of claims 1 to 30 or 35, the method further comprising signaling in the bit stream a second indication indicating that position-related weighting is applied to the TIMD modes, wherein determining the orientation of each TIMD mode in the TIMD modes is further based on the second indication.

37. The method according to any one of claims 1 to 30 or 35 to 36, the method further comprising: The residual is determined based on the predicted block and the current block; as well as The residual is encoded into the bit stream.

38. The method of claim 37, wherein the residual is determined as the difference between the current block and the predicted block.

39. An encoder, the encoder comprising: One or more processors; as well as A memory for storing instructions that, when executed by the one or more processors, cause the encoder to perform the method as described in any one of claims 1 to 30 or 35 to 38.

40. A decoder, the decoder comprising: One or more processors; as well as A memory for storing instructions that, when executed by the one or more processors, cause the decoder to perform the method as described in any one of claims 1 to 34.

41. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a device, cause the device to perform the method as described in any one of claims 1 to 38.